Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

DeepSeek v4.1 Flash vs GLM 5.3 Flash I asked both models to recreate the Psychonauts 2 and Metroid Prime game menus as closely as possible to the originals. Attached are full-sized videos from both models for both games. All were done in a single shot, both at Max thinking,...

23,835 Aufrufe • vor 10 Tagen •via X (Twitter)

41 Kommentare

Profilbild von Ivan Fioravanti
Ivan Fioravantivor 10 Tagen

Yes GLM 5.3 Flash by far!

Profilbild von Mia
Miavor 10 Tagen

Yep, not even close !

Profilbild von Mia
Miavor 10 Tagen

The reference videos I gave both models:

Profilbild von Herlon Aguiar
Herlon Aguiarvor 10 Tagen

Slowly we will see the hype of DS4.1F go away 🙂

Profilbild von filipe
filipevor 10 Tagen

i'm surprised, ds4.1 flash is actually good on this benchmark

Profilbild von Mia
Miavor 10 Tagen

done terribly on this one

Profilbild von Dan Harper
Dan Harpervor 10 Tagen

GLM 5.3 Flash way better. Will have to try it

Profilbild von zoltan
zoltanvor 10 Tagen

Deepseek is always so overly verbose. 1M ctx window on ds is equivalent to ~300k on glm. I love both, ds is way more fun to talk to but i like glm for execution.

Profilbild von Fatih Yeşilova
Fatih Yeşilovavor 10 Tagen

Both models run on 4xspark? or for testing on cloud?

Profilbild von Mia
Miavor 10 Tagen

GLM on 2 sparks, DS on 3

Profilbild von Christopher Owen
Christopher Owenvor 10 Tagen

Have you tried the harnesses from the model providers? maybe there is some secret sauce in zcode or deepcode/harness?

Profilbild von J A Z I I
J A Z I Ivor 10 Tagen

jeez why is v4.1 flash using more tokens? it's cheap after all right?

Profilbild von Mia
Miavor 10 Tagen

It thinks way too much and over complicates things, and in the end the results are worse.

Profilbild von J A Z I I
J A Z I Ivor 10 Tagen

yeah that's the worst thing about v4.1

Profilbild von Young Sneed
Young Sneedvor 10 Tagen

its jus ta game menu

Profilbild von Vassio
Vassiovor 10 Tagen

A quick question, what harnes did you use?

Profilbild von Mia
Miavor 10 Tagen

My own harness. Unreleased.

Profilbild von Atma Blabber
Atma Blabbervor 10 Tagen

such test results these days are pretty much useless if you dont use the HARNESS that it is meant to be used with. glm with Zcode ds4.1 with DSH These results are with ur custom harness, and therefore are not only inaccurate but useless

Profilbild von Belcebuu
Belcebuuvor 10 Tagen

Not very impressed by 4.1 at all such a disappointment

Profilbild von Mia
Miavor 10 Tagen

On these cases, very disappointed

Profilbild von Belcebuu
Belcebuuvor 10 Tagen

It doubled the size and I don't see a big jump vs much smaller size ones like GLM 5.3 flash or Qwen3.8 next flash

Profilbild von Uncle cat bangkok
Uncle cat bangkokvor 10 Tagen

Thats amazing!

Profilbild von hi
hivor 10 Tagen

Hory chet

Profilbild von Random Libertarian Tech Lead
Random Libertarian Tech Leadvor 10 Tagen

I don't know, but I also know that this is a pretty worthless benchmark. Can we get serious about real useful benchmarks, rather than these worthless "do a bad one-shot of something useless, testing arbitrary trivial world knowledge" types of tests?

Profilbild von Inference User
Inference Uservor 10 Tagen

Putting both recreations side by side at max thinking is a clean comparison; identical prompts and retry counts would make the result even easier to reproduce.

Profilbild von PS5_BRO
PS5_BROvor 10 Tagen

GLM 5.3 Flash FTW

Profilbild von Egret
Egretvor 10 Tagen

whats the harness

Profilbild von Paul
Paulvor 10 Tagen

Have you tried the same prompt with qwen3.8 flash next? I gave it the task to make the koi-pond, using your prompt, and it did quite well.

Profilbild von Leo 的游戏人生
Leo 的游戏人生vor 10 Tagen

Token count is a cost, not a score. Burning 3–4x to land closer to the original isn't losing — it's paying more. Fidelity needs a rating, not a tally.Ran DS4.1F on real coding work this week. The token burn is real — but so is what comes out. Efficiency and capability aren't the same axis.

Profilbild von Kennan
Kennanvor 10 Tagen

Sure DSF4.1 used far more tokens, but it got far better looking results in my opinion(havn't played the games).

Profilbild von MASA
MASAvor 10 Tagen

Easily the lighter one did the other just overthink everything?

Profilbild von Paul Cleverly
Paul Cleverlyvor 10 Tagen

This type of example really brings into question just chasing the t/sec. Glm might be slower but like faster even at lower t/sec should probably include time taken with these comparisons.

Profilbild von Webster | JARVIS
Webster | JARVISvor 10 Tagen

GLM’s 3-4x token win matters. But menus are visual state machines, not text blobs. If 134k nails hover/transition fidelity, that’s a real architecture edge, not just brevity.

Profilbild von Carlo Pires
Carlo Piresvor 10 Tagen

That's my perception for complex coding tasks too.

Profilbild von jeremy.og
jeremy.ogvor 10 Tagen

Better than qwen3.8-flash-next?

Profilbild von momo 🇺🇸
momo 🇺🇸vor 10 Tagen

Both look really bad

Profilbild von basedcapital
basedcapitalvor 10 Tagen

a menu is all values, so every thinking step is another chance to swap one it saw for one it expected. whoever invented less won.

Profilbild von Dan C.
Dan C.vor 10 Tagen

I'm starting to think DS results are better with supervision/well defined tasks and less so for open ended problems.

Profilbild von AtxSal
AtxSalvor 10 Tagen

I think what I seen was leave my GLM build alone and wait for DS F 4.2.

Profilbild von 0x7z7z
0x7z7zvor 10 Tagen

I love these comparisons please make more of these!

Profilbild von Thomas Klein
Thomas Kleinvor 9 Tagen

I love seeing this stuff - but 90% of real world coding is fixing bugs / looking through repos, etc. not One shotting random websites and games. no real world engineer has stuff like this - they deal with edge cases, etc. thats what I am intersted in

Ähnliche Videos

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 Aufrufe • vor 29 Tagen

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,997 Aufrufe • vor 1 Monat

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,272 Aufrufe • vor 1 Monat

UPDATE: Charlie Kirk 🚨 Muzzle Flash: Second Shooter Location, Reflection, or Both? This video captures a flash on a window right as Charlie Kirk was shot. Let’s break this down… follow the numbers on the videos. #1. This video shows what appears to be a muzzle flash, just as Charlie is shot, leading people to believe the shooter was to Charlie’s right.. Notate the people to the left of the flash #2. This photo shows the location of the flash. Notate people to the left, on a middle-landing on a staircase. #3. This photo shows a clearer image showing both the middle-landing and the same location of the flash. It is a window. #4 . This video shows the building behind Charlie, which is a long hallway. This debunks any claims the shot came from here. — I zoom in where Charlie was — I zoom in on staircase — I zoom in on the window/flash — I zoom in on where suspected shooter was #5. I notate the flash reflection angles. — Video angle #1 notated — two reflection paths notated in relation to Video angle #1. 1. Reflection angle shows where we were told the shooter was. 2. Reflection angle shows mirrored angle, obviously where there is no reports of a shooter being in. 🔻 Final conclusion: This to me is very likely a muzzle flash reflection from a far off location. I am unsure of the reflection angle however, but this can be 100% proven if someone were to recreate the angles… maybe with a flash camera. If anyone is willing to, or has the means to do this (safely), this will either prove the shot came where the FBI said it came from, or it will prove there was a second shooter, in the direction of the mirrored angle.

MJTruthUltra

10,665,164 Aufrufe • vor 1 Jahr