正在加载视频...
视频加载失败
DeepSeek v4.1 Flash vs GLM 5.3 Flash I asked both models to recreate the Psychonauts 2 and Metroid Prime game menus as closely as possible to the originals. Attached are full-sized videos from both models for both games. All were done in a single shot, both at Max thinking,... show more
23,835 次观看 • 10 天前 •via X (Twitter)
41 条评论

Yes GLM 5.3 Flash by far!

Yep, not even close !

The reference videos I gave both models:

Slowly we will see the hype of DS4.1F go away 🙂

i'm surprised, ds4.1 flash is actually good on this benchmark

done terribly on this one

GLM 5.3 Flash way better. Will have to try it

Deepseek is always so overly verbose. 1M ctx window on ds is equivalent to ~300k on glm. I love both, ds is way more fun to talk to but i like glm for execution.

Both models run on 4xspark? or for testing on cloud?

GLM on 2 sparks, DS on 3

Have you tried the harnesses from the model providers? maybe there is some secret sauce in zcode or deepcode/harness?

jeez why is v4.1 flash using more tokens? it's cheap after all right?

It thinks way too much and over complicates things, and in the end the results are worse.

yeah that's the worst thing about v4.1

its jus ta game menu

A quick question, what harnes did you use?

My own harness. Unreleased.

such test results these days are pretty much useless if you dont use the HARNESS that it is meant to be used with. glm with Zcode ds4.1 with DSH These results are with ur custom harness, and therefore are not only inaccurate but useless

Not very impressed by 4.1 at all such a disappointment

On these cases, very disappointed

It doubled the size and I don't see a big jump vs much smaller size ones like GLM 5.3 flash or Qwen3.8 next flash

Thats amazing!

Hory chet

I don't know, but I also know that this is a pretty worthless benchmark. Can we get serious about real useful benchmarks, rather than these worthless "do a bad one-shot of something useless, testing arbitrary trivial world knowledge" types of tests?

Putting both recreations side by side at max thinking is a clean comparison; identical prompts and retry counts would make the result even easier to reproduce.

GLM 5.3 Flash FTW

whats the harness

Have you tried the same prompt with qwen3.8 flash next? I gave it the task to make the koi-pond, using your prompt, and it did quite well.

Token count is a cost, not a score. Burning 3–4x to land closer to the original isn't losing — it's paying more. Fidelity needs a rating, not a tally.Ran DS4.1F on real coding work this week. The token burn is real — but so is what comes out. Efficiency and capability aren't the same axis.

Sure DSF4.1 used far more tokens, but it got far better looking results in my opinion(havn't played the games).

Easily the lighter one did the other just overthink everything?

This type of example really brings into question just chasing the t/sec. Glm might be slower but like faster even at lower t/sec should probably include time taken with these comparisons.

GLM’s 3-4x token win matters. But menus are visual state machines, not text blobs. If 134k nails hover/transition fidelity, that’s a real architecture edge, not just brevity.

That's my perception for complex coding tasks too.

Better than qwen3.8-flash-next?

Both look really bad

a menu is all values, so every thinking step is another chance to swap one it saw for one it expected. whoever invented less won.

I'm starting to think DS results are better with supervision/well defined tasks and less so for open ended problems.

I think what I seen was leave my GLM build alone and wait for DS F 4.2.

I love these comparisons please make more of these!

I love seeing this stuff - but 90% of real world coding is fixing bugs / looking through repos, etc. not One shotting random websites and games. no real world engineer has stuff like this - they deal with edge cases, etc. thats what I am intersted in
