Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Ran FlappyBench on DeepSeek V4.1-Flash, GLM 5.3, and Kimi K3 with the same /design prompt. 🔹 DSV4.1-Flash: 9/10 · $0.0089 one-shot with the lowest cost 🔹 GLM-5.3: 9.5/10 · $0.0184 · nailed the UI but cost 2x more 🔹 Kimi K3: 8/10 · $0.0740 · hard-coded gameplay and 8x more

18,397 Aufrufe • vor 2 Tagen •via X (Twitter)

17 Kommentare

Profilbild von Command Code
Command Codevor 2 Tagen

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here:

Profilbild von Naveen Saradhi
Naveen Saradhivor 2 Tagen

Chinese model are doing great

Profilbild von Rick A.F.
Rick A.F.vor 2 Tagen

I remember when the cost was in full $. Now it’s so little we have to use different currency soon 😂

Profilbild von Wahed Sikder
Wahed Sikdervor 2 Tagen

GLM looks sharp

Profilbild von Naymur Rahman
Naymur Rahmanvor 2 Tagen

Fastest and most efficient model so far by @deepseek_ai

Profilbild von shahx
shahxvor 2 Tagen

n=10, one prompt: 9 vs 9.5 is inside noise, so the real result is 'all three pass'. And $0.0089 is one-shot spend, not cost per accepted run - log the failed attempt too. Retries are where the 8x gap shrinks. My 85-task run: $9 base -> $13 with 30% retries.

Profilbild von Maedah Batool
Maedah Batoolvor 2 Tagen

Our very own FlappyBench 😉

Profilbild von Paul · SpellWright
Paul · SpellWrightvor 2 Tagen

The cost gap is the interesting result, but I’d want to see the same prompt rerun across seeds before calling it a win. Did you keep the tool and model settings identical besides the model?

Profilbild von Pablo
Pablovor 2 Tagen

I have the GOAT plan, I ran out of weekly credits, I bought additional credits, but now I realize I can't use them until the end of the month... Why? It would be more practical to be able to use them when I need them, like today...

Profilbild von balega_dev
balega_devvor 2 Tagen

I honestly think this is one of the most useless bench, but one thing is clear, kimi somehow is uniquely good in every benchmark and its impressive

Profilbild von san zhang
san zhangvor 2 Tagen

看来gemini连flash模型都要out了

Profilbild von Webster | JARVIS
Webster | JARVISvor 2 Tagen

DSV4.1-Flash is the rational pick: 9/10 at $0.0089. GLM-5.3's +0.5 at 2x cost is demo polish, not production value. Kimi's hard-coded gameplay at 8x is the tell.

Profilbild von Usman
Usmanvor 2 Tagen

please improve your permissions settings and auto approve settings. ........................................... very problematic

Profilbild von Sabda
Sabdavor 2 Tagen

why is ds v4.1 flash not available on /model list?

Profilbild von Command Code
Command Codevor 2 Tagen

Update CLI

Profilbild von Guacamole
Guacamolevor 2 Tagen

请修复issue#829

Profilbild von chance
chancevor 2 Tagen

Same prompt, published cost. 9/10 at $0.009 is the routing argument — leaderboards without unit economics are vibes.

Ähnliche Videos

glm 5.3 flash is 7.5x cheaper, but 3.4x slower than gemini 3.7 flash Z.ai glm 5.3 flash – shipped aug 26, $0.07/$0.25 per 1m Google DeepMind gemini 3.7 flash – shipped aug 13, $0.38/$1.88 per 1m we put the two models on one job: write one html file that draws an animated 3d scene in the browser. no images, no downloads, and it has to look the same on every load. the setup: three scenes – a glass aquarium in a lit room, the solar system, a night city under a thunderstorm. identical brief word for word, reasoning effort high, 64k output cap. the numbers below are not the whole run. they cover the three scenes we kept – the best one per task from each model, the ones in the video. - total generation time for the three scenes #1 gemini 3.7 flash – 10m 36s #2 glm 5.3 flash – 36m 30s - tokens spent on those three scenes #1 glm 5.3 flash – 110k #2 gemini 3.7 flash – 111k - cost of those three scenes #1 glm 5.3 flash – $0.027 #2 gemini 3.7 flash – $0.202 observations: • glm's first 10 attempts: 7 blank pages. it kept inventing short random helpers and forgetting to define one of them. the fix was one line in the brief: use exactly one random helper, named rand(), and don't invent shorthands next to it. next 12 attempts: 11 alive, 0 crashes. • glm spends 66% of its output on reasoning, gemini 57%. that is the whole speed gap. • gemini's storm came back as a black rectangle in 4 of 6 runs. glm's best storm has a branching bolt, lit rain and wet asphalt – for $0.01. conclusion: same three scenes, same token spend – glm 5.3 flash billed $0.027 and took 36m 30s, gemini 3.7 flash billed $0.202 and took 10m 36s. glm wins gemini on price and made the best storm of the whole run follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

15,983 Aufrufe • vor 18 Tagen

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 Aufrufe • vor 16 Tagen