Loading video...

Video Failed to Load

Go Home

DeepSeek V4 Flash 0731 vs Opus 5. Only 7 days separated their releases. While Opus 5 accomplished in 1 shot what took DS 3, DeepSeek V4 Flash did it in ~900 lines vs ~3000 lines for Opus. Cost delta is stark. DeepSeek cost 1 cent. DS V4 Flash 0731...

43,797 views • 1 month ago •via X (Twitter)

18 Comments

sxcpc's profile picture
sxcpc1 month ago

Looks like an ad for opus despite the cost, game looks that clean

Blue - Klarname (Olaf Merz)'s profile picture
Blue - Klarname (Olaf Merz)1 month ago

And most important. You can run DSV 4 Flash 0731 @Home.

AI Mastery Guide's profile picture
AI Mastery Guide1 month ago

900 lines versus 3000 for the same result is a massive efficiency gap.

Zarko Rashev's profile picture
Zarko Rashev1 month ago

Wait for Pro

Maximus's profile picture
Maximus1 month ago

the real question is whether the 900-line solution actually handles edge cases or just looks cleaner on a happy path

Jeremy Chone's profile picture
Jeremy Chone1 month ago

The apples are Luna vs v4 flash

Tushar Koshti's profile picture
Tushar Koshti1 month ago

Who wins over here #Opus or #deepseek

Victor's profile picture
Victor1 month ago

Deepseek lagging like crazy.

David Hendrickson's profile picture
David Hendrickson1 month ago

Actually it wasn’t but DeepSeek V4 added an intentional shudder to the screen when an asteroid was hit so the play looked jittery

Victor's profile picture
Victor1 month ago

Most vibe coded games look like they run on 5-10 fps.

Vito Botta's profile picture
Vito Botta1 month ago

I'd want to inspect both outputs before calling a winner. Did the shorter one stay readable after follow-up edits?

Eric Kang's profile picture
Eric Kang1 month ago

One-shot success is meaningful, but line count alone can mislead. The Responses API walkthrough:

Barrios's profile picture
Barrios1 month ago

amazing! deepseek v4 flash!

Rusti's profile picture
Rusti1 month ago

So what were all the extra lines of code that Opus added?

Sergey Stolyarov's profile picture
Sergey Stolyarov1 month ago

I won’t be using it even for free.. hardly worth risking breaches, hacks ., who will be responsible ? Better say where

Jay L. Kerschner's profile picture
Jay L. Kerschner1 month ago

time is money friends

Oleksandr's profile picture
Oleksandr1 month ago

efficiency win. 900 lines at $0.01 beats 3000 lines at $1.00 value per line skyrockets

Loong🐉's profile picture
Loong🐉1 month ago

The 7-day gap between V4 Flash and Opus 5 is the real story. 1 cent vs 5x that means every closed-model vendor is now selling a marketing premium, not a math one. Either match cents-per-task or pivot to features the open stack can't ship.

Related Videos

Dostlar yeni LLM Coding Benchmark çıktılarımız hazır. 🎉 Bu kez yine zorlu alanları bir araya getiren Trail ile Çift sarkaç: Euler vs RK4 entegrasyonu task'ını test ettim.Bu task ile modellerin hibrit; algoritma, matematik, coding ve Frontend yeteneklerini ölçümledim. Tüm modelleri opencode kullanarak kodlama yaptırdım. Aynı zamanda hepsinde en üst thinking eforu kullandım. Bu kez skorlamada Fiyat/Performans ve Kalite/Performans skalasında yaptım. Değerlendirmeyi modellerin isimlerini görmeden GPT 5.6 Pro yaptı. Testte yer alan Modeller; - DeepSeek V4 Flash 0731 - Gemini 3.6 Flash - Grok 4.5 - Sonnet 5 - GPT5.6-Luna - GPT5.6-Sol - Opus 5 - Kimi K3 - Qwen 3.8 Kalite/Performans - Genel sıralama - Opus 5 — 96.32/100 - $1.4660 - GPT5.6-Sol — 93.85/100 - $0.3627 - GPT5.6-Luna — 92.15/100 - $0.2373 - Kimi K3 — 91.70/100 - $0.3312 - Qwen3.8 — 89.95/100 - $0.2807 - Sonnet 5 — 88.16/100 - $0.4429 - DeepSeek V4 Flash 0731 — 85.12/100 - $0.0994 - Gemini 3.6 Flash — 83.24/100 - $0.0637 - Grok 4.5 — 79.82/100 - $0.4768 Fiyat/performans - Genel sıralama - DeepSeek V4 Flash - 88.00 - GPT5.6-Luna - 87.04 - Qwen 3.8 - 85.75 - Kimi K3 - 84.91 - Gemini 3.6 Flash - 83.57 - GPT5.6-Sol - 82.91 - Grok 4.5 - 81.87 - Sonnet 5 - 80.66 - Opus 5 - 80.22 Genel Değerlendirme: - Prod fiyat/performans kazananı: Qwen3.8. - Maksimum kalite, maliyet önemsiz: Opus 5. - En düşük bütçede yeterli ve tam özellikli çıktı: DeepSeek V4 Flash - En dengeli orta nokta: GPT5.6-Luna ve Kimi K3 Teknik Alan Spesifik Başarı: - En iyi saf numerical integrator: Opus 5 = GPT5.6-Sol - En doğru built-in validator: GPT5.6-Luna - En iyi standart runtime performansı: DeepSeek V4 Flash ve Kimi K3 - En iyi görsel kalite: Opus 5 - En iyi UI/ürün: GPT5.6-Sol - En iyi mimari: Opus 5 - En iyi ekonomik F/P: DeepSeek V4 Flash - En iyi production F/P: GPT5.6-Sol - En pahalı marjinal kalite artışı: Opus 5 - En ciddi validator hatası: Grok 4.5 - En yüksek hot-path/GC riski: Gemini 3.6 Flash

Alican Kiraz

15,102 views • 1 month ago

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,470 views • 1 month ago

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 views • 1 month ago