Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

DeepSeek V4 Flash 0731 vs Opus 5. Only 7 days separated their releases. While Opus 5 accomplished in 1 shot what took DS 3, DeepSeek V4 Flash did it in ~900 lines vs ~3000 lines for Opus. Cost delta is stark. DeepSeek cost 1 cent. DS V4 Flash 0731...

43,797 Aufrufe • vor 1 Monat •via X (Twitter)

18 Kommentare

Profilbild von sxcpc
sxcpcvor 1 Monat

Looks like an ad for opus despite the cost, game looks that clean

Profilbild von Blue - Klarname (Olaf Merz)
Blue - Klarname (Olaf Merz)vor 1 Monat

And most important. You can run DSV 4 Flash 0731 @Home.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

900 lines versus 3000 for the same result is a massive efficiency gap.

Profilbild von Zarko Rashev
Zarko Rashevvor 1 Monat

Wait for Pro

Profilbild von Maximus
Maximusvor 1 Monat

the real question is whether the 900-line solution actually handles edge cases or just looks cleaner on a happy path

Profilbild von Jeremy Chone
Jeremy Chonevor 1 Monat

The apples are Luna vs v4 flash

Profilbild von Tushar Koshti
Tushar Koshtivor 1 Monat

Who wins over here #Opus or #deepseek

Profilbild von Victor
Victorvor 1 Monat

Deepseek lagging like crazy.

Profilbild von David Hendrickson
David Hendricksonvor 1 Monat

Actually it wasn’t but DeepSeek V4 added an intentional shudder to the screen when an asteroid was hit so the play looked jittery

Profilbild von Victor
Victorvor 1 Monat

Most vibe coded games look like they run on 5-10 fps.

Profilbild von Vito Botta
Vito Bottavor 1 Monat

I'd want to inspect both outputs before calling a winner. Did the shorter one stay readable after follow-up edits?

Profilbild von Eric Kang
Eric Kangvor 1 Monat

One-shot success is meaningful, but line count alone can mislead. The Responses API walkthrough:

Profilbild von Barrios
Barriosvor 1 Monat

amazing! deepseek v4 flash!

Profilbild von Rusti
Rustivor 1 Monat

So what were all the extra lines of code that Opus added?

Profilbild von Sergey Stolyarov
Sergey Stolyarovvor 1 Monat

I won’t be using it even for free.. hardly worth risking breaches, hacks ., who will be responsible ? Better say where

Profilbild von Jay L. Kerschner
Jay L. Kerschnervor 1 Monat

time is money friends

Profilbild von Oleksandr
Oleksandrvor 1 Monat

efficiency win. 900 lines at $0.01 beats 3000 lines at $1.00 value per line skyrockets

Profilbild von Loong🐉
Loong🐉vor 1 Monat

The 7-day gap between V4 Flash and Opus 5 is the real story. 1 cent vs 5x that means every closed-model vendor is now selling a marketing premium, not a math one. Either match cents-per-task or pivot to features the open stack can't ship.

Ähnliche Videos

Dostlar yeni LLM Coding Benchmark çıktılarımız hazır. 🎉 Bu kez yine zorlu alanları bir araya getiren Trail ile Çift sarkaç: Euler vs RK4 entegrasyonu task'ını test ettim.Bu task ile modellerin hibrit; algoritma, matematik, coding ve Frontend yeteneklerini ölçümledim. Tüm modelleri opencode kullanarak kodlama yaptırdım. Aynı zamanda hepsinde en üst thinking eforu kullandım. Bu kez skorlamada Fiyat/Performans ve Kalite/Performans skalasında yaptım. Değerlendirmeyi modellerin isimlerini görmeden GPT 5.6 Pro yaptı. Testte yer alan Modeller; - DeepSeek V4 Flash 0731 - Gemini 3.6 Flash - Grok 4.5 - Sonnet 5 - GPT5.6-Luna - GPT5.6-Sol - Opus 5 - Kimi K3 - Qwen 3.8 Kalite/Performans - Genel sıralama - Opus 5 — 96.32/100 - $1.4660 - GPT5.6-Sol — 93.85/100 - $0.3627 - GPT5.6-Luna — 92.15/100 - $0.2373 - Kimi K3 — 91.70/100 - $0.3312 - Qwen3.8 — 89.95/100 - $0.2807 - Sonnet 5 — 88.16/100 - $0.4429 - DeepSeek V4 Flash 0731 — 85.12/100 - $0.0994 - Gemini 3.6 Flash — 83.24/100 - $0.0637 - Grok 4.5 — 79.82/100 - $0.4768 Fiyat/performans - Genel sıralama - DeepSeek V4 Flash - 88.00 - GPT5.6-Luna - 87.04 - Qwen 3.8 - 85.75 - Kimi K3 - 84.91 - Gemini 3.6 Flash - 83.57 - GPT5.6-Sol - 82.91 - Grok 4.5 - 81.87 - Sonnet 5 - 80.66 - Opus 5 - 80.22 Genel Değerlendirme: - Prod fiyat/performans kazananı: Qwen3.8. - Maksimum kalite, maliyet önemsiz: Opus 5. - En düşük bütçede yeterli ve tam özellikli çıktı: DeepSeek V4 Flash - En dengeli orta nokta: GPT5.6-Luna ve Kimi K3 Teknik Alan Spesifik Başarı: - En iyi saf numerical integrator: Opus 5 = GPT5.6-Sol - En doğru built-in validator: GPT5.6-Luna - En iyi standart runtime performansı: DeepSeek V4 Flash ve Kimi K3 - En iyi görsel kalite: Opus 5 - En iyi UI/ürün: GPT5.6-Sol - En iyi mimari: Opus 5 - En iyi ekonomik F/P: DeepSeek V4 Flash - En iyi production F/P: GPT5.6-Sol - En pahalı marjinal kalite artışı: Opus 5 - En ciddi validator hatası: Grok 4.5 - En yüksek hot-path/GC riski: Gemini 3.6 Flash

Alican Kiraz

15,102 Aufrufe • vor 1 Monat

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,470 Aufrufe • vor 1 Monat

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 Aufrufe • vor 1 Monat