Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

DeepSeek V4 Flash 0731 vs Opus 5. Only 7 days separated their releases. While Opus 5 accomplished in 1 shot what took DS 3, DeepSeek V4 Flash did it in ~900 lines vs ~3000 lines for Opus. Cost delta is stark. DeepSeek cost 1 cent. DS V4 Flash 0731...

41,980 görüntüleme • 27 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Dostlar yeni LLM Coding Benchmark çıktılarımız hazır. 🎉 Bu kez yine zorlu alanları bir araya getiren Trail ile Çift sarkaç: Euler vs RK4 entegrasyonu task'ını test ettim.Bu task ile modellerin hibrit; algoritma, matematik, coding ve Frontend yeteneklerini ölçümledim. Tüm modelleri opencode kullanarak kodlama yaptırdım. Aynı zamanda hepsinde en üst thinking eforu kullandım. Bu kez skorlamada Fiyat/Performans ve Kalite/Performans skalasında yaptım. Değerlendirmeyi modellerin isimlerini görmeden GPT 5.6 Pro yaptı. Testte yer alan Modeller; - DeepSeek V4 Flash 0731 - Gemini 3.6 Flash - Grok 4.5 - Sonnet 5 - GPT5.6-Luna - GPT5.6-Sol - Opus 5 - Kimi K3 - Qwen 3.8 Kalite/Performans - Genel sıralama - Opus 5 — 96.32/100 - $1.4660 - GPT5.6-Sol — 93.85/100 - $0.3627 - GPT5.6-Luna — 92.15/100 - $0.2373 - Kimi K3 — 91.70/100 - $0.3312 - Qwen3.8 — 89.95/100 - $0.2807 - Sonnet 5 — 88.16/100 - $0.4429 - DeepSeek V4 Flash 0731 — 85.12/100 - $0.0994 - Gemini 3.6 Flash — 83.24/100 - $0.0637 - Grok 4.5 — 79.82/100 - $0.4768 Fiyat/performans - Genel sıralama - DeepSeek V4 Flash - 88.00 - GPT5.6-Luna - 87.04 - Qwen 3.8 - 85.75 - Kimi K3 - 84.91 - Gemini 3.6 Flash - 83.57 - GPT5.6-Sol - 82.91 - Grok 4.5 - 81.87 - Sonnet 5 - 80.66 - Opus 5 - 80.22 Genel Değerlendirme: - Prod fiyat/performans kazananı: Qwen3.8. - Maksimum kalite, maliyet önemsiz: Opus 5. - En düşük bütçede yeterli ve tam özellikli çıktı: DeepSeek V4 Flash - En dengeli orta nokta: GPT5.6-Luna ve Kimi K3 Teknik Alan Spesifik Başarı: - En iyi saf numerical integrator: Opus 5 = GPT5.6-Sol - En doğru built-in validator: GPT5.6-Luna - En iyi standart runtime performansı: DeepSeek V4 Flash ve Kimi K3 - En iyi görsel kalite: Opus 5 - En iyi UI/ürün: GPT5.6-Sol - En iyi mimari: Opus 5 - En iyi ekonomik F/P: DeepSeek V4 Flash - En iyi production F/P: GPT5.6-Sol - En pahalı marjinal kalite artışı: Opus 5 - En ciddi validator hatası: Grok 4.5 - En yüksek hot-path/GC riski: Gemini 3.6 Flash

Alican Kiraz

14,814 görüntüleme • 26 gün önce

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,595 görüntüleme • 24 gün önce

gemini 3.7 flash vs deepseek v4 pro 0813 vs muse spark 1.2 – on voxel city dioramas three models each built three crossy road-style 3d scenes – a construction site, a nyc intersection, a river with a drawbridge – as single self-contained html files the setup: Nous Research's hermes agent cli on OpenRouter, three.js skills preloaded, identical prompts per scene tasks: 1. construction site – tower crane on a working lift loop, paver laying fresh road, roller compacting it behind 2. nyc crossing – four-way intersection with a traffic light state machine, queuing cars, pedestrians crossing on the walk signal 3. river drawbridge – double-leaf bascule that lifts for tall boats, cars queuing at the barriers, animated water every scene: Three.js r185, box geometry only, a locked 20-color palette, four camera presets, and a day/night mode with bloom. one file, no build step, no assets models: Google DeepMind gemini 3.7 flash, DeepSeek v4 pro 0813, AI at Meta muse spark 1.2 muse and gemini finished every scene in two to three minutes. deepseek took 15 to 41 minutes per scene - build time, all three scenes #1 gemini 3.7 flash – 6m 43s #2 muse spark 1.2 – 7m 20s #3 deepseek v4 pro – 91m 25s - total tokens #1 muse spark 1.2 – 440,279 #2 gemini 3.7 flash – 713,855 #3 deepseek v4 pro – 20,957,568 - total price #1 muse spark 1.2 – $0.53 #2 gemini 3.7 flash – $0.56 #3 deepseek v4 pro – $4.57 - agent calls across the three builds #1 muse spark 1.2 – 12 #2 gemini 3.7 flash – 18 #3 deepseek v4 pro – 143 observations: • muse won two of the three scenes on looks with the smallest files in the test – 887 to 1,042 lines against gemini's 1,934 to 2,377. cheapest, fastest to a good frame, and shortest turned out to be the same column • deepseek burned 20.96m tokens – 29x gemini, 48x muse – across 143 agent calls. prompt caching is the only reason that cost $4.57: the cache discount absorbed roughly $30 of resent context • gemini was the only model whose files needed zero fixes to render – and the only one whose night mode is cosmetic. the sky never darkens and one camera button does nothing. clean code for a scene it never looked at follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

28,407 görüntüleme • 10 gün önce

Advanced Frontier LLM Coding Benchmark 13.08.2026 Modeller: - Grok 4.6 Extra High - Cursor - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - Kimi K3 Max - Kimi-Code (App) - GLM 5.2 Max - Zcode (App) - GPT 5.6 Sol Very High - Codex (App) - Deepseek-v4-flash-0731 - Qwen-Code (CLI) - Cursor Auto - Cursor (App) - Deepseek-v4-pro-0813 - Opencode (CLI) Görev: Ferrofluid: Metaball yüzeyi + mıknatıs imlecine doğru yükselen Rosensweig dikenleri. Ferrofluid'in zorluğu, "imlece doğru akan sıvı" görünümünün arkasındaki fiziğin bir eşik/bifurkasyon fenomeni olması modellerin %90'ının ıskaladığı nokta. Not: Bu test oldukça zorlu bir yapıda ve frontier benchmark temeldedir. Zorluk katmanları: - Fiziksel modeli doğru anlayıp denklemlere dönüştürme - Kararlı bir SPH/PBF çözücü geliştirme - Manyetik alan, yüzey gerilimi ve histerezis davranışını modelleme - Metaball ve marching-squares ile yüzey çıkarımı - Ölçülebilir ve dürüst bir self-grading sistemi oluşturma - Tarayıcıda yeterli performans sağlama - Görsel olarak anlaşılır bir ürün yüzeyi hazırlama Modellerin Kullanım miktarları ve süreleri: - GPT 5.6 - 57m 6s - $82.1 - Opus 5 - 1s 36m - $106.6 - GLM 5.2 - 1h 2m - $35.7 - Kimi K3 - 1h 23m - $57.3 - Qwen3.8 - 48m 1s - $18 - Deepseek-v4-flash-0731 - 1h 25d - $0.79 - Cursor Auto - 25m 15s - $8-9 - Grok 4.6 - 39m - $5.1 - Deepseek-v4-pro-0813 - 2h 51m - $1.92 Modellerin Performans Tablosu: - Maliyet + Performans Kazanı: Qwen3.8 - Saf teknik ve fiziksel kalite kazananı: Opus 5 - En iyi ultra-düşük maliyetli mühendislik çekirdeği: DeepSeek V4 Pro 0813 - En iyi frontend, ürünleştirme ve gözlemlenebilirlik katmanı: GPT-5.6 Sol - En iyi düşük maliyetli matematiksel eleştirmen: Grok 4.6 - Performans ve sadeleştirme için en uygun yardımcı model: GLM 5.2 Bağımsız saf teknik başarı sıralaması: - Opus 5: 86/100 - Düzenli ve gerçek spike trenine en yakın sonuç. - Qwen3.8: 79/100 - Spike oluşuyor fakat yüksek alanda parçalanma artıyor. - DeepSeek V4 Pro-0813: 64/100 - Gelişmiş çekirdek, fakat ciddi indeksleme riski ve parçalanma - GPT-5.6 Sol: 63/100 - En iyi ürün yüzeyi; fizik çekirdeğinde damlacıklaşma - GLM 5.2: 57/100 - Çok kararlı ve hızlı, fakat tek geniş kubbe üretiyor - Grok 4.6: 53/100 - Matematiksel olarak iyi düşünülmüş, tarayıcı performansı çok düşük - Cursor Auto: 47/100 - Hızlı üretim; garip ince kolonlar, parçalanma ve düşük runtime FPS - Kimi K3: 42/100 - Özellik kapsamı geniş; fiziksel sonuç ve skor güvenilirliği zayıf - DeepSeek V4 Flash-0731: 32/100 - En ucuz; varsayılan görüntü kullanılamaz ve ölçümler yanıltıcı Nasıl kullanmalıyız: - Tek model seçeceksem: Qwen3.8 - Bütçe önemsiz ve en iyi teknik sonucu istiyorsam: Opus 5 - Çok düşük bütçeyle güçlü bir çekirdek istiyorsam: DeepSeek V4 Pro - Sadece frontend ve ürün kalitesi istiyorsam: GPT-5.6 Sol - Formül ve algoritma eleştirmeni istiyorsam: Grok 4.6 - Optimization/refactor istiyorsam: GLM 5.2 - Visual QA istiyorsam: Kimi K3 - Cursor Auto’yu model olarak değil, sabit modelleri yöneten agent harness olarak kullanırım - DeepSeek Flash’ı zorlu caseler dışında, harness yaparak çözeceksem ve disposable boilerplate ile test üretiminde kullanırım.

Alican Kiraz

12,564 görüntüleme • 15 gün önce