Загрузка видео...

Не удалось загрузить видео

На главную

Deepseek-V4-Pro-0813 vs Grok-4.6 Flappybird comparison test. DS : 22,848tok, $0.019 total Grok : 5,211tok, $0.030 total Which one do you think is better? 🤔

92,106 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

ox alpha vs deepseek v4 flash vision vs grok 4.6 vs gemini 3.7 flash vs – on photo-to-3d four vision models got one photograph each and had to rebuild the place inside it as a Three.js scene. twelve scenes, twelve first-try runs, zero console errors the setup: one reference photo per scene, sent as an image on OpenRouter. the prompt never says what is in the picture – no "motel", no "bar", no "gas station". the model has to read the photo and rebuild it: layout, materials, hour of the day, and whatever is around the corner that the frame does not show tasks – three photographs of early-2000s america: 1. a motel at night, neon pylon lit, snow on the ground 2. an old new york tavern interior, tin ceiling, tiled floor 3. an abandoned service station in the california desert, midday sun each scene ships as one self-contained html file, procedural geometry and canvas textures only, no downloads. three timed camera shots, and shot 1 has to reproduce the framing of the reference photo models: xAI grok 4.6, Google DeepMind gemini 3.7 flash, DeepSeek deepseek v4 flash vision exp, and ox alpha – a stealth model on openrouter, free, no lab attached to it yet results: - wall clock, three scenes #1 gemini 3.7 flash – 11m 12s #2 deepseek v4 flash – 15m 20s #3 grok 4.6 – 28m 11s #4 ox alpha – 38m 54s - output tokens #1 gemini 3.7 flash – 77,396 #2 ox alpha – 87,613 #3 grok 4.6 – 105,687 #4 deepseek v4 flash – 127,884 - lines of code shipped #1 ox alpha – 2,090 #2 deepseek v4 flash – 2,291 #3 grok 4.6 – 3,529 #4 gemini 3.7 flash – 3,989 - total price #1 ox alpha – $0.000 #2 deepseek v4 flash – $0.091 #3 gemini 3.7 flash – $0.136 #4 grok 4.6 – $0.697 observations: • grok is 7.7x the price of deepseek. it is the only model that read the light – low sun, real shadows on the station, a cold night on the motel • gemini is the fastest and the least deliberate. 17,158 reasoning tokens against deepseek's 99,172, and it still shipped the most code – 3,989 lines • deepseek thought hardest and rendered plainest. 99,172 reasoning tokens, 5.8x gemini's, spent on layout rather than on light. its motel is the second best in the set for $0.030 • ox alpha is free and reads a photo as well as anything here – it lifted "family units / kitchenettes" off the pylon and redrew it in canvas conclusion: twelve scenes, four models, zero fixes, and the whole run cost $0.924! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

24,202 просмотров • 22 дней назад

Advanced Frontier LLM Coding Benchmark 13.08.2026 Modeller: - Grok 4.6 Extra High - Cursor - Opus 5 Max - Claude Code (App) - Qwen 3.8 Max - Qwen-Code (CLI) - Kimi K3 Max - Kimi-Code (App) - GLM 5.2 Max - Zcode (App) - GPT 5.6 Sol Very High - Codex (App) - Deepseek-v4-flash-0731 - Qwen-Code (CLI) - Cursor Auto - Cursor (App) - Deepseek-v4-pro-0813 - Opencode (CLI) Görev: Ferrofluid: Metaball yüzeyi + mıknatıs imlecine doğru yükselen Rosensweig dikenleri. Ferrofluid'in zorluğu, "imlece doğru akan sıvı" görünümünün arkasındaki fiziğin bir eşik/bifurkasyon fenomeni olması modellerin %90'ının ıskaladığı nokta. Not: Bu test oldukça zorlu bir yapıda ve frontier benchmark temeldedir. Zorluk katmanları: - Fiziksel modeli doğru anlayıp denklemlere dönüştürme - Kararlı bir SPH/PBF çözücü geliştirme - Manyetik alan, yüzey gerilimi ve histerezis davranışını modelleme - Metaball ve marching-squares ile yüzey çıkarımı - Ölçülebilir ve dürüst bir self-grading sistemi oluşturma - Tarayıcıda yeterli performans sağlama - Görsel olarak anlaşılır bir ürün yüzeyi hazırlama Modellerin Kullanım miktarları ve süreleri: - GPT 5.6 - 57m 6s - $82.1 - Opus 5 - 1s 36m - $106.6 - GLM 5.2 - 1h 2m - $35.7 - Kimi K3 - 1h 23m - $57.3 - Qwen3.8 - 48m 1s - $18 - Deepseek-v4-flash-0731 - 1h 25d - $0.79 - Cursor Auto - 25m 15s - $8-9 - Grok 4.6 - 39m - $5.1 - Deepseek-v4-pro-0813 - 2h 51m - $1.92 Modellerin Performans Tablosu: - Maliyet + Performans Kazanı: Qwen3.8 - Saf teknik ve fiziksel kalite kazananı: Opus 5 - En iyi ultra-düşük maliyetli mühendislik çekirdeği: DeepSeek V4 Pro 0813 - En iyi frontend, ürünleştirme ve gözlemlenebilirlik katmanı: GPT-5.6 Sol - En iyi düşük maliyetli matematiksel eleştirmen: Grok 4.6 - Performans ve sadeleştirme için en uygun yardımcı model: GLM 5.2 Bağımsız saf teknik başarı sıralaması: - Opus 5: 86/100 - Düzenli ve gerçek spike trenine en yakın sonuç. - Qwen3.8: 79/100 - Spike oluşuyor fakat yüksek alanda parçalanma artıyor. - DeepSeek V4 Pro-0813: 64/100 - Gelişmiş çekirdek, fakat ciddi indeksleme riski ve parçalanma - GPT-5.6 Sol: 63/100 - En iyi ürün yüzeyi; fizik çekirdeğinde damlacıklaşma - GLM 5.2: 57/100 - Çok kararlı ve hızlı, fakat tek geniş kubbe üretiyor - Grok 4.6: 53/100 - Matematiksel olarak iyi düşünülmüş, tarayıcı performansı çok düşük - Cursor Auto: 47/100 - Hızlı üretim; garip ince kolonlar, parçalanma ve düşük runtime FPS - Kimi K3: 42/100 - Özellik kapsamı geniş; fiziksel sonuç ve skor güvenilirliği zayıf - DeepSeek V4 Flash-0731: 32/100 - En ucuz; varsayılan görüntü kullanılamaz ve ölçümler yanıltıcı Nasıl kullanmalıyız: - Tek model seçeceksem: Qwen3.8 - Bütçe önemsiz ve en iyi teknik sonucu istiyorsam: Opus 5 - Çok düşük bütçeyle güçlü bir çekirdek istiyorsam: DeepSeek V4 Pro - Sadece frontend ve ürün kalitesi istiyorsam: GPT-5.6 Sol - Formül ve algoritma eleştirmeni istiyorsam: Grok 4.6 - Optimization/refactor istiyorsam: GLM 5.2 - Visual QA istiyorsam: Kimi K3 - Cursor Auto’yu model olarak değil, sabit modelleri yöneten agent harness olarak kullanırım - DeepSeek Flash’ı zorlu caseler dışında, harness yaparak çözeceksem ve disposable boilerplate ile test üretiminde kullanırım.

Alican Kiraz

12,765 просмотров • 1 месяц назад