Loading video...

Video Failed to Load

Go Home

Latest models on the self-portrait test 👀 Gemini 3.8 Flash, Muse Spark 1.3, Claude Fable 5.1 and Tencent Hy4. Gemini 3.8 Flash looks cool 😎

25,774 views • 7 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Advanced Frontier LLM Benchmark - Üç Cisim & Gravity Choreography 🚀 - Modeller: Grok 4.6 / AI at Meta Muse Spark 1.3 / Claude Fable 5.1 / Google Gemini Gemini 3.8 Flash - Effort: Max & Very high - Agent: Cursor & Opencode Task: Görev, modele tek bir HTML dosyası içinde, dış kaynak kullanmadan, Canvas 2D ile düzlemsel N-cisim kütleçekimi simülasyonu yazdırmak; bu simülasyon açıldığı anda deterministik bir sinematik gösteriyle başlıyor (Kepler yörüngesi → hiyerarşik yıldız sistemi → figure-8 koreografisi → 12 kopyalı kaotik ayrışma → uzun pozlama), ayrıca ~40 saniyelik video kaydı için ayrı bir demo reel modu, gizli fizik self-check'i, korunum diagnostikleri, sürüklenebilir cisimler ve kompakt bir arayüz içeriyor. 🎉 Fiyat/Performans Kazanan: Grok 4.6🎉 🎉 Core Teknik Kazanan: Fable 5.1 🎉 Sebep: Fable, bu benchmark'ın iki tarafını en iyi birlikte çözüyor: solver ve cinematic renderer. Pairwise Newtonian çekirdek, softened potential, conservation diagnostics, Verlet/Yoshida yolu ve validator mimarisi güçlü; runtime self-check sonuçları da dört model içindeki en iyi seviyede. Daha önemlisi, model yalnızca ekrana “PASS” yazmıyor; hidden systems gerçek integrator yoluyla kademeli çalıştırılıyor ve ölçüm sonucundan verdict üretiliyor. Maliyet: - Grok 4.6: $1.10 - Muse Spark 1.3: $1.12 - Gemini 3.8 Flash: $1.3 - Fable 5.1: $6.2 Teknik skor: - Fable 5.1 - 96.4/100 - Grok 4.6 - 92.1/100 - Muse Spark 1.3 - 89.6/100 - Gemini 3.8 Flash - 81.0/100 Performans / $ - Grok 4.6 - 83.7/100 - Muse Spark 1.3 - 80.0/100 - Gemini 3.8 Flash - 62.3/100 - Fable 5.1 - 15.6/100 Teknik İnceleme Metrikleri: 1) Prompt uyumu: Fable'ın substep genişletmesi ve Muse'un bazı spesifikasyonları sadeleştirmesi nedeniyle kimseye 10 vermiyorum. - Fable 5.1: 9.6/10 - Grok 4.6: 9.3/10 - Gemini 3.8 Flash: 8.8/10 - Muse Spark 1.3: 8.2/10 2) Newtonian physics + conservation: Dört modelin de bu kategori şaşırtıcı derecede güçlü. Muse dahil pairwise equal-and-opposite acceleration ve aynı softened potential kernel'ini kullanıyor. - Fable 5.1: 9.9/10 - Grok 4.6: 9.8/10 - Gemini 3.8 Flash: 9.8/10 - Muse Spark 1.3: 9.6/10 3) Symplectic integrasyon + fixed timestep: Gemini'nin Yoshida formülü yanlış değil; puan kaybı scheduler/backlog yönetiminden geliyor. Numerical method doğru, execution model chaos altında sorunlu. - Fable 5.1: 9.6/10 - Grok 4.6: 9.5/10 - Muse Spark 1.3: 9.4/10 - Gemini 3.8 Flash: 8.0/10 4) Self-check / validator doğruluğu: Burada Fable açık lider. Gemini çok yakın; dtExact metodolojik küçük pürüz. Grok/Muse PASS ama validator resolution daha kaba. - Fable 5.1: 10.0/10 - Gemini 3.8 Flash: 9.4/10 - Grok 4.6: 8.6/10 - Muse Spark 1.3: 8.0/10 5) Görsel kalite: Gemini'nin celestial-body shading'i ve UI'sı çok iyi. Fable ise kompozisyon ve “görsel gürültüyü bastırma” açısından biraz daha dengeli. Muse özellikle büyük title choreography ile etkileyici. Grok en restrained olanı. - Fable 5.1: 9.6/10 - Gemini 3.8 Flash: 9.5/10 - Grok 4.6: 9.2/10 - Muse Spark 1.3: 9.1/10 6) Trail rendering: Gemini'nin single-system trail'i güzel; fakat chaos trail timestamp kaynağındaki hata bu kategori için ciddi. - Fable 5.1: 9.7/10 - Grok 4.6: 9.3/10 - Muse Spark 1.3: 9.0/10 - Gemini 3.8 Flash: 7.0/10 Alternatif winner'lar: - En iyi saf numerical correctness: Fable 5.1 - En iyi self-validator: Fable 5.1 - En iyi görsel kalite: Fable 5.1 ≈ Gemini 3.8 Flash - En iyi chaos throughput: Muse Spark 1.3 - En iyi genel mimari: Fable 5.1 - En iyi fiyat/performans: Grok 4.6 - En düşük maliyet: Grok 4.6 - Sadece physics/integrator kodu puanlansaydı: Fable ≈ Grok > Gemini ≈ Muse - Maliyet de kararın temel kriteri olsaydı: Grok > Muse > Gemini >>> Fable

Alican Kiraz

18,781 views • 8 days ago

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,250 views • 14 days ago