正在加载视频...

视频加载失败

Grok 4.5 vs Meta Muse 1.1 I think both cooked really well!! WOW! Left: Grok 4.5 Right: Meta Muse 1.1

23,291 次观看 • 2 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

meta muse spark 1.1 vs gpt 5.6 sol vs fable 5 vs grok 4.5 meta recently dropped muse spark 1.1 – a multimodal reasoning model from meta superintelligence labs built for agentic tasks. key facts: • 1m token context with active self-management – the model compacts its own history and keeps only the steps needed for later work • trained to orchestrate multi-agent systems: as main agent it plans and delegates to parallel subagents, as subagent it sticks to its job and knows when to escalate back • computer use trained to pick between scripting and clicking – writes automation when it's faster, clicks when it's simpler, batches actions per step • first public api from meta: the meta model api is now in preview • benchmarks: sweeps the agent column – mcp atlas 88.1 (opus 4.8: 82.2), jobbench 54.7 (opus: 48.4), humanity's last exam 62.1 (1st). loses coding – deepswe 1.1 53.3 vs gpt 5.5's 67.0, swe bench pro 61.5 vs opus's 69.2 our test – 3 prompts, single-file html, three.js, fully procedural, no assets: 1. norwegian house cantilevered over a fjord in a snowstorm – transmissive glass wall, fully modelled interior 2. beijing siheyuan courtyard house in dawn fog – instanced roof tiles, dougong brackets, glowing paper windows 3. new mexico adobe pueblo in an approaching dust storm – deep window reveals, windward grit accumulation we ran the test on AI/ML API platform results: - cost #1 muse spark 1.1 – $0.20 #2 grok 4.5 – $0.51 #3 gpt 5.6 sol – $1.93 #4 fable 5 – ~$5.20 - output tokens #1 muse spark 1.1 – 41,868 #2 gpt 5.6 sol – 49,139 #3 grok 4.5 – 64,954 #4 fable 5 – 81,849 - lines of code #1 muse spark 1.1 – 1,799 #2 gpt 5.6 sol – 2,377 #3 fable 5 – 3,088 #4 grok 4.5 – 4,216 observations: • muse spark is the cheapest of the four by a wide margin – 2.5x under grok, ~26x under fable per run. output quality tracks the price • only 7.4% of its output tokens are reasoning (3,104 of 41,868) – the model barely thinks before writing. economic, not pedantic: it commits to the first plan and ships it • the low loc is not compression, it's omission – all three prompts demanded instancing, muse spark delivered it in one muse spark's code quality – reviewed by fable 5: upsides: 1. all three files run 2. the adobe grit effect is legit – shader injection via onbeforecompile, windward faces detect storm direction through a normal-dot-wind term and darken procedurally 3. the fjord glass is real meshphysicalmaterial with transmission and ior, not a transparent quad 4. the siheyuan properly instances barrel tiles, dougong blocks and courtyard pavers downsides: 1. in the fjord file the strafe vector is negated – press a, you move right; press d, you move left. exactly the key mix-up we kept hitting with this model 2. all three files ship the model's self-doubt as comments: "// actually yaw orientation: need correct" sits above a direction vector that gets computed, abandoned and recomputed – dead vectors allocated every frame, 60 times a second 3. the siheyuan registers two separate keydown listeners, one containing an empty if-block 4. snow "accumulation" on the norway roof is a sine wobble on a scale value, not accumulation 5. "instanced snow" became 3,500 plain points. zero dispose calls anywhere pattern: minimal reasoning, minimal code, minimal price. it nails the flashy requirements – shaders, transmissive glass – and quietly drops the boring ones: instancing, controls, cleanup. you get a demo that mostly runs and a control scheme you can't trust follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

135,556 次观看 • 2 个月前

Advanced Frontier LLM Benchmark - Üç Cisim & Gravity Choreography 🚀 - Modeller: Grok 4.6 / AI at Meta Muse Spark 1.3 / Claude Fable 5.1 / Google Gemini Gemini 3.8 Flash - Effort: Max & Very high - Agent: Cursor & Opencode Task: Görev, modele tek bir HTML dosyası içinde, dış kaynak kullanmadan, Canvas 2D ile düzlemsel N-cisim kütleçekimi simülasyonu yazdırmak; bu simülasyon açıldığı anda deterministik bir sinematik gösteriyle başlıyor (Kepler yörüngesi → hiyerarşik yıldız sistemi → figure-8 koreografisi → 12 kopyalı kaotik ayrışma → uzun pozlama), ayrıca ~40 saniyelik video kaydı için ayrı bir demo reel modu, gizli fizik self-check'i, korunum diagnostikleri, sürüklenebilir cisimler ve kompakt bir arayüz içeriyor. 🎉 Fiyat/Performans Kazanan: Grok 4.6🎉 🎉 Core Teknik Kazanan: Fable 5.1 🎉 Sebep: Fable, bu benchmark'ın iki tarafını en iyi birlikte çözüyor: solver ve cinematic renderer. Pairwise Newtonian çekirdek, softened potential, conservation diagnostics, Verlet/Yoshida yolu ve validator mimarisi güçlü; runtime self-check sonuçları da dört model içindeki en iyi seviyede. Daha önemlisi, model yalnızca ekrana “PASS” yazmıyor; hidden systems gerçek integrator yoluyla kademeli çalıştırılıyor ve ölçüm sonucundan verdict üretiliyor. Maliyet: - Grok 4.6: $1.10 - Muse Spark 1.3: $1.12 - Gemini 3.8 Flash: $1.3 - Fable 5.1: $6.2 Teknik skor: - Fable 5.1 - 96.4/100 - Grok 4.6 - 92.1/100 - Muse Spark 1.3 - 89.6/100 - Gemini 3.8 Flash - 81.0/100 Performans / $ - Grok 4.6 - 83.7/100 - Muse Spark 1.3 - 80.0/100 - Gemini 3.8 Flash - 62.3/100 - Fable 5.1 - 15.6/100 Teknik İnceleme Metrikleri: 1) Prompt uyumu: Fable'ın substep genişletmesi ve Muse'un bazı spesifikasyonları sadeleştirmesi nedeniyle kimseye 10 vermiyorum. - Fable 5.1: 9.6/10 - Grok 4.6: 9.3/10 - Gemini 3.8 Flash: 8.8/10 - Muse Spark 1.3: 8.2/10 2) Newtonian physics + conservation: Dört modelin de bu kategori şaşırtıcı derecede güçlü. Muse dahil pairwise equal-and-opposite acceleration ve aynı softened potential kernel'ini kullanıyor. - Fable 5.1: 9.9/10 - Grok 4.6: 9.8/10 - Gemini 3.8 Flash: 9.8/10 - Muse Spark 1.3: 9.6/10 3) Symplectic integrasyon + fixed timestep: Gemini'nin Yoshida formülü yanlış değil; puan kaybı scheduler/backlog yönetiminden geliyor. Numerical method doğru, execution model chaos altında sorunlu. - Fable 5.1: 9.6/10 - Grok 4.6: 9.5/10 - Muse Spark 1.3: 9.4/10 - Gemini 3.8 Flash: 8.0/10 4) Self-check / validator doğruluğu: Burada Fable açık lider. Gemini çok yakın; dtExact metodolojik küçük pürüz. Grok/Muse PASS ama validator resolution daha kaba. - Fable 5.1: 10.0/10 - Gemini 3.8 Flash: 9.4/10 - Grok 4.6: 8.6/10 - Muse Spark 1.3: 8.0/10 5) Görsel kalite: Gemini'nin celestial-body shading'i ve UI'sı çok iyi. Fable ise kompozisyon ve “görsel gürültüyü bastırma” açısından biraz daha dengeli. Muse özellikle büyük title choreography ile etkileyici. Grok en restrained olanı. - Fable 5.1: 9.6/10 - Gemini 3.8 Flash: 9.5/10 - Grok 4.6: 9.2/10 - Muse Spark 1.3: 9.1/10 6) Trail rendering: Gemini'nin single-system trail'i güzel; fakat chaos trail timestamp kaynağındaki hata bu kategori için ciddi. - Fable 5.1: 9.7/10 - Grok 4.6: 9.3/10 - Muse Spark 1.3: 9.0/10 - Gemini 3.8 Flash: 7.0/10 Alternatif winner'lar: - En iyi saf numerical correctness: Fable 5.1 - En iyi self-validator: Fable 5.1 - En iyi görsel kalite: Fable 5.1 ≈ Gemini 3.8 Flash - En iyi chaos throughput: Muse Spark 1.3 - En iyi genel mimari: Fable 5.1 - En iyi fiyat/performans: Grok 4.6 - En düşük maliyet: Grok 4.6 - Sadece physics/integrator kodu puanlansaydı: Fable ≈ Grok > Gemini ≈ Muse - Maliyet de kararın temel kriteri olsaydı: Grok > Muse > Gemini >>> Fable

Alican Kiraz

18,781 次观看 • 12 天前