Загрузка видео...
Не удалось загрузить видео
New Opus 4.8 crashed Opus 4.7 at physics on canvas! We gave both models the same three prompts: simulate a real physics phenomenon on raw HTML5 canvas. Prompt 1: "A triple pendulum swings into chaos and paints glowing trails with its tip" Prompt 2: "A 1 kg block bounces... show more
638,609 просмотров • 4 месяцев назад •via X (Twitter)
Комментарии: 39

thanks I’ll buy Opus 4.8 to pass my physics exam

@the_great_hug0 the right decision!

This is the result from sonnet with medium thinking

Sonnet did solid job here

I think it is better then 4.7

Now gives us examples about something that we would use in real life and not for just make an X post.

Skill Issue

Fun exercise. Thought I'd have my agent Sirius give it a shot. What I think he did pretty good. @OpenAIDevs

@OpenAIDevs sirius did pretty well! Appreciate seeing builders run the same experiment with their own agents :))

Impressive

nice

The pendulum test is the best physics benchmark I've seen. Canvas code that produces realistic chaos requires actually modeling the system, not just pattern-matching from training data.

exactly! and fake motion vs real chaotic dynamics shows up immediately

For these comparisons you need to launch opus 4.7 2x times vs opus 4.8 2x. You could be just looking at run to run variations of how these LLMs work.

Very interesting comparison visualizations. It clearly shows the differences between the models.

We're slowly reaching the point where model upgrades feel less like "better chat" and more like software evolution.

the pendulum one is wild, i use Opus daily for coding and the jump from 4.7 to 4.8 on complex stuff is pretty noticeable

Have these prompts run on each model’s release date? Or have you used a today’s 4.7 (crippled) version?

opus 4.7 is crippled:))

4.8 is wild 🔥 Built Claude Pulse for exactly this — live KV cache timer + usage right on claude .ai.

tbh the triple pendulum render looks slick but a canvas physics demo doesnt really tell me if the model holds up on actual coding work

Nice. The pi-digits collision one is a good test because it needs exact physics, not just something that looks plausible. Curious whether it holds up on messier problems like fluids or soft-body collisions, where the math gets harder to fake visually.

992 isn't a random number. The Galperin billiard: collisions ≈ π × √(M/m). With a 100,000:1 mass ratio, that's π × 316 ≈ 993. Opus 4.8 essentially solved the physics. 4.7 got 241 not even close. The benchmark is the whole point.

That's impressive! It's fascinating to see how Opus 4.8 handles complex physics simulations. Looking forward to more insights from your comparisons.

@ai_rohitt glad you enjoyed it! More side by side runs on the way

did either use a fixed timestep because variable dt usually makes canvas physics demos lie

opus 4.8 math model is so great. only the limit sucks 😔

Sorry, but what as to do Opus with local AI models? It can only run remotely on Anthropic servers via APIs.

CC @RayKurz

4.8 crashing on physics implies an over-eager alignment filter. Complex simulations can hit safety boundaries if intermediate steps are misinterpreted as harmful or exploit prompts.

本当に面白そう!バージョン4.8がバージョン4.7を打ち負かしたんだね。どんな物理現象が再現されているんだろう?楽しみ!

i want to use this to improve my memory system! can i use atomic chat to benchmark my self improving memory system?

Should it be Opus 4.8 or 4.6? Opus 4.7 has received mixed reviews.

the detail in that pendulum prompt is actually a solid benchmark for chaos visualization

Yeeah, exactly why we picked it!

wild how clearly opus 4.8 dominates 4.7 here

壁を通りぬけるのは面白いんじゃ🤣

We definitely should see more hard, thick balls hitting the wall more often

Snowdog has been watching this for 3 hours. 🐾❄️ Not the physics. Not the AI benchmark. Just the bouncing balls. 🟡 Same energy as staring at snow falling at 3am. ❄️🍯 $SD #Snowdog

