正在加载视频...

视频加载失败

New Opus 4.8 crashed Opus 4.7 at physics on canvas! We gave both models the same three prompts: simulate a real physics phenomenon on raw HTML5 canvas. Prompt 1: "A triple pendulum swings into chaos and paints glowing trails with its tip" Prompt 2: "A 1 kg block bounces...

638,609 次观看 • 4 个月前 •via X (Twitter)

39 条评论

Hugo Plat 的头像
Hugo Plat4 个月前

thanks I’ll buy Opus 4.8 to pass my physics exam

atomic.chat 的头像
atomic.chat4 个月前

@the_great_hug0 the right decision!

Chandraprakash Darji 的头像
Chandraprakash Darji4 个月前

This is the result from sonnet with medium thinking

atomic.chat 的头像
atomic.chat4 个月前

Sonnet did solid job here

Chandraprakash Darji 的头像
Chandraprakash Darji4 个月前

I think it is better then 4.7

LiberalSismo 的头像
LiberalSismo4 个月前

Now gives us examples about something that we would use in real life and not for just make an X post.

Shounak Das 的头像
Shounak Das4 个月前

Skill Issue

Hutch, D.O. 的头像
Hutch, D.O.4 个月前

Fun exercise. Thought I'd have my agent Sirius give it a shot. What I think he did pretty good. @OpenAIDevs

atomic.chat 的头像
atomic.chat4 个月前

@OpenAIDevs sirius did pretty well! Appreciate seeing builders run the same experiment with their own agents :))

iksammy1 的头像
iksammy14 个月前

Impressive

cws test 的头像
cws test4 个月前

nice

Alexander Benz 的头像
Alexander Benz4 个月前

The pendulum test is the best physics benchmark I've seen. Canvas code that produces realistic chaos requires actually modeling the system, not just pattern-matching from training data.

atomic.chat 的头像
atomic.chat4 个月前

exactly! and fake motion vs real chaotic dynamics shows up immediately

Venkatesh 的头像
Venkatesh4 个月前

For these comparisons you need to launch opus 4.7 2x times vs opus 4.8 2x. You could be just looking at run to run variations of how these LLMs work.

insightformer 的头像
insightformer4 个月前

Very interesting comparison visualizations. It clearly shows the differences between the models.

WeSee 的头像
WeSee4 个月前

We're slowly reaching the point where model upgrades feel less like "better chat" and more like software evolution.

Tech With Matteo 的头像
Tech With Matteo4 个月前

the pendulum one is wild, i use Opus daily for coding and the jump from 4.7 to 4.8 on complex stuff is pretty noticeable

rpello 的头像
rpello4 个月前

Have these prompts run on each model’s release date? Or have you used a today’s 4.7 (crippled) version?

atomic.chat 的头像
atomic.chat4 个月前

opus 4.7 is crippled:))

Samir Patil 的头像
Samir Patil4 个月前

4.8 is wild 🔥 Built Claude Pulse for exactly this — live KV cache timer + usage right on claude .ai.

BruzWJ 的头像
BruzWJ4 个月前

tbh the triple pendulum render looks slick but a canvas physics demo doesnt really tell me if the model holds up on actual coding work

Cristobal Santana 的头像
Cristobal Santana4 个月前

Nice. The pi-digits collision one is a good test because it needs exact physics, not just something that looks plausible. Curious whether it holds up on messier problems like fluids or soft-body collisions, where the math gets harder to fake visually.

Sathish Harry 的头像
Sathish Harry4 个月前

992 isn't a random number. The Galperin billiard: collisions ≈ π × √(M/m). With a 100,000:1 mass ratio, that's π × 316 ≈ 993. Opus 4.8 essentially solved the physics. 4.7 got 241 not even close. The benchmark is the whole point.

Rohit 的头像
Rohit4 个月前

That's impressive! It's fascinating to see how Opus 4.8 handles complex physics simulations. Looking forward to more insights from your comparisons.

atomic.chat 的头像
atomic.chat4 个月前

@ai_rohitt glad you enjoyed it! More side by side runs on the way

Sebastian Buzdugan 的头像
Sebastian Buzdugan4 个月前

did either use a fixed timestep because variable dt usually makes canvas physics demos lie

Build with Jeje 的头像
Build with Jeje4 个月前

opus 4.8 math model is so great. only the limit sucks 😔

GiovanniArturi 的头像
GiovanniArturi3 个月前

Sorry, but what as to do Opus with local AI models? It can only run remotely on Anthropic servers via APIs.

🐺 steve 🐺 的头像
🐺 steve 🐺4 个月前

CC @RayKurz

Noureddine 的头像
Noureddine4 个月前

4.8 crashing on physics implies an over-eager alignment filter. Complex simulations can hit safety boundaries if intermediate steps are misinterpreted as harmful or exploit prompts.

Gia huy 的头像
Gia huy4 个月前

本当に面白そう!バージョン4.8がバージョン4.7を打ち負かしたんだね。どんな物理現象が再現されているんだろう?楽しみ!

Ex0 Byte 的头像
Ex0 Byte4 个月前

i want to use this to improve my memory system! can i use atomic chat to benchmark my self improving memory system?

Ziv Chen 的头像
Ziv Chen4 个月前

Should it be Opus 4.8 or 4.6? Opus 4.7 has received mixed reviews.

Life Around Earth 的头像
Life Around Earth4 个月前

the detail in that pendulum prompt is actually a solid benchmark for chaos visualization

atomic.chat 的头像
atomic.chat4 个月前

Yeeah, exactly why we picked it!

Blum 的头像
Blum4 个月前

wild how clearly opus 4.8 dominates 4.7 here

AIじいさん 的头像
AIじいさん4 个月前

壁を通りぬけるのは面白いんじゃ🤣

Theoden 的头像
Theoden4 个月前

We definitely should see more hard, thick balls hitting the wall more often

Snowdog_base 的头像
Snowdog_base3 个月前

Snowdog has been watching this for 3 hours. 🐾❄️ Not the physics. Not the AI benchmark. Just the bouncing balls. 🟡 Same energy as staring at snow falling at 3am. ❄️🍯 $SD #Snowdog

相关视频

Gemini 2.5 Flash demolishes my Galton Board test, I could not get 4omini, 4o mini high, or 03 to produce this. I found that Gemini 2.5 Flash understands my intents almost instantly, code produced is tight and neat. The prompt is a merging of various steps. It took me 5 steps to achieve this in Gemini 2.5 Flash, I gave up on OpenAI models after about half an hour. My iterations are obviously not exact. But people can test with this one prompt for more objective comparison. Please try this prompt on your end to confirm: -------------------------------------------------- Create a self-contained HTML file for a Galton board simulation using client-side JavaScript and a 2D physics engine (like Matter.js, included via CDN). The simulation should be rendered on an HTML5 canvas and meet the following criteria: 1. **Single File:** All necessary HTML, CSS, and JavaScript code must be within this single `.html` file. 2. **Canvas Size:** The overall simulation area (canvas) should be reasonably sized to fit on a standard screen without requiring extensive scrolling or zooming (e.g., around 500x700 pixels). 3. **Physics:** Utilize a 2D rigid body physics engine for realistic ball-peg and ball-wall interactions. 4. **Obstacles (Pegs):** Create static, circular pegs arranged in full-width horizontal rows extending across the usable width of the board (not just a triangle). The pegs should be small enough and spaced appropriately for balls to navigate and bounce between them. 5. **Containment:** * Include static, sufficiently thick side walls and a ground at the bottom to contain the balls within the board. * Implement *physical* static dividers between the collection bins at the bottom. These dividers must be thick enough to prevent balls from passing through them, ensuring accurate accumulation in each bin. 6. **Ball Dropping:** Balls should be dropped from a controlled, narrow area near the horizontal center at the top of the board to ensure they enter the peg field consistently. 7. **Bins:** The collection area at the bottom should be divided into distinct bins by the physical dividers. The height of the bins should be sufficient to clearly visualize the accumulation of balls. 8. **Visualization:** Use a high-contrast color scheme to clearly distinguish between elements. Specifically, use yellow for the structural elements (walls, top guides, physical bin dividers, ground), a contrasting color (like red) for the pegs, and a highly contrasting color (like dark grey or black) for the balls. 9. **Demonstration:** The simulation should visually demonstrate the formation of the normal (or binomial) distribution as multiple balls fall through the pegs and collect in the bins. Ensure the physics parameters (restitution, friction, density) and ball drop rate are tuned for a smooth and clear demonstration of the distribution. #OpenAI Sam Altman Greg Brockman AshutoshShrivastava Aidan McLaughlin

RameshR

247,923 次观看 • 1 年前

So Runway Gen 4.5 finally adds image-to-video, the workflow most pros rely on for consistency. We put it head to head with Kling AI and Google Flow VEO using the same reference images and prompts (below) to evaluate motion quality, stability, and cinematic realism. 1. Action/WaterPhysics Test Prompt: Cinematic, wide-shot of a man running in a shallow river. The camera is tracking the man from behind as he runs up the river. Handheld camera shake as the camera follows the man. 2. Fire Physics Test Prompt: Cinematic, wide-shot of terrified woman running towards her burning barn. She abruptly stops, and puts in hands on her head as she watches her barn burn down. 3. VFX test prompt: Cinematic, wide-shot of a hooded figure. Flashes of purple magic and smoke whirl around the figure. The figure lifts its arms as the purple magic and smoke intensifies. 4. 2D Animation Test Prompt: 2D animated shot of a waiting at a bus stop in a thunderstorm. The man turns, walks to the bench, and sits down. 5. 3D Animation Test Prompt: 3D animated shot of an octopus. The octopus reaches into a coral and picks up a glowing white gem. 6. Conversation Test Prompt: slow camera push-in as two friends are having a conversation at a coffee shop Overall verdict: Despite the “world’s best” claim, Runway Gen 4.5 is not there yet. Prompt adherence is solid, but motion, physics, and cinematic realism still lag behind tools like Kling and VEO. Great platform, mid-tier model for now.

Curious Refuge

25,712 次观看 • 8 个月前