Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

New Opus 4.8 crashed Opus 4.7 at physics on canvas! We gave both models the same three prompts: simulate a real physics phenomenon on raw HTML5 canvas. Prompt 1: "A triple pendulum swings into chaos and paints glowing trails with its tip" Prompt 2: "A 1 kg block bounces...

638,609 görüntüleme • 4 ay önce •via X (Twitter)

39 Yorum

Hugo Plat profil fotoğrafı
Hugo Plat4 ay önce

thanks I’ll buy Opus 4.8 to pass my physics exam

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

@the_great_hug0 the right decision!

Chandraprakash Darji profil fotoğrafı
Chandraprakash Darji4 ay önce

This is the result from sonnet with medium thinking

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

Sonnet did solid job here

Chandraprakash Darji profil fotoğrafı
Chandraprakash Darji4 ay önce

I think it is better then 4.7

LiberalSismo profil fotoğrafı
LiberalSismo4 ay önce

Now gives us examples about something that we would use in real life and not for just make an X post.

Shounak Das profil fotoğrafı
Shounak Das4 ay önce

Skill Issue

Hutch, D.O. profil fotoğrafı
Hutch, D.O.4 ay önce

Fun exercise. Thought I'd have my agent Sirius give it a shot. What I think he did pretty good. @OpenAIDevs

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

@OpenAIDevs sirius did pretty well! Appreciate seeing builders run the same experiment with their own agents :))

iksammy1 profil fotoğrafı
iksammy14 ay önce

Impressive

cws test profil fotoğrafı
cws test4 ay önce

nice

Alexander Benz profil fotoğrafı
Alexander Benz4 ay önce

The pendulum test is the best physics benchmark I've seen. Canvas code that produces realistic chaos requires actually modeling the system, not just pattern-matching from training data.

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

exactly! and fake motion vs real chaotic dynamics shows up immediately

Venkatesh profil fotoğrafı
Venkatesh4 ay önce

For these comparisons you need to launch opus 4.7 2x times vs opus 4.8 2x. You could be just looking at run to run variations of how these LLMs work.

insightformer profil fotoğrafı
insightformer4 ay önce

Very interesting comparison visualizations. It clearly shows the differences between the models.

WeSee profil fotoğrafı
WeSee4 ay önce

We're slowly reaching the point where model upgrades feel less like "better chat" and more like software evolution.

Tech With Matteo profil fotoğrafı
Tech With Matteo4 ay önce

the pendulum one is wild, i use Opus daily for coding and the jump from 4.7 to 4.8 on complex stuff is pretty noticeable

rpello profil fotoğrafı
rpello4 ay önce

Have these prompts run on each model’s release date? Or have you used a today’s 4.7 (crippled) version?

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

opus 4.7 is crippled:))

Samir Patil profil fotoğrafı
Samir Patil4 ay önce

4.8 is wild 🔥 Built Claude Pulse for exactly this — live KV cache timer + usage right on claude .ai.

BruzWJ profil fotoğrafı
BruzWJ4 ay önce

tbh the triple pendulum render looks slick but a canvas physics demo doesnt really tell me if the model holds up on actual coding work

Cristobal Santana profil fotoğrafı
Cristobal Santana4 ay önce

Nice. The pi-digits collision one is a good test because it needs exact physics, not just something that looks plausible. Curious whether it holds up on messier problems like fluids or soft-body collisions, where the math gets harder to fake visually.

Sathish Harry profil fotoğrafı
Sathish Harry4 ay önce

992 isn't a random number. The Galperin billiard: collisions ≈ π × √(M/m). With a 100,000:1 mass ratio, that's π × 316 ≈ 993. Opus 4.8 essentially solved the physics. 4.7 got 241 not even close. The benchmark is the whole point.

Rohit profil fotoğrafı
Rohit4 ay önce

That's impressive! It's fascinating to see how Opus 4.8 handles complex physics simulations. Looking forward to more insights from your comparisons.

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

@ai_rohitt glad you enjoyed it! More side by side runs on the way

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan4 ay önce

did either use a fixed timestep because variable dt usually makes canvas physics demos lie

Build with Jeje profil fotoğrafı
Build with Jeje4 ay önce

opus 4.8 math model is so great. only the limit sucks 😔

GiovanniArturi profil fotoğrafı
GiovanniArturi3 ay önce

Sorry, but what as to do Opus with local AI models? It can only run remotely on Anthropic servers via APIs.

🐺 steve 🐺 profil fotoğrafı
🐺 steve 🐺4 ay önce

CC @RayKurz

Noureddine profil fotoğrafı
Noureddine4 ay önce

4.8 crashing on physics implies an over-eager alignment filter. Complex simulations can hit safety boundaries if intermediate steps are misinterpreted as harmful or exploit prompts.

Gia huy profil fotoğrafı
Gia huy4 ay önce

本当に面白そう!バージョン4.8がバージョン4.7を打ち負かしたんだね。どんな物理現象が再現されているんだろう?楽しみ!

Ex0 Byte profil fotoğrafı
Ex0 Byte4 ay önce

i want to use this to improve my memory system! can i use atomic chat to benchmark my self improving memory system?

Ziv Chen profil fotoğrafı
Ziv Chen4 ay önce

Should it be Opus 4.8 or 4.6? Opus 4.7 has received mixed reviews.

Life Around Earth profil fotoğrafı
Life Around Earth4 ay önce

the detail in that pendulum prompt is actually a solid benchmark for chaos visualization

atomic.chat profil fotoğrafı
atomic.chat4 ay önce

Yeeah, exactly why we picked it!

Blum profil fotoğrafı
Blum4 ay önce

wild how clearly opus 4.8 dominates 4.7 here

AIじいさん profil fotoğrafı
AIじいさん4 ay önce

壁を通りぬけるのは面白いんじゃ🤣

Theoden profil fotoğrafı
Theoden4 ay önce

We definitely should see more hard, thick balls hitting the wall more often

Snowdog_base profil fotoğrafı
Snowdog_base3 ay önce

Snowdog has been watching this for 3 hours. 🐾❄️ Not the physics. Not the AI benchmark. Just the bouncing balls. 🟡 Same energy as staring at snow falling at 3am. ❄️🍯 $SD #Snowdog

Benzer Videolar

Gemini 2.5 Flash demolishes my Galton Board test, I could not get 4omini, 4o mini high, or 03 to produce this. I found that Gemini 2.5 Flash understands my intents almost instantly, code produced is tight and neat. The prompt is a merging of various steps. It took me 5 steps to achieve this in Gemini 2.5 Flash, I gave up on OpenAI models after about half an hour. My iterations are obviously not exact. But people can test with this one prompt for more objective comparison. Please try this prompt on your end to confirm: -------------------------------------------------- Create a self-contained HTML file for a Galton board simulation using client-side JavaScript and a 2D physics engine (like Matter.js, included via CDN). The simulation should be rendered on an HTML5 canvas and meet the following criteria: 1. **Single File:** All necessary HTML, CSS, and JavaScript code must be within this single `.html` file. 2. **Canvas Size:** The overall simulation area (canvas) should be reasonably sized to fit on a standard screen without requiring extensive scrolling or zooming (e.g., around 500x700 pixels). 3. **Physics:** Utilize a 2D rigid body physics engine for realistic ball-peg and ball-wall interactions. 4. **Obstacles (Pegs):** Create static, circular pegs arranged in full-width horizontal rows extending across the usable width of the board (not just a triangle). The pegs should be small enough and spaced appropriately for balls to navigate and bounce between them. 5. **Containment:** * Include static, sufficiently thick side walls and a ground at the bottom to contain the balls within the board. * Implement *physical* static dividers between the collection bins at the bottom. These dividers must be thick enough to prevent balls from passing through them, ensuring accurate accumulation in each bin. 6. **Ball Dropping:** Balls should be dropped from a controlled, narrow area near the horizontal center at the top of the board to ensure they enter the peg field consistently. 7. **Bins:** The collection area at the bottom should be divided into distinct bins by the physical dividers. The height of the bins should be sufficient to clearly visualize the accumulation of balls. 8. **Visualization:** Use a high-contrast color scheme to clearly distinguish between elements. Specifically, use yellow for the structural elements (walls, top guides, physical bin dividers, ground), a contrasting color (like red) for the pegs, and a highly contrasting color (like dark grey or black) for the balls. 9. **Demonstration:** The simulation should visually demonstrate the formation of the normal (or binomial) distribution as multiple balls fall through the pegs and collect in the bins. Ensure the physics parameters (restitution, friction, density) and ball drop rate are tuned for a smooth and clear demonstration of the distribution. #OpenAI Sam Altman Greg Brockman AshutoshShrivastava Aidan McLaughlin

RameshR

247,923 görüntüleme • 1 yıl önce

So Runway Gen 4.5 finally adds image-to-video, the workflow most pros rely on for consistency. We put it head to head with Kling AI and Google Flow VEO using the same reference images and prompts (below) to evaluate motion quality, stability, and cinematic realism. 1. Action/WaterPhysics Test Prompt: Cinematic, wide-shot of a man running in a shallow river. The camera is tracking the man from behind as he runs up the river. Handheld camera shake as the camera follows the man. 2. Fire Physics Test Prompt: Cinematic, wide-shot of terrified woman running towards her burning barn. She abruptly stops, and puts in hands on her head as she watches her barn burn down. 3. VFX test prompt: Cinematic, wide-shot of a hooded figure. Flashes of purple magic and smoke whirl around the figure. The figure lifts its arms as the purple magic and smoke intensifies. 4. 2D Animation Test Prompt: 2D animated shot of a waiting at a bus stop in a thunderstorm. The man turns, walks to the bench, and sits down. 5. 3D Animation Test Prompt: 3D animated shot of an octopus. The octopus reaches into a coral and picks up a glowing white gem. 6. Conversation Test Prompt: slow camera push-in as two friends are having a conversation at a coffee shop Overall verdict: Despite the “world’s best” claim, Runway Gen 4.5 is not there yet. Prompt adherence is solid, but motion, physics, and cinematic realism still lag behind tools like Kling and VEO. Great platform, mid-tier model for now.

Curious Refuge

25,712 görüntüleme • 8 ay önce