Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

1/ What do video models know about high school physics? Less than you'd think. Led by Varun Varma Thozhiyoor and Shivam Tripathi with Venkatesh Babu, we built Principia: relational physics tests for video models. Eight laws, 500+ real scenes, one simple idea. paper:

11,198 Aufrufe • vor 15 Tagen •via X (Twitter)

7 Kommentare

Profilbild von Ayush Tewari
Ayush Tewarivor 14 Tagen

@varunvarmat @shivam_tr @rvbabuiisc Congrats, Anand. Looks great! We also looked into this in VDAWorld ( check out HSPBench (

Profilbild von Anand Bhattad
Anand Bhattadvor 14 Tagen

Thanks Ayush for the pointer! HSPBench is quite related to our work. After skimming through it, I see that VDAWorld fills the physics gaps using generative models with physics engines. Our Principia paper provides a formal calibration-free way to benchmark these relational failures directly from generated videos. We will discuss and cite this in our updated paper. For additional context, here’s the first version of our results (#CVPR2026 Findings): Thanks!

Profilbild von Anand Bhattad
Anand Bhattadvor 15 Tagen

2/ How do we verify physics accuracy in generated videos? The main problem is that we cannot determine absolute time simply from the number of generated frames. Next, we often do not know the spatial scale at which the generative models are operating, the camera’s coordinates, or the mass of the objects involved. To address this, our approach shifts away from measuring single objects in isolation, focusing instead on comparing multiple objects to evaluate their relative consistency. For example, two blocks of different masses placed on identical inclines must reach the bottom at the same time. By checking for this relative consistency, we can completely circumvent the need for absolute measurements or other calibration.

Profilbild von Anand Bhattad
Anand Bhattadvor 15 Tagen

3/ For our experiment, we tested six different generators, four vision-language models, several hundred manually recorded and curated scenes, and thousands of hours of compute spent simply generating many videos for testing purposes. None of the models we tested performed well on the Principia benchmark. Summary: Video models are still far from generating physically consistent video. Principia provides an objective evaluation protocol for testing high school physics in generators -- a step towards making physics consistency a go-to evaluation for video models.

Profilbild von Anand Bhattad
Anand Bhattadvor 15 Tagen

4/ Several interesting videos are on our website: paper: PS: Principia is the latest step in a series of work from our group on exploring what generative models actually know (and don't) about the physical world.

Profilbild von crisiumnih
crisiumnihvor 13 Tagen

@varunvarmat @shivam_tr @rvbabuiisc coooooll

Profilbild von Isabelle E-Z
Isabelle E-Zvor 14 Tagen

@varunvarmat @shivam_tr @rvbabuiisc The gap says it all: 0.8 on VBench, under 0.42 on Principia. Rendered motion without Newtonian consistency won't get us reliable world models. These generators make things look right, not be right. 💻

Ähnliche Videos

Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,821,691 Aufrufe • vor 1 Jahr