Loading video...

Video Failed to Load

Go Home

I will defend my PhD thesis, “Physics-Constrained Generative Models for Computational Design,” on Tuesday, September 8, 2026, at PM ET. Location: MIT MIT CSAIL 32-G449 Zoom: Committee: Kaiming He, Bill Freeman, Wojciech Matusik Generative models can now propose candidate designs at high throughput, yet the physical world remains the...

78,288 views • 11 days ago •via X (Twitter)

36 Comments

Robert Scoble's profile picture
Robert Scoble11 days ago

@MIT_CSAIL Your work caught my eye. Thank you and knock it out of the park!

Minghao Guo's profile picture
Minghao Guo11 days ago

Relying on downstream validation creates a fundamental trade-off between throughput and physical realizability: generating a candidate may take seconds, while validating it can take hours to months. Physical realizability is a scaling problem, not just a quality problem. (2/7)

Minghao Guo's profile picture
Minghao Guo11 days ago

Inspired by Marr’s separation of levels of analysis, we organize generative computational design into three levels: physical, computational, and generative. Because each level abstracts the one below, physical requirements can be weakened or lost at the interfaces between them. (3/7)

Minghao Guo's profile picture
Minghao Guo11 days ago

This thesis treats physical requirements as first-class citizens of the generative system. The central idea is to carry those requirements bottom-up through explicit interfaces, from the physical world through computation and into generation. We formulate computational design as constrained optimization. (4/7)

Minghao Guo's profile picture
Minghao Guo11 days ago

We realize these interfaces through three core elements: representation, constraints, and objective evaluation. At the representation interface, procedural molecular grammars, medial skeletons, and tetrahedral sphere primitives define candidate spaces that are valid by construction. At the constraint interface, differentiable projection modules enforce physical feasibility during generation. (5/7)

Minghao Guo's profile picture
Minghao Guo11 days ago

At the objective interface, unified simulators evaluate design behavior with scientific fidelity at generative throughput, spanning physical regimes from molecular dynamics to aerodynamics. Together, these contributions lay the foundation for universal precision design: general-purpose systems that tailor physically realizable designs to the specific requirements of a problem, individual, or environment, at scale. (6/7)

Minghao Guo's profile picture
Minghao Guo11 days ago

A research map of selected publications across my PhD study: #GenerativeAI #AIforScience #ComputationalDesign #ScientificMachineLearning #PhDDefense #worldmodel

Minghao Guo's profile picture
Minghao Guo11 days ago

Missing the time haha! It’s 2-3pm ET!

중심's profile picture
중심11 days ago

@MIT_CSAIL at what time?

BrenJ's profile picture
BrenJ11 days ago

@MIT_CSAIL @evelovesolive @beffjezos felt like this was up y'alls alley.

DrKnowItAll's profile picture
DrKnowItAll11 days ago

@MIT_CSAIL Great work! This looks really impressive.

rob's profile picture
rob11 days ago

@MIT_CSAIL Great stuff here.

samir sawarkar's profile picture
samir sawarkar11 days ago

@MIT_CSAIL The next leap in generative AI isn’t generating more designs. It’s generating designs that cannot violate the physics. That’s the shift from “AI can imagine it” to “AI can engineer it.”

Aman Chourasia's profile picture
Aman Chourasia11 days ago

@MIT_CSAIL This is great!

Daniel Samanez's profile picture
Daniel Samanez11 days ago

@MIT_CSAIL 👍 at what time is it?

Ralf Sigmund🇺🇦🇪🇺 @sistar-hh.bsky.social's profile picture
Ralf Sigmund🇺🇦🇪🇺 @sistar-hh.bsky.social10 days ago

@MIT_CSAIL Congrats. It was a honor to listen to You leading us through Your work. All the best!

Xenonostra''s profile picture
Xenonostra'11 days ago

@MIT_CSAIL I wish you the best of luck I'm sure you will do well

GSC's profile picture
GSC11 days ago

@MIT_CSAIL

Zhiyuan's profile picture
Zhiyuan11 days ago

@MIT_CSAIL solid work! may you good luck!

Yuxi Xiao's profile picture
Yuxi Xiao11 days ago

@MIT_CSAIL Wow nice video for the phd thesis!!! congrats

Big data annotation's profile picture
Big data annotation11 days ago

@MIT_CSAIL 大佬,牛逼

Raj Kachhadiya (KARAM)'s profile picture
Raj Kachhadiya (KARAM)11 days ago

@MIT_CSAIL Interesting, congratulations!

Anand Kumar's profile picture
Anand Kumar11 days ago

@MIT_CSAIL Interesting. Looking forward to it.

Neil's profile picture
Neil8 days ago

@MIT_CSAIL Wow this is such an interesting research direction, do you think the future of 3D is headed this way?

Yossi Eliaz's profile picture
Yossi Eliaz11 days ago

@MIT_CSAIL loved the video!

likhesh's profile picture
likhesh11 days ago

@MIT_CSAIL Amazing, what time exactly?

Mandar Wagh's profile picture
Mandar Wagh11 days ago

congrats, and that is a committee. the split i keep getting stuck on: physics constraints are mostly differentiable, so they go straight into a loss. manufacturing constraints mostly are not — tool reach, support-free overhang, demoldability, minimum wall are combinatorial, and you cannot gradient your way to a draft angle. does the thesis treat those as one kind of constraint, or as two different problems? genuinely curious which way you landed.

the guy's profile picture
the guy11 days ago

@MIT_CSAIL 🧡

Jesse Jr Lim (林振燊)'s profile picture
Jesse Jr Lim (林振燊)11 days ago

@MIT_CSAIL wow very cool.. paused the vid alot to see whats going on

X.M.'s profile picture
X.M.11 days ago

@MIT_CSAIL all the best king

Volatile Markets's profile picture
Volatile Markets11 days ago

@MIT_CSAIL Love this!

Verdi's profile picture
Verdi11 days ago

@MIT_CSAIL congrats!! you didnt share a time tho? > "on Tuesday, September 8, 2026, at PM ET."

Jacques Mₜ=∏ₛ((1−λₛ)+λₛEₛ)'s profile picture
Jacques Mₜ=∏ₛ((1−λₛ)+λₛEₛ)11 days ago

@MIT_CSAIL PhaGGOT

Curtis's profile picture
Curtis10 days ago

@MIT_CSAIL Grest work, great slides, great marketing video!

Black Canvas's profile picture
Black Canvas11 days ago

@MIT_CSAIL Wish you the best and interesting work!

Supratik Bhattacharya's profile picture
Supratik Bhattacharya11 days ago

@MIT_CSAIL Wow! this is super interesting. all the best!

Related Videos

Excited to show some surprising inventions on generative multiplayer games we made at Google with Stanford. We call the work MultiGen. I've always been inspired by early studios like id Software with Doom or Blizzard with Warcraft bringing networked video games to the next level. We are at the point in history where we can make strides like them, but for generative games. It's a strange feeling to be in the age of generative video games while still discovering how exactly to train the models and design the tools that make them useful. All of the tools that have been invented for classic game engines need to be redesigned for generative games. For example level and world design is not entirely possible with existing technology. We introduce editable memory to diffusion game engines that allow for design of new levels via a minimap. But we can easily imagine how this can be expanded with different creation tools. The end goal of this research direction is to allow game designers to be able to guide the generation process of their world, at the granularity that they prefer. Editable memory also allows us to add multiplayer to Generative Doom. We were amazed when we saw GameNGen some years ago, and now you can play it live with friends in real-time, on your couch or even online. Shared representations like our editable memory seem like the future for this type of experience. Models are, in some cases, expensive and approximate encoders but great interpolators and extrapolators. Leveraging their strengths lets you have completely new experiences that can be realized now and not in the distant future. This work was started at my previous team and continued in collaboration with Stanford. Congratulations to all for the discoveries.

Nataniel Ruiz

104,932 views • 6 months ago

Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,821,580 views • 1 year ago

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,079,815 views • 3 months ago

Real-time world models represent a fundamental shift in AI. reactor is building the platform for real-time generative video infrastructure, supporting developers who need the tech for use across entertainment, physical AI, and robotics. Co-founders Alberto and Bryce Schmidtchen joined us last week on The Investment Memo, hosted by Partners Bucky Moore and Amber Yang, to talk about the era of world models. The conversation centered around the infrastructure Reactor is building, why real-time models are the edge right now, and current use cases for the product. Alberto and Bryce agreed that world models are shaping the way simulations are created, and that developers need a streamlined platform that can support their ideas. We believe Reactor is positioned to be at the frontier of research into real-time generative models. We look forward to seeing how these models apply across industries. Chapters 00:00 Introduction & Overview of Reactor 01:08 Meet the Hosts & Founders 02:18 The Origin Story: From 3D Assets to World Models 05:07 Real-Time Video Applications Across Industries 06:55 The Open Source World Model Explosion 07:23 Why Infrastructure Is the Opportunity 08:42 Parallels to Past Technology Waves 09:51 Bridging the Research-to-Production Gap 13:13 What Developers Are Building with World Models 16:41 Lessons from Luma AI 18:23 What Apple Vision Pro Taught Bryce About Real-Time Systems 20:48 Company Values & Team Culture 22:40 Series A: What the Capital Unlocks 24:13 Reactor's Five-Year Vision 26:09 Closing Remarks

Lightspeed

144,942 views • 3 months ago

Demis Hassabis on the limit in today’s AI: language can describe the world, but it cannot contain it - and why "World Models" are his "longest standing passion". Language models absorbed far more structure about reality from text than many researchers expected, because human language quietly carries physics, psychology, culture, tools, plans, and cause-and-effect. But text is still a compressed residue of experience, not experience itself. A sentence can say a cup falls from a table, yet it does not fully encode weight, grip, balance, friction, timing, sound, surprise, or the tiny motor corrections a body makes before it even notices them. The world is not only made of facts that can be named; it is made of constraints that have to be lived through, touched, predicted, violated, and repaired. That is why world models matter. They aim to learn the hidden grammar of physical reality: how objects persist, how forces unfold, how space changes when an agent moves, and how action creates feedback. Language models can often reason about the world because people have written so much about it. World models try to learn what the world is like before it becomes words. The difference is exactly what matters because intelligence is not just answering well; it is knowing what would happen next if you moved, reached, pushed, smelled, slipped, or failed. A mind trained only on descriptions may become brilliant at explanation. A mind trained on experience may become better at consequence. --- Full video from "Google DeepMind" and "Hannah Fry" YT channel (link in comment)

Rohan Paul

49,938 views • 3 months ago