Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I will defend my PhD thesis, “Physics-Constrained Generative Models for Computational Design,” on Tuesday, September 8, 2026, at PM ET. Location: MIT MIT CSAIL 32-G449 Zoom: Committee: Kaiming He, Bill Freeman, Wojciech Matusik Generative models can now propose candidate designs at high throughput, yet the physical world remains the...

78,288 görüntüleme • 11 gün önce •via X (Twitter)

36 Yorum

Robert Scoble profil fotoğrafı
Robert Scoble11 gün önce

@MIT_CSAIL Your work caught my eye. Thank you and knock it out of the park!

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

Relying on downstream validation creates a fundamental trade-off between throughput and physical realizability: generating a candidate may take seconds, while validating it can take hours to months. Physical realizability is a scaling problem, not just a quality problem. (2/7)

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

Inspired by Marr’s separation of levels of analysis, we organize generative computational design into three levels: physical, computational, and generative. Because each level abstracts the one below, physical requirements can be weakened or lost at the interfaces between them. (3/7)

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

This thesis treats physical requirements as first-class citizens of the generative system. The central idea is to carry those requirements bottom-up through explicit interfaces, from the physical world through computation and into generation. We formulate computational design as constrained optimization. (4/7)

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

We realize these interfaces through three core elements: representation, constraints, and objective evaluation. At the representation interface, procedural molecular grammars, medial skeletons, and tetrahedral sphere primitives define candidate spaces that are valid by construction. At the constraint interface, differentiable projection modules enforce physical feasibility during generation. (5/7)

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

At the objective interface, unified simulators evaluate design behavior with scientific fidelity at generative throughput, spanning physical regimes from molecular dynamics to aerodynamics. Together, these contributions lay the foundation for universal precision design: general-purpose systems that tailor physically realizable designs to the specific requirements of a problem, individual, or environment, at scale. (6/7)

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

A research map of selected publications across my PhD study: #GenerativeAI #AIforScience #ComputationalDesign #ScientificMachineLearning #PhDDefense #worldmodel

Minghao Guo profil fotoğrafı
Minghao Guo11 gün önce

Missing the time haha! It’s 2-3pm ET!

중심 profil fotoğrafı
중심11 gün önce

@MIT_CSAIL at what time?

BrenJ profil fotoğrafı
BrenJ11 gün önce

@MIT_CSAIL @evelovesolive @beffjezos felt like this was up y'alls alley.

DrKnowItAll profil fotoğrafı
DrKnowItAll11 gün önce

@MIT_CSAIL Great work! This looks really impressive.

rob profil fotoğrafı
rob11 gün önce

@MIT_CSAIL Great stuff here.

samir sawarkar profil fotoğrafı
samir sawarkar11 gün önce

@MIT_CSAIL The next leap in generative AI isn’t generating more designs. It’s generating designs that cannot violate the physics. That’s the shift from “AI can imagine it” to “AI can engineer it.”

Aman Chourasia profil fotoğrafı
Aman Chourasia11 gün önce

@MIT_CSAIL This is great!

Daniel Samanez profil fotoğrafı
Daniel Samanez11 gün önce

@MIT_CSAIL 👍 at what time is it?

Ralf Sigmund🇺🇦🇪🇺 @sistar-hh.bsky.social profil fotoğrafı
Ralf Sigmund🇺🇦🇪🇺 @sistar-hh.bsky.social10 gün önce

@MIT_CSAIL Congrats. It was a honor to listen to You leading us through Your work. All the best!

Xenonostra' profil fotoğrafı
Xenonostra'11 gün önce

@MIT_CSAIL I wish you the best of luck I'm sure you will do well

GSC profil fotoğrafı
GSC11 gün önce

@MIT_CSAIL

Zhiyuan profil fotoğrafı
Zhiyuan11 gün önce

@MIT_CSAIL solid work! may you good luck!

Yuxi Xiao profil fotoğrafı
Yuxi Xiao11 gün önce

@MIT_CSAIL Wow nice video for the phd thesis!!! congrats

Big data annotation profil fotoğrafı
Big data annotation11 gün önce

@MIT_CSAIL 大佬,牛逼

Raj Kachhadiya (KARAM) profil fotoğrafı
Raj Kachhadiya (KARAM)11 gün önce

@MIT_CSAIL Interesting, congratulations!

Anand Kumar profil fotoğrafı
Anand Kumar11 gün önce

@MIT_CSAIL Interesting. Looking forward to it.

Neil profil fotoğrafı
Neil8 gün önce

@MIT_CSAIL Wow this is such an interesting research direction, do you think the future of 3D is headed this way?

Yossi Eliaz profil fotoğrafı
Yossi Eliaz11 gün önce

@MIT_CSAIL loved the video!

likhesh profil fotoğrafı
likhesh11 gün önce

@MIT_CSAIL Amazing, what time exactly?

Mandar Wagh profil fotoğrafı
Mandar Wagh10 gün önce

congrats, and that is a committee. the split i keep getting stuck on: physics constraints are mostly differentiable, so they go straight into a loss. manufacturing constraints mostly are not — tool reach, support-free overhang, demoldability, minimum wall are combinatorial, and you cannot gradient your way to a draft angle. does the thesis treat those as one kind of constraint, or as two different problems? genuinely curious which way you landed.

the guy profil fotoğrafı
the guy11 gün önce

@MIT_CSAIL 🧡

Jesse Jr Lim (林振燊) profil fotoğrafı
Jesse Jr Lim (林振燊)11 gün önce

@MIT_CSAIL wow very cool.. paused the vid alot to see whats going on

X.M. profil fotoğrafı
X.M.11 gün önce

@MIT_CSAIL all the best king

Volatile Markets profil fotoğrafı
Volatile Markets11 gün önce

@MIT_CSAIL Love this!

Verdi profil fotoğrafı
Verdi11 gün önce

@MIT_CSAIL congrats!! you didnt share a time tho? > "on Tuesday, September 8, 2026, at PM ET."

Jacques Mₜ=∏ₛ((1−λₛ)+λₛEₛ) profil fotoğrafı
Jacques Mₜ=∏ₛ((1−λₛ)+λₛEₛ)11 gün önce

@MIT_CSAIL PhaGGOT

Curtis profil fotoğrafı
Curtis10 gün önce

@MIT_CSAIL Grest work, great slides, great marketing video!

Black Canvas profil fotoğrafı
Black Canvas11 gün önce

@MIT_CSAIL Wish you the best and interesting work!

Supratik Bhattacharya profil fotoğrafı
Supratik Bhattacharya11 gün önce

@MIT_CSAIL Wow! this is super interesting. all the best!

Benzer Videolar

Excited to show some surprising inventions on generative multiplayer games we made at Google with Stanford. We call the work MultiGen. I've always been inspired by early studios like id Software with Doom or Blizzard with Warcraft bringing networked video games to the next level. We are at the point in history where we can make strides like them, but for generative games. It's a strange feeling to be in the age of generative video games while still discovering how exactly to train the models and design the tools that make them useful. All of the tools that have been invented for classic game engines need to be redesigned for generative games. For example level and world design is not entirely possible with existing technology. We introduce editable memory to diffusion game engines that allow for design of new levels via a minimap. But we can easily imagine how this can be expanded with different creation tools. The end goal of this research direction is to allow game designers to be able to guide the generation process of their world, at the granularity that they prefer. Editable memory also allows us to add multiplayer to Generative Doom. We were amazed when we saw GameNGen some years ago, and now you can play it live with friends in real-time, on your couch or even online. Shared representations like our editable memory seem like the future for this type of experience. Models are, in some cases, expensive and approximate encoders but great interpolators and extrapolators. Leveraging their strengths lets you have completely new experiences that can be realized now and not in the distant future. This work was started at my previous team and continued in collaboration with Stanford. Congratulations to all for the discoveries.

Nataniel Ruiz

104,932 görüntüleme • 6 ay önce

Everything you love about generative models — now powered by real physics! Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,821,580 görüntüleme • 1 yıl önce

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,079,815 görüntüleme • 3 ay önce

Real-time world models represent a fundamental shift in AI. reactor is building the platform for real-time generative video infrastructure, supporting developers who need the tech for use across entertainment, physical AI, and robotics. Co-founders Alberto and Bryce Schmidtchen joined us last week on The Investment Memo, hosted by Partners Bucky Moore and Amber Yang, to talk about the era of world models. The conversation centered around the infrastructure Reactor is building, why real-time models are the edge right now, and current use cases for the product. Alberto and Bryce agreed that world models are shaping the way simulations are created, and that developers need a streamlined platform that can support their ideas. We believe Reactor is positioned to be at the frontier of research into real-time generative models. We look forward to seeing how these models apply across industries. Chapters 00:00 Introduction & Overview of Reactor 01:08 Meet the Hosts & Founders 02:18 The Origin Story: From 3D Assets to World Models 05:07 Real-Time Video Applications Across Industries 06:55 The Open Source World Model Explosion 07:23 Why Infrastructure Is the Opportunity 08:42 Parallels to Past Technology Waves 09:51 Bridging the Research-to-Production Gap 13:13 What Developers Are Building with World Models 16:41 Lessons from Luma AI 18:23 What Apple Vision Pro Taught Bryce About Real-Time Systems 20:48 Company Values & Team Culture 22:40 Series A: What the Capital Unlocks 24:13 Reactor's Five-Year Vision 26:09 Closing Remarks

Lightspeed

144,942 görüntüleme • 3 ay önce

Demis Hassabis on the limit in today’s AI: language can describe the world, but it cannot contain it - and why "World Models" are his "longest standing passion". Language models absorbed far more structure about reality from text than many researchers expected, because human language quietly carries physics, psychology, culture, tools, plans, and cause-and-effect. But text is still a compressed residue of experience, not experience itself. A sentence can say a cup falls from a table, yet it does not fully encode weight, grip, balance, friction, timing, sound, surprise, or the tiny motor corrections a body makes before it even notices them. The world is not only made of facts that can be named; it is made of constraints that have to be lived through, touched, predicted, violated, and repaired. That is why world models matter. They aim to learn the hidden grammar of physical reality: how objects persist, how forces unfold, how space changes when an agent moves, and how action creates feedback. Language models can often reason about the world because people have written so much about it. World models try to learn what the world is like before it becomes words. The difference is exactly what matters because intelligence is not just answering well; it is knowing what would happen next if you moved, reached, pushed, smelled, slipped, or failed. A mind trained only on descriptions may become brilliant at explanation. A mind trained on experience may become better at consequence. --- Full video from "Google DeepMind" and "Hannah Fry" YT channel (link in comment)

Rohan Paul

49,938 görüntüleme • 3 ay önce