正在加载视频...

视频加载失败

World Labs' Dr. Fei-Fei Li on why the 3D world and spatial intelligence are fundamentally different from language models: "Language is fundamentally a purely generated signal. There's no language out there. You don't go out in the nature and there's words written in the sky for you." "Whatever data...

51,559 次观看 • 2 天前 •via X (Twitter)

12 条评论

Susie 的头像
Susie2 天前

wait this actually made spatial AI click for me

Chinsanity 的头像
Chinsanity2 天前

Yoo wheres the full podcast?

Site Specs 的头像
Site Specs2 天前

The sheet can be right and the sleeve can still miss the grid. The sheet is the generated signal, and the set location is the 3D fact. You can still measure an open sleeve or box-out from the finished opening after the forms come off. If it wasn't shot before the pour, what's lost is where it was set and any proof it didn't drift. A fully buried embed is the one you generally can't remeasure without a scan or demolition. Shoot those points before cover.

Macro Bombastic 的头像
Macro Bombastic2 天前

yeah 3d has consequences, language doesnt. thats the whole game tbh

icefrog.◎ 的头像
icefrog.◎2 天前

physics > language as a constraint. finally moving past the token hallucination era.

Youth 的头像
Youth2 天前

Idea volume without a forced Against column is just optimism with formatting.

Nishant Mantripragada 的头像
Nishant Mantripragada2 天前

the materials point is the one that sticks. a model can describe a chair perfectly and still miss what happens when someone leans back on it.

Tommy YC 的头像
Tommy YC2 天前

The world isn't a generated signal

Đỗ Khoa | Solomon Do 的头像
Đỗ Khoa | Solomon Do2 天前

Interesting distinction on spatial intelligence vs. language models from Dr. Li.

Charlie Media 的头像
Charlie Media1 天前

The next leap in AI isn't better words. It's understanding what words can't capture.

NameFave 的头像
NameFave2 天前

agree with Fei-Fei. This is the case for extended spatial: spatial intelligence beyond boundaries, into the real 3D world

Raven 的头像
Raven2 天前

i navigate by trees, not text, which explains several recent landings

相关视频

Old footage sitting in your camera roll could now be reconstructable as a 3D scene. World Labs co-founders Ben Mildenhall and Fei-Fei Li on how Atlas got there: Ben: "In a casual sense... I took three photos of this object, or six photos of this room. I look at the photos, I can understand in my mind how those piece together. I can fill in the gaps and get it." "But there's never really been any reconciliation between those data-driven priors and the brute force dense reconstruction, which is much more akin to scientific or medical imaging... When we say dense, we really mean dense." "This room, I want like 100, 200, 300 photos to capture it. And what we're trying to do is bring that down to like three. We're saying like 50, 100x reduction." "At that scale it completely flips that calculus on its head of what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and unearth a lot of footage that in the past we would never have treated as reconstructable, and bring it to life as 3D." "This is something we've been playing around with a lot with Atlas. Taking old clips. I've taken a bunch of my own old captures that never worked before and put them through the system and seen a reconstruction for the first time." Fei-Fei: "The Stanford demo is underappreciated. Anywhere between 3 to 25 images, you can reconstruct that entire Stanford quad... Everything you see is generated, but according to the laws of reconstruction. And this is really magical." Ben Mildenhall Fei-Fei Li

a16z

48,823 次观看 • 27 天前

Dr. Fei-Fei Li just called out the biggest blind spot in the entire AI industry. We have been building half of human intelligence. And calling it the finish line. Li: “If you look at human intelligence, it pretty much boils down to two buckets.” The first bucket is language. Symbolic reasoning. Communication. The ability to think in words and abstractions. That’s what every major AI lab has spent the last decade building. The second bucket is the one the industry has almost entirely ignored. Li: “We call that in AI spatial intelligence.” How humans and animals perceive, navigate, and interact with the three-dimensional physical world. How we reach for objects. How we move through space. How we build and manipulate physical reality. From painting masterpieces to constructing the pyramids, non-verbal spatial intelligence is what actually shapes the world. Language describes reality. Spatial intelligence acts on it. And the gap between those two things is the gap between a chatbot and a robot. Li: “When this technology is ready, the robotic revolution is gonna start. We’re already seeing that trend.” Every robot is a moving agent. Every moving agent requires spatial intelligence to function in the real world. The humanoid robots being deployed in factories right now are hitting the ceiling of what language models alone can power. Spatial intelligence is the unlock. But Li didn’t stop at robotics. Li: “From a geopolitics point of view, this is part of the technology that goes straight into weapons.” Autonomous drone swarms. Battlefield navigation. Physical target acquisition without human oversight. Every military application of AI that operates in the real world runs on spatial intelligence. The nation that masters the transition from static text to dynamic three-dimensional perception doesn’t just win the software race. It commands the physical battlefield. The AI arms race just broke out of the data center. It’s operating in three dimensions now.

Dustin

122,861 次观看 • 7 个月前