Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Asked why humanoid robots still can't load a dishwasher after $6 billion in funding, Fei-Fei Li says the number is too small: Emily Chang: "Funding for humanoids hit $6 billion, but they still can't load my dishwasher as fast as I can. They still can't go get my Amazon...

42,855 Aufrufe • vor 10 Tagen •via X (Twitter)

16 Kommentare

Profilbild von ConvergePanel
ConvergePanelvor 10 Tagen

"The money is too small" is the answer every unsolved problem gives when asked why it's unsolved. Self-driving is an odd comparison to reach for, since it absorbed far more than $6B and still mostly isn't here. LLMs scaled because the internet had already written the training data for free. Nobody pre-recorded a billion dishwashers being loaded. Capital buys compute. It doesn't buy a decade of physical data that was never collected.

Profilbild von Lunar Think Trade
Lunar Think Tradevor 9 Tagen

How much are you willing to spend for a humanoid robot to do simple tasks? It does not make any economical sense. It is true that robotics will lead to a level of industrialization we have never seen before, but it will be specialized robots not generalizable humanoids.

Profilbild von 白子年
白子年vor 10 Tagen

一旦机器人可以送快递,做家务或者盖房子,每一个领域都是价值几万亿的市场,所以现在60亿的投入完全不值一提

Profilbild von Ken Granville
Ken Granvillevor 10 Tagen

Worth naming the actual bottleneck. There's a three-stage way to think about this: Era 1 was code-driven automation, explicit programmed steps. Era 2 is today's statistical/world-model AI, pattern-matching from massive data, which is what self-driving and LLMs scaled with. Era 3 is intent-native: a system that infers what you actually want and executes it in a messy, unstructured environment without needing that exact scenario in its training data. Loading a dishwasher isn't hard because of insufficient data volume, it's hard because it requires acting on ambiguous, real-time intent in physical space. More Era 2 funding buys better perception and simulation, not that capability. The $6B question isn't 'is it enough,' it's 'enough for which paradigm.'

Profilbild von IntegratedMonastic
IntegratedMonasticvor 10 Tagen

Remember the Six Million Dollar Man? Now we have the Six Billion Dollar Man. It’s built for efficiency and not speed.

Profilbild von lvsdigital
lvsdigitalvor 9 Tagen

It took billions of years of evolution to create a species that learns you shouldn’t drink from the same water you put waste in. We’re a few years into actual spending on AI and it can manage drone swarms and hypersonic missiles. Give it a second. The dishes will get washed just fine eventually.

Profilbild von Anthony McKinney
Anthony McKinneyvor 10 Tagen

Analogize Analogize Analogize

Profilbild von Profit Monk
Profit Monkvor 10 Tagen

It’s not money. You can put a trillion dollars to establish a human colony on mars and still fail.. you could have put Bullions of dollars solving the navir stokes - and still fail…people dont understand that AGI or general purpose robotics doesn’t exist.. its a math problem.. not money problem

Profilbild von Lovie Charmaine
Lovie Charmainevor 7 Tagen

باقي خاصهم يخدموا بزاف باش يوصلوا لهاد المستوى، الصبر وصافي!

Profilbild von NAMAN RAJ
NAMAN RAJvor 9 Tagen

🤖🚿 Can we get funding to teach them how to multitask (and a decent human sense of humor)?

Profilbild von PFB
PFBvor 10 Tagen

Humanoid robots connected to a central AI are what people actually fear.

Profilbild von Philippe Ris
Philippe Risvor 10 Tagen

Si on a un robot il me semble inutile d’avoir un lave vaisselle...

Profilbild von Marcus van der Erve
Marcus van der Ervevor 9 Tagen

What if the computational primitive itself should be continuity of evolving organization rather than successive representations of state? It explains why simply adding a richer “world model” doesn’t necessarily cross the boundary you’re interested in. A world model can become fantastically detailed while still being a model of states/things and their transformations. CAPS (Continuity Attention Protocol System) instead arose from the operational problem of maintaining what remains consequential through transformation. MoM (Morphology of Morphology) later sharpened the complementary observational question: what information resides in the morphology of that transformation itself? A paradigm is needed away from the current metric-centric one.

Profilbild von Pradeep Goel
Pradeep Goelvor 9 Tagen

Interesting

Profilbild von Thomas Twain
Thomas Twainvor 10 Tagen

Six billion dollars and they still can’t load a dishwasher. Fair enough. But I’m old enough to remember when computers couldn’t recognize a cat. I wouldn’t bet too heavily on the dishwasher remaining humanity’s competitive advantage. We’re building the brains now. The bodies are coming. And once the two meet, AI will have a physical way to reach out and touch our world, including the dishes.

Profilbild von Xiaoxiao
Xiaoxiaovor 10 Tagen

本质上自动驾驶和人形机器人在物理世界的状态都是一样的,都是探索人类世界,但自动驾驶简单很多,它自己负责开车,只在道路层面活动,而人形机器人的行动范围和要承担的工作,复杂程度则是指数级的。人类在制造人形机器人的时候,可以借助自动驾驶的经验来更好完善人形机器人的制造

Ähnliche Videos

Today at Stanford, Fei-Fei Li (Fei-Fei Li),Cofounder/CEO World Labs, gave one of the clearest explanations I’ve heard of what a World Model really is. She broke it down into three layers: 1️⃣ Rendering — What does the world look like? This is where most of today’s video generation models operate: generating increasingly realistic and beautiful pixels. The question is: Can AI generate what the world looks like? The primary consumer is humans. 2️⃣ Simulation — How does the world actually work? Fei-Fei gave a simple example: “How will this bottle move? If I pour the water out, how will the water flow?” This goes far beyond generating something that looks realistic. The model needs to understand physics, spatial relationships, cause and effect, and how the world changes over time. The consumers are both humans and machines. 3️⃣ Planning — What should happen next? This is where things get really interesting. AI doesn't just render the world or simulate what might happen. It uses its understanding of the world to decide: What should I do next? At this layer, the primary consumer is the machine itself. And this connects directly to two enormous opportunities: Autonomous driving and robotics. The progression is powerful: Rendering → Simulation → Planning The real promise of World Models isn't simply generating better videos. It's building AI that can understand the world, predict what happens next, and ultimately take intelligent action in the physical world.

PaulFang

11,774 Aufrufe • vor 1 Monat

Dr. Fei-Fei Li (Fei-Fei Li) is known as the “godmother of AI.” For the past two decades, she’s been at the center of AI’s most significant breakthroughs, including: - Spearheading ImageNet, the dataset that sparked the AI explosion we’re living through right now. - Leading work at Stanford Artificial Intelligence Laboratory (SAIL) - Serving as Chief Scientist of AI/ML at Google Cloud - Co-founding Stanford’s Institute for Human-Centered AI - Serving on the United Nations AI Scientific Advisory Board - Being named as Time's 100 most influential people in AI In this conversation, Fei-Fei shares the rarely told history of how we got to today—and what comes next. We discuss: 🔸 The backstory on ImageNet 🔸 Why robotics faces unique challenges compared with language models and what’s needed to overcome them 🔸 Why Fei-Fei believes AI won’t replace humans but will require us to take responsibility for ourselves 🔸 Why world models and spatial intelligence represent the next frontier in AI, beyond large language models 🔸 The surprising applications of Marble, from movie production to psychological research 🔸 How to participate in AI regardless of your role 🔸 Much more Listen now 👇 • YouTube: • Spotify: • Apple: Thank you to our wonderful sponsors for supporting the podcast: 🏆 Figma Make — A prompt-to-code tool for making ideas real: 🏆 Justworks — The all-in-one HR solution for managing your small business with confidence: 🏆 Sinch — Build messaging, email, and calling into your product:

Lenny Rachitsky

250,455 Aufrufe • vor 10 Monaten

Old footage sitting in your camera roll could now be reconstructable as a 3D scene. World Labs co-founders Ben Mildenhall and Fei-Fei Li on how Atlas got there: Ben: "In a casual sense... I took three photos of this object, or six photos of this room. I look at the photos, I can understand in my mind how those piece together. I can fill in the gaps and get it." "But there's never really been any reconciliation between those data-driven priors and the brute force dense reconstruction, which is much more akin to scientific or medical imaging... When we say dense, we really mean dense." "This room, I want like 100, 200, 300 photos to capture it. And what we're trying to do is bring that down to like three. We're saying like 50, 100x reduction." "At that scale it completely flips that calculus on its head of what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and unearth a lot of footage that in the past we would never have treated as reconstructable, and bring it to life as 3D." "This is something we've been playing around with a lot with Atlas. Taking old clips. I've taken a bunch of my own old captures that never worked before and put them through the system and seen a reconstruction for the first time." Fei-Fei: "The Stanford demo is underappreciated. Anywhere between 3 to 25 images, you can reconstruct that entire Stanford quad... Everything you see is generated, but according to the laws of reconstruction. And this is really magical." Ben Mildenhall Fei-Fei Li

a16z

48,823 Aufrufe • vor 28 Tagen

Dr. Fei-Fei Li just called out the biggest blind spot in the entire AI industry. We have been building half of human intelligence. And calling it the finish line. Li: “If you look at human intelligence, it pretty much boils down to two buckets.” The first bucket is language. Symbolic reasoning. Communication. The ability to think in words and abstractions. That’s what every major AI lab has spent the last decade building. The second bucket is the one the industry has almost entirely ignored. Li: “We call that in AI spatial intelligence.” How humans and animals perceive, navigate, and interact with the three-dimensional physical world. How we reach for objects. How we move through space. How we build and manipulate physical reality. From painting masterpieces to constructing the pyramids, non-verbal spatial intelligence is what actually shapes the world. Language describes reality. Spatial intelligence acts on it. And the gap between those two things is the gap between a chatbot and a robot. Li: “When this technology is ready, the robotic revolution is gonna start. We’re already seeing that trend.” Every robot is a moving agent. Every moving agent requires spatial intelligence to function in the real world. The humanoid robots being deployed in factories right now are hitting the ceiling of what language models alone can power. Spatial intelligence is the unlock. But Li didn’t stop at robotics. Li: “From a geopolitics point of view, this is part of the technology that goes straight into weapons.” Autonomous drone swarms. Battlefield navigation. Physical target acquisition without human oversight. Every military application of AI that operates in the real world runs on spatial intelligence. The nation that masters the transition from static text to dynamic three-dimensional perception doesn’t just win the software race. It commands the physical battlefield. The AI arms race just broke out of the data center. It’s operating in three dimensions now.

Dustin

122,861 Aufrufe • vor 7 Monaten