Loading video...
Video Failed to Load
Robots can already fold laundry, make espresso, clean kitchens, and assemble things. The harder problem is getting them to do those tasks reliably, for long periods of time, without a human babysitting them. At Startup School 2026, Physical Intelligence cofounder Chelsea Finn explains what it takes to build general-purpose... show more
112,717 views • 1 month ago •via X (Twitter)
27 Comments

Tune in: Transcript:

@physical_int @chelseabfinn Love that they used laundry as the test. Growing up in dry cleaning, the hard part was doing it right every day. Reliability is the product.

@physical_int @chelseabfinn Love seeing thoughtful work like this.

@physical_int @chelseabfinn Organizations adopting reinforcement learning typically succeed when they integrate it into broader strategic goals in real-world applications.

@physical_int @chelseabfinn exactly. a 30 second demo proves basically nothing

@physical_int @chelseabfinn Reliability is the moat here. A robot that folds laundry once is a demo. One that runs 13 hours without anyone babysitting it is a business.

@physical_int @chelseabfinn Robotics are in our near future. We need to make sure we pay attention to safety, including transparency, and our ability to audit all the important or mundane things we will have robots doing for us when we’re out having fun. @OfficialXYO

@physical_int @chelseabfinn The reliability problem is where AI and robotics become truly interesting. Moving from specialized systems to general-purpose AI that can operate reliably for hours in real-world environments is a much bigger leap than simply making robots perform more tasks.

"Without a human babysitting them" is the whole problem, and it shows up in software agents too. I gave mine 53 tools instead of 3. It called zero of them, answered from memory, wrong, on 4 runs out of 4. Not one error raised. Reliability isn't the capability. It's whether the failure is visible.

@garrytan imo you guys should invite @Lindon_Gao and @YorkYang5050 to give a talk - YC 16 alumni and knowing what I know I think they’re gonna do well long term

@physical_int @chelseabfinn Reliability is where demos become businesses. The commercial threshold is not whether a robot can do the task once, but whether operators can trust the exception rate and recovery path.

@physical_int @chelseabfinn Reliability at scale is where the physical world meets the rigor of software. In Japan, we’re seeing a fascinating shift: blending precision robotics with generative AI to lower that "babysitting" threshold. It's the new frontier for autonomy.

@physical_int @chelseabfinn most people don't realize the real challenge isn't programming robots to perform tasks, but making them do so in complex environments withou

@physical_int @chelseabfinn “Can it do the task?” is becoming the easy question. I care more about: how often does it need rescuing, does it know when it’s outside its bounds, and what happens when nobody is watching? That’s the difference between automation and babysitting.

@physical_int @chelseabfinn That reliability gap has a direct organizational parallel. Capability gets attention, but repeatable performance comes from feedback loops, memory and learning from failure. That’s usually what separates a demo from an operating system.

@physical_int @chelseabfinn the reliability problem is the real robotics problem

Reliability in production is never about executing the happy path, it is about how gracefully the system handles exceptions. Whether it is a CI pipeline, a distributed cloud database, or a physical robot, the real breakthrough is self-healing. True autonomy starts when the system learns to recover from its own failures without triggering an alert for a human.

@physical_int @chelseabfinn YC, you’ll have to watch this speech. YC and Steve Hilton @SteveHiltonx should have a talk about the future of California. He’s got all the right ideas and he’s calling for a decade of building for California in this great speech.

Le problème avec ce type de présentation n’est pas que Physical Intelligence mente sur ses résultats. Les démonstrations existent : leurs robots peuvent réellement plier du linge, préparer des expressos, nettoyer une cuisine, assembler des objets et, dans certains cas, exécuter ces tâches de manière autonome pendant plusieurs heures. Ce qui est survendu, c’est la portée de ces résultats. Ces capacités sont obtenues au terme d’un processus d’entraînement lourd et coûteux : collecte massive de données robotiques, téléopération et démonstrations humaines, pré-entraînement, fine-tuning, reinforcement learning, corrections et répétitions. Le reinforcement learning peut ensuite améliorer fortement la robustesse et le débit d’une politique déjà entraînée. Mais exécuter avec fiabilité une tâche réelle dans un contexte pour lequel le système a été entraîné n’est pas disposer d’une compétence générale transposable dans le monde réel. Le problème apparaît lorsqu’on change véritablement le contexte : autre environnement, disposition différente, objets inconnus, morphologie ou dynamique du robot différente, nouvelles contraintes physiques, événements imprévus ou tâche absente de l’entraînement. Les capacités acquises ne se transposent pas automatiquement. Il faut généralement de nouvelles données, de l’adaptation, du fine-tuning ou un nouvel entraînement. Et encore une fois, c'est très coûteux. C’est une distinction essentielle lorsqu’on parle de « monde réel ». Un vrai robot manipulant de vrais objets n’est pas, à lui seul, une démonstration de généralisation au monde réel. Le monde réel signifie précisément l’ouverture du contexte : variations non prévues, événements rares, changements de friction ou de charge, objets inconnus, erreurs qui s’accumulent, interruptions, interactions humaines et situations qui ne figuraient pas dans les données d’entraînement. Il faut donc distinguer : Exécuter une tâche physique réelle après un entraînement massif et transposer spontanément cette compétence dans un contexte réellement différent faire face à une situation imprévue et y acquérir une nouvelle compétence intégrer durablement cette nouvelle expérience sans perdre les compétences précédemment acquises continuer ainsi pendant des jours, des mois ou des années en accumulant une histoire propre. C’est ici qu’apparaît une autre limite fondamentale : l’apprentissage continu. Les architectures neuronales actuelles restent confrontées à l’oubli catastrophique. Lorsqu’on modifie leurs paramètres pour apprendre successivement de nouvelles tâches ou distributions, les nouvelles acquisitions peuvent dégrader les anciennes. Replay, régularisation, architectures modulaires ou mémoires externes permettent de réduire le problème, mais il n’existe pas aujourd’hui de solution générale permettant à un robot d’apprendre continuellement tout au long de son existence, sans réentraînement lourd et sans dégradation de ses acquis. Enfin, eux n'ont pas la solution. Même l’affirmation selon laquelle ces robots peuvent fonctionner « autonomously for hours » doit donc être correctement interprétée. Faire fonctionner pendant dix heures une politique devenue robuste dans son domaine appris est une remarquable démonstration de robustesse d’exécution. Ce n’est pas démontrer qu’un robot peut s’adapter et apprendre pendant dix heures dans un monde ouvert qui change. Le passage vers des systèmes plus généralistes est réel. Les modèles peuvent couvrir davantage de tâches, de robots et de situations qu’auparavant. Mais plus généraliste ne signifie pas général.

@physical_int @chelseabfinn Reliable performance in the real world is very different from a controlled demo. The interesting part is not only what a robot can do once, but how consistently it can do it, recover from failure and be maintained economically.

@physical_int @chelseabfinn After artificial intelligence, it is an era of physical intelligence it seems

@physical_int @chelseabfinn A robot that works for five minutes is a demo. One that works for five hours without intervention is a product.

"Without a human babysitting them" changes the problem from capability to failure recovery. A demo needs one successful run. Eight hours unattended means handling every abnormal state the robot itself creates — a sleeve caught on a drawer handle, a portafilter seated wrong, a counter that no longer resembles anything in training because a previous grasp slipped. The scarce data isn't successful trajectories, it's recovery-from-arbitrary-failure trajectories, and that's the hardest kind to collect: you have to make it fail first, and fail in enough different ways to matter.

@physical_int @chelseabfinn one bad grasp near a hot stove and 13 unsupervised hours is the headline. software agents fail for free. robots don't get that deal.

@physical_int @chelseabfinn Great overview! 🤖

@physical_int @chelseabfinn Reliability is a different problem from capability, and it gets budgeted last. A demo shows the best run; deployment is priced on the worst one. What decides it is recovery — whether the robot notices it failed and carries on without a human standing by.

@physical_int @chelseabfinn

