Loading video...

Video Failed to Load

Go Home

Anyone skeptical about AI’s ability to perform diverse physical tasks should watch this. Physical Intelligence's π₀ model in action: 18 minutes of bimanual robots autonomously handling complex, dexterous chores.

210,507 views • 1 year ago •via X (Twitter)

11 Comments

Victor's profile picture
Victor1 year ago

watching robots do chores is the new asmr... kinda therapeutic

David Cantone's profile picture
David Cantone1 year ago

It may not be the best right now, but at least it can work 24 hours a day.

The Humanoid Hub's profile picture
The Humanoid Hub1 year ago

This is only the initial version with a small team working for just 8 months. The proof of existence is real. The next couple of years will be crazy exciting for foundation models.

Larry Panozzo's profile picture
Larry Panozzo1 year ago

It’s trying to use one heuristic and it fails many times over and over, stuck making no progress. If the heuristic works then yeah it folds nicely.

The Humanoid Hub's profile picture
The Humanoid Hub1 year ago

Corrections and recoveries are an emergent feature of pre-training on a large, diverse dataset.

Shane McGrath's profile picture
Shane McGrath1 year ago

Leaders in the field

MR MR's profile picture
MR MR1 year ago

Looks teleoperated

Nicholas Shanks's profile picture
Nicholas Shanks1 year ago

This looks teleoperated by a child.

Robert Newton @robertnewton.bsky.social's profile picture
Robert Newton @robertnewton.bsky.social1 year ago

I don't understand your point. the robot did a terrible job of folding those clothes.

Vlad's profile picture
Vlad1 year ago

it's tele operated obviously

The Humanoid Hub's profile picture
The Humanoid Hub1 year ago

@vzhidkovUSA It's all autonomous. The model is fine-tuned with carefully curated high-quality examples of humans operating the robot in the same settings. This induces the desired pattern of movements and dexterity similar to those in the human operated dataset.

Related Videos

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

466,226 views • 1 year ago