Loading video...

Video Failed to Load

Go Home

LLM post-training used to mean fine-tuning to a downstream task Robotics has been stuck in this setting, needing task-specific fine-tuning for best performance π07 changes this: It works out of the box & outperforms fine-tuned specialists Details:

77,108 views • 5 months ago •via X (Twitter)

32 Comments

Chelsea Finn's profile picture
Chelsea Finn5 months ago

A few highlights of what makes π0.7 special: 1. It achieves dexterity and precision of fine-tuned models. Check out this 1x speed video of a sub-mm precision arm assembly subtask.

Chelsea Finn's profile picture
Chelsea Finn5 months ago

2. It achieves zero-shot cross-embodiment transfer, across drastically different robot platforms. No training data for folding was collected on this robot platform.

Chelsea Finn's profile picture
Chelsea Finn5 months ago

3. It generalizes compositionally to new tasks, like interacting with appliances that are barely represented in the pre-training data

Chelsea Finn's profile picture
Chelsea Finn5 months ago

4. Out of the box, it achieves reliability & throughput that matches or exceeds that of pi*06. This is without any fine-tuning.

Chelsea Finn's profile picture
Chelsea Finn5 months ago

We share many details & experimental results in the blog post and paper! Blog post: Paper:

Reppo's profile picture
Reppo5 months ago

Very cool! We bring prediction markets to the fix to provide real time feedback on specialized tasks.

Erik Schluntz's profile picture
Erik Schluntz5 months ago

I'm glad the knife is tied to the table!

Elodie 3DFontaine's profile picture
Elodie 3DFontaine5 months ago

The knife mounting is giving me flashbacks to thermistor soldering, but the generalization is solid. End-to-end over hand-coded trajectories.

andrew's profile picture
andrew5 months ago

this video is actually wildly impressive. that's such a hard task for a robot.

Utkarsh's profile picture
Utkarsh5 months ago

Honestly, out of the box beating specialists is the threshold that actually matters. Really curious to see how it holds up outside the training distribution, because that’s where most robotics models still quietly fall apart.

Josef Pa's profile picture
Josef Pa5 months ago

Why is the knife tied to the table 🤨

Aryan Dhawan's profile picture
Aryan Dhawan27 days ago

Out-of-the-box transfer is the part I care about. A robot at home can’t realistically keep a separate fine-tuned specialist per chore.

Adam Koszek's profile picture
Adam Koszek4 months ago

It will be cool to see new range of products that get better for people just because robots require them. Looking at this video, first of all, I'm pretty impressed by what the robot can do. Finally, it's something fairly practical. But also the cutting board could benefit from more stability.

Priyesh Gandhi's profile picture
Priyesh Gandhi5 months ago

Zero-shot task transfer is the holy grail and π07 finally shows it's tractable. Huge work. The next frontier: scaling the pretraining data distribution itself. Most of what models haven't seen isn't exotic — it's the messy long-tail of non-Western kitchens, tools, and object variations that don't exist in any lab dataset. We're collecting that distribution shift — 5,000+ hrs across diverse Indian households. Would be curious how π07 performs on it.

Alexis's profile picture
Alexis1 month ago

instructions unclear, stabbing the robot owner

just10101's profile picture
just101015 months ago

Shameless @danfei_xu Stop stealing from students/ interviewees Shame on @gtcomputing

Danielius Stasiulis's profile picture
Danielius Stasiulis5 months ago

wow, this is cool!

Chriminal's profile picture
Chriminal5 months ago

Microplastics

Amit's profile picture
Amit5 months ago

True progress in robotics comes when models transcend brittle task-specific fine-tuning and instead internalize a generalized world model. π07 signals a shift toward foundation models that actually encode transferable priors about physics and interaction. This is moving from brittle scripts to robust competence—finally, robots starting to show up ready for the open world rather than the sandbox.

AiGentsy's profile picture
AiGentsy3 months ago

This is really interesting because if π07-style generalization keeps working, the constraint shifts. It’s less “can we fine tune a robot for this task?” and more “when should a generalist robot policy be allowed to create real-world consequence?” That’s the problem I’ve been building around with AiGentsy. Acceptance gates for autonomous work. For robotics, a trace is not enough. You need proof of what happened, who/what accepted it, what was refused, and why the action was allowed to move downstream. Would be very interested in how you think about acceptance/rejection gates around generalist robot policies.

Nishanth Rajkumar's profile picture
Nishanth Rajkumar27 days ago

zero shot scam

냉동참치/冷凍ツナ's profile picture
냉동참치/冷凍ツナ1 month ago

우왕!! 디게 싱기하다!!! 제일 대단하다고 생각하는건 칼질이 아니라 칼을 자기가 직접 꼽아 넣는단 거임! 써는건 어떻게든 되지만 칼을 자기가 집는게 중요한거임!!

Arjun Raj Jain's profile picture
Arjun Raj Jain2 months ago

Coding is oddly still stuck in the pre-pi07 setting, general-purpose models still lean on task-specific fine-tunes for real production repos, since the training data for that is scarce and mostly synthetic. Curious if the same out-of-the-box generalization is coming for coding.

Simon Manning's profile picture
Simon Manning5 months ago

Yes but its not even as good as FSD version 11.1

Prakhar Kaushik ✈️ ECCV's profile picture
Prakhar Kaushik ✈️ ECCV5 months ago

What does out of the box mean exactly? Has there it seen no data of cutting a similar vegetable with a similar knife? Is there no error correction or finetuning with generative data?

John Wei's profile picture
John Wei1 month ago

task specific fine tuning might finally be on the way out

Steven Cheng's profile picture
Steven Cheng5 months ago

终于等到开箱即用的机器人模型了!我们实验室那台delta臂还在为每个新任务重训,泪目 😅

ar0cket1's profile picture
ar0cket15 months ago

its out, no way

Srijan's profile picture
Srijan25 days ago

The jump to generalist models shifts the data bottleneck too -- less about task variety, more about environmental diversity at scale. The same policy needs real-world variation across geographies and conditions to generalize. That coverage is the hard collection problem.

XENON's profile picture
XENON3 months ago

This is precision autonomy in action 🤖 Fine-tuned manipulation, adaptive control, and real-world task execution at 5x speed , exactly where embodied AI is heading. For brands building in robotics, fine-tuned autonomous systems, or adaptive workload execution , offers a strategic fit. USPTO-verified, built on a Sensor-Wide-Adaptive-Workload-Relay framework. The name carries deep linguistic roots Proto-Germanic, Hebrew, and Arabic origins meaning oath, protection, and purposeful reach a compact, globally pronounceable identity engineered for next-generation autonomy. For more information, visit

PsudoMike 🇨🇦's profile picture
PsudoMike 🇨🇦5 months ago

Same shift NLP went through around 2019 when transfer learning killed task specific fine tuning. Robotics getting there is bigger because physical tasks have way less data than text. Curious how it handles novel objects and environments it never saw in training.

Kamesh 🇺🇸's profile picture
Kamesh 🇺🇸5 months ago

This is where AI stops being software and starts behaving like a system that can act in the real world without constant retraining.

Related Videos

New Course: Reinforcement Fine-Tuning LLMs with GRPO! Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with Predibase by Rubrik, and taught by Travis Addair, its Co-Founder and CTO, and Arnav Garg, its Senior Engineer and Machine Learning Lead. Reasoning models have been one of the most important developments in LLMs. Reinforcement Fine-Tuning (RFT) uses rewards to encourage LLMs to find solutions to multi-step reasoning tasks such as solving math problems and debugging code - without needing pre-existing training examples like in traditional supervised fine-tuning. Group Relative Policy Optimization (GRPO) is a reinforcement fine-tuning algorithm gaining rapid adoption. Developed by the DeepSeek team and used to train the R1 reasoning model, GRPO uses reward functions that you can write in Python to assign rewards to model responses. It’s beneficial for tasks with verifiable outcomes and can work well even with fewer than 100 training examples. It can also significantly improve the reasoning ability of smaller LLMs, making applications faster and more cost effective. In this course, you’ll take a technical deep dive into RFT with GRPO. You’ll learn to build reward functions that you can use in the GRPO training process to guide an LLM toward better performance on multi-step reasoning tasks. In detail, you’ll: - Learn when reinforcement fine-tuning is a better fit than supervised fine-tuning, especially for tasks involving multi-step reasoning or limited labeled data. - Understand how GRPO uses programmable reward functions as a more scalable alternative to the human feedback required for other reinforcement learning algorithms, such as RLHF and DPO. - Frame the Wordle game as a reinforcement fine-tuning problem and see how an LLM can learn to plan, analyze feedback, and improve its strategy over time. - Design reward functions that power the reinforcement fine-tuning process. - Learn techniques for evaluating more subjective tasks, such as rating the quality of a text summary, using an LLM as a judge. - Understand why reward hacking happens and how to avoid it by adding penalty functions to discourage undesirable behaviors. - Learn the four key components of the loss calculation in the GRPO algorithm: token probability distribution ratios, advantages, clipping, and KL-divergence. - Launch reinforcement fine-tuning jobs using Predibase’s hosted training services. By the end of this course, you’ll be able to build and fine-tune LLMs using reinforcement learning to improve reasoning without relying on large labeled datasets or subjective human feedback. Please sign up here:

Andrew Ng

86,697 views • 1 year ago