Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

LLM post-training used to mean fine-tuning to a downstream task Robotics has been stuck in this setting, needing task-specific fine-tuning for best performance π07 changes this: It works out of the box & outperforms fine-tuned specialists Details:

77,108 Aufrufe • vor 5 Monaten •via X (Twitter)

32 Kommentare

Profilbild von Chelsea Finn
Chelsea Finnvor 5 Monaten

A few highlights of what makes π0.7 special: 1. It achieves dexterity and precision of fine-tuned models. Check out this 1x speed video of a sub-mm precision arm assembly subtask.

Profilbild von Chelsea Finn
Chelsea Finnvor 5 Monaten

2. It achieves zero-shot cross-embodiment transfer, across drastically different robot platforms. No training data for folding was collected on this robot platform.

Profilbild von Chelsea Finn
Chelsea Finnvor 5 Monaten

3. It generalizes compositionally to new tasks, like interacting with appliances that are barely represented in the pre-training data

Profilbild von Chelsea Finn
Chelsea Finnvor 5 Monaten

4. Out of the box, it achieves reliability & throughput that matches or exceeds that of pi*06. This is without any fine-tuning.

Profilbild von Chelsea Finn
Chelsea Finnvor 5 Monaten

We share many details & experimental results in the blog post and paper! Blog post: Paper:

Profilbild von Reppo
Reppovor 5 Monaten

Very cool! We bring prediction markets to the fix to provide real time feedback on specialized tasks.

Profilbild von Erik Schluntz
Erik Schluntzvor 5 Monaten

I'm glad the knife is tied to the table!

Profilbild von Elodie 3DFontaine
Elodie 3DFontainevor 5 Monaten

The knife mounting is giving me flashbacks to thermistor soldering, but the generalization is solid. End-to-end over hand-coded trajectories.

Profilbild von andrew
andrewvor 5 Monaten

this video is actually wildly impressive. that's such a hard task for a robot.

Profilbild von Utkarsh
Utkarshvor 5 Monaten

Honestly, out of the box beating specialists is the threshold that actually matters. Really curious to see how it holds up outside the training distribution, because that’s where most robotics models still quietly fall apart.

Profilbild von Josef Pa
Josef Pavor 5 Monaten

Why is the knife tied to the table 🤨

Profilbild von Aryan Dhawan
Aryan Dhawanvor 28 Tagen

Out-of-the-box transfer is the part I care about. A robot at home can’t realistically keep a separate fine-tuned specialist per chore.

Profilbild von Adam Koszek
Adam Koszekvor 4 Monaten

It will be cool to see new range of products that get better for people just because robots require them. Looking at this video, first of all, I'm pretty impressed by what the robot can do. Finally, it's something fairly practical. But also the cutting board could benefit from more stability.

Profilbild von Priyesh Gandhi
Priyesh Gandhivor 5 Monaten

Zero-shot task transfer is the holy grail and π07 finally shows it's tractable. Huge work. The next frontier: scaling the pretraining data distribution itself. Most of what models haven't seen isn't exotic — it's the messy long-tail of non-Western kitchens, tools, and object variations that don't exist in any lab dataset. We're collecting that distribution shift — 5,000+ hrs across diverse Indian households. Would be curious how π07 performs on it.

Profilbild von Alexis
Alexisvor 1 Monat

instructions unclear, stabbing the robot owner

Profilbild von just10101
just10101vor 5 Monaten

Shameless @danfei_xu Stop stealing from students/ interviewees Shame on @gtcomputing

Profilbild von Danielius Stasiulis
Danielius Stasiulisvor 5 Monaten

wow, this is cool!

Profilbild von Chriminal
Chriminalvor 5 Monaten

Microplastics

Profilbild von Amit
Amitvor 5 Monaten

True progress in robotics comes when models transcend brittle task-specific fine-tuning and instead internalize a generalized world model. π07 signals a shift toward foundation models that actually encode transferable priors about physics and interaction. This is moving from brittle scripts to robust competence—finally, robots starting to show up ready for the open world rather than the sandbox.

Profilbild von AiGentsy
AiGentsyvor 3 Monaten

This is really interesting because if π07-style generalization keeps working, the constraint shifts. It’s less “can we fine tune a robot for this task?” and more “when should a generalist robot policy be allowed to create real-world consequence?” That’s the problem I’ve been building around with AiGentsy. Acceptance gates for autonomous work. For robotics, a trace is not enough. You need proof of what happened, who/what accepted it, what was refused, and why the action was allowed to move downstream. Would be very interested in how you think about acceptance/rejection gates around generalist robot policies.

Profilbild von Nishanth Rajkumar
Nishanth Rajkumarvor 28 Tagen

zero shot scam

Profilbild von 냉동참치/冷凍ツナ
냉동참치/冷凍ツナvor 1 Monat

우왕!! 디게 싱기하다!!! 제일 대단하다고 생각하는건 칼질이 아니라 칼을 자기가 직접 꼽아 넣는단 거임! 써는건 어떻게든 되지만 칼을 자기가 집는게 중요한거임!!

Profilbild von Arjun Raj Jain
Arjun Raj Jainvor 2 Monaten

Coding is oddly still stuck in the pre-pi07 setting, general-purpose models still lean on task-specific fine-tunes for real production repos, since the training data for that is scarce and mostly synthetic. Curious if the same out-of-the-box generalization is coming for coding.

Profilbild von Simon Manning
Simon Manningvor 5 Monaten

Yes but its not even as good as FSD version 11.1

Profilbild von Prakhar Kaushik ✈️ ECCV
Prakhar Kaushik ✈️ ECCVvor 5 Monaten

What does out of the box mean exactly? Has there it seen no data of cutting a similar vegetable with a similar knife? Is there no error correction or finetuning with generative data?

Profilbild von John Wei
John Weivor 1 Monat

task specific fine tuning might finally be on the way out

Profilbild von Steven Cheng
Steven Chengvor 5 Monaten

终于等到开箱即用的机器人模型了!我们实验室那台delta臂还在为每个新任务重训,泪目 😅

Profilbild von ar0cket1
ar0cket1vor 5 Monaten

its out, no way

Profilbild von Srijan
Srijanvor 25 Tagen

The jump to generalist models shifts the data bottleneck too -- less about task variety, more about environmental diversity at scale. The same policy needs real-world variation across geographies and conditions to generalize. That coverage is the hard collection problem.

Profilbild von XENON
XENONvor 3 Monaten

This is precision autonomy in action 🤖 Fine-tuned manipulation, adaptive control, and real-world task execution at 5x speed , exactly where embodied AI is heading. For brands building in robotics, fine-tuned autonomous systems, or adaptive workload execution , offers a strategic fit. USPTO-verified, built on a Sensor-Wide-Adaptive-Workload-Relay framework. The name carries deep linguistic roots Proto-Germanic, Hebrew, and Arabic origins meaning oath, protection, and purposeful reach a compact, globally pronounceable identity engineered for next-generation autonomy. For more information, visit

Profilbild von PsudoMike 🇨🇦
PsudoMike 🇨🇦vor 5 Monaten

Same shift NLP went through around 2019 when transfer learning killed task specific fine tuning. Robotics getting there is bigger because physical tasks have way less data than text. Curious how it handles novel objects and environments it never saw in training.

Profilbild von Kamesh 🇺🇸
Kamesh 🇺🇸vor 5 Monaten

This is where AI stops being software and starts behaving like a system that can act in the real world without constant retraining.

Ähnliche Videos

New Course: Reinforcement Fine-Tuning LLMs with GRPO! Learn to use reinforcement learning to improve your LLM performance in this short course, built in collaboration with Predibase by Rubrik, and taught by Travis Addair, its Co-Founder and CTO, and Arnav Garg, its Senior Engineer and Machine Learning Lead. Reasoning models have been one of the most important developments in LLMs. Reinforcement Fine-Tuning (RFT) uses rewards to encourage LLMs to find solutions to multi-step reasoning tasks such as solving math problems and debugging code - without needing pre-existing training examples like in traditional supervised fine-tuning. Group Relative Policy Optimization (GRPO) is a reinforcement fine-tuning algorithm gaining rapid adoption. Developed by the DeepSeek team and used to train the R1 reasoning model, GRPO uses reward functions that you can write in Python to assign rewards to model responses. It’s beneficial for tasks with verifiable outcomes and can work well even with fewer than 100 training examples. It can also significantly improve the reasoning ability of smaller LLMs, making applications faster and more cost effective. In this course, you’ll take a technical deep dive into RFT with GRPO. You’ll learn to build reward functions that you can use in the GRPO training process to guide an LLM toward better performance on multi-step reasoning tasks. In detail, you’ll: - Learn when reinforcement fine-tuning is a better fit than supervised fine-tuning, especially for tasks involving multi-step reasoning or limited labeled data. - Understand how GRPO uses programmable reward functions as a more scalable alternative to the human feedback required for other reinforcement learning algorithms, such as RLHF and DPO. - Frame the Wordle game as a reinforcement fine-tuning problem and see how an LLM can learn to plan, analyze feedback, and improve its strategy over time. - Design reward functions that power the reinforcement fine-tuning process. - Learn techniques for evaluating more subjective tasks, such as rating the quality of a text summary, using an LLM as a judge. - Understand why reward hacking happens and how to avoid it by adding penalty functions to discourage undesirable behaviors. - Learn the four key components of the loss calculation in the GRPO algorithm: token probability distribution ratios, advantages, clipping, and KL-divergence. - Launch reinforcement fine-tuning jobs using Predibase’s hosted training services. By the end of this course, you’ll be able to build and fine-tune LLMs using reinforcement learning to improve reasoning without relying on large labeled datasets or subjective human feedback. Please sign up here:

Andrew Ng

86,697 Aufrufe • vor 1 Jahr