Loading video...
Video Failed to Load
Introducing EXPO-FT – Efficient, Reliable & Open-Source VLA Finetuning! EXPO-FT unlocks π0.5 for challenging manipulation tasks: Routing string lights & inserting the power connector to illuminate them Striking pool ball into pocket Inserting flower into wine bottle (1/5)
79,711 views • 4 months ago •via X (Twitter)
17 Comments

Unlike prior works that only train lightweight policies, rely on latent space prediction, or decoupled RL training, EXPO-FT fully finetunes the VLA with RL — leveraging EXPO's stability and sample efficiency, augmented by human-in-the-loop corrections (2/5)

Using an average of 19.1 minutes of online robot data, EXPO-FT reaches 30/30 on these tasks: Routing/powering string lights Striking a pool ball into a pocket Inserting flower into a wine bottle Scooping candy Picking up cube from large initial states Flipping egg (3/5)

EXPO-FT outperforms highly performant prior RL approaches by fully leveraging the VLA for online learning (4/5)

Project with @khhung906, @TianGao_19, @DorsaSadigh, @chelseabfinn Website: Paper: Open-source code coming soon! (5/5)

Really cool work :)

Cool work! I'm looking into exactly this atm. What are tasks you tried the method on where it didn't work well enough and that you didn't include in the paper?

Cool idea congrats!!!

Incredible work!

This is what I was looking for. Awesome work!

I don’t get it, this isn’t really finetuning a vla right? You’re finetuning an edit policy that updates the vla output? Eitherway good work but was just confused, what’s the improvement on prior work? Or is the only change is the fact that the policy is a vla

The VLA becomes the base policy in EXPO and is fine-tuned

awsome

Impressive results – 30/30 on 8 tasks with only 19 min of RL data. Quick question: these results are tested on visual variations within the same lab setup. How does EXPO-FT perform when the environment itself changes – different kitchen layout, unfamiliar objects, or a completely new cultural context? Is that a data problem or a method problem?

amazing

The most interesting AI progress isn't happening on screens anymore. It's happening in the physical world. Getting a robot to reliably plug in a connector or insert a flower into a bottle is often harder than generating a thousand lines of code.

can we say this is "fully finetuning" the VLA with RL? seems that you freeze original VLA during RL process and just modify the edit residuals actor and Q-function

Next you need to try with a thread and a needle😅 Looking forward to the open source too
