Loading video...

Video Failed to Load

Go Home

How can robots acquire fine-grained manipulation skills? Introducing ACT: Action Chunking with Transformers 🤖 Key idea: Imitation, but predict actions in chunks instead of one at a time. Here are results with only ~15min of demonstrations, running on low-cost arms:

247,890 views • 3 years ago •via X (Twitter)

11 Comments

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

In case you missed ALOHA 🏖, the hardware we use for all these experiments, here is the thread!

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

Fine manipulation is difficult: either from RL, Sim2Real, or Imitation. - Hard exploration and sparse reward - Large Sim2Real gap - Compounding error for BC - No large dataset We introduce three important design choices behind ACT, an efficient imitation learning method:

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

(1) Predict action sequence Standard BC predicts one action at a time, while a fine manipulation task can have >1000 steps easily. Predicting action in chunks slows down compounding error, and can better model non-stationary human behavior.

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

(2) Generative model policy The policy is trained as the decoder of a VAE, reconstructing action chunks from latent z, 4 RGB images, and proprioception. Intuitively, z extracts the “style” of the action chunk. This is crucial when learning from human demos.

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

(3) Transformer We modernize the VAE by using a BERT-like encoder and a DETR-like decoder, training end-to-end from scratch. This transformer architecture benefits more from chunking than ConvNets and non-parametric methods.

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

With all above, ACT obtains 64%, 96%, 84%, 92% success for 4 tasks shown, with objects randomized along the 15 cm line. It does not just memorize the training data, and is able to react to external disturbances:

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

It is also robust to a certain level of distractor objects:

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

Similar to ALOHA, we open source ACT together with 2 simulated environments for reproducibility. You can find it in the project website: We hope ALOHA+ACT would be a helpful resource towards advancing fine-grained manipulation!

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

Personally, this is a challenging project to work on, spanning from hardware to ML. It would certainly not be possible without my amazing advisor @chelseabfinn and collaboration from @svlevine @Vikashplus!

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

Here are some really cool related works you should also know about! Chopstick-holding cherry-picking robot from @xkelym, trained with RL in the real world. The motion is very reactive and precise!

Tony Z. Zhao's profile picture
Tony Z. Zhao3 years ago

Diffusion policy from @chichengcc: also uses a generative model for policy. Great for fitting multi-modal data and made large progress on the RoboMimic benchmark. Also very impressive real-world experiments!

Related Videos

Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with NVIDIA, we also integrated FLUX 3 Action natively into Hugging Face's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).

Black Forest Labs

177,726 views • 11 days ago