ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

How can robots acquire fine-grained manipulation skills? Introducing ACT: Action Chunking with Transformers ๐Ÿค– Key idea: Imitation, but predict actions in chunks instead of one at a time. Here are results with only ~15min of demonstrations, running on low-cost arms:

247,890 ๆฌก่ง‚็œ‹ โ€ข 3 ๅนดๅ‰ โ€ขvia X (Twitter)

11 ๆก่ฏ„่ฎบ

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

In case you missed ALOHA ๐Ÿ–, the hardware we use for all these experiments, here is the thread!

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

Fine manipulation is difficult: either from RL, Sim2Real, or Imitation. - Hard exploration and sparse reward - Large Sim2Real gap - Compounding error for BC - No large dataset We introduce three important design choices behind ACT, an efficient imitation learning method:

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

(1) Predict action sequence Standard BC predicts one action at a time, while a fine manipulation task can have >1000 steps easily. Predicting action in chunks slows down compounding error, and can better model non-stationary human behavior.

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

(2) Generative model policy The policy is trained as the decoder of a VAE, reconstructing action chunks from latent z, 4 RGB images, and proprioception. Intuitively, z extracts the โ€œstyleโ€ of the action chunk. This is crucial when learning from human demos.

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

(3) Transformer We modernize the VAE by using a BERT-like encoder and a DETR-like decoder, training end-to-end from scratch. This transformer architecture benefits more from chunking than ConvNets and non-parametric methods.

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

With all above, ACT obtains 64%, 96%, 84%, 92% success for 4 tasks shown, with objects randomized along the 15 cm line. It does not just memorize the training data, and is able to react to external disturbances:

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

It is also robust to a certain level of distractor objects:

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

Similar to ALOHA, we open source ACT together with 2 simulated environments for reproducibility. You can find it in the project website: We hope ALOHA+ACT would be a helpful resource towards advancing fine-grained manipulation!

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

Personally, this is a challenging project to work on, spanning from hardware to ML. It would certainly not be possible without my amazing advisor @chelseabfinn and collaboration from @svlevine @Vikashplus!

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

Here are some really cool related works you should also know about! Chopstick-holding cherry-picking robot from @xkelym, trained with RL in the real world. The motion is very reactive and precise!

Tony Z. Zhao ็š„ๅคดๅƒ
Tony Z. Zhao3 ๅนดๅ‰

Diffusion policy from @chichengcc: also uses a generative model for policy. Great for fitting multi-modal data and made large progress on the RoboMimic benchmark. Also very impressive real-world experiments!

็›ธๅ…ณ่ง†้ข‘

Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.โ โ  FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with NVIDIA, we also integrated FLUX 3 Action natively into Hugging Face's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, weโ€™re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. Weโ€™re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).

Black Forest Labs

177,726 ๆฌก่ง‚็œ‹ โ€ข 12 ๅคฉๅ‰