Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

How can robots acquire fine-grained manipulation skills? Introducing ACT: Action Chunking with Transformers 🤖 Key idea: Imitation, but predict actions in chunks instead of one at a time. Here are results with only ~15min of demonstrations, running on low-cost arms:

247,890 Aufrufe • vor 3 Jahren •via X (Twitter)

11 Kommentare

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

In case you missed ALOHA 🏖, the hardware we use for all these experiments, here is the thread!

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

Fine manipulation is difficult: either from RL, Sim2Real, or Imitation. - Hard exploration and sparse reward - Large Sim2Real gap - Compounding error for BC - No large dataset We introduce three important design choices behind ACT, an efficient imitation learning method:

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

(1) Predict action sequence Standard BC predicts one action at a time, while a fine manipulation task can have >1000 steps easily. Predicting action in chunks slows down compounding error, and can better model non-stationary human behavior.

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

(2) Generative model policy The policy is trained as the decoder of a VAE, reconstructing action chunks from latent z, 4 RGB images, and proprioception. Intuitively, z extracts the “style” of the action chunk. This is crucial when learning from human demos.

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

(3) Transformer We modernize the VAE by using a BERT-like encoder and a DETR-like decoder, training end-to-end from scratch. This transformer architecture benefits more from chunking than ConvNets and non-parametric methods.

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

With all above, ACT obtains 64%, 96%, 84%, 92% success for 4 tasks shown, with objects randomized along the 15 cm line. It does not just memorize the training data, and is able to react to external disturbances:

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

It is also robust to a certain level of distractor objects:

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

Similar to ALOHA, we open source ACT together with 2 simulated environments for reproducibility. You can find it in the project website: We hope ALOHA+ACT would be a helpful resource towards advancing fine-grained manipulation!

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

Personally, this is a challenging project to work on, spanning from hardware to ML. It would certainly not be possible without my amazing advisor @chelseabfinn and collaboration from @svlevine @Vikashplus!

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

Here are some really cool related works you should also know about! Chopstick-holding cherry-picking robot from @xkelym, trained with RL in the real world. The motion is very reactive and precise!

Profilbild von Tony Z. Zhao
Tony Z. Zhaovor 3 Jahren

Diffusion policy from @chichengcc: also uses a generative model for policy. Great for fitting multi-modal data and made large progress on the RoboMimic benchmark. Also very impressive real-world experiments!

Ähnliche Videos

Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with NVIDIA, we also integrated FLUX 3 Action natively into Hugging Face's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).

Black Forest Labs

177,726 Aufrufe • vor 11 Tagen