Loading video...

Video Failed to Load

Go Home

Going back to training single task models with a focus on improving success rates - here’s a longer, unedited real-time rollout of an ACT model trained with 40 episodes. Explanation in 🧵

11,843 views • 1 year ago •via X (Twitter)

11 Comments

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

The robot freezing initially was due to a false positive from my safety system. If you can see a regular pattern of fast movements, this is when the action chunk border chances. I used longer action chunks which makes the robot less responsive but more “confident” in movements

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

My previous data was recorded in 50Hz, while data for this model (and the inference) was in 20Hz. I also controlled the robot in a slower and more careful manner. In AI, garbage in -> garbage out so data quality really matters

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

Why 20Hz and not 50Hz? I noticed my cheap cameras were lagging, especially the wrist camera during movements, so many frames would contain an outdated view of the scene. This did not eliminate the issues and buying better cameras is high on the priority list but 20Hz helps

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

The wrist camera’s shakiness is a likely reason why the robot can pick capsules in front of the main camera but struggles with one’s further away, or ones that get occluded by the arm. Perhaps masking one of the cameras (or patches of them) at times during training would help?

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

The data I collected is all from the same table and from 2 different coffee cups - I’ll probably add to the dataset by recording a few episodes from other places in my flat and a few other cups as well. Data diversity matters. A lot.

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

I also want to augment my data from human corrections during inference, dAgger style. This will help the robot align itself back on track, as the original training episodes don’t have many mistakes.

Dan Peña's profile picture
Dan Peña3 years ago

29 Years of proven track record and creating generational wealth! If you want to learn what it takes to becomes super successful then take the success test now!

Shreyas Dixit's profile picture
Shreyas Dixit1 year ago

Question: Why ACT why not Diffusion Policy? I had tried ACT with around 100 samples and lot of epochs (I mean a lot) but it still failed to perform good.

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

I might try diffusion (and VQ-BeT) as well to compare. But ACT is a good baseline and the model I have the most experience using. These experiments exist to see how different techniques can boost success rates, so you could start with any baseline model that's quick to train

Julian Fried's profile picture
Julian Fried1 year ago

So strange to watch

Ville Kuosmanen's profile picture
Ville Kuosmanen1 year ago

one day it will look indistinguishable from a human operator controlling the robot robot turing test 🤖

Related Videos

A team tested Pi0, Pi0 Fast, Gr00t, and ACT on real robot arms in manufacturing tasks. (🔖 Bookmark this for later!) The task was precise: place thin rectangular frames from a messy stack into a holder. The team fine-tuned each model on 100 real trajectories and compared training time, inference speed, motion quality, and success rates. ⬇️ Here’s a breakdown of what they found Pi0 (Original) ✅ Strongest overall performance in precise pick-and-place ✅ High success rate even in edge cases ✅ Longest training time (~11 hours, ~$30 per run) ✅ Inference time of 80 ms causes short pauses between actions Despite delays, it handles complex scenarios well… solid for high-precision tasks, but slow to train. Gr00t ✅ Trains fast (~2 hours, ~$5 per run) ✅ Performs almost as well as Pi0 on large-object tasks ✅ Struggles with fine precision; random movement in some trials ✅ More training didn’t fix jitter or random offsets Best suited for tasks where exact precision isn’t critical. Not ready for manufacturing-grade accuracy without more tuning. Pi0 Fast ✅ Promised faster training, but results were underwhelming ✅ Training at 6 hours still showed low success rates ✅ Inference was slower than expected ✅ Not reliable for generalizing even slightly new tasks Currently too unstable for real-world deployment. Doesn’t live up to the “Fast” name yet. ACT (Baseline) ✅ 200MB model—lightweight, but limited ✅ Struggles with stacked objects or ambiguous scenes ✅ Success rates around 70% in best-case setups ✅ Can’t match newer models on precision or generalization Still a solid baseline, but clearly a generation behind in robustness. 🚨 Extra Notes All newer models share a common issue: •Inference takes longer than a frame (80 ms vs 33 ms), so robots “pause” between chunks. •This results in jittery movements, but not a dealbreaker unless tasks are time-sensitive. Language-conditioned tasks also fell short: after training on two labeled tasks, the model couldn’t generalize to a third unseen combination using only text prompts. ✅ The good news? These models adapt well to new robot arms with quick fine-tuning. ❌ The bad news? There’s still no plug-and-play solution for improving performance after deployment. Reinforcement learning or DAgger-style data collection during real-world operation may be the next big step, something many teams in robotics are actively working on.

Ilir Aliu

21,844 views • 1 year ago