Loading video...
Video Failed to Load
Argus: an open-source robotics data annotation and quality pipeline. With new frontier VLMs like GPT-6 Astra, we can generate rich, high-quality annotations for robotics data for both training and dataset analysis. Argus delivers detailed, timestamp-level annotations while catching issues like mislabeled instructions, sped-up recordings, swapped camera streams, and unflagged... show more
75,142 views • 2 days ago •via X (Twitter)
33 Comments

Argus takes in any episode in its raw format and outputs the following: - A densely annotated timeline with all robot or gripper actions, state changes, task progression percentages, and more - A record for every operator mistake, along with notes on the timing of the failure, the nature of the failure, the operator’s recovery response, and the ideal fix - Any metadata mistakes - A severity rating for each issue based on the predicted impact on training - Any subgoals and the completion outcome - The goal frame if it is achieved - Overall grade on completion along with a short rationale review

Argus can also be used on inference and tele-op trajectories captured in production. Here’s an example of a fully annotated trajectory from one of our production facilities.

Argus is a model-agnostic pipeline and can run on any sufficiently capable VLM. In our test comparisons, Astra and GPT-6.1 Sol produced the densest annotations, which we’ve found is highly correlated with overall accuracy.

Of the 3,546 episodes we labelled across nine datasets, 27% have at least one issue of medium severity or above, meaning part or all of the episode teaches the model something wrong. We’re releasing the full audit and more than 150K annotations at

Read the full analysis at and run Argus on your data at For collaborations or questions, contact us at [email protected]. Credit to @ericli_ for undertaking this project.

@PantheonInc Eric 🚀

@PantheonInc 🚀

@PantheonInc Incredible

@PantheonInc Pushing the frontier! VLMs are improving so fast, habit. Great harnesses and clarity for training ready data is key to extract signals from raw data and train the right policies!

@PantheonInc chef Li @ericli_

@PantheonInc 🚢

@PantheonInc Eric li rolling the robotics boulder up the hill

@PantheonInc @ericli_ 🚀🚀🚀

@PantheonInc 🔥

@PantheonInc Mr Li

@PantheonInc So cool

@PantheonInc very very cool stuff @calixo888 @nettedjay 😎

@PantheonInc Very cool.

@PantheonInc 🚢🚢🚢

@PantheonInc Fire contributions for the community

@PantheonInc beautiful

@PantheonInc 🧑🍳

@PantheonInc Swapped camera streams and sped up recordings silently ruin a training run. Good that this is open source.

@PantheonInc damn, this is amazing.

@PantheonInc Congrats Calix!

@PantheonInc 27% of episodes having training-affecting issues is a wild number for datasets people assumed were clean

@PantheonInc cooking

@PantheonInc catching swapped camera streams and mislabeled instructions before training saves more than another round of fine tuning, factory data has the same mess

@PantheonInc Super cool ! Have you guys looked at how much variation there is in SR of VLA style models when there is variance in the task description?

@PantheonInc That's amazing!

Generated annotations need a sampling audit or you inherit the model's blind spots at full scale. The failure mode is not random noise, it is consistent error, which looks exactly like clean data until something downstream learns the same mistake. Worth holding back a human labelled slice purely to measure drift against.

@PantheonInc sharing with my friend in robotics!

@PantheonInc WOAHH

