Loading video...
Video Failed to Load
Need help in the real world? RoboVQA can guide robots and humans through long-horizon tasks on a phone via Google Meet. We release a dataset of 800k (video, question/answer) with robots & humans doing various long-horizon tasks. Data: Google DeepMind
64,369 views • 2 years ago •via X (Twitter)
9 Comments

The video-conditioned RoboVQA model can instruct a robot through long-horizon tasks (teleoperated here). A human supervisor speaks corrections which are automatically quantified as intervention rate. This allows performing tasks to completion in supervised real-world deployment

We combine RoboVQA with RT-2 in a system1-system2 fashion so that the entire system runs autonomously (except for high-level human interventions via chat). This run completed with a 12.5% intervention rate (25% cognitive intervention, 0% physical intervention).

Our small video model (380M) outperforms a large model by ~2x (zero-shot PaLM-E 12B). While not intended as a fair comparison, it highlights models trained on existing datasets are still not enough. Scalable data collection like ours remains critical for real-world deployment.

Collection throughput boosts: A key advantage of collecting continuous long-horizon episodes is the 2.2x throughput gain compared to collecting individual steps one at a time. We also get accuracy gains from cross-embodiment transfer from human data (14x faster to collect).

Accuracy boost from text augmentation: From instructions, we automatically generate question-answer pairs of type: success, affordance, planning, past description, future prediction. Training on all augmentations yields better results than training on a single question type.

Accuracy boost from video: The model is 20% more accurate when using 16 frames as input compared to a single frame.

Authors: @psermanet, @TianliDing, @jezhao, @xf1280, @debidatta, @keerthanpg, Christine Chan, @gabepsilon, Sharath Maddineni, @nikhil_j_joshi, @peteflorence, Wei Han, Robert Baruch, Yao Lu, @suvir_m, Peng Xu, @pannag_, @hausman_k, Izhak Shafran, @brian_ichter, @caoyuan33

@GoogleDeepMind Wow, this is amazing! RoboVQA seems like a game-changer for guiding robots and humans in long-horizon tasks. The dataset release of 800k (video, question/answer) is impressive. Can't wait to explore it! Thanks for sharing, Pierre!

@GoogleDeepMind This could be incredibly helpful for people with severe depression, adhd or other issues that overload the brain when there are multiple „where to start“ options. Having third person decisions for those choices could massively improve self-organization skills

