Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Need help in the real world? RoboVQA can guide robots and humans through long-horizon tasks on a phone via Google Meet. We release a dataset of 800k (video, question/answer) with robots & humans doing various long-horizon tasks. Data: Google DeepMind

64,369 Aufrufe • vor 2 Jahren •via X (Twitter)

9 Kommentare

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

The video-conditioned RoboVQA model can instruct a robot through long-horizon tasks (teleoperated here). A human supervisor speaks corrections which are automatically quantified as intervention rate. This allows performing tasks to completion in supervised real-world deployment

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

We combine RoboVQA with RT-2 in a system1-system2 fashion so that the entire system runs autonomously (except for high-level human interventions via chat). This run completed with a 12.5% intervention rate (25% cognitive intervention, 0% physical intervention).

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

Our small video model (380M) outperforms a large model by ~2x (zero-shot PaLM-E 12B). While not intended as a fair comparison, it highlights models trained on existing datasets are still not enough. Scalable data collection like ours remains critical for real-world deployment.

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

Collection throughput boosts: A key advantage of collecting continuous long-horizon episodes is the 2.2x throughput gain compared to collecting individual steps one at a time. We also get accuracy gains from cross-embodiment transfer from human data (14x faster to collect).

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

Accuracy boost from text augmentation: From instructions, we automatically generate question-answer pairs of type: success, affordance, planning, past description, future prediction. Training on all augmentations yields better results than training on a single question type.

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

Accuracy boost from video: The model is 20% more accurate when using 16 frames as input compared to a single frame.

Profilbild von Pierre Sermanet
Pierre Sermanetvor 2 Jahren

Authors: @psermanet, @TianliDing, @jezhao, @xf1280, @debidatta, @keerthanpg, Christine Chan, @gabepsilon, Sharath Maddineni, @nikhil_j_joshi, @peteflorence, Wei Han, Robert Baruch, Yao Lu, @suvir_m, Peng Xu, @pannag_, @hausman_k, Izhak Shafran, @brian_ichter, @caoyuan33

Profilbild von TadasG | topyappers.com
TadasG | topyappers.comvor 2 Jahren

@GoogleDeepMind Wow, this is amazing! RoboVQA seems like a game-changer for guiding robots and humans in long-horizon tasks. The dataset release of 800k (video, question/answer) is impressive. Can't wait to explore it! Thanks for sharing, Pierre!

Profilbild von JRaw
JRawvor 2 Jahren

@GoogleDeepMind This could be incredibly helpful for people with severe depression, adhd or other issues that overload the brain when there are multiple „where to start“ options. Having third person decisions for those choices could massively improve self-organization skills

Ähnliche Videos