Загрузка видео...

Не удалось загрузить видео

На главную

Need help in the real world? RoboVQA can guide robots and humans through long-horizon tasks on a phone via Google Meet. We release a dataset of 800k (video, question/answer) with robots & humans doing various long-horizon tasks. Data: Google DeepMind

64,369 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 9

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

The video-conditioned RoboVQA model can instruct a robot through long-horizon tasks (teleoperated here). A human supervisor speaks corrections which are automatically quantified as intervention rate. This allows performing tasks to completion in supervised real-world deployment

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

We combine RoboVQA with RT-2 in a system1-system2 fashion so that the entire system runs autonomously (except for high-level human interventions via chat). This run completed with a 12.5% intervention rate (25% cognitive intervention, 0% physical intervention).

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

Our small video model (380M) outperforms a large model by ~2x (zero-shot PaLM-E 12B). While not intended as a fair comparison, it highlights models trained on existing datasets are still not enough. Scalable data collection like ours remains critical for real-world deployment.

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

Collection throughput boosts: A key advantage of collecting continuous long-horizon episodes is the 2.2x throughput gain compared to collecting individual steps one at a time. We also get accuracy gains from cross-embodiment transfer from human data (14x faster to collect).

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

Accuracy boost from text augmentation: From instructions, we automatically generate question-answer pairs of type: success, affordance, planning, past description, future prediction. Training on all augmentations yields better results than training on a single question type.

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

Accuracy boost from video: The model is 20% more accurate when using 16 frames as input compared to a single frame.

Фото профиля Pierre Sermanet
Pierre Sermanet2 лет назад

Authors: @psermanet, @TianliDing, @jezhao, @xf1280, @debidatta, @keerthanpg, Christine Chan, @gabepsilon, Sharath Maddineni, @nikhil_j_joshi, @peteflorence, Wei Han, Robert Baruch, Yao Lu, @suvir_m, Peng Xu, @pannag_, @hausman_k, Izhak Shafran, @brian_ichter, @caoyuan33

Фото профиля TadasG | topyappers.com
TadasG | topyappers.com2 лет назад

@GoogleDeepMind Wow, this is amazing! RoboVQA seems like a game-changer for guiding robots and humans in long-horizon tasks. The dataset release of 800k (video, question/answer) is impressive. Can't wait to explore it! Thanks for sharing, Pierre!

Фото профиля JRaw
JRaw2 лет назад

@GoogleDeepMind This could be incredibly helpful for people with severe depression, adhd or other issues that overload the brain when there are multiple „where to start“ options. Having third person decisions for those choices could massively improve self-organization skills

Похожие видео