Video yükleniyor...
Video Yüklenemedi
Combining real-time interactivity, task understanding, and full-body action prediction on a humanoid is so, so hard. Here's an example where we bring all of these together in Gemini Robotics 2 🤖🧠
41,579 görüntüleme • 1 ay önce •via X (Twitter)
12 Yorum

Under the hood, this requires: - Action model that coordinates whole-body movement and fine motor skills simultaneously, while following instructions and recovering on-the-fly - Embodied reasoning (ER) agent with SOTA video understanding, long-horizon reasoning about user intent, and interruption handling

This was a huge team effort @GoogleDeepMind! hard to fathom the amount of work that goes into making full systems like this, but it's incredibly rewarding when things come together, and it's so fun to see how responsive the models are :)

this brings me back to when you interrupted my toy cleanup session in toyon and i started crashing out. huge moment.

@DvijKalaria

The response feels too slow. If you begin streaming tokens as soon as you start generating, and flush the first speakable clause as early as possible, the user experience would improve significantly—even though it has no real impact on the underlying task.

Apollo never hesitated, great!

i want Apollo 🥰

You sure do look nervous... suppose, it's normal with beta versions 👉 keep us updated! 🍀

All three on a humanoid in real-time? No joke.

@StanfordAILab Yeah! But the passion to make it happen is the thing that move people forward

The interactivity part is what stands out. Most demos show a robot executing, this one shows it adjusting to a person mid-task. Curious how much of that came from data with humans actually in the loop, versus scripted rollouts.

I think having a human test subject do the same task, and while they are performing the task, ask them to describe why, and how, they are making the decisions, and the movements, would help go a long way for training AI/Robotics. The more detailed the human subject gets and explains their through process helps train the AI/Robot to do the same thing. AI/Robotics need to know more than how to do something, they need to know why as well. There are many ways to perform a task, first train them on the different approaches with extractable expirations, then focus optimization through efficiency and effectiveness. The why is a very crucial part of learning.
