Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing ♾OmniJARVIS, our latest venture to #AgentGPT, or vision-language-action (VLA) models for open-world instruction-following agents 🦾🕹️ tuning in 👉 by Team CraftJarvis, 🤿⏬

32,562 Aufrufe • vor 2 Jahren •via X (Twitter)

8 Kommentare

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

🌟Our key insights into building more versatile VLAs: 💡♾OmniJARVIS emits **behavior token**, a skill representation that emerged from **self-supervised learning (SSL)** on massive replay (gaming, robotics, driving, etc) data. A de-tokenizer policy jointly learned with SSL will turn it into low-level actions. We choose to stay in between VLA down to producing motor commands directly (RT series @xiao_ted @QuanVng and folks, STEVE-1 @Shalev_lif), and VLA remains a task planner (SayCan, DEPS/JARVIS-1). ➡️ It offers a more compact and more situated/semantically richer interface between🧠"agent brain" (reasoning and planning) and the agent's 🏃"motor neurons" while assuring high-frequency motor control(~20Hz). ➡️ It echos Dieter Fox's (@NVIDIARobotics) recent talk at @CMU_Robotics. Yes, ♾OmniJARVIS is precisely the "middle ground" you hinted 🤗🤗

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

💡♾OmniJARVIS seeks to do native modeling on "VLA data". The idea is just intuitive: upon the multimodal dialogues being modeled by VLMs/MLMs, we further loop in 🏃**actions**(behaviors) interleaved within these dialogues -- we call this "multimodal interaction data". Technically, this adds a new modality "Action" over vision and language. Why interaction data modeling? This allows the unification of reasoning(self-talk), chat, and control, but the ultimate goal is enabling in-context agent learning😈😈, we're still working on it.

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

Just another word for the self-supervised behavior tokenizer/de-tokenizer. It inherits our SSL policy learning framework GROOT ( but adopts a discrete latent space (Finite-Scalar Quantizer @mentzer_f). Checkout my summary on GROOT

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

Very quick heads up on the results: 1⃣ Unifying reasoning, chat, and control enable better agent performance, especially on open-world tasks. 2⃣ FSQ-based behavior tokenizer is great! Consistent behaviours✅ Stable training ✅ Fast decoding ✅ 3⃣ ♾OmniJARVIS adheres to scaling law (first time in VLA?)

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

Wrapping up here. Calling in the Team CraftJarvis at and @PKU1898: @RealZihaoWang @liu_anji @YitaoLiang @ShaofeiCai. Checkout our series of generalist agent work below JARVIS-1

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

LEO

Profilbild von Xiaojian Ma
Xiaojian Mavor 2 Jahren

@AdeenaY8 guess you may be interested! We’ve uploaded it to daily papers 🤩

Profilbild von Martin Andrews
Martin Andrewsvor 2 Jahren

I'll be talking about this (briefly) at tonight's Machine Learning Singapore MeetUp ( One small thing I noticed : The Instruction Following table on the Project Page shows 'higher is better' (where it should be 'lower is better')

Ähnliche Videos