Loading video...

Video Failed to Load

Go Home

Sharing some exciting DYNA-1 result: zero-shot environment generalization We put DYNA-1 under test in a completely different environment from our training distribution – with an entirely different background (Dyna Robotics banner) and metal table. The table has a reflective and smooth surface, creating a wildly different visual appearance as...

49,008 views • 1 year ago •via X (Twitter)

10 Comments

Jason Ma's profile picture
Jason Ma1 year ago

Original thread:

MrRobotics's profile picture
MrRobotics1 year ago

@DynaRobotics why is it 10X and doesn't look impressive as the first demo?

Hiro Protagonist's profile picture
Hiro Protagonist1 year ago

@DynaRobotics 10x speed so it's still basically useless for real life applications. Robotics amateur hour.

pfung's profile picture
pfung1 year ago

@DynaRobotics nice robustness!

Ming Qin's profile picture
Ming Qin1 year ago

@DynaRobotics The table can even reflect the robot itself, and it doesn’t seem surprised at all

atharva's profile picture
atharva1 year ago

@DynaRobotics insane

Max von Wolff's profile picture
Max von Wolff1 year ago

@DynaRobotics Impressive!

Humanoid Pulse's profile picture
Humanoid Pulse1 year ago

@DynaRobotics wow- there is something sacinating in that - can watch for hours🤓

Peter Christie's profile picture
Peter Christie1 year ago

@DynaRobotics Nope….

Michael Cho - Rbt/Acc's profile picture
Michael Cho - Rbt/Acc1 year ago

@DynaRobotics Impressive stuff!

Related Videos

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 views • 1 year ago