Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Another day, another humanoid robot from china AGIBOT introduces GO-1, a generalist foundation model that integrates a vision-language model with a latent planner for enhanced long-horizon and dexterous manipulation.

23,513 Aufrufe • vor 1 Jahr •via X (Twitter)

8 Kommentare

Profilbild von @letsleepingfoxeslie
@letsleepingfoxeslievor 1 Jahr

Damn, China has been cooking. And they seem to be going hard. I heard some top Chinese universities plan on enrolling more students than usual in frontier fields like AI.

Profilbild von VistaShares ETFs
VistaShares ETFsvor 1 Jahr

From semiconductors to data centers, AIS targets the critical components behind AI's exponential growth. Capture potential returns from this transformative technology sector.

Profilbild von 33 - 11:11
33 - 11:11vor 1 Jahr

Most EV influences have had to flight to china to keep rolling their content. Will it happen the same to you?

Profilbild von AI Theology
AI Theologyvor 1 Jahr

Yeah, but who put all that shit on the table 🤷🏽‍♂️🤷🏽‍♂️

Profilbild von S
Svor 1 Jahr

i'll wait until they release the GO-1 slim

Profilbild von Van Liuten
Van Liutenvor 1 Jahr

This is how you scale

Profilbild von Michael
Michaelvor 1 Jahr

I love seeing the advancements in humanoid robots but, the hands throw me off. Still not sure if I want to see human looking hands or not but this is not it. The future is so weird sometimes.

Profilbild von Aurora
Auroravor 1 Jahr

🤖🌏🔄

Ähnliche Videos

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,296 Aufrufe • vor 1 Monat