Video yükleniyor...
Video Yüklenemedi
JoyIn unveiled Aether: a 4B-param energy-based model for robots, trained on ~200 hrs of human video, zero real-robot data. Not VLA (perception+language to action) or world models. It's an "action-energy" model: actions emerge from minimizing a learned energy function over perception, memory and dynamics. No planner mapping instructions to... show more
12,885 görüntüleme • 8 gün önce •via X (Twitter)
6 Yorum

The architecture

About 200 hours of human video and zero real-robot data is a strong claim about where skill lives. If actions come from that footage alone, the scarce work is still episode quality: task, scene, and clean cuts before any robot body is in the loop.

200 hz means the energy search has a few milliseconds at most i am probably missing something

Interesting approach for learning robot behavior from human video alone

액션-에너지 모델은 처음 들어보네요 인간 영상만으로 어디까지 가능할지도 궁금합니다

The action-energy framing is more interesting than the label. For robots, the useful question is whether the learned objective stays stable when perception is noisy and the scene changes. Recovery behavior will matter as much as the demo.
