正在加载视频...

视频加载失败

Introducing LDA, a latent world action foundation model that, for the first time, unifies the utilization of heterogeneous embodied data across simulation and reality, humans and robots, and varying levels of action quality and annotation. By breaking long-standing data silos in embodied intelligence, LDA enables the field, much like...

40,883 次观看 • 5 个月前 •via X (Twitter)

5 条评论

Jiangran Lyu 的头像
Jiangran Lyu5 个月前

We have open-sourced the model and code! Project Page: arXiv:

Sharpa 的头像
Sharpa5 个月前

Nothing better than Sharpa hands to cook a burger just right! Joke aside, it's great to see the ecosystem coming up with new recipes to scale robot learning👍

abdel 的头像
abdel5 个月前

oh no why do you release it now.. i just finished my essay on world models haha and I am not sure yet in which category your world model would fit in. seems exciting, great job.

Tradeye 的头像
Tradeye5 个月前

LDA just dropped the mic @GalbotRobotics 🔥 Unifying sim + real, human + robot data? Absolute game changer! Tradeye is out here supplying the messy real-world chaos you need: narrated POV video from licensed HVAC/plumbing/electrical crews in actual tight attics, crawlspaces & lived-in homes (C20 licensed of course). DM me if you want some authentic trades data to supercharge the model? Let’s make robots unstoppable together 🚀

Akito 的头像
Akito5 个月前

Can consumers-captured video serve as training data for the model??

相关视频

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,416 次观看 • 1 年前