Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing LDA, a latent world action foundation model that, for the first time, unifies the utilization of heterogeneous embodied data across simulation and reality, humans and robots, and varying levels of action quality and annotation. By breaking long-standing data silos in embodied intelligence, LDA enables the field, much like...

40,883 Aufrufe • vor 5 Monaten •via X (Twitter)

5 Kommentare

Profilbild von Jiangran Lyu
Jiangran Lyuvor 5 Monaten

We have open-sourced the model and code! Project Page: arXiv:

Profilbild von Sharpa
Sharpavor 5 Monaten

Nothing better than Sharpa hands to cook a burger just right! Joke aside, it's great to see the ecosystem coming up with new recipes to scale robot learning👍

Profilbild von abdel
abdelvor 5 Monaten

oh no why do you release it now.. i just finished my essay on world models haha and I am not sure yet in which category your world model would fit in. seems exciting, great job.

Profilbild von Tradeye
Tradeyevor 5 Monaten

LDA just dropped the mic @GalbotRobotics 🔥 Unifying sim + real, human + robot data? Absolute game changer! Tradeye is out here supplying the messy real-world chaos you need: narrated POV video from licensed HVAC/plumbing/electrical crews in actual tight attics, crawlspaces & lived-in homes (C20 licensed of course). DM me if you want some authentic trades data to supercharge the model? Let’s make robots unstoppable together 🚀

Profilbild von Akito
Akitovor 5 Monaten

Can consumers-captured video serve as training data for the model??

Ähnliche Videos

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,416 Aufrufe • vor 1 Jahr