Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

What state representation should robots have? 🤖 I’m thrilled to present an Any-point Trajectory Model (ATM), which models physical motions from videos without additional assumptions and shows significant positive transfer from cross-embodiment human and robot videos! 🧵👇

124,318 Aufrufe • vor 2 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

1/5 Our goal is to improve policy learning from video data, a rich and scalable source. Since videos lack explicit actions, we focus on learning to predict the future trajectories of any set of particles based on their initial 2D positions, circumventing the need for actions

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

2/5 Once the trajectory model is trained, we learn trajectory-guided policies. We simply look at the trajectories of points from a fixed grid. We do not assume any calibration and our model utilizes cameras of different viewpoints.

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

3/5 By modeling the low-level particle trajectories, we find significant positive transfer from videos of humans or from a different robot! Our current model is trained from relatively in-domain videos. Stay tuned for developments on a more generalized model!

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

4/5 Our work is enabled by recent advances in video tracking. We build on top of the great works from CoTracker (@n_karaev @chrirupp ) and Tracking-Any-Point by @CarlDoersch et al.

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

@chrirupp @CarlDoersch 5/5 Work done with great collaborators @ChuanWen15, @johnrso_ , Kai Chen, Qi Dou, Yang Gao, and @pabbeel ! Project website: Paper:

Profilbild von Jiawei Yang
Jiawei Yangvor 2 Jahren

@yuewang314 @JunjieYe9 @PointsCoder

Profilbild von Aleksandr Kovalev
Aleksandr Kovalevvor 2 Jahren

Looks fascinating! Does it work in real time on the CPU?

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

On GPU the model runs in real time. But running on CPU can be much slower.

Profilbild von Juan Stoppa
Juan Stoppavor 2 Jahren

very interesting! when are you planning to make the code available?

Profilbild von Xingyu Lin
Xingyu Linvor 2 Jahren

Will release the code soon. Stay tuned!

Ähnliche Videos

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

466,148 Aufrufe • vor 1 Jahr

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,049 Aufrufe • vor 1 Monat

DisCo: Disentangled Control for Referring Human Dance Generation in Real World paper page: Generative AI has made significant strides in computer vision, particularly in image/video synthesis conditioned on text descriptions. Despite the advancements, it remains challenging especially in the generation of human-centric content such as dance synthesis. Existing dance synthesis methods struggle with the gap between synthesized content and real-world dance scenarios. In this paper, we define a new problem setting: Referring Human Dance Generation, which focuses on real-world dance scenarios with three important properties: (i) Faithfulness: the synthesis should retain the appearance of both human subject foreground and background from the reference image, and precisely follow the target pose; (ii) Generalizability: the model should generalize to unseen human subjects, backgrounds, and poses; (iii) Compositionality: it should allow for composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce a novel approach, DISCO, which includes a novel model architecture with disentangled control to improve the faithfulness and compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions.

AK

161,453 Aufrufe • vor 3 Jahren