Загрузка видео...

Не удалось загрузить видео

На главную

What state representation should robots have? 🤖 I’m thrilled to present an Any-point Trajectory Model (ATM), which models physical motions from videos without additional assumptions and shows significant positive transfer from cross-embodiment human and robot videos! 🧵👇

124,318 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

1/5 Our goal is to improve policy learning from video data, a rich and scalable source. Since videos lack explicit actions, we focus on learning to predict the future trajectories of any set of particles based on their initial 2D positions, circumventing the need for actions

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

2/5 Once the trajectory model is trained, we learn trajectory-guided policies. We simply look at the trajectories of points from a fixed grid. We do not assume any calibration and our model utilizes cameras of different viewpoints.

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

3/5 By modeling the low-level particle trajectories, we find significant positive transfer from videos of humans or from a different robot! Our current model is trained from relatively in-domain videos. Stay tuned for developments on a more generalized model!

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

4/5 Our work is enabled by recent advances in video tracking. We build on top of the great works from CoTracker (@n_karaev @chrirupp ) and Tracking-Any-Point by @CarlDoersch et al.

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

@chrirupp @CarlDoersch 5/5 Work done with great collaborators @ChuanWen15, @johnrso_ , Kai Chen, Qi Dou, Yang Gao, and @pabbeel ! Project website: Paper:

Фото профиля Jiawei Yang
Jiawei Yang2 лет назад

@yuewang314 @JunjieYe9 @PointsCoder

Фото профиля Aleksandr Kovalev
Aleksandr Kovalev2 лет назад

Looks fascinating! Does it work in real time on the CPU?

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

On GPU the model runs in real time. But running on CPU can be much slower.

Фото профиля Juan Stoppa
Juan Stoppa2 лет назад

very interesting! when are you planning to make the code available?

Фото профиля Xingyu Lin
Xingyu Lin2 лет назад

Will release the code soon. Stay tuned!

Похожие видео

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

466,333 просмотров • 1 год назад

🇨🇳 It has started. A new home service in China pairs human cleaners with autonomous AI robots to tackle household chores. Residents in Shenzhen can now book a service where a human professional and an autonomous robot arrive together to clean their home. Real houses present a chaotic mess of dropped toys and random furniture that confuse traditional machines. X Square Robot and a major service platform named 58[.]com decided to tackle this chaos by launching China's first robot cleaner service in March-26. Customers use an application to hire a cleaning crew that consists of 1 human worker and 1 robot. The human takes care of the tricky chores that require complex judgment. The robot handles the repetitive physical work like picking up trash and wiping down flat surfaces. This machine runs on a system called WALL-A, which acts as a single continuous AI brain rather than a list of pre-written rules. They built this AI foundation model to perceive its surroundings and make its own decisions without human guidance. It processes visual data and plans multi-step actions. And deploying these robots into actual homes now provides the massive amounts of extremely important training data to improve it continuously. Alibaba and ByteDance backed this project. IMO, if the foundational model behind it figures out how to navigate a messy living room without getting stuck, it can learn to operate in almost any other physical environment.

Rohan Paul

52,996 просмотров • 4 месяцев назад

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,282 просмотров • 2 месяцев назад

DisCo: Disentangled Control for Referring Human Dance Generation in Real World paper page: Generative AI has made significant strides in computer vision, particularly in image/video synthesis conditioned on text descriptions. Despite the advancements, it remains challenging especially in the generation of human-centric content such as dance synthesis. Existing dance synthesis methods struggle with the gap between synthesized content and real-world dance scenarios. In this paper, we define a new problem setting: Referring Human Dance Generation, which focuses on real-world dance scenarios with three important properties: (i) Faithfulness: the synthesis should retain the appearance of both human subject foreground and background from the reference image, and precisely follow the target pose; (ii) Generalizability: the model should generalize to unseen human subjects, backgrounds, and poses; (iii) Compositionality: it should allow for composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce a novel approach, DISCO, which includes a novel model architecture with disentangled control to improve the faithfulness and compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions.

AK

161,478 просмотров • 3 лет назад