Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Gen2Act: Casting language-conditioned manipulation as *human video generation* followed by *closed-loop policy execution conditioned on the generated video* enables solving diverse real-world tasks unseen in the robot dataset! 1/n

71,154 Aufrufe • vor 1 Jahr •via X (Twitter)

10 Kommentare

Profilbild von Homanga Bharadhwaj
Homanga Bharadhwajvor 1 Jahr

We opt for generating human videos because we find that current best video models (e.g. VideoPoet) are already good at generating human videos *zero-shot* given an image of a scene and a language description of a task. This doesn't require any fine-tuning/adaption! 2/n

Profilbild von Homanga Bharadhwaj
Homanga Bharadhwajvor 1 Jahr

The video model generalizes well to new scenarios by virtue of web-scale training The policy also generalizes to tasks beyond that in the robot data as it is tasked with a much simpler job of translating the generated video to actions by following motion cues from the video 3/n

Profilbild von Homanga Bharadhwaj
Homanga Bharadhwajvor 1 Jahr

We can also chain Gen2Act for long-horizon activities with multiple tasks by sequentially rolling out video generation and policy execution conditioned on the generated video. 4/n

Profilbild von Homanga Bharadhwaj
Homanga Bharadhwajvor 1 Jahr

Following prior works, we categorize results with respect to different levels of generalization. Gen2Act achieves non-trivial success rates (30-60%) for even the challenging categories of motion-type and object-type generalization 5/n

Profilbild von Homanga Bharadhwaj
Homanga Bharadhwajvor 1 Jahr

This was a fun project w/ @debidatta @gupta_abhinav_ @shubhtuls @CarlDoersch @shahdhruv_ @xiao_ted @SeanKirmani @xf1280 @DorsaSadigh @GoogleDeepMind @CMU_Robotics @StanfordAILab More details: Video: n/n

Profilbild von Samarth Sinha
Samarth Sinhavor 1 Jahr

Congrats Homanga!!

Profilbild von Jason Ma
Jason Mavor 1 Jahr

Excited to see this out, congrats Homanga!

Profilbild von Paweł Budzianowski
Paweł Budzianowskivor 1 Jahr

Great to see first video-based model employed! This opens up completely new category of possibilities!

Profilbild von Jay Vakil
Jay Vakilvor 1 Jahr

Amazing work @mangahomanga

Profilbild von Rui Chen
Rui Chenvor 1 Jahr

Great work! Human video is a useful and unlimited source for mainpulation.

Ähnliche Videos

DisCo: Disentangled Control for Referring Human Dance Generation in Real World paper page: Generative AI has made significant strides in computer vision, particularly in image/video synthesis conditioned on text descriptions. Despite the advancements, it remains challenging especially in the generation of human-centric content such as dance synthesis. Existing dance synthesis methods struggle with the gap between synthesized content and real-world dance scenarios. In this paper, we define a new problem setting: Referring Human Dance Generation, which focuses on real-world dance scenarios with three important properties: (i) Faithfulness: the synthesis should retain the appearance of both human subject foreground and background from the reference image, and precisely follow the target pose; (ii) Generalizability: the model should generalize to unseen human subjects, backgrounds, and poses; (iii) Compositionality: it should allow for composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce a novel approach, DISCO, which includes a novel model architecture with disentangled control to improve the faithfulness and compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions.

AK

161,453 Aufrufe • vor 3 Jahren