Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Alert🚨 This could be by far the most impressive video about 4D human motion generation! 🔥🔥🔥 We present TRUMANS, a pioneering and solid effort that scales up 4D human-scene interactions with both static and dynamic objects. With our model, we could eventually generate diverse and physically plausible human motions...

13,689 görüntüleme • 2 yıl önce •via X (Twitter)

7 Yorum

Siyuan Huang profil fotoğrafı
Siyuan Huang2 yıl önce

Hi @_akhaliq, please check out our full video here

Jingbo Wang profil fotoğrafı
Jingbo Wang2 yıl önce

Looks great!

Siyuan Huang profil fotoğrafı
Siyuan Huang2 yıl önce

thanks!

German Barquero profil fotoğrafı
German Barquero2 yıl önce

Impressive work! Congrats!

David Ferrera profil fotoğrafı
David Ferrera2 yıl önce

I'm primarily focused on 3D for VFX, but I'm curious, do these models come with a rig, for instance, to facilitate animation retargeting?

LEO LI profil fotoğrafı
LEO LI2 yıl önce

I think your work could lead to some innovative visual effects, some amazing effects, and of course great use in games like Max Payne's bullet Time

Siyuan Huang profil fotoğrafı
Siyuan Huang2 yıl önce

Thank you!

Benzer Videolar

.President Donald J. Trump: "In 250 Years, the Free People of this land have accomplished more with our Liberty than any other society has accomplished even in thousands and thousands of years... What our critics will never understand is that America is not the sum of its mistakes. Our mistakes make us human—our achievements make us American... We are the nation that dreamed and created the modern world—we laid the railroads, we raised up those big, beautiful skyscrapers, harnessed electricity, and invented the light bulb, the telephone, the airplane, the assembly line, the television, the microchip, the personal computer, the internet, GPS, the smartphone, and almost everything else that has ever been invented—including... a thing called air conditioning. We charted the human genome to cure diseases, we powered entire cities by splitting single atoms, and planted our flag on the Moon. Americans filled the airwaves of the planet with our music and our culture. We invented baseball, basketball, football, volleyball, NASCAR, and the Rodeos of the West. Americans have won the most Olympic medals of any country in the world by far, the most Nobel Prizes... and the most world records. We publish by far the most patents, we produce the best movies, we make the best music, and we raise up the greatest entertainers and strongest athletes the world has ever seen. We built the biggest and most dynamic economy... with, as of last week, 19.2 trillion dollars pouring into the United States from all over the world... And thanks to our great election win and tariffs, plants and factories are being built all over the United States right now... We created the strongest and most powerful military. We won Two World Wars, the Cold War, and left America’s enemies in the depths of history."

Rapid Response 47

274,275 görüntüleme • 2 ay önce

Exciting updates on Project GR00T! We discover a systematic way to scale up robot data, tackling the most painful pain point in robotics. The idea is simple: human collects demonstration on a real robot, and we multiply that data 1000x or more in simulation. Let’s break it down: 1. We use Apple Vision Pro (yes!!) to give the human operator first person control of the humanoid. Vision Pro parses human hand pose and retargets the motion to the robot hand, all in real time. From the human’s point of view, they are immersed in another body like the Avatar. Teleoperation is slow and time-consuming, but we can afford to collect a small amount of data. 2. We use RoboCasa, a generative simulation framework, to multiply the demonstration data by varying the visual appearance and layout of the environment. In Jensen’s keynote video below, the humanoid is now placing the cup in hundreds of kitchens with a huge diversity of textures, furniture, and object placement. We only have 1 physical kitchen at the GEAR Lab in NVIDIA HQ, but we can conjure up infinite ones in simulation. 3. Finally, we apply MimicGen, a technique to multiply the above data even more by varying the *motion* of the robot. MimicGen generates vast number of new action trajectories based on the original human data, and filters out failed ones (e.g. those that drop the cup) to form a much larger dataset. To sum up, given 1 human trajectory with Vision Pro -> RoboCasa produces N (varying visuals) -> MimicGen further augments to NxM (varying motions). This is the way to trade compute for expensive human data by GPU-accelerated simulation. A while ago, I mentioned that teleoperation is fundamentally not scalable, because we are always limited by 24 hrs/robot/day in the world of atoms. Our new GR00T synthetic data pipeline breaks this barrier in the world of bits. Scaling has been so much fun for LLMs, and it's finally our turn to have fun in robotics! We are building tools to enable everyone in the ecosystem to scale up with us. Links in thread:

Jim Fan

364,670 görüntüleme • 2 yıl önce

DisCo: Disentangled Control for Referring Human Dance Generation in Real World paper page: Generative AI has made significant strides in computer vision, particularly in image/video synthesis conditioned on text descriptions. Despite the advancements, it remains challenging especially in the generation of human-centric content such as dance synthesis. Existing dance synthesis methods struggle with the gap between synthesized content and real-world dance scenarios. In this paper, we define a new problem setting: Referring Human Dance Generation, which focuses on real-world dance scenarios with three important properties: (i) Faithfulness: the synthesis should retain the appearance of both human subject foreground and background from the reference image, and precisely follow the target pose; (ii) Generalizability: the model should generalize to unseen human subjects, backgrounds, and poses; (iii) Compositionality: it should allow for composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce a novel approach, DISCO, which includes a novel model architecture with disentangled control to improve the faithfulness and compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions.

AK

161,479 görüntüleme • 3 yıl önce