Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We present MotionStream — real-time, long-duration video generation that you can interactively control just by dragging your mouse. All videos here are raw, real-time screen captures without any post-processing. Model runs on a single H100 at 29 FPS and 0.4s latency.

99,145 görüntüleme • 10 ay önce •via X (Twitter)

38 Yorum

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

It feels totally different when video models actually become real-time interactive. Drag the mouse, and the cup instantly moves with it, and the water follows. You’re not just watching a video anymore, you’re playing with it.

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

We achieve this by first fine-tuning a video diffusion model conditioned on point track motions. We then apply Self-Forcing style causal distillation to post-train it into a real-time interactive model. To reduce error accumulation of long video generation, we apply attention sink during the training and inference. For maximum speed in real-time scenario, we incorporate several additional design optimizations: - Efficient encoding of track conditions via simple sinusoidal encoding and input-channel concatenation - Local, sliding-window causal attention - Lightweight VAE decoder

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

Using point tracks as a unified motion representation allows us to control not only the object motion but also the camera, all in real time. We use monocular depth estimation to lift the input image into 3D and uniformly sample 3D points. Then, as the virtual camera moves, we project each 3D point into the new view to create 2D motion trajectories.

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

It supports real-time video editing. Given an input video stream, we first edit the initial frame. We can estimates point tracks from the incoming video in real time. MotionStream then generates the output on the fly, conditioned on the edited initial frame and tracked points.

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

Work led by Joonghyuk Shin (@jhshin2000), and done with wonderful collaborators @zhengqi_li @rzhang88 @junyanz89 Jaesik Park @elishechtman Arxiv: Check out our project page for more details!

Learn to use AI profil fotoğrafı
Learn to use AI10 ay önce

is open source, github ?

davidsong profil fotoğrafı
davidsong10 ay önce

This is phenomenal. @twominutepapers you've gotta see this if you haven't already

Tianpei Gu profil fotoğrafı
Tianpei Gu10 ay önce

Does attention sink solve the memory issue in long videos? Like if the same object is completely out of frame then come back later does the model can still generate the correct object?

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

The model does not really have long-term memory because we are using local attention. Objects in the middle will not be remembered. The architecture design is more suitable for manipulating the same object/scene than world exploration.

Tianpei Gu profil fotoğrafı
Tianpei Gu10 ay önce

Got it, thanks for the reply! Attention sink looks really interesting here so I just wonder if dynamic attention sink would work, like discussed in the future work. This work is really cool!

Zhengzhong Tu profil fotoğrafı
Zhengzhong Tu10 ay önce

when to release🤨

deathbycasio profil fotoğrafı
deathbycasio10 ay önce

One issue I can see straight away is the off screen aspects, like hands in this example, are not 'generated' That would be a bit of an issue.

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

Right I think this is mostly limited by tha capability of the base model - below is another failure case. But we do observe that better base model can fix it, and models are progressing fast!

Rhumple profil fotoğrafı
Rhumple10 ay önce

Found a typo 👍

H5 profil fotoğrafı
H510 ay önce

That's insanely cool!

MetaDJ profil fotoğrafı
MetaDJ10 ay önce

Awesome ✨

Avery Superpowers profil fotoğrafı
Avery Superpowers10 ay önce

29 FPS on a single H100 is efficient. What’s the resolution limit for real-time playback?

DON'T BAN ME profil fotoğrafı
DON'T BAN ME10 ay önce

直接控制图像,这不就离用AI玩游戏越来越近了嘛

Dennis, Founder @Tixu 🎟️ | profil fotoğrafı
Dennis, Founder @Tixu 🎟️ |10 ay önce

Could I generate an Avatar out of it that respondieren to Text, and add an lypsinc Model over it?

bald spamton real profil fotoğrafı
bald spamton real10 ay önce

what the shit

Saleh Abdulaziz profil fotoğrafı
Saleh Abdulaziz10 ay önce

Good stuff

FCG Studio profil fotoğrafı
FCG Studio10 ay önce

Oh that's fancy!! Nice

RyanOnTheInside profil fotoğrafı
RyanOnTheInside10 ay önce

lets gooo! this is awesome

Philip Bankier profil fotoğrafı
Philip Bankier10 ay önce

very cool!

database__error profil fotoğrafı
database__error10 ay önce

@TomLikesRobots A link to site would be really useful rather than reading multiple threads and unable to find but interesting reading

Zefan Cai profil fotoğrafı
Zefan Cai10 ay önce

Any plan to open-source?

Frank “Mr. Four” Tarangelo profil fotoğrafı
Frank “Mr. Four” Tarangelo10 ay önce

Unbelievable work

たくろう profil fotoğrafı
たくろう10 ay önce

Borderline sci‑fi, can't wait to try it

timedoctor.eth profil fotoğrafı
timedoctor.eth10 ay önce

Looks promising!

Eric profil fotoğrafı
Eric10 ay önce

Congratulations! Great to see this work being published.

Mykhailo Sorochuk profil fotoğrafı
Mykhailo Sorochuk10 ay önce

This takes interactivity to another level. Instant feedback with real-time control is a game changer for video generation.

DON'T BAN ME profil fotoğrafı
DON'T BAN ME10 ay önce

这不就离用AI玩游戏越来越近了嘛

vfxpro88.Ξth profil fotoğrafı
vfxpro88.Ξth10 ay önce

Excellent work!

Mattias Johansson profil fotoğrafı
Mattias Johansson10 ay önce

C an you have image input too?

Xun Huang profil fotoğrafı
Xun Huang10 ay önce

Yes it's already conditioned on the initial frame.

Alexi Derkatsch profil fotoğrafı
Alexi Derkatsch10 ay önce

Animation will become a bespoke talent of the old world

Ben profil fotoğrafı
Ben10 ay önce

Incredible work and awesome presentation!

Mattias Johansson profil fotoğrafı
Mattias Johansson10 ay önce

Wow 🤯

Benzer Videolar