Loading video...

Video Failed to Load

Go Home

Another beast from Alibaba. Live Avatar: real-time, infinite-length streaming avatar generator; - 20 FPS, based on a 14B WanS2V trained with DMD. - locks identity; They plan to: add TTS, low VRAM, 3 steps, SVD quantz...

21,313 views • 8 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Want to create an avatar from a single image? FlexAvatar is a transformer model that creates full 360°, high-quality, and expressive 3D head avatar from just a single portrait image in minutes. Real-time Demo: FlexAvatar's lightweight architecture allows both animation and rendering in real-time, enabling interactive user experiences. To create a new 3D head avatar, only one image is required, e.g., from a webcam. The final avatar is ready after 2 minutes. Architecture: Under the hood, FlexAvatar adopts a transformer-based encoder-decoder design. The encoder maps the input image onto a latent avatar space, while the decoder produces 3D Gaussian attribute maps by incorporating the animation signal via cross-attention. The model learns all facial animations directly from the data without relying on pre-built 3D face models. This equips the avatars with realistic facial expressions. The internal avatar latent space can be conveniently used to integrate additional observations of a person via fitting. This enables use-cases where more than one image of a person is available, e.g., from a phone scan of the person. We train jointly on 2D monocular videos and multi-view data. However, in monocular videos, the animation signal leaks the target viewpoint, causing the model to produce incomplete 3D heads. We call this phenomenon entanglement of driving signal and target viewpoint. To prevent entanglement, we introduce bias sinks. These are learnable tokens that indicate whether a training sample stems from a monocular or a multi-view dataset. During training, the model learns to produce incomplete 3D heads only when the monocular token is present. During inference, FlexAvatar then always uses the multi-view token for which the model has learned to produce complete 3D heads. This simple design allows to combine the generalizability from monocular data with the quality of multi-view data. FlexAvatar summary: - Input: Single-image, phone scan, or monocular video - Output: Full 360° head avatar - Expressive animations - Real-time rendering and animation - Generalization to any portrait - Create a new avatar in 2 minutes - Use bias sinks to combine 2D and 3D data 🏠 🌍 🎥 Great work by Tobias Kirschstein and Simon Giebenhain!

Matthias Niessner

96,186 views • 7 months ago

I made a digital twin of myself from 10 seconds of video. In the clip: left is the real me, middle is a leading avatar model, right is Mirage Avatar X. Watch the eyes. The difference is not subtle. I have been testing AI avatar models since my first clone in 2023. Every one of them was impressive for about 30 seconds, then your brain caught up. Still eyes. One polite expression. A mouth doing all the work. Avatar X is the first model where that moment never came. Here is what makes it different: It is trained on you. Avatar X preserves your identity. Most avatar models can copy your appearance. Avatar X captures the subtle details that make you you. The way you move, the way you express yourself, and the way you naturally deliver speech. It looks like you. It moves like you. It sounds like you. It understands non-verbal performance Laughing, crying, yawning, sighing. These are the moments where most avatar models fall apart, trying to lip-sync through sounds that aren't words. Avatar X responds naturally, generating realistic facial expressions and micro-expressions instead of forcing every sound into speech. The expression goes beyond the lips Expressions are driven by the audio, through the whole face and body. Ask a question and it furrows its brows and shrugs on the tone. No other model does this to this degree. No quality degradation The first second and the last second look the same. Other models lose quality the longer the video runs. 10 seconds of input That is the entire requirement. Other models need 15 seconds, some even 1 to five minutes. Three years ago my AI clone was a party trick. This one can carry my face, my expressions and my delivery without me in the room. The bar for AI avatars just moved. Avatar X is live today. → Try it here:

Linus ✦ Ekenstam

20,752 views • 17 days ago

Physics-based Motion Retargeting from Sparse Inputs paper page: Avatars are important to create interactive and immersive experiences in virtual worlds. One challenge in animating these characters to mimic a user's motion is that commercial AR/VR products consist only of a headset and controllers, providing very limited sensor data of the user's pose. Another challenge is that an avatar might have a different skeleton structure than a human and the mapping between them is unclear. In this work we address both of these challenges. We introduce a method to retarget motions in real-time from sparse human sensor data to characters of various morphologies. Our method uses reinforcement learning to train a policy to control characters in a physics simulator. We only require human motion capture data for training, without relying on artist-generated animations for each avatar. This allows us to use large motion capture datasets to train general policies that can track unseen users from real and sparse data in real-time. We demonstrate the feasibility of our approach on three characters with different skeleton structure: a dinosaur, a mouse-like creature and a human. We show that the avatar poses often match the user surprisingly well, despite having no sensor information of the lower body available. We discuss and ablate the important components in our framework, specifically the kinematic retargeting step, the imitation, contact and action reward as well as our asymmetric actor-critic observations. We further explore the robustness of our method in a variety of settings including unbalancing, dancing and sports motions.

AK

106,527 views • 3 years ago