Loading video...

Video Failed to Load

Go Home

Meet LongCat-Video-Avatar 1.5🐱—our upgraded, open-source digital human framework. Built for real production, not just short demos. What's New: 🔹 Upgraded Audio Encoder: Replaces Wav2Vec2 with Whisper-Large, yielding significantly smoother and more natural lip dynamics. 🔹 Production-Ready Stability: Achieves accurate lip-synchronization, full-body temporal stability, and robust long-video generation with strict...

31,524 views • 3 months ago •via X (Twitter)

11 Comments

Garry's profile picture
Garry3 months ago

How much time it will take to generate 60 seconds of avatar video (720p) with audio upload?

ZIAISTAN's profile picture
ZIAISTAN3 months ago

Amazing

blankbrain's profile picture
blankbrain3 months ago

oh finally , some good open sauce stuff from good plebs

Alice The Ai Expert's profile picture
Alice The Ai Expert3 months ago

Open digital humans leveled up

T1000's profile picture
T10003 months ago

@toyxyz3 This is great.

Ruzaina's profile picture
Ruzaina3 months ago

Open-source avatars are leveling up

.'s profile picture
.3 months ago

the question is: it is better then ltx 2.3? its faster? if dont, it is useless

𝘿𝙖𝙫𝙞𝙙 ✦ 𝙈𝙂𝙏's profile picture
𝘿𝙖𝙫𝙞𝙙 ✦ 𝙈𝙂𝙏3 months ago

whisper swap is the right call, wav2vec2 was always the weak link for lip sync. Curious what your real-time latency looks like end-to-end on a production load

Sani Ai Tech's profile picture
Sani Ai Tech3 months ago

Open-source digital humans with 8-step inference and solid lip sync is a big deal

chener's profile picture
chener3 months ago

@HeyGen

Artyom's profile picture
Artyom3 months ago

Is this based on wan 2.2 video model?

Related Videos