正在加载视频...

视频加载失败

Meet LongCat-Video-Avatar 1.5🐱—our upgraded, open-source digital human framework. Built for real production, not just short demos. What's New: 🔹 Upgraded Audio Encoder: Replaces Wav2Vec2 with Whisper-Large, yielding significantly smoother and more natural lip dynamics. 🔹 Production-Ready Stability: Achieves accurate lip-synchronization, full-body temporal stability, and robust long-video generation with strict...

31,524 次观看 • 3 个月前 •via X (Twitter)

11 条评论

Garry 的头像
Garry3 个月前

How much time it will take to generate 60 seconds of avatar video (720p) with audio upload?

ZIAISTAN 的头像
ZIAISTAN3 个月前

Amazing

blankbrain 的头像
blankbrain3 个月前

oh finally , some good open sauce stuff from good plebs

Alice The Ai Expert 的头像
Alice The Ai Expert3 个月前

Open digital humans leveled up

T1000 的头像
T10003 个月前

@toyxyz3 This is great.

Ruzaina 的头像
Ruzaina3 个月前

Open-source avatars are leveling up

. 的头像
.3 个月前

the question is: it is better then ltx 2.3? its faster? if dont, it is useless

𝘿𝙖𝙫𝙞𝙙 ✦ 𝙈𝙂𝙏 的头像
𝘿𝙖𝙫𝙞𝙙 ✦ 𝙈𝙂𝙏3 个月前

whisper swap is the right call, wav2vec2 was always the weak link for lip sync. Curious what your real-time latency looks like end-to-end on a production load

Sani Ai Tech 的头像
Sani Ai Tech3 个月前

Open-source digital humans with 8-step inference and solid lip sync is a big deal

chener 的头像
chener3 个月前

@HeyGen

Artyom 的头像
Artyom3 个月前

Is this based on wan 2.2 video model?

相关视频