Video wird geladen...
Video konnte nicht geladen werden
Meet LongCat-Video-Avatar 1.5🐱—our upgraded, open-source digital human framework. Built for real production, not just short demos. What's New: 🔹 Upgraded Audio Encoder: Replaces Wav2Vec2 with Whisper-Large, yielding significantly smoother and more natural lip dynamics. 🔹 Production-Ready Stability: Achieves accurate lip-synchronization, full-body temporal stability, and robust long-video generation with strict... show more
31,524 Aufrufe • vor 3 Monaten •via X (Twitter)
11 Kommentare

Garryvor 3 Monaten
How much time it will take to generate 60 seconds of avatar video (720p) with audio upload?

ZIAISTANvor 3 Monaten
Amazing

blankbrainvor 3 Monaten
oh finally , some good open sauce stuff from good plebs

Alice The Ai Expertvor 3 Monaten
Open digital humans leveled up

T1000vor 3 Monaten
@toyxyz3 This is great.

Ruzainavor 3 Monaten
Open-source avatars are leveling up

.vor 3 Monaten
the question is: it is better then ltx 2.3? its faster? if dont, it is useless

𝘿𝙖𝙫𝙞𝙙 ✦ 𝙈𝙂𝙏vor 3 Monaten
whisper swap is the right call, wav2vec2 was always the weak link for lip sync. Curious what your real-time latency looks like end-to-end on a production load

Sani Ai Techvor 3 Monaten
Open-source digital humans with 8-step inference and solid lip sync is a big deal

chenervor 3 Monaten
@HeyGen

Artyomvor 3 Monaten
Is this based on wan 2.2 video model?

