Загрузка видео...

Не удалось загрузить видео

На главную

📦 Fully open-source under Apache 2.0 — the FULL stack: ✅ Base model (256p) ✅ Distilled model ✅ Super-resolution model (540p & 1080p) ✅ Inference code 🤗 Models: 🎮 Demo: 💻 Code: 📄 Paper: More videos below! Go build something amazing 🚀

12,163 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 6

Фото профиля Pengfei Liu
Pengfei Liu5 месяцев назад

📊 Benchmarks — 2,000 pairwise human comparisons: 🥊 vs Ovi 1.1 → 80.0% win rate 💥 🥊 vs LTX 2.3 → 60.9% win rate 🔥 🎨 Visual Quality: 4.80 (highest) 📝 Text Alignment: 4.18 (highest) 🎙️ WER: 14.60% (lowest) Opensource SOTA across the board 👑

Фото профиля Pengfei Liu
Pengfei Liu5 месяцев назад

🧑 Human-centric quality that hits different: 😊 Expressive facial performance 🗣️ Natural speech-expression coordination 🕺 Realistic body motion 🎵 Accurate audio-video sync 🌍 Multilingual: Chinese (Mandarin + Cantonese), English, Japanese, Korean, German, French 📉 WER of just 14.6% — lowest among all competitors

Фото профиля Pengfei Liu
Pengfei Liu5 месяцев назад

🏎️ Why it's SO fast: ⚡ Latent-space super-res — upscale in latent space, and skip the extra VAE decode-encode round trip 🔄 Turbo VAE Decoder — lightweight retrained decoder slashes decoding overhead 🔧 MagiCompiler — full-graph compilation fuses ops across layers (~1.2x speedup) 🎯 DMD-2 distillation → only 8 denoising steps, no CFG 256p in 2s ⚡ 540p in 8s ⚡ 1080p in 38s All on a single H100 🤯

Фото профиля Pengfei Liu
Pengfei Liu5 месяцев назад

🏗️ Architecture deep dive: A single 15B-param, 40-layer Transformer processes text, video, and audio tokens through self-attention only. 🥪 "Sandwich" design: first & last 4 layers use modality-specific projections, middle 32 layers share params across all modalities. 🚫 No timestep embeddings — the model infers denoising state directly from input latents.

Фото профиля Pengfei Liu
Pengfei Liu5 месяцев назад

Seedance 2.0 is impressive. But it's closed-source! Introducing our daVinci-MagiHuman — a single-stream 15B Transformer trained from scratch that jointly generates video + audio. No cross-attention. No multi-stream branches. Just self-attention. ⚡ 5s 1080p video in 38s on a single H100 🏆 80% win rate vs Ovi 1.1 | 60.9% vs LTX 2.3 (2,000 human comparisons) 🌍 6 languages 📦 Fully open-source Speed by simplicity. By @SII_GAIR × @SandAI_HQ 📄 💻 🤗

Фото профиля Surendra Anubolu
Surendra Anubolu5 месяцев назад

Impressive! Has anyone tried this on 5090? How long on 5090 for a 5s video

Похожие видео