Video wird geladen...
Video konnte nicht geladen werden
๐ฆ Fully open-source under Apache 2.0 โ the FULL stack: โ Base model (256p) โ Distilled model โ Super-resolution model (540p & 1080p) โ Inference code ๐ค Models: ๐ฎ Demo: ๐ป Code: ๐ Paper: More videos below! Go build something amazing ๐
12,163 Aufrufe โข vor 5 Monaten โขvia X (Twitter)
6 Kommentare

๐ Benchmarks โ 2,000 pairwise human comparisons: ๐ฅ vs Ovi 1.1 โ 80.0% win rate ๐ฅ ๐ฅ vs LTX 2.3 โ 60.9% win rate ๐ฅ ๐จ Visual Quality: 4.80 (highest) ๐ Text Alignment: 4.18 (highest) ๐๏ธ WER: 14.60% (lowest) Opensource SOTA across the board ๐

๐ง Human-centric quality that hits different: ๐ Expressive facial performance ๐ฃ๏ธ Natural speech-expression coordination ๐บ Realistic body motion ๐ต Accurate audio-video sync ๐ Multilingual: Chinese (Mandarin + Cantonese), English, Japanese, Korean, German, French ๐ WER of just 14.6% โ lowest among all competitors

๐๏ธ Why it's SO fast: โก Latent-space super-res โ upscale in latent space, and skip the extra VAE decode-encode round trip ๐ Turbo VAE Decoder โ lightweight retrained decoder slashes decoding overhead ๐ง MagiCompiler โ full-graph compilation fuses ops across layers (~1.2x speedup) ๐ฏ DMD-2 distillation โ only 8 denoising steps, no CFG 256p in 2s โก 540p in 8s โก 1080p in 38s All on a single H100 ๐คฏ

๐๏ธ Architecture deep dive: A single 15B-param, 40-layer Transformer processes text, video, and audio tokens through self-attention only. ๐ฅช "Sandwich" design: first & last 4 layers use modality-specific projections, middle 32 layers share params across all modalities. ๐ซ No timestep embeddings โ the model infers denoising state directly from input latents.

Seedance 2.0 is impressive. But it's closed-source! Introducing our daVinci-MagiHuman โ a single-stream 15B Transformer trained from scratch that jointly generates video + audio. No cross-attention. No multi-stream branches. Just self-attention. โก 5s 1080p video in 38s on a single H100 ๐ 80% win rate vs Ovi 1.1 | 60.9% vs LTX 2.3 (2,000 human comparisons) ๐ 6 languages ๐ฆ Fully open-source Speed by simplicity. By @SII_GAIR ร @SandAI_HQ ๐ ๐ป ๐ค

Impressive! Has anyone tried this on 5090? How long on 5090 for a 5s video

