Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐Ÿ“ฆ Fully open-source under Apache 2.0 โ€” the FULL stack: โœ… Base model (256p) โœ… Distilled model โœ… Super-resolution model (540p & 1080p) โœ… Inference code ๐Ÿค— Models: ๐ŸŽฎ Demo: ๐Ÿ’ป Code: ๐Ÿ“„ Paper: More videos below! Go build something amazing ๐Ÿš€

12,163 Aufrufe โ€ข vor 5 Monaten โ€ขvia X (Twitter)

6 Kommentare

Profilbild von Pengfei Liu
Pengfei Liuvor 5 Monaten

๐Ÿ“Š Benchmarks โ€” 2,000 pairwise human comparisons: ๐ŸฅŠ vs Ovi 1.1 โ†’ 80.0% win rate ๐Ÿ’ฅ ๐ŸฅŠ vs LTX 2.3 โ†’ 60.9% win rate ๐Ÿ”ฅ ๐ŸŽจ Visual Quality: 4.80 (highest) ๐Ÿ“ Text Alignment: 4.18 (highest) ๐ŸŽ™๏ธ WER: 14.60% (lowest) Opensource SOTA across the board ๐Ÿ‘‘

Profilbild von Pengfei Liu
Pengfei Liuvor 5 Monaten

๐Ÿง‘ Human-centric quality that hits different: ๐Ÿ˜Š Expressive facial performance ๐Ÿ—ฃ๏ธ Natural speech-expression coordination ๐Ÿ•บ Realistic body motion ๐ŸŽต Accurate audio-video sync ๐ŸŒ Multilingual: Chinese (Mandarin + Cantonese), English, Japanese, Korean, German, French ๐Ÿ“‰ WER of just 14.6% โ€” lowest among all competitors

Profilbild von Pengfei Liu
Pengfei Liuvor 5 Monaten

๐ŸŽ๏ธ Why it's SO fast: โšก Latent-space super-res โ€” upscale in latent space, and skip the extra VAE decode-encode round trip ๐Ÿ”„ Turbo VAE Decoder โ€” lightweight retrained decoder slashes decoding overhead ๐Ÿ”ง MagiCompiler โ€” full-graph compilation fuses ops across layers (~1.2x speedup) ๐ŸŽฏ DMD-2 distillation โ†’ only 8 denoising steps, no CFG 256p in 2s โšก 540p in 8s โšก 1080p in 38s All on a single H100 ๐Ÿคฏ

Profilbild von Pengfei Liu
Pengfei Liuvor 5 Monaten

๐Ÿ—๏ธ Architecture deep dive: A single 15B-param, 40-layer Transformer processes text, video, and audio tokens through self-attention only. ๐Ÿฅช "Sandwich" design: first & last 4 layers use modality-specific projections, middle 32 layers share params across all modalities. ๐Ÿšซ No timestep embeddings โ€” the model infers denoising state directly from input latents.

Profilbild von Pengfei Liu
Pengfei Liuvor 5 Monaten

Seedance 2.0 is impressive. But it's closed-source! Introducing our daVinci-MagiHuman โ€” a single-stream 15B Transformer trained from scratch that jointly generates video + audio. No cross-attention. No multi-stream branches. Just self-attention. โšก 5s 1080p video in 38s on a single H100 ๐Ÿ† 80% win rate vs Ovi 1.1 | 60.9% vs LTX 2.3 (2,000 human comparisons) ๐ŸŒ 6 languages ๐Ÿ“ฆ Fully open-source Speed by simplicity. By @SII_GAIR ร— @SandAI_HQ ๐Ÿ“„ ๐Ÿ’ป ๐Ÿค—

Profilbild von Surendra Anubolu
Surendra Anuboluvor 5 Monaten

Impressive! Has anyone tried this on 5090? How long on 5090 for a 5s video

ร„hnliche Videos