
Enze Xie
@xieenze_jr • 3,074 subscribers
Tech Lead & Staff Research Scientist @ NVIDIA, Efficient VideoGen / SANA / Sol-Engine , CS PhD from HKU MMLab.
Shorts
Videos

🚀 Sol-H3: MiniMax (official) H3 Video Generation Faster Than Playback 🤩 Five seconds of world. 1.653 seconds to infer. We’re releasing Sol-H3, our fastest end-to-end MiniMax-H3 inference stack yet. On one 8× NVIDIA B300 Blackwell system, it generates five seconds of 1344×768 video with stereo audio in 1.653 seconds. Across 1×, 4×, and 8× B300, Sol-H3 reaches up to a 15.54× speedup versus Base H3. Compared with 50-step Base H3 Dense on the same 8× B300 system, the four-step Sol-H3 profile delivers: • 5s: 18.250s → 1.653s (11.04×) • 10s: 50.660s → 3.732s (13.57×) • 15s: 99.513s → 6.612s (15.05×) Sol-H3 also scales across GPU counts: • 4× B300: 2.918s / 6.993s / 12.542s for 5s / 10s / 15s (12.11–15.54×) • 1× B300: 13.745s / 37.813s / 52.260s for 5s / 10s / 15s (9.45–14.29×) All figures are medians of three measured runs after one warmup at 1344×768 and 24 FPS with stereo audio. Base H3 uses 50 scheduler points (49 DiT forwards); Sol-H3 uses four DiT forwards, so this is a full-profile comparison—not an attention-only runtime change. Sol-H3 uses Dense attention on 1× B300 and SOL with INT8 QKV / FP8 output transport on 4× / 8×. Timing includes text encoding, DiT denoising, and video/audio VAE decoding; model loading, compilation warmup, and final MP4 encoding are excluded. Sol-H3 brings Sol-Engine × Sol-Attn into one full-stack runtime: • dynamic sparse attention with no retraining • fused norm, RoPE, MLP, and sparse-attention setup • fused INT8 QKV / FP8 output communication across 8 GPUs • parallel, batched VAE decoding • precomputed AdaLN caching Inside the stack: • sparse-attention setup: 1.206 → 0.285 ms (−76.4%) • VAE decode: 7.55 → 0.602 s • ~24 GB memory freed per GPU Any MiniMax-H3 few-step LoRA can plug into the same engine, and the code is deployment-friendly under Apache 2.0. For us, the bigger milestone is crossing from “fast generation” into “faster than playback.” That opens the path toward continuous 24 FPS generation and truly interactive video systems. We’re excited to partner with reactor to release Sol-H3 and make it available as an API day-0. Try it now on Reactor: 🔗 Amazing team effort—full credits in the blog. LoveSy Junsong_Chen yitong li Haopeng Li Haocheng Xi Song Han
Enze Xie202,121 Aufrufe • vor 8 Tagen

🚀 Sol-H3 on DGX Spark: 768p in Under a Minute 🤩 Monday: 8×B300, 5s 768p in 1.65s — faster than playback. Today: the same stack on one desktop Spark — about 56s hot E2E. Five seconds of 1344×768 video at 24 FPS with stereo audio, on a single NVIDIA DGX Spark (GB10). Two-stage, not the datacenter profile: 384p H3 draft → latent ×2 → H3-to-LTX VAE adapter → 768p LTX refine → VAE decode No decode/re-encode between stages. Stage 2 is conditioned on the draft latent, so Gemma stays off the box. Quantized weights stay resident; sparse attention cuts the rest. Stage 1 takes any MiniMax-H3 few-step LoRA. Timing is hot E2E (encode → both stages → video/audio VAE); cold start and MP4 mux are separate. Apache 2.0. Server was realtime. Edge is one box, under a minute. 🔗 Amazing team effort—full credits in the blog. Haopeng Li Junsong_Chen yitong li LoveSy ,Jingyu Xin, Haocheng Xi Song Han
Enze Xie91,620 Aufrufe • vor 4 Tagen

🚀 SANA-Video 2.0 is here! A full-stack optimized video model designed for efficiency — while still delivering high quality. We combine a hybrid architecture closely related to the recent Kimi K3 design, adopt Self-Flow from FLUX 3, and further accelerate it with our own Sol-Engine. Key technical ingredients: 🧠 Hybrid Linear–Softmax Attention 3 gated linear layers + 1 gated-softmax anchor (75% linear / 25% softmax) 🧱 Block Attention Residuals (AttnRes) Boosts deep-layer effective rank by ~12% 🏗️ Unified 5B & 14B models Trained from scratch under very limited resources: • 5B → only 16 nodes of H100 • 14B → 48 nodes of B200 Results 📊 • 84.30 VBench Total • 3.2× faster DiT forward than matched full-softmax at 720p/60s • 720p/5s in 13.06s on a single H100 with Sol-Engine • 120× faster than Wan 2.2-A14B under the same one-H100 setup SANA-Video 2.0 shows that high-quality 720p video generation can be both efficient and practical on a single GPU. 🎬 Project: 📄 Paper: 💻 Code: Proud of the team! 🎉 More details below 🧵
Enze Xie25,558 Aufrufe • vor 1 Monat

Fast-dLLM accelerates multi-modal diffusion VLM LLaDA-V 10 times! 🚀
Enze Xie11,314 Aufrufe • vor 1 Jahr
Keine weiteren Inhalte verfügbar