Loading video...
Video Failed to Load
🚀 SANA-Video 2.0 is here! A full-stack optimized video model designed for efficiency — while still delivering high quality. We combine a hybrid architecture closely related to the recent Kimi K3 design, adopt Self-Flow from FLUX 3, and further accelerate it with our own Sol-Engine. Key technical ingredients: 🧠... show more
25,538 views • 1 month ago •via X (Twitter)
24 Comments

1/3 Motivation: Why hybrid attention? 🤔 Full 3D softmax attention captures rich spatial–temporal interactions, but its O(N²) cost grows rapidly with video resolution and duration. Pure linear attention scales as O(N), but compressing token interactions into a fixed-size state limits rank and expressiveness. SANA-Video 2.0 keeps most layers linear, then inserts sparse softmax anchors to recover the interactions that pure linear attention misses—combining long-sequence efficiency with softmax-level quality.

2/3 🧬 From Frontier LLMs to Video Generation SANA-Video 2.0 follows the same high-level direction as Kimi K3: a mostly-linear hybrid-attention backbone paired with cross-layer Block AttnRes. Each 4-layer cycle contains: • 3 gated linear-attention layers for efficient global mixing • 1 gated-softmax anchor for periodic full-rank refresh • Attention Residual for recovery of low-rank representations in linear attention A from-scratch ratio sweep identifies this 75% linear / 25% softmax design as the practical quality–efficiency Pareto knee. The softmax anchors restore richer token interactions, while Block AttnRes carries these refreshed features into deeper linear layers—improving their effective rank by 11.7%. We adapt this frontier LLM recipe to bidirectional video diffusion and massive spatial–temporal token sequences. The backbone is trained from scratch, without post-hoc linearization or a pretrained image-generation prior.

3/3 ⚙️ Full-Stack Training & Deployment We also disclose the complete training recipe: • Pre-training → continual training → 720p SFT • Raw-video processing, cleaning, filtering, and multi-axis scoring • Flow matching with content- and token-aware timestep sampling • Self-Flow feature distillation during pre-training and continual training • DPO and online ReFL for preference alignment For deployment, the convolution-free SwiGLU backbone maps cleanly to fused kernels. Sol-Engine further combines kernel fusion, caching, and sparse attention—bringing 720p/5s generation to 13.06s on a single H100. Efficient architecture, efficient training, and efficient inference—optimized end to end.

Looks pretty good! Do you plan to share the model weights?

yes, will open sourse asap

Awesome!

so fast

Exciting updates, especially the Self-Flow integration sounds promising!

Fast and high quality is a great combination.

Efficiency beast 720p video on a single H100 in 13s and 120x faster than Wan SANA Video 2.0 just made high quality video generation actually practical

720p in 13 seconds on one H100 is genuinely impressive.

Efficiency-first architecture is the move now. Hybrid design, sparse activation, every layer optimized—Kimi K3 proved it works. Curious what the token-per-dollar math looks like with your accelerations.

This is the real democratization of video AI: 720p/5s on a single H100 at 13 seconds isn't just an optimization—it's the threshold where research labs and indie devs finally get to play in the same sandbox as the giants. 🚀5

Impressive engineering work pushing efficient AI video generation forward

Very interesting work from NVlabs. Hybrid linear-softmax is the right approach. Pure linear attention loses too much expressiveness at deeper layers.

Why choose the LTX VAE vs. others?

because ltx vae has high compression ratio and good reconstruction capabilities

solid. what'd break first when you pushed kimi past the demo?

Nice but seems like dataset and training duration are insufficient

open source coming soon?

The quality looks good but the slow motion bias is terrible :(

open source?

sure thing

honest take: if kimi only works on a clean demo tree I'm out. you run this on real work yet?
