Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

MiniMax (official) H3 is live in SGLang Diffusion, with day-0 serving support 🎬 This open model matches Seedance 2.0 at 1/3 the cost, or $0 if you run it locally on 2x 5090 or 1 RTX 6000. With SGLang Diffusion, you can build visual concepts, motion design, e-commerce creatives,...

100,189 görüntüleme • 1 ay önce •via X (Twitter)

18 Yorum

LMSYS Org profil fotoğrafı
LMSYS Org1 ay önce

Thanks @lena_z01 and @yesand_ai for the great prompts we use in our model comparison 🪄

RyanLee profil fotoğrafı
RyanLee1 ay önce

@MiniMax_AI Thanks SGLang day 0 support! The best partner

Leanna Ren profil fotoğrafı
Leanna Ren1 ay önce

@MiniMax_AI Love you all! ❤️ Love how we keep growing and progressing together

LMSYS Org profil fotoğrafı
LMSYS Org1 ay önce

Run MiniMax-H3 with SGLang Diffusion:

Kai profil fotoğrafı
Kai1 ay önce

@MiniMax_AI Chose Open. Chose @MiniMax_AI. Chose @sgl_project 🏃‍♀️

Scribblewick profil fotoğrafı
Scribblewick1 ay önce

@MiniMax_AI It doesn't match it yet, but it will once the community works on it. Look what they did with LTX!

DegenApeDev profil fotoğrafı
DegenApeDev1 ay önce

@MiniMax_AI I did this one is 14min on a RTX3090... Sure it's not a datacenter but it's still local AI! MiniMax H3

Markets & Mayhem profil fotoğrafı
Markets & Mayhem1 ay önce

@MiniMax_AI 👀

泰多宝 profil fotoğrafı
泰多宝1 ay önce

@MiniMax_AI The "$0 if you run it locally" line is the one that changes behavior. "Water finds the low ground and fills it." — Chinese proverb

IpezyGJ profil fotoğrafı
IpezyGJ1 ay önce

@MiniMax_AI Matching Seedance 2.0 at a third of the cost is the claim that will travel, so the instrument is worth naming. Arena already controls presentation order, which most video comparisons skip; we measured the second-shown option taking 50.6 to 49.4. Did H3 get that treatment?

lsm_ profil fotoğrafı
lsm_1 ay önce

@MiniMax_AI @grok what is the model size ? are they doing block/layer level offloading in sglang?

Rain Miao profil fotoğrafı
Rain Miao1 ay önce

@MiniMax_AI what does inference cost per task look like against seedance?

AlphaRomeoSierra profil fotoğrafı
AlphaRomeoSierra1 ay önce

@MiniMax_AI A RTX 3060 is enough as MiniMax H3 optimized for VRAM offloading. It'll be slow but works.

Alice The Ai Expert profil fotoğrafı
Alice The Ai Expert1 ay önce

@MiniMax_AI Amazing bMiniMax H3 + SGLang Diffusion pro video creation, zero barriers.

Layla CryptoWhiz profil fotoğrafı
Layla CryptoWhiz1 ay önce

@MiniMax_AI keep the cloud fees... we buidl different

AgentSparko 💥 profil fotoğrafı
AgentSparko 💥1 ay önce

@MiniMax_AI If anyone has some performance numbers for DGX Spark with SGLang please let me know too. My Spark is working hard at the project I posted about.

Kreta HB profil fotoğrafı
Kreta HB1 ay önce

@MiniMax_AI Superb launch video. Minimax H3 has bring hope to world that quality is just a matter of smartness to bring in models. I found a really cool prompt guide on minimax h3 video.

John Doe profil fotoğrafı
John Doe1 ay önce

@MiniMax_AI 2 5090s? 😢

Benzer Videolar

🚀 Sol-H3: MiniMax (official) H3 Video Generation Faster Than Playback 🤩 Five seconds of world. 1.653 seconds to infer. We’re releasing Sol-H3, our fastest end-to-end MiniMax-H3 inference stack yet. On one 8× NVIDIA B300 Blackwell system, it generates five seconds of 1344×768 video with stereo audio in 1.653 seconds. Across 1×, 4×, and 8× B300, Sol-H3 reaches up to a 15.54× speedup versus Base H3. Compared with 50-step Base H3 Dense on the same 8× B300 system, the four-step Sol-H3 profile delivers: • 5s: 18.250s → 1.653s (11.04×) • 10s: 50.660s → 3.732s (13.57×) • 15s: 99.513s → 6.612s (15.05×) Sol-H3 also scales across GPU counts: • 4× B300: 2.918s / 6.993s / 12.542s for 5s / 10s / 15s (12.11–15.54×) • 1× B300: 13.745s / 37.813s / 52.260s for 5s / 10s / 15s (9.45–14.29×) All figures are medians of three measured runs after one warmup at 1344×768 and 24 FPS with stereo audio. Base H3 uses 50 scheduler points (49 DiT forwards); Sol-H3 uses four DiT forwards, so this is a full-profile comparison—not an attention-only runtime change. Sol-H3 uses Dense attention on 1× B300 and SOL with INT8 QKV / FP8 output transport on 4× / 8×. Timing includes text encoding, DiT denoising, and video/audio VAE decoding; model loading, compilation warmup, and final MP4 encoding are excluded. Sol-H3 brings Sol-Engine × Sol-Attn into one full-stack runtime: • dynamic sparse attention with no retraining • fused norm, RoPE, MLP, and sparse-attention setup • fused INT8 QKV / FP8 output communication across 8 GPUs • parallel, batched VAE decoding • precomputed AdaLN caching Inside the stack: • sparse-attention setup: 1.206 → 0.285 ms (−76.4%) • VAE decode: 7.55 → 0.602 s • ~24 GB memory freed per GPU Any MiniMax-H3 few-step LoRA can plug into the same engine, and the code is deployment-friendly under Apache 2.0. For us, the bigger milestone is crossing from “fast generation” into “faster than playback.” That opens the path toward continuous 24 FPS generation and truly interactive video systems. We’re excited to partner with reactor to release Sol-H3 and make it available as an API day-0. Try it now on Reactor: 🔗 Amazing team effort—full credits in the blog. LoveSy Junsong_Chen yitong li Haopeng Li Haocheng Xi Song Han

Enze Xie

204,399 görüntüleme • 10 gün önce