Loading video...

Video Failed to Load

Go Home

Mercury 2 is live 🚀🚀 The world’s first reasoning diffusion LLM, delivering 5x faster performance than leading speed-optimized LLMs. Watching the team turn years of research into a real product never gets old, and I’m incredibly proud of what we’ve built. We’re just getting started on what diffusion can...

1,050,489 views • 7 months ago •via X (Twitter)

67 Comments

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Check out more in our blog

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Full story by @dinabass in @business

NVIDIA AI's profile picture
NVIDIA AI7 months ago

Huge congrats to the @_inception_ai team on Mercury 2 — blazing-fast diffusion reasoning on NVIDIA Blackwell GPUs and an exciting leap for real-time AI.

Stefano Ermon's profile picture
Stefano Ermon7 months ago

@_inception_ai Thank you @NVIDIAAI! Couldn't have done it without the infrastructure your team has built🤝

Volodymyr Kuleshov 🇺🇦's profile picture
Volodymyr Kuleshov 🇺🇦7 months ago

🚨Hot off the presses: the official Artificial Analysis benchmarking results are in! 🚀Mercury morels set a new frontier of speed and agentic quality

Beff (e/acc)'s profile picture
Beff (e/acc)7 months ago

very exciting

Stefano Ermon's profile picture
Stefano Ermon7 months ago

🚀

YEETUS DELETUS's profile picture
YEETUS DELETUS7 months ago

Mate, if it's already 1009 Tok/s on standard hardware from Nvidia, imagine if it was running on Groq or Cerebras hardware? Heck, is it even possible running it on Taalas hardware? Exciting times for Mercury 2, hoping to test it out in heavy duty reasoning tasks!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Exciting times indeed!

Yash Kr Gupta's profile picture
Yash Kr Gupta7 months ago

1009 tokens per second on standard GPUs. Parallel diffusion instead of sequential decoding. This is a fundamentally different architecture not just an optimization.

Stefano Ermon's profile picture
Stefano Ermon7 months ago

We have been working hard!

Volodymyr Kuleshov 🇺🇦's profile picture
Volodymyr Kuleshov 🇺🇦7 months ago

Breaking the 1000/tok sec barrier at the intelligence level of GPT5-mini, Flash, Haiku 🚀🚀🚀

Aditya Grover's profile picture
Aditya Grover7 months ago

Reasoning quality at non-reasoning speeds, super excited to see this out @_inception_ai!

Krishna Kaasyap's profile picture
Krishna Kaasyap7 months ago

Great product. Congratulations on shipping a better model, and even faster than Google (not sure what happened to Gemini diffusion). Small typo though: there is no GPT-5.2 Mini, just GPT-5 Mini!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks!

Vishnu Vardhan's profile picture
Vishnu Vardhan7 months ago

Super cool, Can’t wait to try it out guys, we are trying to build something like openclaw but with end to end voice pipeline , this could really make the difference, anyway we could get api access soon?

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Glad to hear your interest! Sign up for access here. We are rolling out access quickly.

Vishnu Vardhan's profile picture
Vishnu Vardhan7 months ago

Done!

AVB's profile picture
AVB7 months ago

For people who are wondering how dLLMs work: If you've seen Arrival, it's exactly that. Bidirectional attention, start with noise & iteratively denoise into actual text. Autoregressive LLMs iterate on tokens per step Diffusion LLMs iterate on revisions of large chunks per step

Samar Khanna's profile picture
Samar Khanna7 months ago

Super excited to share the fastest reasoning LLM, built with diffusion! Mercury 2 is even faster and much, MUCH better than Mercury 1

Mike Knoop's profile picture
Mike Knoop7 months ago

Congrats on the launch!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks for your support!

m's profile picture
m7 months ago

Excited to support u guys 🚀

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks! We are just getting started.

Jeffrey Emanuel's profile picture
Jeffrey Emanuel7 months ago

Very cool. I’m curious how the speed changes as the model size grows. For a Sonnet-class model (say, Sonnet 4), would it only be 3x faster? Is it currently capable of performing at that cognitive level at any speed?

Stefano Ermon's profile picture
Stefano Ermon7 months ago

We're actively scaling our models. Speed declines for larger models, but the same way that it does for a traditional LLM.

Boyuan (Nemo) Chen's profile picture
Boyuan (Nemo) Chen7 months ago

5x faster is wild but does diffusion-based decoding degrade gracefully under long context? autoregressive models at least fail predictably. curious what happens when you push Mercury to 128k tokens

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks! Mercury 2 supports up to 128k context length!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

@boyuan_chen Not much published on this

Xen's profile picture
Xen7 months ago

Some answers to expected questions about Mercury 2 right now: Can I run it locally? → No. It's API-only for now (closed weights, forget about local/RTX 4090 via LM Studio option) - their playground or API is the only option. How fast & how much? → Up to 1,196 tok/s output (world's fastest reasoning model via diffusion parallel gen) → $0.25/M input | $0.75/M output (cheaper blended than many other models and you get 10M free tokens to start) How does it look vs MiniMax M2.5 benchmarks? → MiniMax wins raw intelligence (~42 AA Index vs 33 for Mercury2), coding strength, 205k context → Mercury 2 crushes speed/throughput + cost for real-time agents/voice/inference Major plus: → Just because this reminds me of Freddie, I'll try it. Bottom line: Mercury 2 seems a valid option for "instant everything and fast".

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thank you!

VioP's profile picture
VioP7 months ago

you guys arr back congrats!!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks! It's live in the playground, give it a try at

mrkelly's profile picture
mrkelly7 months ago

Speed is a commodity. Reasoning depth is the only moat left.

Stefano Ermon's profile picture
Stefano Ermon7 months ago

We are excited to see the new use cases that our speed enables!

Julia Turc's profile picture
Julia Turc7 months ago

Congrats on the launch!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks! Been great collaborating recently.

Misha Laskin's profile picture
Misha Laskin7 months ago

Impressive! Congratulations

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks!

Jonathan Rosenberg's profile picture
Jonathan Rosenberg7 months ago

Been advising Inception. Mercury 2 is live: first reasoning diffusion LLM, ~1,000 tokens/sec, 5x faster. Not a tuning trick, a different architecture. @stefanoermon helped invent diffusion. His team made it work at production scale. Pay attention.

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks for your support!

Davis Treybig's profile picture
Davis Treybig7 months ago

So fast!!!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

@TreybigDavis ⚡️

Michele Catasta's profile picture
Michele Catasta7 months ago

Exceptional work, Stefano & team!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks! We are excited for the use cases that our speed will enable.

Saurabh Shukla's profile picture
Saurabh Shukla7 months ago

Amazing. I learned so much from your course on diffusion. all the best. This looks impressive. 🚀

Sarah Catanzaro's profile picture
Sarah Catanzaro7 months ago

Mercury can code apps faster than you can scroll through code. So cool.

Stefano Ermon's profile picture
Stefano Ermon7 months ago

⚡️🚀

Victor M's profile picture
Victor M7 months ago

any plan to open source it? or a small variant so we can get a taste 👀

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Not at the moment. But you can get a taste right now at

himanshu's profile picture
himanshu7 months ago

what are the params?

Yinjie Wang's profile picture
Yinjie Wang7 months ago

dLLM is good base model for action modeling, e.g. computer-use agents. You don't wanna wait for few seconds before a single click or one line of code in terminal.

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Agreed!

Christian T White's profile picture
Christian T White7 months ago

Well done and love the opening denoising of your video

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Glad you liked it!

Amit Verma's profile picture
Amit Verma7 months ago

Amazing milestone Mercury 2: "Delivering 5x faster reasoning." 🧠⚡️ Me: Currently spending 4 hours trying to logically reason with my Python environment because it refuses to run a basic script. 🐍💻 If this AI can politely explain my syntax errors to me 5x faster, I am completely sold.

sankalp's profile picture
sankalp7 months ago

where can we see benchmarks

Stefano Ermon's profile picture
Stefano Ermon7 months ago

See benchmarks in our blog post here:

Shivendra Soni's profile picture
Shivendra Soni7 months ago

How does the cost compare to SOTA models? Like sonnet of 5.2codex. This would be an amazing thing to have for voice ai agents

Stefano Ermon's profile picture
Stefano Ermon7 months ago

See our pricing here!

Mathias LÜCK's profile picture
Mathias LÜCK7 months ago

@oyegpt How well does it perform?

Stefano Ermon's profile picture
Stefano Ermon7 months ago

@oyegpt See Artificial Analysis's report here!

AVB's profile picture
AVB7 months ago

Great! Is there a technical paper? Or a technical blogpost? Would love to know more, mainly to create some content around this!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

See our research here!

AVB's profile picture
AVB7 months ago

Thanks! 🙏🏼

Marco Mascorro's profile picture
Marco Mascorro7 months ago

This is awesome. Congrats @StefanoErmon & team!

Stefano Ermon's profile picture
Stefano Ermon7 months ago

Thanks!

Related Videos