Video wird geladen...
Video konnte nicht geladen werden
Mercury 2 is live 🚀🚀 The world’s first reasoning diffusion LLM, delivering 5x faster performance than leading speed-optimized LLMs. Watching the team turn years of research into a real product never gets old, and I’m incredibly proud of what we’ve built. We’re just getting started on what diffusion can... show more
1,050,489 Aufrufe • vor 7 Monaten •via X (Twitter)
67 Kommentare

Check out more in our blog

Full story by @dinabass in @business

Huge congrats to the @_inception_ai team on Mercury 2 — blazing-fast diffusion reasoning on NVIDIA Blackwell GPUs and an exciting leap for real-time AI.

@_inception_ai Thank you @NVIDIAAI! Couldn't have done it without the infrastructure your team has built🤝

🚨Hot off the presses: the official Artificial Analysis benchmarking results are in! 🚀Mercury morels set a new frontier of speed and agentic quality

very exciting

🚀

Mate, if it's already 1009 Tok/s on standard hardware from Nvidia, imagine if it was running on Groq or Cerebras hardware? Heck, is it even possible running it on Taalas hardware? Exciting times for Mercury 2, hoping to test it out in heavy duty reasoning tasks!

Exciting times indeed!

1009 tokens per second on standard GPUs. Parallel diffusion instead of sequential decoding. This is a fundamentally different architecture not just an optimization.

We have been working hard!

Breaking the 1000/tok sec barrier at the intelligence level of GPT5-mini, Flash, Haiku 🚀🚀🚀

Reasoning quality at non-reasoning speeds, super excited to see this out @_inception_ai!

Great product. Congratulations on shipping a better model, and even faster than Google (not sure what happened to Gemini diffusion). Small typo though: there is no GPT-5.2 Mini, just GPT-5 Mini!

Thanks!

Super cool, Can’t wait to try it out guys, we are trying to build something like openclaw but with end to end voice pipeline , this could really make the difference, anyway we could get api access soon?

Glad to hear your interest! Sign up for access here. We are rolling out access quickly.

Done!

For people who are wondering how dLLMs work: If you've seen Arrival, it's exactly that. Bidirectional attention, start with noise & iteratively denoise into actual text. Autoregressive LLMs iterate on tokens per step Diffusion LLMs iterate on revisions of large chunks per step

Super excited to share the fastest reasoning LLM, built with diffusion! Mercury 2 is even faster and much, MUCH better than Mercury 1

Congrats on the launch!

Thanks for your support!

Excited to support u guys 🚀

Thanks! We are just getting started.

Very cool. I’m curious how the speed changes as the model size grows. For a Sonnet-class model (say, Sonnet 4), would it only be 3x faster? Is it currently capable of performing at that cognitive level at any speed?

We're actively scaling our models. Speed declines for larger models, but the same way that it does for a traditional LLM.

5x faster is wild but does diffusion-based decoding degrade gracefully under long context? autoregressive models at least fail predictably. curious what happens when you push Mercury to 128k tokens

Thanks! Mercury 2 supports up to 128k context length!

@boyuan_chen Not much published on this

Some answers to expected questions about Mercury 2 right now: Can I run it locally? → No. It's API-only for now (closed weights, forget about local/RTX 4090 via LM Studio option) - their playground or API is the only option. How fast & how much? → Up to 1,196 tok/s output (world's fastest reasoning model via diffusion parallel gen) → $0.25/M input | $0.75/M output (cheaper blended than many other models and you get 10M free tokens to start) How does it look vs MiniMax M2.5 benchmarks? → MiniMax wins raw intelligence (~42 AA Index vs 33 for Mercury2), coding strength, 205k context → Mercury 2 crushes speed/throughput + cost for real-time agents/voice/inference Major plus: → Just because this reminds me of Freddie, I'll try it. Bottom line: Mercury 2 seems a valid option for "instant everything and fast".

Thank you!

you guys arr back congrats!!

Thanks! It's live in the playground, give it a try at

Speed is a commodity. Reasoning depth is the only moat left.

We are excited to see the new use cases that our speed enables!

Congrats on the launch!

Thanks! Been great collaborating recently.

Impressive! Congratulations

Thanks!

Been advising Inception. Mercury 2 is live: first reasoning diffusion LLM, ~1,000 tokens/sec, 5x faster. Not a tuning trick, a different architecture. @stefanoermon helped invent diffusion. His team made it work at production scale. Pay attention.

Thanks for your support!

So fast!!!

@TreybigDavis ⚡️

Exceptional work, Stefano & team!

Thanks! We are excited for the use cases that our speed will enable.

Amazing. I learned so much from your course on diffusion. all the best. This looks impressive. 🚀

Mercury can code apps faster than you can scroll through code. So cool.

⚡️🚀

any plan to open source it? or a small variant so we can get a taste 👀

Not at the moment. But you can get a taste right now at

what are the params?

dLLM is good base model for action modeling, e.g. computer-use agents. You don't wanna wait for few seconds before a single click or one line of code in terminal.

Agreed!

Well done and love the opening denoising of your video

Glad you liked it!

Amazing milestone Mercury 2: "Delivering 5x faster reasoning." 🧠⚡️ Me: Currently spending 4 hours trying to logically reason with my Python environment because it refuses to run a basic script. 🐍💻 If this AI can politely explain my syntax errors to me 5x faster, I am completely sold.

where can we see benchmarks

See benchmarks in our blog post here:

How does the cost compare to SOTA models? Like sonnet of 5.2codex. This would be an amazing thing to have for voice ai agents

See our pricing here!

@oyegpt How well does it perform?

@oyegpt See Artificial Analysis's report here!

Great! Is there a technical paper? Or a technical blogpost? Would love to know more, mainly to create some content around this!

See our research here!

Thanks! 🙏🏼

This is awesome. Congrats @StefanoErmon & team!

Thanks!
