Loading video...
Video Failed to Load
Today we're excited to announce Mercury 2.5 It’s the most capable diffusion LLM on the market. It is a 40% jump in intelligence over Mercury 2 and runs at over 1,100 tokens/sec on widely available NVIDIA AI GPUs.
143,813 views • 16 days ago •via X (Twitter)
41 Comments

Mercury powers real-time systems with the tightest latency budgets. @OpenCall_AI uses Mercury to run voice agents on live patient calls. After switching from an AI inference chip provider, p50 latency fell below 200 ms and p99 fell from several minutes to one second. This keeps multi-step reasoning inside the latency budget of a live call.

Mercury 2.5 is available through our API and on @OpenRouter and @Baseten. New accounts include 100M free tokens. At launch, Mercury 2.5 is 80% off at $0.04/M input and $0.15/M output. Try Mercury 2.5:

Congrats to the team Stefano!

@NVIDIAAI Mercury 2.5 continues to redraw the performance-throughput pareto frontier for LLMs

@NVIDIAAI Thanks Deedy!

@NVIDIAAI Congrats on the launch! Excited for developers to get their hands on Mercury 2.5 and see what they build.

@NVIDIAAI let's goooo!!! congratulations @StefanoErmon @volokuleshov @adityagrover_ and team Inception!!!

@NVIDIAAI @volokuleshov @adityagrover_ Thanks!

@NVIDIAAI i love the idea of diffusion LLMs but i just wish you've benchmarked it against popular modals too and also mention which TerminalBench version you've tested

@NVIDIAAI ⚡️

@NVIDIAAI 🧠⚡️📈

@NVIDIAAI let's goooo!! best team 😄

@NVIDIAAI Thanks Neal!

@NVIDIAAI And we’re hiring 😏

@NVIDIAAI 🧨🧨🧨

@NVIDIAAI Congrats to the whole Inception team on Mercury 2.5! Great to see diffusion models keep proving out in production.

@NVIDIAAI This is great, congrats!

@NVIDIAAI Thank you Azalia!

@NVIDIAAI So fast!!

@NVIDIAAI Congrats on the launch!

@NVIDIAAI omg I'm so excited to try this out, thank you!

@NVIDIAAI Thank you!

@NVIDIAAI Wow this is speedy!!

@NVIDIAAI congrats on the launch @StefanoErmon @volokuleshov !

@NVIDIAAI Congrats on the launch, team! Excited to see what gets built with Mercury 2.5 🚀

@NVIDIAAI very fast

@NVIDIAAI Thank you!

@NVIDIAAI What's the 40% measured on? Genuinely asking the speed numbers are easy to verify but "intelligence" needs an index name to mean anything. Happy to be pointed at it.

@NVIDIAAI Lets go! Congrats on the launch @StefanoErmon @volokuleshov and team @_inception_ai !

@NVIDIAAI This is blazing fast. congrats.

@NVIDIAAI wow sick

@NVIDIAAI I am really hyped up for this. Building one myself right now.

@NVIDIAAI HELLLL YAHHHH upgrading my app

@NVIDIAAI excited for diffusion language models.

@NVIDIAAI congrats

@NVIDIAAI Congrats!

@NVIDIAAI Congratulations @StefanoErmon — great to see progress in text diffusion at scale

@NVIDIAAI Reading this made me realize I don't know shit about how diffusion decoding actually works. Guess I have homework this weekend.

@NVIDIAAI Comparable quality to: Haiku and Gemini Flash-light 😭 But im glad someone at least continues work on different architectures.

@NVIDIAAI Congrats! Really impressive to see the team keep shipping Mercury 2.5

@NVIDIAAI Congrats!
