Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Mercury 2 is live 🚀🚀 The world’s first reasoning diffusion LLM, delivering 5x faster performance than leading speed-optimized LLMs. Watching the team turn years of research into a real product never gets old, and I’m incredibly proud of what we’ve built. We’re just getting started on what diffusion can...

1,050,489 görüntüleme • 7 ay önce •via X (Twitter)

67 Yorum

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Check out more in our blog

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Full story by @dinabass in @business

NVIDIA AI profil fotoğrafı
NVIDIA AI7 ay önce

Huge congrats to the @_inception_ai team on Mercury 2 — blazing-fast diffusion reasoning on NVIDIA Blackwell GPUs and an exciting leap for real-time AI.

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

@_inception_ai Thank you @NVIDIAAI! Couldn't have done it without the infrastructure your team has built🤝

Volodymyr Kuleshov 🇺🇦 profil fotoğrafı
Volodymyr Kuleshov 🇺🇦7 ay önce

🚨Hot off the presses: the official Artificial Analysis benchmarking results are in! 🚀Mercury morels set a new frontier of speed and agentic quality

Beff (e/acc) profil fotoğrafı
Beff (e/acc)7 ay önce

very exciting

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

🚀

YEETUS DELETUS profil fotoğrafı
YEETUS DELETUS7 ay önce

Mate, if it's already 1009 Tok/s on standard hardware from Nvidia, imagine if it was running on Groq or Cerebras hardware? Heck, is it even possible running it on Taalas hardware? Exciting times for Mercury 2, hoping to test it out in heavy duty reasoning tasks!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Exciting times indeed!

Yash Kr Gupta profil fotoğrafı
Yash Kr Gupta7 ay önce

1009 tokens per second on standard GPUs. Parallel diffusion instead of sequential decoding. This is a fundamentally different architecture not just an optimization.

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

We have been working hard!

Volodymyr Kuleshov 🇺🇦 profil fotoğrafı
Volodymyr Kuleshov 🇺🇦7 ay önce

Breaking the 1000/tok sec barrier at the intelligence level of GPT5-mini, Flash, Haiku 🚀🚀🚀

Aditya Grover profil fotoğrafı
Aditya Grover7 ay önce

Reasoning quality at non-reasoning speeds, super excited to see this out @_inception_ai!

Krishna Kaasyap profil fotoğrafı
Krishna Kaasyap7 ay önce

Great product. Congratulations on shipping a better model, and even faster than Google (not sure what happened to Gemini diffusion). Small typo though: there is no GPT-5.2 Mini, just GPT-5 Mini!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks!

Vishnu Vardhan profil fotoğrafı
Vishnu Vardhan7 ay önce

Super cool, Can’t wait to try it out guys, we are trying to build something like openclaw but with end to end voice pipeline , this could really make the difference, anyway we could get api access soon?

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Glad to hear your interest! Sign up for access here. We are rolling out access quickly.

Vishnu Vardhan profil fotoğrafı
Vishnu Vardhan7 ay önce

Done!

AVB profil fotoğrafı
AVB7 ay önce

For people who are wondering how dLLMs work: If you've seen Arrival, it's exactly that. Bidirectional attention, start with noise & iteratively denoise into actual text. Autoregressive LLMs iterate on tokens per step Diffusion LLMs iterate on revisions of large chunks per step

Samar Khanna profil fotoğrafı
Samar Khanna7 ay önce

Super excited to share the fastest reasoning LLM, built with diffusion! Mercury 2 is even faster and much, MUCH better than Mercury 1

Mike Knoop profil fotoğrafı
Mike Knoop7 ay önce

Congrats on the launch!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks for your support!

m profil fotoğrafı
m7 ay önce

Excited to support u guys 🚀

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks! We are just getting started.

Jeffrey Emanuel profil fotoğrafı
Jeffrey Emanuel7 ay önce

Very cool. I’m curious how the speed changes as the model size grows. For a Sonnet-class model (say, Sonnet 4), would it only be 3x faster? Is it currently capable of performing at that cognitive level at any speed?

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

We're actively scaling our models. Speed declines for larger models, but the same way that it does for a traditional LLM.

Boyuan (Nemo) Chen profil fotoğrafı
Boyuan (Nemo) Chen7 ay önce

5x faster is wild but does diffusion-based decoding degrade gracefully under long context? autoregressive models at least fail predictably. curious what happens when you push Mercury to 128k tokens

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks! Mercury 2 supports up to 128k context length!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

@boyuan_chen Not much published on this

Xen profil fotoğrafı
Xen7 ay önce

Some answers to expected questions about Mercury 2 right now: Can I run it locally? → No. It's API-only for now (closed weights, forget about local/RTX 4090 via LM Studio option) - their playground or API is the only option. How fast & how much? → Up to 1,196 tok/s output (world's fastest reasoning model via diffusion parallel gen) → $0.25/M input | $0.75/M output (cheaper blended than many other models and you get 10M free tokens to start) How does it look vs MiniMax M2.5 benchmarks? → MiniMax wins raw intelligence (~42 AA Index vs 33 for Mercury2), coding strength, 205k context → Mercury 2 crushes speed/throughput + cost for real-time agents/voice/inference Major plus: → Just because this reminds me of Freddie, I'll try it. Bottom line: Mercury 2 seems a valid option for "instant everything and fast".

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thank you!

VioP profil fotoğrafı
VioP7 ay önce

you guys arr back congrats!!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks! It's live in the playground, give it a try at

mrkelly profil fotoğrafı
mrkelly7 ay önce

Speed is a commodity. Reasoning depth is the only moat left.

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

We are excited to see the new use cases that our speed enables!

Julia Turc profil fotoğrafı
Julia Turc7 ay önce

Congrats on the launch!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks! Been great collaborating recently.

Misha Laskin profil fotoğrafı
Misha Laskin7 ay önce

Impressive! Congratulations

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks!

Jonathan Rosenberg profil fotoğrafı
Jonathan Rosenberg7 ay önce

Been advising Inception. Mercury 2 is live: first reasoning diffusion LLM, ~1,000 tokens/sec, 5x faster. Not a tuning trick, a different architecture. @stefanoermon helped invent diffusion. His team made it work at production scale. Pay attention.

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks for your support!

Davis Treybig profil fotoğrafı
Davis Treybig7 ay önce

So fast!!!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

@TreybigDavis ⚡️

Michele Catasta profil fotoğrafı
Michele Catasta7 ay önce

Exceptional work, Stefano & team!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks! We are excited for the use cases that our speed will enable.

Saurabh Shukla profil fotoğrafı
Saurabh Shukla7 ay önce

Amazing. I learned so much from your course on diffusion. all the best. This looks impressive. 🚀

Sarah Catanzaro profil fotoğrafı
Sarah Catanzaro7 ay önce

Mercury can code apps faster than you can scroll through code. So cool.

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

⚡️🚀

Victor M profil fotoğrafı
Victor M7 ay önce

any plan to open source it? or a small variant so we can get a taste 👀

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Not at the moment. But you can get a taste right now at

himanshu profil fotoğrafı
himanshu7 ay önce

what are the params?

Yinjie Wang profil fotoğrafı
Yinjie Wang7 ay önce

dLLM is good base model for action modeling, e.g. computer-use agents. You don't wanna wait for few seconds before a single click or one line of code in terminal.

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Agreed!

Christian T White profil fotoğrafı
Christian T White7 ay önce

Well done and love the opening denoising of your video

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Glad you liked it!

Amit Verma profil fotoğrafı
Amit Verma7 ay önce

Amazing milestone Mercury 2: "Delivering 5x faster reasoning." 🧠⚡️ Me: Currently spending 4 hours trying to logically reason with my Python environment because it refuses to run a basic script. 🐍💻 If this AI can politely explain my syntax errors to me 5x faster, I am completely sold.

sankalp profil fotoğrafı
sankalp7 ay önce

where can we see benchmarks

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

See benchmarks in our blog post here:

Shivendra Soni profil fotoğrafı
Shivendra Soni7 ay önce

How does the cost compare to SOTA models? Like sonnet of 5.2codex. This would be an amazing thing to have for voice ai agents

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

See our pricing here!

Mathias LÜCK profil fotoğrafı
Mathias LÜCK7 ay önce

@oyegpt How well does it perform?

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

@oyegpt See Artificial Analysis's report here!

AVB profil fotoğrafı
AVB7 ay önce

Great! Is there a technical paper? Or a technical blogpost? Would love to know more, mainly to create some content around this!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

See our research here!

AVB profil fotoğrafı
AVB7 ay önce

Thanks! 🙏🏼

Marco Mascorro profil fotoğrafı
Marco Mascorro7 ay önce

This is awesome. Congrats @StefanoErmon & team!

Stefano Ermon profil fotoğrafı
Stefano Ermon7 ay önce

Thanks!

Benzer Videolar