Загрузка видео...

Не удалось загрузить видео

На главную

Mercury 2 is live 🚀🚀 The world’s first reasoning diffusion LLM, delivering 5x faster performance than leading speed-optimized LLMs. Watching the team turn years of research into a real product never gets old, and I’m incredibly proud of what we’ve built. We’re just getting started on what diffusion can...

1,050,489 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 67

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Check out more in our blog

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Full story by @dinabass in @business

Фото профиля NVIDIA AI
NVIDIA AI7 месяцев назад

Huge congrats to the @_inception_ai team on Mercury 2 — blazing-fast diffusion reasoning on NVIDIA Blackwell GPUs and an exciting leap for real-time AI.

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

@_inception_ai Thank you @NVIDIAAI! Couldn't have done it without the infrastructure your team has built🤝

Фото профиля Volodymyr Kuleshov 🇺🇦
Volodymyr Kuleshov 🇺🇦7 месяцев назад

🚨Hot off the presses: the official Artificial Analysis benchmarking results are in! 🚀Mercury morels set a new frontier of speed and agentic quality

Фото профиля Beff (e/acc)
Beff (e/acc)7 месяцев назад

very exciting

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

🚀

Фото профиля YEETUS DELETUS
YEETUS DELETUS7 месяцев назад

Mate, if it's already 1009 Tok/s on standard hardware from Nvidia, imagine if it was running on Groq or Cerebras hardware? Heck, is it even possible running it on Taalas hardware? Exciting times for Mercury 2, hoping to test it out in heavy duty reasoning tasks!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Exciting times indeed!

Фото профиля Yash Kr Gupta
Yash Kr Gupta7 месяцев назад

1009 tokens per second on standard GPUs. Parallel diffusion instead of sequential decoding. This is a fundamentally different architecture not just an optimization.

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

We have been working hard!

Фото профиля Volodymyr Kuleshov 🇺🇦
Volodymyr Kuleshov 🇺🇦7 месяцев назад

Breaking the 1000/tok sec barrier at the intelligence level of GPT5-mini, Flash, Haiku 🚀🚀🚀

Фото профиля Aditya Grover
Aditya Grover7 месяцев назад

Reasoning quality at non-reasoning speeds, super excited to see this out @_inception_ai!

Фото профиля Krishna Kaasyap
Krishna Kaasyap7 месяцев назад

Great product. Congratulations on shipping a better model, and even faster than Google (not sure what happened to Gemini diffusion). Small typo though: there is no GPT-5.2 Mini, just GPT-5 Mini!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks!

Фото профиля Vishnu Vardhan
Vishnu Vardhan7 месяцев назад

Super cool, Can’t wait to try it out guys, we are trying to build something like openclaw but with end to end voice pipeline , this could really make the difference, anyway we could get api access soon?

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Glad to hear your interest! Sign up for access here. We are rolling out access quickly.

Фото профиля Vishnu Vardhan
Vishnu Vardhan7 месяцев назад

Done!

Фото профиля AVB
AVB7 месяцев назад

For people who are wondering how dLLMs work: If you've seen Arrival, it's exactly that. Bidirectional attention, start with noise & iteratively denoise into actual text. Autoregressive LLMs iterate on tokens per step Diffusion LLMs iterate on revisions of large chunks per step

Фото профиля Samar Khanna
Samar Khanna7 месяцев назад

Super excited to share the fastest reasoning LLM, built with diffusion! Mercury 2 is even faster and much, MUCH better than Mercury 1

Фото профиля Mike Knoop
Mike Knoop7 месяцев назад

Congrats on the launch!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks for your support!

Фото профиля m
m7 месяцев назад

Excited to support u guys 🚀

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks! We are just getting started.

Фото профиля Jeffrey Emanuel
Jeffrey Emanuel7 месяцев назад

Very cool. I’m curious how the speed changes as the model size grows. For a Sonnet-class model (say, Sonnet 4), would it only be 3x faster? Is it currently capable of performing at that cognitive level at any speed?

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

We're actively scaling our models. Speed declines for larger models, but the same way that it does for a traditional LLM.

Фото профиля Boyuan (Nemo) Chen
Boyuan (Nemo) Chen7 месяцев назад

5x faster is wild but does diffusion-based decoding degrade gracefully under long context? autoregressive models at least fail predictably. curious what happens when you push Mercury to 128k tokens

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks! Mercury 2 supports up to 128k context length!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

@boyuan_chen Not much published on this

Фото профиля Xen
Xen7 месяцев назад

Some answers to expected questions about Mercury 2 right now: Can I run it locally? → No. It's API-only for now (closed weights, forget about local/RTX 4090 via LM Studio option) - their playground or API is the only option. How fast & how much? → Up to 1,196 tok/s output (world's fastest reasoning model via diffusion parallel gen) → $0.25/M input | $0.75/M output (cheaper blended than many other models and you get 10M free tokens to start) How does it look vs MiniMax M2.5 benchmarks? → MiniMax wins raw intelligence (~42 AA Index vs 33 for Mercury2), coding strength, 205k context → Mercury 2 crushes speed/throughput + cost for real-time agents/voice/inference Major plus: → Just because this reminds me of Freddie, I'll try it. Bottom line: Mercury 2 seems a valid option for "instant everything and fast".

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thank you!

Фото профиля VioP
VioP7 месяцев назад

you guys arr back congrats!!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks! It's live in the playground, give it a try at

Фото профиля mrkelly
mrkelly7 месяцев назад

Speed is a commodity. Reasoning depth is the only moat left.

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

We are excited to see the new use cases that our speed enables!

Фото профиля Julia Turc
Julia Turc7 месяцев назад

Congrats on the launch!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks! Been great collaborating recently.

Фото профиля Misha Laskin
Misha Laskin7 месяцев назад

Impressive! Congratulations

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks!

Фото профиля Jonathan Rosenberg
Jonathan Rosenberg7 месяцев назад

Been advising Inception. Mercury 2 is live: first reasoning diffusion LLM, ~1,000 tokens/sec, 5x faster. Not a tuning trick, a different architecture. @stefanoermon helped invent diffusion. His team made it work at production scale. Pay attention.

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks for your support!

Фото профиля Davis Treybig
Davis Treybig7 месяцев назад

So fast!!!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

@TreybigDavis ⚡️

Фото профиля Michele Catasta
Michele Catasta7 месяцев назад

Exceptional work, Stefano & team!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks! We are excited for the use cases that our speed will enable.

Фото профиля Saurabh Shukla
Saurabh Shukla7 месяцев назад

Amazing. I learned so much from your course on diffusion. all the best. This looks impressive. 🚀

Фото профиля Sarah Catanzaro
Sarah Catanzaro7 месяцев назад

Mercury can code apps faster than you can scroll through code. So cool.

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

⚡️🚀

Фото профиля Victor M
Victor M7 месяцев назад

any plan to open source it? or a small variant so we can get a taste 👀

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Not at the moment. But you can get a taste right now at

Фото профиля himanshu
himanshu7 месяцев назад

what are the params?

Фото профиля Yinjie Wang
Yinjie Wang7 месяцев назад

dLLM is good base model for action modeling, e.g. computer-use agents. You don't wanna wait for few seconds before a single click or one line of code in terminal.

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Agreed!

Фото профиля Christian T White
Christian T White7 месяцев назад

Well done and love the opening denoising of your video

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Glad you liked it!

Фото профиля Amit Verma
Amit Verma7 месяцев назад

Amazing milestone Mercury 2: "Delivering 5x faster reasoning." 🧠⚡️ Me: Currently spending 4 hours trying to logically reason with my Python environment because it refuses to run a basic script. 🐍💻 If this AI can politely explain my syntax errors to me 5x faster, I am completely sold.

Фото профиля sankalp
sankalp7 месяцев назад

where can we see benchmarks

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

See benchmarks in our blog post here:

Фото профиля Shivendra Soni
Shivendra Soni7 месяцев назад

How does the cost compare to SOTA models? Like sonnet of 5.2codex. This would be an amazing thing to have for voice ai agents

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

See our pricing here!

Фото профиля Mathias LÜCK
Mathias LÜCK7 месяцев назад

@oyegpt How well does it perform?

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

@oyegpt See Artificial Analysis's report here!

Фото профиля AVB
AVB7 месяцев назад

Great! Is there a technical paper? Or a technical blogpost? Would love to know more, mainly to create some content around this!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

See our research here!

Фото профиля AVB
AVB7 месяцев назад

Thanks! 🙏🏼

Фото профиля Marco Mascorro
Marco Mascorro7 месяцев назад

This is awesome. Congrats @StefanoErmon & team!

Фото профиля Stefano Ermon
Stefano Ermon7 месяцев назад

Thanks!

Похожие видео