Video wird geladen...
Video konnte nicht geladen werden
Introducing NVIDIA Nemotron 3 Super 🎉 Open 120B-parameter (12B active) hybrid Mamba-Transformer MoE model Native 1M-token context Built for compute-efficient, high-accuracy multi-agent applications Plus, fully open weights, datasets and recipes for easy customization and deployment. 🧵
135,620 Aufrufe • vor 7 Monaten •via X (Twitter)
40 Kommentare

This latest addition to the Nemotron family isn't just a bigger Nano. ✅ Up to 5x higher throughput and 2x accuracy than the previous version ✅ Latent MoE that calls 4x as many expert specialists for the same inference cost ✅ Multi-token prediction that dramatically reduces generation time ✅ Hybrid Mamba-Transformer backbone delivers 4x improved memory and compute efficiency ✅ Native NVFP4 pretraining optimized for NVIDIA Blackwell Check out the deep dive into the architectural decisions and training methods behind the model 👇

🦞These innovations come together to create a model that is well suited for long-running autonomous agents. On PinchBench—a benchmark for evaluating LLMs as @OpenClaw coding agents—Nemotron 3 Super scores 85.6% across the full test suite, making it the best open model in its class.

“NVIDIA Nemotron 3 Super: The new leader in open, efficient intelligence”

✨ Nemotron 3 Super is now available to @Perplexity_ai Pro and Max subscribers in the model selector drop-down. It can also be used through the Agent API and Perplexity Computer.

Ready to get started? Nemotron 3 Super supports deployment across environments, from workstations to the cloud, and can be accessed through API, OpenRouter, or It is now live and available on major inference platforms, packaged as NVIDIA NIM: 📥 Download the weights from @HuggingFace, launch an optimized instance through NVIDIA NIM, fine-tune with @UnslothAI, or start with the cookbooks from @lmsysorg and @vllm_project to get running in minutes. Super is also available through @baseten, @Cloudflare, @deepinfra, @FireworksAI_HQ, @friendliai, @LightningAI, and @modal. 📗Read the Nemotron 3 Super technical report for the full details

Running this on my DGX Spark right now. It’s absolutely incredible. Bravo @NVIDIAAI

@NVIDIAAI 🙌

Try it for free in Kilo Code!

lets gooo

🙌

Given the focus on multi-agent applications, will Nemotron 3 prioritize emergent behavior understanding, or is that still largely a black box? 🤔

fine tune it with trl 💚!

🕸🤖🔥🛢🔥🤖🕸

Really like this model! Using it in this demo to act as a fraud advisor, it's very capable 🙌

is that @llm_wizard? 🧙

120 billion parameters but only 12 billion active at once. Open weights. 1 million token context. Nvidia just made every closed source AI company nervous with a single drop. The moat isn't the model anymore. It's the hardware that runs it.

Some demo applications I created:

wow Nemotron 3 Super looks next level!

spinning up 100 of these models for AI subminds with tool use will be crazy after a fine-tune, ablation, and retraining.

Here is my take:

Long context windows usually don't mean better accuracy - this is a known problem in current llm architecture. we built Hindsight to fix it

Efficient af. Can't wait for quants, but also hoping Nvidia's next workstation card can fit the full 120b, or maybe support Nvlink? 🤞

Sacrificing quality just to go fast?

@grok how to implement this for openclaw? A plain openrouter api?

@NVIDIA_AI_PC Cool 😎

Multi agent AI getting real

@HannaHajishirzi Good job 👌

Great news for the .1% of companies that can run this, now for consumers?

The interesting question here is not: can it hold 1M tokens? It’s whether long context agent runs stay stable enough to be operationally useful. A lot of systems look great on context length and throughput, then get weird once state, tool use, retries, and long-horizon memory start interacting. That’s where the real architecture test starts.

要点をまとめて日本語で紹介しました Key points summarized in Japanese:

@grok can my GTX 550 handle this?

"Built for multi-agent applications" is the key phrase here. Most models optimize for single-turn accuracy. This optimizes for what agents need: sustained context across long task chains. 12B active / 120B total is the right arch for agents that run hours, not seconds

yall are amazing

native 1m token context is a game changer for multi agent applications, how does it handle context switching between different agents and tasks, any benchmarks on that

where is the spark recipe :(

The future of AI makes me so dang motivated. We live in a blissful time.

Got 100 tokens per second on my RTX Pro 6000!

@NVIDIAAI Looking forward to trying it out

@grok this is exciting. Do you think a Spark will actually do the 1 million context?

1M context with open weights changes the game for small labs overnight.


