Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing NVIDIA Nemotron 3 Super 🎉 Open 120B-parameter (12B active) hybrid Mamba-Transformer MoE model Native 1M-token context Built for compute-efficient, high-accuracy multi-agent applications Plus, fully open weights, datasets and recipes for easy customization and deployment. 🧵

135,620 Aufrufe • vor 7 Monaten •via X (Twitter)

40 Kommentare

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

This latest addition to the Nemotron family isn't just a bigger Nano. ✅ Up to 5x higher throughput and 2x accuracy than the previous version ✅ Latent MoE that calls 4x as many expert specialists for the same inference cost
 ✅ Multi-token prediction that dramatically reduces generation time ✅ Hybrid Mamba-Transformer backbone delivers 4x improved memory and compute efficiency ✅ Native NVFP4 pretraining optimized for NVIDIA Blackwell Check out the deep dive into the architectural decisions and training methods behind the model 👇

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

🦞These innovations come together to create a model that is well suited for long-running autonomous agents. On PinchBench—a benchmark for evaluating LLMs as @OpenClaw coding agents—Nemotron 3 Super scores 85.6% across the full test suite, making it the best open model in its class.

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

“NVIDIA Nemotron 3 Super: The new leader in open, efficient intelligence”

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

✨ Nemotron 3 Super is now available to @Perplexity_ai Pro and Max subscribers in the model selector drop-down. It can also be used through the Agent API and Perplexity Computer.

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

Ready to get started? Nemotron 3 Super supports deployment across environments, from workstations to the cloud, and can be accessed through API, OpenRouter, or It is now live and available on major inference platforms, packaged as NVIDIA NIM: 📥 Download the weights from @HuggingFace, launch an optimized instance through NVIDIA NIM, fine-tune with @UnslothAI, or start with the cookbooks from @lmsysorg and @vllm_project to get running in minutes. Super is also available through @baseten, @Cloudflare, @deepinfra, @FireworksAI_HQ, @friendliai, @LightningAI, and @modal. 📗Read the Nemotron 3 Super technical report for the full details

Profilbild von Alex Finn
Alex Finnvor 7 Monaten

Running this on my DGX Spark right now. It’s absolutely incredible. Bravo @NVIDIAAI

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

@NVIDIAAI 🙌

Profilbild von Kilo (acq. by Anaconda)
Kilo (acq. by Anaconda)vor 7 Monaten

Try it for free in Kilo Code!

Profilbild von Luis Catacora
Luis Catacoravor 7 Monaten

lets gooo

Profilbild von NVIDIA AI Developer
NVIDIA AI Developervor 7 Monaten

🙌

Profilbild von Hüseyin Örskaya
Hüseyin Örskayavor 7 Monaten

Given the focus on multi-agent applications, will Nemotron 3 prioritize emergent behavior understanding, or is that still largely a black box? 🤔

Profilbild von Sergio Paniego
Sergio Paniegovor 7 Monaten

fine tune it with trl 💚!

Profilbild von Joe Stevens
Joe Stevensvor 7 Monaten

🕸🤖🔥🛢🔥🤖🕸

Profilbild von Vision Agents
Vision Agentsvor 6 Monaten

Really like this model! Using it in this demo to act as a fraud advisor, it's very capable 🙌

Profilbild von Maziyar PANAHI
Maziyar PANAHIvor 7 Monaten

is that @llm_wizard? 🧙

Profilbild von Leo Sukkar
Leo Sukkarvor 7 Monaten

120 billion parameters but only 12 billion active at once. Open weights. 1 million token context. Nvidia just made every closed source AI company nervous with a single drop. The moat isn't the model anymore. It's the hardware that runs it.

Profilbild von Cobus Greyling
Cobus Greylingvor 7 Monaten

Some demo applications I created:

Profilbild von Elena🌸
Elena🌸vor 7 Monaten

wow Nemotron 3 Super looks next level!

Profilbild von teromee
teromeevor 7 Monaten

spinning up 100 of these models for AI subminds with tool use will be crazy after a fine-tune, ablation, and retraining.

Profilbild von Cobus Greyling
Cobus Greylingvor 7 Monaten

Here is my take:

Profilbild von Nicolò Boschi
Nicolò Boschivor 7 Monaten

Long context windows usually don't mean better accuracy - this is a known problem in current llm architecture. we built Hindsight to fix it

Profilbild von Slop-Swap
Slop-Swapvor 7 Monaten

Efficient af. Can't wait for quants, but also hoping Nvidia's next workstation card can fit the full 120b, or maybe support Nvlink? 🤞

Profilbild von Debdoot Ghosh
Debdoot Ghoshvor 7 Monaten

Sacrificing quality just to go fast?

Profilbild von Fahd.
Fahd.vor 7 Monaten

@grok how to implement this for openclaw? A plain openrouter api?

Profilbild von BraveTom
BraveTomvor 7 Monaten

@NVIDIA_AI_PC Cool 😎

Profilbild von Kyriakos
Kyriakosvor 7 Monaten

Multi agent AI getting real

Profilbild von Götz-Henrik
Götz-Henrikvor 7 Monaten

@HannaHajishirzi Good job 👌

Profilbild von Eric ⚡️ Building...
Eric ⚡️ Building...vor 7 Monaten

Great news for the .1% of companies that can run this, now for consumers?

Profilbild von Ali Noori
Ali Noorivor 6 Monaten

The interesting question here is not: can it hold 1M tokens? It’s whether long context agent runs stay stable enough to be operationally useful. A lot of systems look great on context length and throughput, then get weird once state, tool use, retries, and long-horizon memory start interacting. That’s where the real architecture test starts.

Profilbild von あいり|海外AIニュースを毎日届ける人
あいり|海外AIニュースを毎日届ける人vor 6 Monaten

要点をまとめて日本語で紹介しました Key points summarized in Japanese:

Profilbild von Cameron Taylor
Cameron Taylorvor 7 Monaten

@grok can my GTX 550 handle this?

Profilbild von Schelling Protocol
Schelling Protocolvor 7 Monaten

"Built for multi-agent applications" is the key phrase here. Most models optimize for single-turn accuracy. This optimizes for what agents need: sustained context across long task chains. 12B active / 120B total is the right arch for agents that run hours, not seconds

Profilbild von matthew stevick
matthew stevickvor 7 Monaten

yall are amazing

Profilbild von Youth
Youthvor 7 Monaten

native 1m token context is a game changer for multi agent applications, how does it handle context switching between different agents and tasks, any benchmarks on that

Profilbild von Jeffrey 杰弗瑞
Jeffrey 杰弗瑞vor 7 Monaten

where is the spark recipe :(

Profilbild von tylerdotai
tylerdotaivor 7 Monaten

The future of AI makes me so dang motivated. We live in a blissful time.

Profilbild von Dr. Tristan Behrens
Dr. Tristan Behrensvor 7 Monaten

Got 100 tokens per second on my RTX Pro 6000!

Profilbild von NowDigi
NowDigivor 7 Monaten

@NVIDIAAI Looking forward to trying it out

Profilbild von Kevin Faircloth
Kevin Fairclothvor 7 Monaten

@grok this is exciting. Do you think a Spark will actually do the 1 million context?

Profilbild von Martin Szerment | Practical AI
Martin Szerment | Practical AIvor 7 Monaten

1M context with open weights changes the game for small labs overnight.

Ähnliche Videos

Inside Nemotron and NVIDIA's AI lab: my conversation with Bryan Catanzaro (Bryan Catanzaro). NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? We go deep into the Nemotron models, what it takes to build a top AI lab, and the future of frontier AI. 01:33 - Is open source AI catching the frontier? 05:29 - Do closed labs blocking distillation slow open source down? 07:42 - Is the US falling behind China? 10:30 - Why companies actually choose open models 12:39 - A "crazy" 2008 bet: machine learning on GPUs 15:33 - Working with Andrew Ng and Dario Amodei at Baidu 17:41 - Coming back to NVIDIA: DLSS and the birth of Megatron 21:55 - The real reason NVIDIA builds its own models 24:28 - Is Moore's Law really dead? 33:37 - The Nemotron family: Nano, Super, Ultra 35:09 - Built for agents: why NVIDIA bets on speed 36:02 - How you train a 550B model in 4 bits 39:25 - Hybrid Mamba-Transformer, explained simply 42:31 - Mixture of experts, and why NVIDIA built NVL72 around it 47:26 - Why a 1-million-token context window matters 49:26 - Multi-token prediction: how the model predicts 5 tokens at once 52:47 - Multi-teacher distillation: teaching one model from many 58:01 - Where reinforcement learning goes next 01:00:16 - Inside NVIDIA's research org: "the mission is the boss" 01:04:03 - How NVIDIA decides who gets the GPUs 01:10:53 - Why NVIDIA still feels entrepreneurial after 33 years 01:12:58 - Why Bryan doesn't believe in the singularity 01:17:50 - The AI backlash 01:19:18 - The controversial case: open AI is safer than closed

Matt Turck

56,954 Aufrufe • vor 3 Monaten

Soofi Consortium Releases Soofi S 30B-A3B: An Open 31.6B Model for German and English Hitting 79.1 German Aggregate With Only 3.2B Active Parameters. Here's how it works. 👇 1. Sparsity in two places at once 52 layers: 23 Mamba-2, 23 granular MoE, 6 Grouped-Query Attention. The MoE router picks 6 of 128 experts per token, plus 2 shared. Mamba-2 carries the sequence mixing with a fixed-size recurrent state, so 46 of 52 layers keep no KV cache at all. → 3.2B of 31.6B parameters active per token 2. Reference architecture on purpose No bespoke backbone. It adopts NVIDIA's Nemotron 3 Nano design without modification — for day-one vLLM kernels, for serving efficiency, and for scientific control. That last one is the real move: Nemotron becomes an architecture-identical baseline, so the data recipe is the only variable left. 3. German as the deliberate variable Three-phase Warmup–Stable–Decay curriculum. Phase 1 is breadth at a 1e-3 plateau, Phase 2 concentrates high-quality data as the LR decays, Phase 3 stretches context to 1M tokens. → ~26.68T consumed tokens → German 7.2% → 15.32% of the mixture, vs ~5% for all non-English in the Nemotron reference → +4.2 German aggregate, +1.8 English, +6.7 held-out English over Nemotron 4. Where the architecture pays: memory bandwidth Every decoded token re-reads the weights and, for a Transformer, the attention cache of every sequence in the batch. Six KV layers instead of 52 keeps that per-sequence state small. Measured on one B200, TP=1, vLLM latency-subtraction. → 8–9× aggregate decode TPS/GPU vs dense 14–24B models at 40K context, batch 32 → decode stays flat from 4K to 256K 5. The numbers (base model, lm-evaluation-harness, 16 open baselines) → 70.1 English aggregate, +2.8 over Olmo 3 32B → 79.1 German aggregate, +6.3 over Apertus 70B → 73.8 HumanEval, 84.2 MBPP-DE, 88.8 GLP-DE, 61.2 INCLUDE-DE Full analysis: Paper: Technical details:

Marktechpost AI

65,637 Aufrufe • vor 2 Monaten