Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing NVIDIA Nemotron 3 Super 🎉 Open 120B-parameter (12B active) hybrid Mamba-Transformer MoE model Native 1M-token context Built for compute-efficient, high-accuracy multi-agent applications Plus, fully open weights, datasets and recipes for easy customization and deployment. 🧵

135,620 görüntüleme • 7 ay önce •via X (Twitter)

40 Yorum

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

This latest addition to the Nemotron family isn't just a bigger Nano. ✅ Up to 5x higher throughput and 2x accuracy than the previous version ✅ Latent MoE that calls 4x as many expert specialists for the same inference cost
 ✅ Multi-token prediction that dramatically reduces generation time ✅ Hybrid Mamba-Transformer backbone delivers 4x improved memory and compute efficiency ✅ Native NVFP4 pretraining optimized for NVIDIA Blackwell Check out the deep dive into the architectural decisions and training methods behind the model 👇

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

🦞These innovations come together to create a model that is well suited for long-running autonomous agents. On PinchBench—a benchmark for evaluating LLMs as @OpenClaw coding agents—Nemotron 3 Super scores 85.6% across the full test suite, making it the best open model in its class.

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

“NVIDIA Nemotron 3 Super: The new leader in open, efficient intelligence”

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

✨ Nemotron 3 Super is now available to @Perplexity_ai Pro and Max subscribers in the model selector drop-down. It can also be used through the Agent API and Perplexity Computer.

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

Ready to get started? Nemotron 3 Super supports deployment across environments, from workstations to the cloud, and can be accessed through API, OpenRouter, or It is now live and available on major inference platforms, packaged as NVIDIA NIM: 📥 Download the weights from @HuggingFace, launch an optimized instance through NVIDIA NIM, fine-tune with @UnslothAI, or start with the cookbooks from @lmsysorg and @vllm_project to get running in minutes. Super is also available through @baseten, @Cloudflare, @deepinfra, @FireworksAI_HQ, @friendliai, @LightningAI, and @modal. 📗Read the Nemotron 3 Super technical report for the full details

Alex Finn profil fotoğrafı
Alex Finn7 ay önce

Running this on my DGX Spark right now. It’s absolutely incredible. Bravo @NVIDIAAI

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

@NVIDIAAI 🙌

Kilo (acq. by Anaconda) profil fotoğrafı
Kilo (acq. by Anaconda)7 ay önce

Try it for free in Kilo Code!

Luis Catacora profil fotoğrafı
Luis Catacora7 ay önce

lets gooo

NVIDIA AI Developer profil fotoğrafı
NVIDIA AI Developer7 ay önce

🙌

Hüseyin Örskaya profil fotoğrafı
Hüseyin Örskaya7 ay önce

Given the focus on multi-agent applications, will Nemotron 3 prioritize emergent behavior understanding, or is that still largely a black box? 🤔

Sergio Paniego profil fotoğrafı
Sergio Paniego7 ay önce

fine tune it with trl 💚!

Joe Stevens profil fotoğrafı
Joe Stevens7 ay önce

🕸🤖🔥🛢🔥🤖🕸

Vision Agents profil fotoğrafı
Vision Agents6 ay önce

Really like this model! Using it in this demo to act as a fraud advisor, it's very capable 🙌

Maziyar PANAHI profil fotoğrafı
Maziyar PANAHI7 ay önce

is that @llm_wizard? 🧙

Leo Sukkar profil fotoğrafı
Leo Sukkar7 ay önce

120 billion parameters but only 12 billion active at once. Open weights. 1 million token context. Nvidia just made every closed source AI company nervous with a single drop. The moat isn't the model anymore. It's the hardware that runs it.

Cobus Greyling profil fotoğrafı
Cobus Greyling7 ay önce

Some demo applications I created:

Elena🌸 profil fotoğrafı
Elena🌸7 ay önce

wow Nemotron 3 Super looks next level!

teromee profil fotoğrafı
teromee7 ay önce

spinning up 100 of these models for AI subminds with tool use will be crazy after a fine-tune, ablation, and retraining.

Cobus Greyling profil fotoğrafı
Cobus Greyling7 ay önce

Here is my take:

Nicolò Boschi profil fotoğrafı
Nicolò Boschi7 ay önce

Long context windows usually don't mean better accuracy - this is a known problem in current llm architecture. we built Hindsight to fix it

Slop-Swap profil fotoğrafı
Slop-Swap7 ay önce

Efficient af. Can't wait for quants, but also hoping Nvidia's next workstation card can fit the full 120b, or maybe support Nvlink? 🤞

Debdoot Ghosh profil fotoğrafı
Debdoot Ghosh7 ay önce

Sacrificing quality just to go fast?

Fahd. profil fotoğrafı
Fahd.7 ay önce

@grok how to implement this for openclaw? A plain openrouter api?

BraveTom profil fotoğrafı
BraveTom7 ay önce

@NVIDIA_AI_PC Cool 😎

Kyriakos profil fotoğrafı
Kyriakos7 ay önce

Multi agent AI getting real

Götz-Henrik profil fotoğrafı
Götz-Henrik7 ay önce

@HannaHajishirzi Good job 👌

Eric ⚡️ Building... profil fotoğrafı
Eric ⚡️ Building...7 ay önce

Great news for the .1% of companies that can run this, now for consumers?

Ali Noori profil fotoğrafı
Ali Noori6 ay önce

The interesting question here is not: can it hold 1M tokens? It’s whether long context agent runs stay stable enough to be operationally useful. A lot of systems look great on context length and throughput, then get weird once state, tool use, retries, and long-horizon memory start interacting. That’s where the real architecture test starts.

あいり|海外AIニュースを毎日届ける人 profil fotoğrafı
あいり|海外AIニュースを毎日届ける人6 ay önce

要点をまとめて日本語で紹介しました Key points summarized in Japanese:

Cameron Taylor profil fotoğrafı
Cameron Taylor7 ay önce

@grok can my GTX 550 handle this?

Schelling Protocol profil fotoğrafı
Schelling Protocol7 ay önce

"Built for multi-agent applications" is the key phrase here. Most models optimize for single-turn accuracy. This optimizes for what agents need: sustained context across long task chains. 12B active / 120B total is the right arch for agents that run hours, not seconds

matthew stevick profil fotoğrafı
matthew stevick7 ay önce

yall are amazing

Youth profil fotoğrafı
Youth7 ay önce

native 1m token context is a game changer for multi agent applications, how does it handle context switching between different agents and tasks, any benchmarks on that

Jeffrey 杰弗瑞 profil fotoğrafı
Jeffrey 杰弗瑞7 ay önce

where is the spark recipe :(

tylerdotai profil fotoğrafı
tylerdotai7 ay önce

The future of AI makes me so dang motivated. We live in a blissful time.

Dr. Tristan Behrens profil fotoğrafı
Dr. Tristan Behrens7 ay önce

Got 100 tokens per second on my RTX Pro 6000!

NowDigi profil fotoğrafı
NowDigi7 ay önce

@NVIDIAAI Looking forward to trying it out

Kevin Faircloth profil fotoğrafı
Kevin Faircloth7 ay önce

@grok this is exciting. Do you think a Spark will actually do the 1 million context?

Martin Szerment | Practical AI profil fotoğrafı
Martin Szerment | Practical AI7 ay önce

1M context with open weights changes the game for small labs overnight.

Benzer Videolar

Inside Nemotron and NVIDIA's AI lab: my conversation with Bryan Catanzaro (Bryan Catanzaro). NVIDIA is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? We go deep into the Nemotron models, what it takes to build a top AI lab, and the future of frontier AI. 01:33 - Is open source AI catching the frontier? 05:29 - Do closed labs blocking distillation slow open source down? 07:42 - Is the US falling behind China? 10:30 - Why companies actually choose open models 12:39 - A "crazy" 2008 bet: machine learning on GPUs 15:33 - Working with Andrew Ng and Dario Amodei at Baidu 17:41 - Coming back to NVIDIA: DLSS and the birth of Megatron 21:55 - The real reason NVIDIA builds its own models 24:28 - Is Moore's Law really dead? 33:37 - The Nemotron family: Nano, Super, Ultra 35:09 - Built for agents: why NVIDIA bets on speed 36:02 - How you train a 550B model in 4 bits 39:25 - Hybrid Mamba-Transformer, explained simply 42:31 - Mixture of experts, and why NVIDIA built NVL72 around it 47:26 - Why a 1-million-token context window matters 49:26 - Multi-token prediction: how the model predicts 5 tokens at once 52:47 - Multi-teacher distillation: teaching one model from many 58:01 - Where reinforcement learning goes next 01:00:16 - Inside NVIDIA's research org: "the mission is the boss" 01:04:03 - How NVIDIA decides who gets the GPUs 01:10:53 - Why NVIDIA still feels entrepreneurial after 33 years 01:12:58 - Why Bryan doesn't believe in the singularity 01:17:50 - The AI backlash 01:19:18 - The controversial case: open AI is safer than closed

Matt Turck

56,954 görüntüleme • 3 ay önce

Soofi Consortium Releases Soofi S 30B-A3B: An Open 31.6B Model for German and English Hitting 79.1 German Aggregate With Only 3.2B Active Parameters. Here's how it works. 👇 1. Sparsity in two places at once 52 layers: 23 Mamba-2, 23 granular MoE, 6 Grouped-Query Attention. The MoE router picks 6 of 128 experts per token, plus 2 shared. Mamba-2 carries the sequence mixing with a fixed-size recurrent state, so 46 of 52 layers keep no KV cache at all. → 3.2B of 31.6B parameters active per token 2. Reference architecture on purpose No bespoke backbone. It adopts NVIDIA's Nemotron 3 Nano design without modification — for day-one vLLM kernels, for serving efficiency, and for scientific control. That last one is the real move: Nemotron becomes an architecture-identical baseline, so the data recipe is the only variable left. 3. German as the deliberate variable Three-phase Warmup–Stable–Decay curriculum. Phase 1 is breadth at a 1e-3 plateau, Phase 2 concentrates high-quality data as the LR decays, Phase 3 stretches context to 1M tokens. → ~26.68T consumed tokens → German 7.2% → 15.32% of the mixture, vs ~5% for all non-English in the Nemotron reference → +4.2 German aggregate, +1.8 English, +6.7 held-out English over Nemotron 4. Where the architecture pays: memory bandwidth Every decoded token re-reads the weights and, for a Transformer, the attention cache of every sequence in the batch. Six KV layers instead of 52 keeps that per-sequence state small. Measured on one B200, TP=1, vLLM latency-subtraction. → 8–9× aggregate decode TPS/GPU vs dense 14–24B models at 40K context, batch 32 → decode stays flat from 4K to 256K 5. The numbers (base model, lm-evaluation-harness, 16 open baselines) → 70.1 English aggregate, +2.8 over Olmo 3 32B → 79.1 German aggregate, +6.3 over Apertus 70B → 73.8 HumanEval, 84.2 MBPP-DE, 88.8 GLP-DE, 61.2 INCLUDE-DE Full analysis: Paper: Technical details:

Marktechpost AI

65,637 görüntüleme • 2 ay önce