Loading video...

Video Failed to Load

Go Home

ICYMI: Not every step in an agent workflow needs the same model. Meet NVIDIA NeMo Switchyard, a new open-source library for model routing. Use frontier models for complex reasoning and planning, and NVIDIA Nemotron Lightning for high-volume, specialized execution.

16,629 views • 1 month ago •via X (Twitter)

15 Comments

NVIDIA AI's profile picture
NVIDIA AI1 month ago

Learn more:

Joe Stevens's profile picture
Joe Stevens1 month ago

🕸🤖🕸

TerraByte AI's profile picture
TerraByte AI1 month ago

Geospatial search is a good routing test. A small model can handle place, date and sensor filters. Only ambiguous queries like “recent burn scars near vineyards” need the expensive model.

无限未来's profile picture
无限未来1 month ago

你好,为什么测试你们平台的模型,几乎全部都连接失败,使用测试成功的模型也经常报错

REVELATOR's profile picture
REVELATOR1 month ago

@grok exine structure closely matched modern maize pollen.

Alex Lee's profile picture
Alex Lee1 month ago

This is the direction agents are heading 👀 Right model, right task, better efficiency 🚀

freesoul's profile picture
freesoul1 month ago

Looking forward to testing the escalation and prefill routers in real workflows.

Patch's profile picture
Patch1 month ago

I thought we already had model routing….?

Strategize Labs's profile picture
Strategize Labs1 month ago

Routing to the right model depending on the job size and complexity. That's how you get real cost compression and value for your bucks.

Gokulan Vikash's profile picture
Gokulan Vikash23 days ago

Can I talk to someone from this team? Why is this not ready/useable for production workload?

Francesco Di Costanzo's profile picture
Francesco Di Costanzo1 month ago

Could my @openclaw pass the message to the router which then dispatches to the right model? That would be cool (provided it works). Equally could Sol be able to dispatch to Terra and Luna via this to optimise token usage?

KVCacheStore LLC's profile picture
KVCacheStore LLC1 month ago

I'm a contributor :)

Haya Innovator's profile picture
Haya Innovator1 month ago

Exactly the kind of routing layer enterprise agents need. I separate complex reasoning from high volume execution across my digital health stacks. Nemotron Lightning is ideal for those inference steps.

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB1 month ago

按任务难度分配模型,确实比全程堆大模型省得多。

Pixel's profile picture
Pixel1 month ago

routing is the missing layer. frontier for thinking, Lightning for volume. this is how agent stacks should work

Related Videos