Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

The right model depends on the task. NVIDIA NeMo Switchyard helps developers route each agent workflow step across a chosen model pool based on their own quality, latency and cost criteria. Kari Briski joins MTS to explain why agent workflows need model routing.

97,218 Aufrufe • vor 1 Monat •via X (Twitter)

33 Kommentare

Profilbild von NVIDIA
NVIDIAvor 1 Monat

Learn more:

Profilbild von Brent Liang
Brent Liangvor 1 Monat

@MTSlive so awesome to have kari on the show!

Profilbild von Denys Khomyn
Denys Khomynvor 1 Monat

@ChrisJBakke @MTSlive NVIDIA monitoring the situation 🤘

Profilbild von DataVLT
DataVLTvor 1 Monat

@MTSlive The smarter the routing, the more important the data behind each decision becomes. Model choice is only one part of the stack.

Profilbild von PixelSilicon
PixelSiliconvor 1 Monat

@MTSlive Model routing by task is becoming essential once agent workflows get multi step. One model for everything is rarely optimal on cost or latency.

Profilbild von KVCacheStore LLC
KVCacheStore LLCvor 1 Monat

@MTSlive I'm a contributor to Switchyard. If anyone wants help setting it up let me know.

Profilbild von Joel Ramirez
Joel Ramirezvor 1 Monat

@MTSlive ⚓️✅❌🚀🎇

Profilbild von Glitch Truth
Glitch Truthvor 1 Monat

@MTSlive Doesn't matter which model wins the routing fight, every single path still ends at an Nvidia GPU.

Profilbild von OGNVDASOL
OGNVDASOLvor 1 Monat

@MTSlive @elonmusk

Profilbild von Zee B-side
Zee B-sidevor 1 Monat

@MTSlive oxi, now proibida no mundo todo

Profilbild von Ce Ce | AI & Systems Infra
Ce Ce | AI & Systems Infravor 1 Monat

@MTSlive model routing based on latency + cost is the part most teams still underestimate

Profilbild von cameron_h23
cameron_h23vor 1 Monat

@MTSlive Model routing is becoming core infrastructure for agent systems—use the expensive reasoning model where it matters, optimize everything else around it.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

@MTSlive Smart routing saves real money

Profilbild von Robert Hero
Robert Herovor 25 Tagen

@MTSlive The right model depends on the task.

Profilbild von Marissa
Marissavor 1 Monat

@MTSlive 👍

Profilbild von J. LEE
J. LEEvor 1 Monat

@MTSlive Love how NeMo Switchyard lets us pick the perfect model for each step—smart, flexible, and cost‑savvy!

Profilbild von MoM
MoMvor 1 Monat

@MTSlive

Profilbild von Abhi
Abhivor 1 Monat

@MTSlive Routing only works after teams save representative workflows and set a quality floor.

Profilbild von Christopher Dean
Christopher Deanvor 1 Monat

@MTSlive Per-step routing is right, and it only works if you have per-step evals: without them you're routing on vibes and the cheap model quietly tanks one node of the workflow.

Profilbild von nock
nockvor 1 Monat

@MTSlive rubin vs tpu pileup this month is the whole story. everyone shipping silicon at once.

Profilbild von Mitchell𖤐⚔️🔺
Mitchell𖤐⚔️🔺vor 1 Monat

@MTSlive ♻️🔰♻️🔰♻️🔰♻️🔰♻️🔰♻️

Profilbild von Cody Gore
Cody Gorevor 1 Monat

@MTSlive

Profilbild von Zentiva AI
Zentiva AIvor 1 Monat

@MTSlive This is a key piece of the agentic AI puzzle. Model routing makes sense when different workflow steps have different requirements. The real challenge is deciding dynamically when quality matters more than latency or cost. That’s where intelligent orchestration becomes critical.

Profilbild von Raul Verdusco
Raul Verduscovor 1 Monat

@MTSlive Smart model routing can make agent workflows far more efficient. Matching each task to the right model helps balance quality, latency, and cost without relying on one model for everything.

Profilbild von seastart
seastartvor 1 Monat

@MTSlive Model routing gets really interesting once you optimize for the whole workflow, not each model in isolation. The best model for one step can be the wrong one for the next.

Profilbild von Greg B
Greg Bvor 1 Monat

@MTSlive Not one person buying this stock. Absolutely makes zero sense why this company is not doubling its stock buyback. Executive leadership screwed investors. This company has an image problem. @JensenHuang

Profilbild von A P E | T H E | N E W S
A P E | T H E | N E W Svor 1 Monat

@MTSlive Routing is becoming part of model performance, not just infra plumbing. A cheaper model on the right step beats pushing a frontier model through every step blindly.

Profilbild von why
whyvor 1 Monat

@MTSlive Switchyard tackles model routing. Good for balancing quality, latency, cost.

Profilbild von 海
海vor 1 Monat

@MTSlive “NVIDIA is amazing. I just hope it keeps pushing a little harder.”

Profilbild von Raghavan
Raghavanvor 1 Monat

@MTSlive Yes sir nvidia has great architecture

Profilbild von Вестник
Вестникvor 1 Monat

@MTSlive Hello @nvidia - why are you helping Russia kill Ukrainians? Do you support mass killing of civilians? In Russian attack drones, Nvidia chips for fully autonomous targeting were found. It was precisely such a drone that killed three civilians in Zaporizhzhia on July 6.

Profilbild von RH Fardin
RH Fardinvor 1 Monat

@MTSlive Quality, latency, and cost routing per workflow step is the right architecture. GPT should get the hard turns.

Profilbild von mia ♡
mia ♡vor 1 Monat

@MTSlive does routing happen per step after failure, or only from the initial plan?

Ähnliche Videos

how to use claude code mods like a top 1% user, step by step: give this to your agent before everyone catches on👇 1. set up Jev connect Jev to the model registry you want to use. check that the connection works and the listed models are available. 2. build your claude code mod open claude code 2.1.287 or later and paste this prompt: “build a mod called run-ledger. load plugin-authoring and use the API types for my installed version. create a dashboard that shows: - estimated cost per run, model, and source plugin where known - which model handles each task - each subagent’s status, latest action, and elapsed time include token counts, cache usage, and reported retries. count each request once. keep background tasks linked to their original run. use dated prices. label costs as API estimates, not subscription charges. show unknown when data is missing. connect the mod to my existing Jev setup. give Jev the task, available models, prices, budget, and relevant past results. ask it to recommend a model and explain why. start with recommendations. make automatic routing optional for eligible subagents. show the recommended model, actual model, result, and cost. include Jev’s own cost. keep state across hot reloads. add details and export. the dashboard itself must make no model calls. validate the plugin. test rendering, costs, attribution, and duplicate counting. give me steps for a live test.” 3. test the setup allow hot reload when prompted. run a simple task, a subagent task, and a mod-triggered model call. check that the dashboard updates, each request counts once, and Jev’s recommended model can actually run the task. 4. install the working mod ask claude to copy it out of the temporary folder and install it as a persistent plugin. see the cost. choose the model. check the result.

Avid

47,276 Aufrufe • vor 6 Tagen