正在加载视频...
视频加载失败
Introducing typesafe/jev-router: a cache-aware model router powered by Jev and TypeSafe AI The Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. Here's how it works 👇🏻
1,297,534 次观看 • 3 天前 •via X (Twitter)
18 条评论

On four agent benchmarks, Jev Router solved 82% more tasks than our Auto Router (237 vs 130 of 423).

On five agent benchmarks, Jev Router had a faster median time to first token than every other router we tested.

Most routers pick a model based on each message. Each switch loses the cached chat, so you pay full price for the new model to reread the whole conversation. Most routers also route on the task type, not the difficulty. Easy and hard coding tasks get the same model.

Jev Router runs on Jev, TypeSafe's first Decision / System One model. Before each turn, Jev reads your prompt and scores it on difficulty and precision. Also checks whether a bigger model or more effort would help, whether a cheaper model is enough, and whether the task changed.

Jev reads the conversation text only to pick the model and effort. It runs under zero data retention (ZDR) terms, so nothing is stored or trained on. Attachments are never sent to Jev, and requests with "zdr: true" work with Jev Router.

Jev Router keeps a model that works for the rest of the session. It can raise or lower effort without switching models. It switches only when the expected gain is larger than the cost, including the cached chat it would lose.

If the Jev call times out or returns invalid output, the request fails instead of falling back to another router. Each response includes routing metadata with the reason for the choice. Try it now using "typesafe/jev-router" in your app, or, see Jev Router's thinking.

You can see Jev Router's thinking process in OpenRouter Chat. The routing insights panel shows the model it picked for each turn and the scores behind it, like task, difficulty, precision, and larger model benefit. Test Jev Router here:

Read more:

@typesafeai is there a directive to limit the worst model it can route to? "Hey Jev don't ever route it to [this dogshit model]"

@typesafeai Are we able to specifically a list of models to choose from?

@typesafeai cache-aware routing should make repeated prompts cheaper without sacrificing quality

@typesafeai the fail-closed call is right, but make the router timeout its own error class: routing unavailable, not a task failure. otherwise every harness counts it against the agent and retries the work when the only thing that needs retrying is the routing decision.

@typesafeai cache-aware routing is just optimizing for what makes them money not what makes your product better

@typesafeai they finally shipped a router that stops my wallet from crying

@typesafeai Önbellek farkındalığı, yönlendirmeyi basit bir maliyet oyunundan çıkarıp gerçek bir mühendislik problemine dönüştürüyor.

@typesafeai And if you use Paseo, I created a TypeSafe's Jev (it also support local Laya as a classifier) model router for free:

This kind of tools will be a game changer once we figure out how to share cache between different models. I currently run something like this for my harness but it's only a single call at the beginning of every conversation. It's so unbeneficial that each call _might_ use the cache

