Loading video...
Video Failed to Load
Jev Engineering is what turns an agent stack into an actual control system and moves the expensive model out of every decision loop. and up to 193x faster and 444x cheaper in tests. the model shouldn’t decide everything. in this setup: request → structured state → Jev router →... show more
175,761 views • 4 days ago •via X (Twitter)
23 Comments

It looks wild Ricker

thanks Morty

This vis perfectly explains Jev Harness

exactly mate

decision layers that don't generate prose is a clean way to frame it

fact Yarchi

wow, that's wild

insane

This architectural shift is huge. Request -> State -> Jev fast decision -> Target Tool means 90% of router ops happen in under 50ms before touching a heavy model. Pure system design elegance.

Route → score → block → approve. That’s the sentence that matters. Prose models shouldn’t sit in every loop. A cheap decider plus a gate is a control system. Still needs a room the worker comes home to — keys, session, human on the last write. Jev Engineering is the reflex. @AgentOS_Tech is the desk around it.

Great one, mate.

Routing system state transitions with LLM prose is paying a high premium for latency and hallucinated control flow.

same. the boring part is what actually ships

очередное дерьмо

The interesting shift is treating the agent stack like a control system instead of making the model responsible for every decision. Cheap, deterministic gates can handle the routine work and reserve expensive reasoning for where it actually adds value.

193x faster and 444x cheaper is wild, and the fact that you measured it makes it even better. What was the baseline: a frontier model in every loop? And did accuracy on the decision layer hold up?

I loved your article btw - I'm going to read through it again but you were def the first (on my feed at least) to share anything about this. Really Solid.

request → state → router → cheap model → gate → tool is the control loop agents have been missing. One caveat: a cheap layer fails more quietly. Without a confidence threshold and escalation to a human or a frontier model, you just take the wrong step faster.

193x faster is seriously crazy

The architecture is right. Worth naming where 193x comes from though, that is TypeSafe running their own workflow eval. The independent number I have seen is nearer 25x across 777 judgments, still enough to change how I build. Quote the one that survives your own data.

Reminds me of how a pattern buffer from a Star Trek teleporter might operate. 🤪

removing LLMs from every loop is literally the only way agents scale

seems like the new paradigm
