Video yükleniyor...
Video Yüklenemedi
I STOPPED LETTING CLAUDE MAKE DECISIONS THE DAY I BUILT THIS JEV AGENT FOLDER I used to let Claude decide, write and act on every single step -> now Claude only writes. Jev makes the calls, code does the acting, and every step leaves a receipt here's what's inside... show more
78,918 görüntüleme • 7 gün önce •via X (Twitter)
19 Yorum

@polydao i hit this wall too. once i started doing weekly reviews, my workflow tightened up a lot.

This is the first control layer I’ve seen that actually logs every decision with receipts. 10k decisions for $0.42 is the number that sticks.

This is the cleanest agent split I’ve seen: LLM only writes, Jev decides, code acts. Receipts on every call + 10k decisions for $0.42 is the part that actually matters.

What's the point of all this?

That folder makes the workflow much cleaner

right split. one gap: if misses a delete, hard_rules never sees it. gate the send/pay/delete tools themselves.

Again job mate! Let's connect

Swapping Claude's judgment for a 100ms gate might miss context. How do you handle ambiguous cases where "wake? safe? good? done?" isn't enough?

Thanks for sharing, much appreciated

i think i should copy what you did i wouldn't have even thought of this

jev is better than claude?

Jev is very interesting

The clean split is the important part. Let the model handle judgment, let code enforce the decision, and keep the expensive generation step out of the loop.

Claude writes, Jev decides, code acts. Saw this in prod: the receipts matter more than the split. My rule: a PostToolUse hook appends {step, input_hash, decided_by, exit_code} to a JSONL file Claude can't write to. If the model can edit the log, it's not a receipt.

The receipt layer is the interesting bit. Are those logs append-only, or can Jev rewrite history when a step fails?

The done check is the part most setups skip. Models agree with whoever spoke last, so asking the same model 'are we done?' gets a yes far too often. A written check it can't argue with is what actually ends the loop. The LLM should never grade its own homework.

How are the decision thresholds updated when new evaluation traces are added to the folder?

treating llms like a middle manager who overthinks every single move is the only way to scale this stuff without going broke

This decoupling architecture is elegant. Each part takes its own responsibility with clear audit trails, solves the old problem of LLM decision-making being untraceable.
