Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I STOPPED LETTING CLAUDE MAKE DECISIONS THE DAY I BUILT THIS JEV AGENT FOLDER I used to let Claude decide, write and act on every single step -> now Claude only writes. Jev makes the calls, code does the acting, and every step leaves a receipt here's what's inside...

78,918 görüntüleme • 7 gün önce •via X (Twitter)

19 Yorum

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack6 gün önce

@polydao i hit this wall too. once i started doing weekly reviews, my workflow tightened up a lot.

Ivana profil fotoğrafı
Ivana7 gün önce

This is the first control layer I’ve seen that actually logs every decision with receipts. 10k decisions for $0.42 is the number that sticks.

Nyrqavel profil fotoğrafı
Nyrqavel6 gün önce

This is the cleanest agent split I’ve seen: LLM only writes, Jev decides, code acts. Receipts on every call + 10k decisions for $0.42 is the part that actually matters.

Horoshi⚡ profil fotoğrafı
Horoshi⚡6 gün önce

What's the point of all this?

magsimich profil fotoğrafı
magsimich7 gün önce

That folder makes the workflow much cleaner

riVeN profil fotoğrafı
riVeN6 gün önce

right split. one gap: if misses a delete, hard_rules never sees it. gate the send/pay/delete tools themselves.

Taqi T| Tech consultant & Sr.Engineer profil fotoğrafı
Taqi T| Tech consultant & Sr.Engineer7 gün önce

Again job mate! Let's connect

sofie unbothered profil fotoğrafı
sofie unbothered7 gün önce

Swapping Claude's judgment for a 100ms gate might miss context. How do you handle ambiguous cases where "wake? safe? good? done?" isn't enough?

Ham profil fotoğrafı
Ham6 gün önce

Thanks for sharing, much appreciated

Archive profil fotoğrafı
Archive6 gün önce

i think i should copy what you did i wouldn't have even thought of this

Iron Mind profil fotoğrafı
Iron Mind7 gün önce

jev is better than claude?

Brian D. Evans profil fotoğrafı
Brian D. Evans6 gün önce

Jev is very interesting

Miano profil fotoğrafı
Miano6 gün önce

The clean split is the important part. Let the model handle judgment, let code enforce the decision, and keep the expensive generation step out of the loop.

Paul profil fotoğrafı
Paul7 gün önce

Claude writes, Jev decides, code acts. Saw this in prod: the receipts matter more than the split. My rule: a PostToolUse hook appends {step, input_hash, decided_by, exit_code} to a JSONL file Claude can't write to. If the model can edit the log, it's not a receipt.

ShadowAguy profil fotoğrafı
ShadowAguy6 gün önce

The receipt layer is the interesting bit. Are those logs append-only, or can Jev rewrite history when a step fails?

Pankaj Kumar profil fotoğrafı
Pankaj Kumar6 gün önce

The done check is the part most setups skip. Models agree with whoever spoke last, so asking the same model 'are we done?' gets a yes far too often. A written check it can't argue with is what actually ends the loop. The LLM should never grade its own homework.

ridhafara profil fotoğrafı
ridhafara7 gün önce

How are the decision thresholds updated when new evaluation traces are added to the folder?

Sael profil fotoğrafı
Sael7 gün önce

treating llms like a middle manager who overthinks every single move is the only way to scale this stuff without going broke

Crio Songo profil fotoğrafı
Crio Songo6 gün önce

This decoupling architecture is elegant. Each part takes its own responsibility with clear audit trails, solves the old problem of LLM decision-making being untraceable.

Benzer Videolar

Jev has been exploding in popularity recently. If you already have access to the Jev API but aren't sure how to start experimenting with it, just copy this checklist: 1. jev-ultrafast Browser Use's fastest agent. Jev decides the next action and which element to click, and a language model is only called when text has to be typed. 2. typesafe-mario Jev plays Super Mario Bros. from structured emulator state, choosing every action from features pulled out of the game. 3. jev-plays-pokemon Reads Pokémon Red's game state as text, answers typed questions each turn, and lets plain code turn the answers into moves. 4. jev-drone A camera-only autonomous drone in MuJoCo with a Jev judgment model sitting in the control loop at 2.5 Hz. 5. robo-harness A real SO-101 robot arm workbench where Jev picks bounded joint steps from typed candidate actions under a spend budget. 6. fast-jev-compaction Claude Code plugin that replaces the compaction summary with Jev decisions, scoring every tool call for whether it is still needed. 7. jev-claude Routes Claude Code's own judgment calls through Jev: typed choices with probabilities at plan approval, on questions, and before risky commands. 8. is-malicious Supply-chain check before you run anything: Jev Noul checks over source and build files, returning the implicated files and lines. 9. sqlite-jev Jev inside SQL. Noul, Choice and Score judgments exposed as SQLite functions, with confidence on every row. 10. jevinci Paints images by having Jev predict every pixel's colour in parallel, with confidence deciding how wide each stroke is drawn. Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓

Hanako

29,629 görüntüleme • 9 gün önce

Claude Code tip: once Opus 5.5 runs your main session, stop leaving Fable 5.1 on the bench put it on call with /advisor run /advisor fable Opus 5.5 keeps writing the code Fable 5.1 reads the whole session, every tool call included, and only speaks up at three moments: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Fable 5.1 reviews. Opus 5.5 ships jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to jev in under half a second, and the big model only sees the ones that split - the full tree > Opus 5.5 on high runs the main session > explorer reads the code > worker edits and runs tests > researcher pulls the docs > all three on medium > Fable 5.1 on call as the advisor paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: opus, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, any variable that stops feature-flag fetching) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

Hanako

51,329 görüntüleme • 3 gün önce

this is f*cking gold 20 GitHub repos with 500K+ combined stars that will level up your JEV workflow AGENTS > jev-ultrafast: Jev picks every click and DOM target, a small LLM only types > hermes-jev-skills: routing, memory, compaction and skill picks in one pack > typesafe-computer-use: OCR reads your Mac screen, Jev picks the next click MEMORY > fast-jev-compaction: scores every tool call keep, truncate or drop instead of summarizing > jevmem: project memory for Claude Code, Cursor and Codex, updated every turn > jev-second-brain: your Obsidian vault, Jev judges which notes duplicate, revise or contradict SAFETY > jev-guard: a gate before every tool call, Jev scores the risk, you set allow, ask or deny TOOLS > skills: the official TypeSafe skill for Claude Code and Codex > system-one-adapter-python: dry-run your questions on an ordinary LLM before you burn a Jev key > jev-mcp: claim checks, screening and ranking as MCP tools > typesafe-mcp: plug Jev into any MCP client > json-render: Vercel's generative UI, where Jev picks the components OPEN MODELS > SemIf-OpenJev: semantic ifs from frozen open models kev: Jev-like models on Qwen that run on your MacBook > laya-mlx: the Laya decision engine on Apple silicon > clm: an open System One model with Choice, Noul and Score > jevlike: train your own Jev-like model TRADING > jev-trader: one buy or sell decision per Monad block START HERE > awesome-jev: the biggest map of everything built on Jev > awesome-jev-by-typesafe: use cases, patterns and starter code bookmark it before your next build

NO1ennn

18,445 görüntüleme • 5 gün önce

Claude Code tip: once Opus 5.5 is your main model, stop leaving Fable 5.1 sitting idle put it on call with /advisor run /advisor fable Opus 5.5 keeps writing the code Fable 5.1 reads the full session, every tool call included, and only speaks up at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Fable 5.1 reviews. Opus 5.5 ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big model only sees the ones that split - the full tree > Opus 5.5 on high runs the main session > explorer reads the code > worker edits and runs tests > researcher pulls the docs > all three on medium > Fable 5.1 on call as the advisor paste the tree and this prompt into Claude Code ↓ "Rebuild my Claude Code setup around this tree: 1. Check ~/.claude/agents and .claude/agents for subagents that already fit explorer, worker and researcher. > Draft new ones only for missing roles > Give each model: opus, effort: medium > Skip any that pin a different model and list them 2. Set the main session to high via effortLevel in ~/.claude/settings.json, and set advisorModel to fable 3. Find anything that keeps the advisor off (CLAUDE_CODE_DISABLE_ADVISOR_TOOL, DISABLE_TELEMETRY, any variable that stops feature-flag fetching) plus CLAUDE_CODE_EFFORT_LEVEL, which overrides subagent effort. Report them, change nothing 4. Add one rule to ~/.claude/CLAUDE.md: consult the advisor before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

692,413 görüntüleme • 5 gün önce

how to use claude code mods like a top 1% user, step by step: give this to your agent before everyone catches on👇 1. set up Jev connect Jev to the model registry you want to use. check that the connection works and the listed models are available. 2. build your claude code mod open claude code 2.1.287 or later and paste this prompt: “build a mod called run-ledger. load plugin-authoring and use the API types for my installed version. create a dashboard that shows: - estimated cost per run, model, and source plugin where known - which model handles each task - each subagent’s status, latest action, and elapsed time include token counts, cache usage, and reported retries. count each request once. keep background tasks linked to their original run. use dated prices. label costs as API estimates, not subscription charges. show unknown when data is missing. connect the mod to my existing Jev setup. give Jev the task, available models, prices, budget, and relevant past results. ask it to recommend a model and explain why. start with recommendations. make automatic routing optional for eligible subagents. show the recommended model, actual model, result, and cost. include Jev’s own cost. keep state across hot reloads. add details and export. the dashboard itself must make no model calls. validate the plugin. test rendering, costs, attribution, and duplicate counting. give me steps for a live test.” 3. test the setup allow hot reload when prompted. run a simple task, a subagent task, and a mod-triggered model call. check that the dashboard updates, each request counts once, and Jev’s recommended model can actually run the task. 4. install the working mod ask claude to copy it out of the temporary folder and install it as a persistent plugin. see the cost. choose the model. check the result.

Avid

45,337 görüntüleme • 1 gün önce

Another insane Jev use case! Jev is making it dramatically cheaper to evaluate what actually happened inside an agent run. And finally, someone open-sourced a self-improving memory layer that can put that signal to work across agent harnesses: - Claude Code - Codex - Cursor - OpenCode, and 20+ more Beacon by Asymptote Labs continuously captures your agent history across harnesses and uses Jev to identify which runs are actually worth learning from. It then turns the highest-signal workflows, corrections, and debugging patterns into reusable skills. GitHub repo: (don’t forget to star it ⭐ ) Beacon preserves the complete session history. But preserving a run and learning from it are two different things. Most coding-agent sessions contain routine exploration, failed commands, and fixes that only apply to one task. The trace can remain available for inspection without turning every detail into guidance for future agents. Jev scores each run for evidence, reuse potential, and human correction signals. An application policy then decides whether to promote, review, or discard it. The recording shows this in action. Claude receives a coding task, modifies the implementation, and runs the tests. I then provide an edge-case correction, so Claude updates the code and adds regression coverage. Beacon automatically captures the complete session. Jev evaluates whether the correction contains a reusable engineering lesson. Once approved, that lesson becomes available to other coding agents working on the project. Since it works across harnesses: - Claude Code sessions can teach Codex. - Cursor debugging can improve OpenCode. So a problem solved by one agent should not need to be learned from scratch by another. If you want to dive deeper into Jev, I also wrote a hands-on guide to building this Jev-style decision path with open models, entirely locally. Read it below.

Avi Chawla

291,153 görüntüleme • 13 gün önce

Jev is cool. So is it's OSS companion, Laya. The Latest Cool Thing In AI™ tends to get a lot of hype, sometimes without everyone even understanding it. So... what is this thing? Jev is an AI model that consumes input and produces output VERY differently than chat, claude, grok. The input is two things: 1) Text state to assess. Email, html, code, whatever. 2) A set of questions which will be asked about the attached state. The canonical example from TypeSafe's docs is to identify the urgency of a support ticket. We pass the model the customer text + a single noul question "is this urgent?". Jev returns a full set of JSON. This JSON is not generated with token-by-token autoregression. Jev is not trained to produce sequences of text tokens, rather to answer questions, and guarantees well-formed responses. In the example below, we see it produces a 0.99 probability (on a 0-1.0 scale) that the answer is "yes." Jev supports exactly three types of questions (seconds example in video): a) Noul: 0–1 probability that the answer to a yes/no question is "yes." b) Choice: Ask question with pre-defined set of answers. Jev chooses the best and assigns probabilities to each. c) Score: Ask question with pre-defined scale of answers. Jev produces a position on the scale. Jev computes answers for all questions in parallel, making responses super fast even for many questions in a single request. This might seem like a narrow set of capabilities, but in the right contexts leads to incredible potential. It also makes for a useful API / primitive for programming, since the outputs are... *ahem*... type-safe and predictable in structure. Jev is not going to replace LLMs for writing your code, auto-generating your docs, or being at the core of an agent harness. But Jev IS incredibly cool, and will be used to build a lot of amazing tech. Hope this helps.

Ben Dicken

40,810 görüntüleme • 14 gün önce

i finally mastered how to maximise my opus 5.5 usage limits... the trick: let jev choose which subagent gets each task and how much effort it should use. [here’s how i’d wire it:] claude breaks the project into tasks. jev selects from predefined worker profiles. claude applies the selected settings and dispatches the work. → main session, medium: clarify the requirements, define what “done” looks like, and prepare the tasks → builder, low: small, clearly defined tasks with existing examples or patterns → builder, medium: tasks that connect multiple parts or need decisions within the approved plan → verifier, high: check requirements, probe edge cases, and report problems for the builder to fix jev gets the task’s scope, what’s uncertain, and the consequences of failure. it chooses from the profiles allowed for that task. your approval checkpoints stay in place. paste this into your next planning session: “use opus 5.5 with jev selecting the worker and effort profile for each task. first, check that a working jev integration is available and that this environment supports separate effort settings for subagents. check for configuration or environment overrides that could prevent those settings from taking effect. if anything is missing, explain what needs wiring before proceeding. break my request into tasks with clear ownership, dependencies, relevant context, and acceptance checks. keep small related tasks together when a separate subagent would add unnecessary overhead. keep the main session at medium effort. offer jev these worker profiles: builder at low effort for small, clearly defined tasks using existing patterns; builder at medium effort for tasks that connect multiple parts or require decisions within the approved plan; verifier at high effort for checking requirements and edge cases. give jev each task’s scope, uncertainties, dependencies, and consequences of failure. only offer profiles appropriate to the current stage. validate its selection before dispatching. use the actual jev integration; don’t simulate its decisions. if it abstains or returns an invalid choice, stop that handoff and ask me. show me the task plan and proposed assignments before starting. after approval, launch the selected workers with their assigned effort settings, relevant context, file ownership, and completion checks. let me review the result before verification. the verifier may add tests but must leave implementation code unchanged. have it report what passed, what failed, and what remains uncertain. send implementation fixes back to the builder, then recheck the affected parts. if a task repeatedly fails, examine the requirements and approach before increasing effort. report available total usage, including jev calls, worker calls, retries, and verification. don’t invent missing data. compare similar completed tasks before claiming savings.” steal this 👇

Avid

31,355 görüntüleme • 7 gün önce

LLMs vs. Jev, clearly explained! LLMs are great, and the ceiling is one you can watch scroll past: an LLM writes the answer one token at a time. give it a failed deploy and four decisions, and it produces a small JSON object where every token depends on the one before it. token nine cannot exist until token eight does, so four decisions that had nothing to do with each other just stood in a queue. then your code parses it, validates the shape, and retries when the shape is wrong. Jev fixes this without being a smaller or faster model: it removes the order. one turn on that deploy has to know: → whether the incident is urgent → which team owns it → whether the next command is risky → whether the task is actually done you declare the questions and the answer type upfront, and all four come back together, typed, with a probability on each. three primitives cover almost every fork in an agent: 1. **Choice** picks one of up to 255 options you define, like engineering, billing or sales. 2. **Score** places the state on an ordered scale you define, like low, medium or high risk. 3. **Noul** returns the probability that a yes-or-no condition is true. here is the sentence that resolves the whole confusion: text is a line you have to walk. an answer space is a room you see all of at once. ↳ generation: one order you cannot change, one string at the end, a shape you hope holds ↳ evaluation: no order at all, typed answers, a probability on every option Prompts → Agents → Loops → Graphs → Jev the probabilities matter more than the answer. ↳ engineering at 0.91 against billing at 0.09 is a route you can automate ↳ 0.52 against 0.46 is a coin flip wearing a label, and the label alone never told you which one you got that last one catches careful people. an LLM would have said "engineering" in a confident sentence and given you no way to know the race was that close. thresholds live in your code, one per action, scaled to what being wrong costs. it works when the options are known and the call depends on meaning. it is not for writing, code, arithmetic, or anything where question two needs the answer to question one. and the one that eats whole nights: type safety prevents malformed output, not incorrect judgment. Jev cannot return an option outside your schema, and it can still pick the wrong valid one with confidence. a schema-valid mistake refunds the wrong customer just as fast. an LLM writes new language when the answer space is open. Jev evaluates known paths when the answer space is closed. below i have quoted my full breakdown on Jev. it covers the three primitives, the parallel battery, the thresholds, and where it does not belong. save this and read it below ↓

Hanako

42,632 görüntüleme • 13 gün önce