Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

JEV ENGINEERING JUST CUT A CLAUDE LOOP FROM $765/MONTH TO $3.02 the expensive part isn’t always writing code. a busy overnight loop can burn ~600 decisions just asking what to do next, whether a command is safe, or if the task is finished. trigger → Jev → allow /...

15,718 Aufrufe • vor 1 Tag •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Jev + SERV is actually insane. We already showed you can increase Jev's performance with SERV Reasoning. Now we're taking it further, bringing Jev-powered Decision nodes into Graph Sharding with the upcoming SERV v3. Here's a breakdown of how it works: Jev is a decision-making model. Given a task and a set of options, it predicts which path is more likely. Think of the octopus that predicted World Cup results. Jev does that for your business, except it's not luck. It weighs every option and tells you how sure it is. It does this by assigning probabilities to outcomes. It doesn't generate text on its own, so you can't expect it to create a new outcome for you. But that's also what enables it to be lightning fast and dirt cheap. For example, in customer service you can ask Jev how to triage an incoming query and route it to the correct department. It can only select from the list of departments you provide it. This also means it can't hallucinate a new outcome outside the options it's given, which makes it incredibly interesting for OpenServ. In Graph Sharding, we take a single system prompt and break it down into multiple LLM steps with deterministic input and output shapes. Some of these steps require an LLM to produce new output, while others are simply decision routers that determine the next possible path. Traditionally, LLMs are slow and expensive. Breaking a single prompt into multiple steps increases accuracy and reliability by a ton, but it also introduces latency. Jev takes on those decision nodes, which are the backbone of a business process and therefore SERV graphs, and makes them super consistent and lightning fast, lowering the overall cost and latency of graph execution. SERV Reasoning on its own is a great force multiplier for Jev because, like all other models, it works by interpreting input instructions. The clearer those instructions are, the better the model performs. That's where SERV Reasoning comes into play. Just like amplifying any other model, we also amplify the accuracy and consistency of Jev's responses. And now we're bringing Jev-powered Decision nodes into Graph Sharding with SERV v3.

Armagan Amcalar

365,108 Aufrufe • vor 5 Tagen

Top 12 agentic use cases for Jev: (bookmark this) Jev handles semantic decisions that ordinary code cannot express reliably. It returns typed answers and probabilities, while code continues to cover the workflow. Here are 12 practical use cases for Jev: 1. Browser next action > Convert the current DOM state into a bounded action such as click, type, or stop. Code executes only valid operation-target pairs. There are already several open-source Jev web agents. 2. Context compaction > Decide which events from a long agent trace should remain. The selected text stays verbatim instead of being replaced with a generated summary. 3. Skill and context loading > Compare the current user turn against the available skills. Load only the instructions needed for that turn instead of filling the context window with every skill. 4. Typed tool-call compilation > Map a natural-language request to a function and fill its typed arguments. Each argument is evaluated separately before code allows execution. 5. Citation verification > Check whether a quoted passage exists and whether the surrounding evidence supports the claim. The output can be supported, unsupported, or contradicted. 6. Extraction verification > Run a cheap extractor first, then use Jev to verify questionable fields. Clean records stay on the fast path while uncertain ones reach a reasoning model. 7. Agent trace evaluation > Turn raw trajectories into queryable labels such as progress and repetition. This avoids asking another LLM to write a full review of every run. 8. Semantic regression tests > Replay a trace suite against a new agent build. Semantic checks can then pass or block prompt, model, tool, and policy changes in CI. 9. Jevgrep code search > Search a codebase by what the code does rather than its exact words. Jev scores candidate snippets and returns the most relevant code first. 10. Entity alignment > Compare two candidate records and decide whether to merge, review, or keep them separate. Candidate generation remains deterministic while Jev handles semantic identity. 11. Retrieval reranking > Let embeddings retrieve a broad candidate set, then use Jev to reorder passages by relevance. The generation model receives the most useful evidence first. 12. Memory promotion gate > Capture a completed agent trace, then judge whether its corrections contain a reusable lesson. Trace-backed lessons can be promoted while task-specific noise is discarded. If you want to see the final pattern in practice, it is already implemented in the Beacon open-source project. Beacon captures full sessions across Claude Code, Codex, Cursor, OpenCode, and 20+ agent harnesses, and then Jev identifies which workflows and corrections are worth learning from, so that a lesson discovered by one agent can become available to the others. GitHub repo: (don’t forget to star it ⭐) If you want to dive deeper, I also wrote about a similar mechanism in a hands-on guide. It covers building a Jev-style decision path with open models, entirely locally. Read it below.

Avi Chawla

97,392 Aufrufe • vor 1 Tag

Jev + Muse is the first AI agent system that actually automate 100% of my life 99% of people pay 200x more for slower AI agents - while 1% run this 2030 setup just 5 min and setup is ready: prompt → Muse → Jev decision → Muse execution → result step 1 → create your Jev API key (typesafe website) step 2 → clone and install the complete router from Github below python3 -m venv .venv && .venv/bin/pip install -r requirements.txt && cp config.example.yaml config.yaml step 3 → export the key before running anything: export TYPESAFE_API_KEY='YOUR_KEY' add the same export to ~/.zshrc or ~/.bashrc if you want it to survive a new terminal session step 4 → give your agent skill/jev-decision-layer.SKILL.md and connect it to src/router.py + recipes/ , raw Jev returns probabilities - the router converts them into executable actions step 5 → test the entire chain, not the raw Jev API: .venv/bin/python -m src.cli '{"goal":"what is 2+2?","kind":"chat"}' the final JSON should contain action, reason, mode, jev_used and confidence details step 6 → keep mode: shadow for 20–50 real decisions: the agent works normally while Jev’s routes are logged and checked; promote only reliable question packs step 7 → switch to mode: active with hard confidence gates: ≥0.80 act automatically, 0.50–0.79 advisory only, <0.50 escalate to the human the result: Jev + Muse is a system that decides what to do, what to skip and when to bring in - I’ve tested it across my daily workflows, and it’s the best setup I’ve found for automating routine Take the exact stack I built, run it yourself from the repo - then read the full Jev architecture behind it ↓

codila

91,306 Aufrufe • vor 6 Tagen

50% cheaper Claude inference with just one line of code change! - Remove → model="claude-opus-4-8" - Add → model="ship-like/claude-opus-4-8" I verified the cost saving in my own terminal by invoking the same Anthropic model with the same prompt. The underlying engineering by Ship is actually interesting, and the patterns can be used in any production LLM stack. Essentially, a trained model is a frozen artifact. Every request performs the same forward-pass, whether it extracts a date or refactors a module, because the compute decision was made at training time, before the request existed. Ship makes that decision at inference time instead. After seeing a request, it searches over executions, involving single models, cascades, ensembles, or harnesses with tools, and serves the cheapest one that will match the reference model's quality. This is not a basic router, because picking a cheaper model per query doesn't ensure the cheaper model preserves the original's behavior, like output shape, tool-call patterns, and refusals. Ship measures this equivalence directly. Outputs stay distributionally indistinguishable from the reference model, not token-identical, since two calls to the same model already differ, but they are indistinguishable in capability and behavior. Of course, some requests execute cheaply and some cost Ship more than the customer pays, but the price per request is still a flat 50% off either way, so the execution-cost variance moves off the application's bill entirely. The video below depicts the cost savings and output in my real invocation, and I partnered with the team to put this together.

Akshay 🚀

64,482 Aufrufe • vor 2 Monaten

New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.

elvis

11,303 Aufrufe • vor 1 Monat

40 hours of human work. That’s what this humanoid could save every month! A construction company is already testing a Unitree G1 on a real job site, using the robot for site inspections, 360° imaging, data collection and progress monitoring. The robot starts at around $13,500, while the company says its deployment can save roughly 40 hours of work every month. That adds up to around 480 hours a year from a machine that costs a fraction of traditional industrial equipment. The interesting part is that this isn't about replacing an entire construction worker. It's about removing hundreds of repetitive hours from the workflow, including walking inspection routes, documenting progress and collecting information across the site. Humans can then spend their time on decisions and tasks that actually require them. This is where humanoid robots become economically interesting. Construction sites are already designed around human movement, so a robot with two arms, two legs and a human-sized body can potentially work in the same spaces without rebuilding the entire environment. Every additional task it learns turns those same hardware costs into more productive hours. And the economics get even more interesting as prices fall and production scales. A robot that saves 480 hours per year doesn't need to be perfect or replace a full-time employee to justify its existence. It just needs to reliably take over the boring, repetitive work that companies are already paying humans to do. 480 hours saved. Thousands of dollars in hardware. One construction site. This is how humanoids will enter the workforce, not by replacing everyone overnight, but by quietly taking over the hours nobody wants to spend.

Future Memo

18,925 Aufrufe • vor 15 Tagen

Don't train the model, evolve the harness. I read a brilliant blog post from Hugging Face where they took a frozen open model scoring 0% on a hard legal agent benchmark, left its weights alone, and let an automated loop rewrite only the code around it. That code layer is the harness, the runtime wrapper that feeds the model context, runs its tool calls, and decides when a run ends. By the time the loop finished, the system had essentially matched Sonnet 4.6 on the benchmark's headline metric, at roughly 7x lower cost per task. Zero weights changed. The gain existed because of where the model was failing. The judge only grades files saved in the right place under the exact requested filename, and the model kept doing the legal analysis correctly, then saving it under the wrong name, dropping it in a scratch folder, or never writing it at all. So the 0% was never measuring legal reasoning. It was measuring the harness. Hand-tuning that layer is slow and model-specific, so they automated it. A Claude proposer adds exactly one mechanism per iteration, and an outer loop keeps it only if it clearly beats the current best, so accepted mechanisms compound. What the loop discovered says a lot about where agents actually fail. → The biggest single gain was file handling, not intelligence. An automatic step that lands the deliverable exactly where the judge expects it beat every prompt change, with zero extra model tokens. → Code fixes transferred across models, prompt playbooks did not. The same harness lifted a smaller model from the same family by 14 points, but the tuned prompts hurt a different model family on tasks it could already finish. → The harness mattered more than anything else. Same model, same judge, same tasks, and five different harnesses scored anywhere between 3.5% and 80.1%. The gains do eventually flatten, and the remaining misses look like real capability gaps. At some point the wrapper runs out of tricks and the model has to carry the work. But the lesson holds. A benchmark score measures the model and its harness together, and until the harness is fixed, it's impossible to know which one failed. I highly recommend reading this: I also wrote a deep dive on agent harness engineering a while back, covering the orchestration loop, tools, memory, context management, and everything that turns a stateless LLM into a capable agent. The article is quoted below.

Akshay 🚀

245,379 Aufrufe • vor 2 Monaten

HydraFusion Explained. Part I: How does the Copilot engine know what to optimize for? Your prompt is evaluated across 4 dimensions: ➡ Does it require deep reasoning? (aka. reasoning depth) ➡ Is it a sophisticated problem? (aka. code generation complexity) ➡ Is it untangling a complicated mess? (aka. debugging difficulty) ➡ Is it dominated by tool-use? (aka. tool orchestration needs) Based on this evaluation, a HyDRA score is assigned to determine the capability profile your task needs the most and to establish a quality bar. Part II: How does it choose a model? Note: It doesn' t pick one model to handle the entire job e2e, (that's Auto mode). Instead, it selects 1 of 3 execution workflows and assigns the best model at different stages based on the HyDRA score: 1️⃣ Single ⚙️ How it works: A single model completes the task from start to finish. ⚖️ Rationale: The task comfortably meets the quality bar with one model. Multi-model orchestration would add latency and cost with no meaningful quality gain. 2️⃣ Cascade ⚙️ How it works: A lightweight, cost-efficient model generates the solution. This draft is evaluated against a quality gate and if it falls short of the quality bar, the entire task escalates to a stronger, frontier model. ⚖️ Rationale: Only bring in the big guns when there is concrete evidence that a lightweight model won't meet the quality threshold. 3️⃣ Critique ⚙️ How it works: A lightweight model drafts the initial code and tool interactions. An independent, read-only frontier model reviews that draft and provides feedback. The original lightweight model then performs any targeted revision(s) before the final response is sent to the user. ⚖️ Rationale: Writing code (output tokens) is expensive while reviewing code (input tokens) is cheap. Instead of incurring the cost of a powerhouse writing hundreds of lines from scratch, a cost-efficient model writes the first draft, and the frontier model just reviews it and points out fixes. HydraFusion is available in experimental preview on the GitHub Copilot CLI: /experimental on, /model and select Hydrafusion (Research Preview)

Julia Muiruri

12,687 Aufrufe • vor 20 Tagen

FIVE LAYERS OF AGENT ENGINEERING, EACH ONE WRAPS THE ONE BELOW IT. IF YOU SKIP LAYER 2, YOUR LAYER 5 WILL LOOK BROKEN WHEN IT IS ACTUALLY JUST STANDING ON NOTHING. for weeks i debated harness vs loop vs graph like they were competing choices. then a stack diagram made the shape obvious. they are not choices. they are floors. 01 | prompt engineering. the message. unit of work: one input. inputs are role, instructions, examples, format. output is a single raw response. 02 | context engineering. the memory. unit of work: what stays in the window. a curator selects, compresses, and drops from query, docs, memory, prior turns, and tool outputs before the prompt runs. 03 | harness engineering. the machine. unit of work: the machine itself. gather (context + prompt) → LLM → tools or sub-agents → verifier → final response. the article calls this the operating environment. 04 | loop engineering. the system. unit of work: the run. goal + success criteria + max iterations + budget + completion check wrap around one harness pass. failed pass appends results to context and retries. 05 | graph engineering. the topology. unit of work: the graph run. goal + nodes + edges + state schema. graph routes to agent nodes, tool nodes, or human approval. a reviewer node with a different model and fresh context checks the final answer. the wrapping is the whole point. layer 5 assumes layer 4 works. layer 4 assumes layer 3 works. skip layer 2 and layer 3's verifier keeps failing without a clear reason. this is why swapping the model is a one-day project and swapping the stack is a quarter. the model is the commodity. the five layers around it are the engineering. full three-layer breakdown of the top of the stack (harness, loop, graph) in the post below.

kocer

31,198 Aufrufe • vor 29 Tagen

a moonshot engineer leaked the benchmark anthropic, openai and xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article below

starmex

33,100 Aufrufe • vor 1 Monat