Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Let me explain the agent loop, simple It's the core of every agentic system, and the part most people overcomplicate It's just this: 1. Send messages to the model 2. Model responds, maybe calls a tool 3. You run the tool 4. Append the result back to messages 5....

12,514 Aufrufe • vor 3 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

this is f*cking beyond comprehension. Google engineers just shipped the entire agent lifecycle in one release: build, scale, govern. and every piece answers a specific way agents die in production > context layers (Static, Turn, User, Cache): you decide what the model carries between turns, so token spend stops being a mystery > a self heal plugin: the agent notices a tool call failed and retries it a different way instead of dying mid run > adk deploy: one command from your laptop to the managed runtime, no packaging, no infra ticket > Go joins Python and Java, with its own A2A SDK then the part nobody builds for themselves: > a dashboard on token consumption, latency, error rates and tool calls: the four things that actually kill an agent > a traces tab that opens the real sequence of actions the agent took, step by step > a playground wired to the deployed agent, past sessions included, so debugging is not a redeploy loop > an Evaluation Layer with a User Simulator, because you cannot unit test a non deterministic system and the part that decides whether it ever ships: > agents get native identities as first class IAM principals: least privilege applies to them like it does to people > Model Armor screens prompt injection, tool calls and responses, inline for Gemini or over REST > Security Command Center inventories every agentic asset and flags data exfiltration by an agent ADK is already at 7 million downloads. the runtime has a free tier, and express mode runs off a Gmail address. the prototype was never the hard part.

NO1ennn

24,666 Aufrufe • vor 18 Tagen

New open-source agent harness just landed! I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.

elvis

11,303 Aufrufe • vor 1 Monat

If you are trying to understand where AI agents are going, learn harness engineering. A capable model is only one part of an agent system. Once the model begins reading files, calling tools, modifying state and working across many steps, the quality of the system depends increasingly on the software around it. Consider a coding agent working through a large repository. The model can decide that it needs to inspect a file, search for a symbol, make an edit or run a test, but those decisions do not execute themselves. The surrounding runtime has to decide which resources are available, whether the requested action is permitted, how the operation should be performed, what result should be retained, and what information should be presented to the model on the next step. This becomes harder as the run gets longer. As history accumulates, replaying everything can become costly and less effective. The harness has to decide what should remain in context, what should be summarized or retrieved later, and what belongs in persistent state outside the context window. Execution has similar problems. A long-running agent may need to survive an interruption, avoid repeating completed work, enforce permissions around consequential actions, and preserve enough history to reconstruct what happened when the final result is wrong. These are harness problems. The harness is the layer that manages context, tools, execution, state, checkpoints, limits and traces around the model. Harness engineering is the work of designing and improving that layer. Engineers inspect execution traces, evaluate agents on representative tasks, look for recurring failure modes, and then change things such as context selection, tool interfaces, state handling or execution controls. That last part matters because agent failures are often not fixed by changing the model. Sometimes the useful change is in what the model sees, how a tool is exposed, what state is preserved, or what the runtime does after a failed step. As agents take on longer tasks, the demands on this surrounding software grow. Model capability remains essential, but harness engineering is what turns that capability into an execution process that can be controlled, inspected, tested and improved.

Tech with Mak

48,318 Aufrufe • vor 8 Tagen

FIVE LAYERS OF AGENT ENGINEERING, EACH ONE WRAPS THE ONE BELOW IT. IF YOU SKIP LAYER 2, YOUR LAYER 5 WILL LOOK BROKEN WHEN IT IS ACTUALLY JUST STANDING ON NOTHING. for weeks i debated harness vs loop vs graph like they were competing choices. then a stack diagram made the shape obvious. they are not choices. they are floors. 01 | prompt engineering. the message. unit of work: one input. inputs are role, instructions, examples, format. output is a single raw response. 02 | context engineering. the memory. unit of work: what stays in the window. a curator selects, compresses, and drops from query, docs, memory, prior turns, and tool outputs before the prompt runs. 03 | harness engineering. the machine. unit of work: the machine itself. gather (context + prompt) → LLM → tools or sub-agents → verifier → final response. the article calls this the operating environment. 04 | loop engineering. the system. unit of work: the run. goal + success criteria + max iterations + budget + completion check wrap around one harness pass. failed pass appends results to context and retries. 05 | graph engineering. the topology. unit of work: the graph run. goal + nodes + edges + state schema. graph routes to agent nodes, tool nodes, or human approval. a reviewer node with a different model and fresh context checks the final answer. the wrapping is the whole point. layer 5 assumes layer 4 works. layer 4 assumes layer 3 works. skip layer 2 and layer 3's verifier keeps failing without a clear reason. this is why swapping the model is a one-day project and swapping the stack is a quarter. the model is the commodity. the five layers around it are the engineering. full three-layer breakdown of the top of the stack (harness, loop, graph) in the post below.

kocer

31,198 Aufrufe • vor 1 Monat

sorry, they just did WHAT someone gave a machine one disease name, the leading cause of blindness in the developed world with 1.5 million americans already in its path, and it came back pointing at a drug that has sat in pharmacies for years under a different label: 551 papers read in 30 minutes against the 294 hours a human would have needed, and the loop that did it is public on GitHub most agent setups answer one question at a time, so the ceiling on the work is the quality of the question you happened to think of this one was handed a single question and wrote the second one itself. turns out that follow-up is where the real find was: a target called ABCA1, upregulated threefold, in an experiment no human ordered i read the whole paper looking for the trick, and the trick is structural. that is the second question, and it is the gap between an assistant and a factory: - hand the loop a field rather than a task: it was given a disease, and choosing the mechanism was part of its job - make it rank before it spends: 151 papers in, ten candidate mechanisms out, scored against each other before anything touched a bench - split reading from judging, so the agent that forms the theory is a different agent from the one grading it - close every cycle on physical reality: the verdict was an experiment, and another model's opinion was never allowed to stand in for one - feed each result back as the next question rather than a log line, which is the step almost nobody builds - search what already passed inspection first: the winner was an approved compound with a safety file already on record - write down what the round learned before opening the next one, so round two starts where round one stopped my read, and i think it is the uncomfortable one: reading was the entire bottleneck in that field, and everybody spent the decade optimising the writing. people ran every physical experiment here, the analysis agent needs a domain expert writing its prompts, and the authors decline to call this the leap it resembles. the thinking got replaced, and the hands did not so the question i cannot answer for my own setup: which step of your loop still stops dead until you sit down and type something bookmark this one. the four parts that turn one model into a line that runs like this, the queue, the rooms, the write permissions and the gate, are built file by file in the piece below ↓

Argona

32,744 Aufrufe • vor 1 Monat

Don't train the model, evolve the harness. I read a brilliant blog post from Hugging Face where they took a frozen open model scoring 0% on a hard legal agent benchmark, left its weights alone, and let an automated loop rewrite only the code around it. That code layer is the harness, the runtime wrapper that feeds the model context, runs its tool calls, and decides when a run ends. By the time the loop finished, the system had essentially matched Sonnet 4.6 on the benchmark's headline metric, at roughly 7x lower cost per task. Zero weights changed. The gain existed because of where the model was failing. The judge only grades files saved in the right place under the exact requested filename, and the model kept doing the legal analysis correctly, then saving it under the wrong name, dropping it in a scratch folder, or never writing it at all. So the 0% was never measuring legal reasoning. It was measuring the harness. Hand-tuning that layer is slow and model-specific, so they automated it. A Claude proposer adds exactly one mechanism per iteration, and an outer loop keeps it only if it clearly beats the current best, so accepted mechanisms compound. What the loop discovered says a lot about where agents actually fail. → The biggest single gain was file handling, not intelligence. An automatic step that lands the deliverable exactly where the judge expects it beat every prompt change, with zero extra model tokens. → Code fixes transferred across models, prompt playbooks did not. The same harness lifted a smaller model from the same family by 14 points, but the tuned prompts hurt a different model family on tasks it could already finish. → The harness mattered more than anything else. Same model, same judge, same tasks, and five different harnesses scored anywhere between 3.5% and 80.1%. The gains do eventually flatten, and the remaining misses look like real capability gaps. At some point the wrapper runs out of tricks and the model has to carry the work. But the lesson holds. A benchmark score measures the model and its harness together, and until the harness is fixed, it's impossible to know which one failed. I highly recommend reading this: I also wrote a deep dive on agent harness engineering a while back, covering the orchestration loop, tools, memory, context management, and everything that turns a stateless LLM into a capable agent. The article is quoted below.

Akshay 🚀

245,379 Aufrufe • vor 3 Monaten