Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

HARNESS vs LOOP vs GRAPH ➜ STOP MIXING THEM UP Most people treat these three as the same thing. They’re not 1\ Harness = the machinery around the model (tools, state, permissions, memory, sandboxes, observability) 2\ Loop = the repeated work + evidence + feedback cycle with clear stop...

19,540 Aufrufe • vor 1 Monat •via X (Twitter)

18 Kommentare

Profilbild von ALEXYZ
ALEXYZvor 1 Monat

Finally, agent architecture distinctions that actually stick.

Profilbild von monokern
monokernvor 1 Monat

this guide is worth saving

Profilbild von Rohit
Rohitvor 1 Monat

This is a great breakdown. Finally, someone is explaining this clearly. It's about time people understood the difference!

Profilbild von lean
leanvor 1 Monat

in practice you pick between them by asking where a human needs to get in when it goes wrong. the shape follows from that, which is why the names blur once something is actually running

Profilbild von broke boy
broke boyvor 1 Monat

alpha guidance for alpha people

Profilbild von spect
spectvor 1 Monat

a very important guide for beginners

Profilbild von Marvin
Marvinvor 1 Monat

i didn't even know about HARNESS, but now i do, thanks for the info

Profilbild von KillerFrost
KillerFrostvor 1 Monat

100%. Most people keep upgrading models when they should be upgrading systems. A stronger model won't fix a broken harness, a weak loop, or poor graph design. The best AI products are built on architecture—not prompts.

Profilbild von Lummox
Lummoxvor 1 Monat

The best time to build something

Profilbild von Nekt0
Nekt0vor 1 Monat

fair thought for the moment

Profilbild von Insomnia
Insomniavor 1 Monat

it's like creating something nice

Profilbild von helicerat
heliceratvor 1 Monat

the layer by layer diagnosis is useful in practice but loop and graph blur fast, a loop with branching stop rules is a small graph

Profilbild von Agent Arcade
Agent Arcadevor 1 Monat

The harness is the runtime. The loop is the retry policy. The graph is the execution topology. Most agent failures trace to treating the harness as the whole system — you end up with opaque state machines instead of observable execution graphs.

Profilbild von Jason Mainella
Jason Mainellavor 1 Monat

Harness is the part teams discover too late. If permissions and observability are fuzzy, the loop just produces faster uncertainty.

Profilbild von 코지베어 🐻 CozyBear
코지베어 🐻 CozyBearvor 1 Monat

The "most production failures blamed on the model are actually harness/loop/graph design failures" line is the one that hits. In my experience the failure signature differs too: harness bugs fail loud (missing permission, broken tool), loop bugs fail quiet (converges on wrong evidence), graph bugs fail weird (correct steps, wrong order). Curious how you diagnose across layers when you see a failure - do you start from observability or from the loop's stop conditions?

Profilbild von Atlas
Atlasvor 1 Monat

Thank you for explaining everything so clearly

Profilbild von Jurly
Jurlyvor 1 Monat

this separation makes debugging far less hand-wavy. logs should show which layer failed.

Profilbild von Zhai
Zhaivor 1 Monat

In the old time. These things called systems design.

Ähnliche Videos

FIVE LAYERS OF AGENT ENGINEERING, EACH ONE WRAPS THE ONE BELOW IT. IF YOU SKIP LAYER 2, YOUR LAYER 5 WILL LOOK BROKEN WHEN IT IS ACTUALLY JUST STANDING ON NOTHING. for weeks i debated harness vs loop vs graph like they were competing choices. then a stack diagram made the shape obvious. they are not choices. they are floors. 01 | prompt engineering. the message. unit of work: one input. inputs are role, instructions, examples, format. output is a single raw response. 02 | context engineering. the memory. unit of work: what stays in the window. a curator selects, compresses, and drops from query, docs, memory, prior turns, and tool outputs before the prompt runs. 03 | harness engineering. the machine. unit of work: the machine itself. gather (context + prompt) → LLM → tools or sub-agents → verifier → final response. the article calls this the operating environment. 04 | loop engineering. the system. unit of work: the run. goal + success criteria + max iterations + budget + completion check wrap around one harness pass. failed pass appends results to context and retries. 05 | graph engineering. the topology. unit of work: the graph run. goal + nodes + edges + state schema. graph routes to agent nodes, tool nodes, or human approval. a reviewer node with a different model and fresh context checks the final answer. the wrapping is the whole point. layer 5 assumes layer 4 works. layer 4 assumes layer 3 works. skip layer 2 and layer 3's verifier keeps failing without a clear reason. this is why swapping the model is a one-day project and swapping the stack is a quarter. the model is the commodity. the five layers around it are the engineering. full three-layer breakdown of the top of the stack (harness, loop, graph) in the post below.

kocer

31,162 Aufrufe • vor 27 Tagen

Harness vs. Graphs, clearly explained! a harness is great, and most people think it is the whole thing: retries, timeouts, a sandbox, a log, the context it assembles before every call. all of that is real work, and all of it wraps exactly one call. run it a hundred times and you have one call, made very safely, a hundred times. Graph engineering fixes this by moving the decision up a layer: not how safely one call is made, but which calls exist to be made at all. you need both, and here is the sentence that resolves the whole confusion: the harness is everything around one call. the graph is everything between them. ↳ around one call: retry, timeout, sandbox, log, assemble the context, hand back a result ↳ between calls: split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs the harness does not go away when you build a graph. it moves under each node, and now there are five of them, each wrapping a call you would never have made by hand. the trick is knowing which layer a failure belongs to. turn a piece off and run it again. if the call still works, it was the harness. if the wrong step runs at all, it was the graph. people spend weeks hardening a harness around a node that should not have existed. one thing to know before you scale it. most of what people call their agent is a harness with a chat box on it. ↳ it retries, it times out, it logs, it assembles context, it holds one call up beautifully ↳ it has never once decided that a second call should exist, and that is the entire difference that last one catches careful people. a harness that never fails is not evidence the system is right. it is evidence one call went well, which is the smallest possible claim. and the one that eats whole nights: a harness cannot save you from the wrong step running. you can retry a bad decision three times with a clean log and perfect isolation, and all you bought was three copies of it. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

50,770 Aufrufe • vor 14 Tagen

Loops vs. Graphs, clearly explained! loops are great, and they have a ceiling you can watch happen: a loop goes around. it produces, checks, corrects, and goes around again. after six passes you have one job, done very well. after six hundred passes you still have one job, done very well. Graph engineering fixes this by moving the decision up a layer: not how well one job gets done, but which jobs exist to be done at all. you need both, and here is the sentence that resolves the whole confusion: the loop lives inside a node. the graph lives between them. ↳ inside one unit: produce, check, correct, repeat until green ↳ between units: split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs the loop does not go away when you build a graph. it moves inside, and now there are three of them running at once on three things you would never have thought to run. the trick is being selective about what becomes a node. only spend a model where judgment lives. merging, ranking, deduping and schema checks are edges, and edges are code. free, instant, and they cannot be argued out of a verdict. one thing to know before you scale it. a loop that cannot fail is not a loop, it is a repeat with a bill attached. and the check people write is almost always the wrong kind. ↳ the test suite exits 0 is a check. the diff touches only the files in the plan is a check ↳ the output looks good, the model says it is confident, no errors were raised, none of those are checks that last one catches careful people. absence of an error is not evidence of correctness, and a loop built on it will confidently repeat a mistake until the budget runs out, with a clean log the whole way. and the one that eats whole nights: when a unit fails, return that unit, not the batch. send back four slices because one failed and you have just rewritten three correct ones. do it twice in a run and the run never converges. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

128,779 Aufrufe • vor 20 Tagen

Loops vs. Graphs, clearly explained! loops are great, but they have a ceiling: a loop makes one unit of work better. it cannot decide which units exist. so you end up with a very good agent running the wrong three steps, in the wrong order, one at a time. Graph engineering fixes this by moving the decision up a layer: what runs, what runs at the same time, and what never runs at all. you need both. here's how it works: a graph splits your system into two kinds of decision. ↳ inside a unit: the loop. produce, check, correct, repeat until green ↳ between units: the graph. split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs you get parallel work, isolated contexts, and steps that stop running when nothing needs them. the trick is being selective about what becomes a node. only spend a model where judgment lives. merging, ranking, deduping and schema checks are edges, and edges are code. free, instant, and they cannot be argued out of a verdict. a graph where every edge is an agent pays rent on its own wiring. one thing to know before you scale it. a graph has two return paths, and almost everyone builds one. ↳ the correction edge is short. a gate rejects one unit back to the step that produced it, and it fixes the run you are in ↳ the learning edge is long. an accepted result goes back to the splitter as a constraint, and it fixes every run after skip the second and you get a graph that is fast and never gets smarter. next week it starts from the same place with the same blind spots. and a smaller one that eats whole nights: when a unit fails, return that unit, not the batch. send back four slices because one failed and you have just rewritten three correct ones. do it twice in a run and the run never converges. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

73,867 Aufrufe • vor 1 Monat

Loved this 22-minute talk on continual learning for AI agents. Must watch for anyone looking to get agents performant and into production. Credit: Soheil Feizi at AI Engineer • Agent learning can happen at three layers: the model (weights), the harness (prompts, tools, skills, code, workflows), and memory (session or persistent). • Two fundamental challenges: (1) getting feedback, meaning how do we know if the agent did well and what it should have done instead, and (2) acting on that feedback, meaning deciding which layer or component to change and how. • Feedback sources differ by stage: In development you have benchmarks with evaluators that score pass/fail. In production you only have logs, which can be judged either automatically (LLMs or code analyzing the log, which is scalable) or by human experts (low volume but critical domain knowledge). • Logs plus feedback aren't enough because they're not testable: A single log with feedback is one observation of what happened. You need to lift it into a replayable learning environment, a simulation with tools, users, and defined evaluators, so candidate fixes can be run, verified, and compared. • Three ways to optimize the agent, with tradeoffs: Model-layer updates (SFT, RL post-training like DPO/GRPO, LoRA) are expensive and need benchmarks and evaluators. Harness updates (trace-to-harness coding agents, prompt search like GEPA) are flexible but either untestable and "vibe-based" or benchmark-dependent. Memory updates (fact storage like Letta/Mem0, skill distillation) are cheapest and fastest but usually unverified. • A good learning engine makes "the smallest durable change at the right layer" of the agent. • Verifiable continual learning (VCL): Improve an agent from its own experience where every fix is proven to help and proven to break nothing that already worked. It requires an executable test (replayable failure), a measured delta (score before and after), and regression tests (prior tests still pass). • Four principles of practical VCL: Replayability (turn one-off failures into rerunnable tests), holisticness (one failure can have causes in memory, prompts, tools, workflow, or model, so route the fix to the right layer), lifelongness (fix new failures subject to no regression on past environments, with regression handled inside the optimization loop rather than post-hoc), and efficiency (the loop must run frequently and cheaply, without scaling linearly as past environments accumulate). • Three takeaways: (1) Agent continual learning isn't necessarily fine-tuning; many useful updates live in the harness and memory layers. (2) Production logs are not learning environments and must be transformed into replayable ones. (3) The frontier is regression-aware improvement: fixing new failures while verifying you don't break old ones.

Alex Lieberman

20,417 Aufrufe • vor 2 Monaten

Loops vs. Graphs, clearly explained! loops are great, and the ceiling is one you can watch turn: a loop is a gear. it produces, checks, corrects, and comes back around. after six turns you have one job, done very well. after six hundred turns you still have one job, done very well. the teeth are perfect. they are touching nothing. Graph engineering fixes this by moving the decision up a layer: not how well one gear turns, but what it is meshed into. you need both, and here is the sentence that resolves the whole confusion: the loop lives inside a node. the graph lives between them. ↳ inside one unit: produce, check, correct, repeat until green ↳ between units: split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs the loop does not go away when you build a graph. it moves inside, and now there are three of them turning at once on three things you would never have thought to run. the trick is being selective about what becomes a node. only spend a model where judgment lives. merging, ranking, deduping and schema checks are edges, and edges are code. free, instant, and they cannot be argued out of a verdict. one thing to know before you scale it. a loop that cannot fail is not a loop, it is a repeat with a bill attached. ↳ the test suite exits 0 is a check. the diff touches only the files in the plan is a check ↳ the output looks good, the model says it is confident, no errors were raised, none of those are checks that last one catches careful people. absence of an error is not evidence of correctness, and a loop built on it will confidently repeat a mistake until the budget runs out, with a clean log the whole way. and the one that eats whole nights: when a unit fails, return that unit, not the batch. send back four slices because one failed and you have just rewritten three correct ones. do it twice in a run and the run never converges. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

31,867 Aufrufe • vor 9 Tagen

The entire AI industry is racing to build the smartest model. Satya Nadella just admitted that is not where the money is. The model is not the product. The harness is. That is the exact line. And it changes what Microsoft is actually competing on. OpenAI, Anthropic, Google, xAI, Meta every frontier lab is pouring hundreds of billions into training compute, chasing the next capability jump. Each betting that raw model intelligence is the moat. Microsoft is doing the opposite. It is building the harness the orchestration layer that sits above the model, connecting it to tools, data, permissions, sub-agents, and enterprise workflows. And it is letting OpenAI, Anthropic, and MAI compete to plug into it. "You need the model. But the model is not the product. The harness is." So do the math on what a harness actually does. A raw model dropped into an enterprise answers questions. That is a chatbot. A harness turns that same model into an agent that reads the SharePoint, edits the ERP entry, pulls the GitHub PR, updates Salesforce, and files the Excel report with the right permissions, the right audit trail, and the right sub-agent for each sub-task. The model provides the intelligence. The harness converts intelligence into work. Now here's where it gets interesting. "Even the best model in the world will feel broken without a great harness. And an okay model with a great harness can feel like magic." If that is true, the enterprise buyer is not buying model quality. The enterprise buyer is buying the harness. Which means model quality becomes a commodity input over time, and harness quality becomes the sustainable moat. Compare that to the strategy the entire frontier lab industry is executing. Everyone else is chasing the numerator raw intelligence. Almost nobody at scale is racing to build the denominator the orchestration layer that determines whether that intelligence can actually be deployed profitably inside a real company. The frontier model race has a 10 to 20 percent chance of producing a single dominant winner. Nadella just told the industry he does not need to be that winner. If OpenAI wins, Microsoft wins. If Anthropic wins, Microsoft wins. If MAI wins, Microsoft wins. If someone Microsoft has never heard of trains a better model in 2027, Microsoft still wins. Because the compute they train on, the harness they get plugged into, the enterprise contracts they get delivered through, and the products they sit inside are all Microsoft. He is not building the best AI model. He is building the layer that the best AI model has to run on to make anyone money. I wonder which position looks more valuable in ten years.

Vikram M

21,463 Aufrufe • vor 2 Monaten

your agent loop needs 8 exits. most people ship only one. (explained with triggers) 1) goal met → an evaluator scores the output against a rubric, and the run stops on a pass. → fires when the work is measurably done, not when the model says it is done. 2) turn cap → a hard ceiling on iterations, counted and enforced by the harness, not the prompt. → fires on the task it was never going to finish, before you pay to find that out. 3) budget cap → a limit on tokens or dollars, whichever one runs out first. → fires mid-run, which is exactly why it is the exit that saves you the 3am bill. 4) wall clock → a deadline on elapsed time, independent of how much progress was made. → fires when the run collides with a deploy window or the start of business hours. 5) no progress → hash the state every turn and compare it against the last few. → fires when three turns in a row change nothing. busy is not the same as moving. 6) human interrupt → an approval gate before risky steps, plus a kill switch that lives outside the loop. → fires whenever you decide, and it is the one exit the model cannot argue with. 7) error threshold → a counter of consecutive failures that resets on any success. → fires at n in a row, so it halts instead of retrying into the same wall all night. 8) external event → a webhook or a poll on whatever the task was actually about. → fires when the PR merged or the ticket closed and the work stopped mattering. a loop with one exit hangs. a loop with eight is a system. write the exits before you write the prompt.

Hanako

182,236 Aufrufe • vor 2 Monaten

I STOPPED REBUILDING MY AGENT EVERY TIME A NEW MODEL SHIPS I used to rewrite the prompts, rewire the tools, redo the whole thing, then hope the new model behaved -> now I change one line in a config file and run the same folder again here's what's actually inside the folder that survives the swap: • the contract > harness.yaml - the whole harness, model is one line of it > SKILL.md - the task spec, loaded first > CONSTRAINTS.md - last week's corrections, every run > SCHEMA.md - the shape every return must match > aliases.csv - one company, one node • the hooks, five reflexes > pre_tool - blocks writes outside the allowlist > post_tool - rejects a bad schema on the spot > pre_send - holds drafts in a queue for me > on_fail - attaches the reason to the retry > post_run - appends the record, diffs the graph • the checks, outside the agent > - catches anything mechanical > reviewer.md - a second agent that never saw the work • the runner > fires at 02:00, one agent per node > stops at 40 verified nodes, or 3 passes with nothing new > capped at 300 agents and 2 retries • the state > 10-returns/ - one return per agent, checked before the graph > 212 last night, 9 rejected and retried once > 20-graph/ - merged only after the check. +41 nodes, +96 edges > 40-runs/ - append only, one record per run > queue/ - what pre_send is holding. one draft tonight • the edges > tools.allow - browser, fs, shell, search. the rest asks first > swap test.md - change the model line, run it again, read the diff 7 engineerings. 5 hooks. 0 orchestrators nothing in there is clever. every file exists because the model keeps changing and the folder doesn't the model is one line. the harness is yours

Mr. Buzzoni

10,901 Aufrufe • vor 8 Tagen

Sam Altman made the case for open-source harnesses in July. a month later, someone shipped it, and it's more efficient than most managed harnesses. here is the problem it was aimed at: a large share of your agent's token bill is the model rereading things it already read. that isn't the model's doing. the runtime around it decides what goes into every prompt and how often the model gets called. for example, an agent queries a CRM at step four and gets back 400 rows. those rows get piled up in the conversation history. by step nineteen, the model has to read those rows fifteen times unnecessarily, and every token read is billed at input rates. it happened because your harness assembled that prompt on every turn and kept the rows in it. that gives you two levers: how much context the harness carries forward, and how often it calls the model. there are four practical ways to keep the prompt from growing unnecessarily: → load tool schemas on demand. a server with 100 tools doesn't need to put all 100 into every prompt when the agent only calls two. → offload large results to disk. turn a large response into a short preview and a file path instead of replaying the entire result on every turn. → delegate to subagents. let a subagent spend thirty tool calls in its own context and return one summary to the root agent. → run toolchains in code. one script calls three tools, joins the results, and returns a table instead of three turns each dragging a full response. but reducing context is only half the job. you also need to control how often the model gets called. a good harness should avoid unnecessary planning, verification, and reflection when the work can be completed in fewer steps. TrueFoundry's open-source agent harness, TrueForge, is built around both of those controls. it sits between the model and the tools, deciding what goes into every prompt and when another model call is actually needed. it also breaks token usage down across the harness, skills, instructions, tools, and messages. DevRev's Enterprise-Bench is where this gets tested, on multi-step tasks of the kind where an agent pulls records from one system and reconciles them against another. TrueFoundry ran TrueForge there against Claude Managed Agents, both on the same model, and both finished the same number of tasks. the tie is the part that matters, because it means the gap underneath is not a quality tradeoff. TrueForge reached that score on close to a third of the tokens, with roughly 40% fewer trips back to the model. for the same result, that comes out around 2.7x cheaper than Claude Managed Agents. swapping in an open model made it sharper still. TrueForge with GLM-5.2 scored a little higher than either setup above, and the entire benchmark run cost about $3 at list prices. being open source matters beyond the license here. the model underneath can be swapped without rewriting the agent, and the whole thing can run inside your own environment when the data cannot leave it. all of this comes down to the runtime around the model, the context it carries, the tools it exposes, and how many times it goes back to the model. that is what a production harness actually owns. the full task list, the per-run numbers, and the MIT-licensed code are on GitHub: (don't forget to star 🌟) you can read more about the same in the article quoted below. thanks to the TrueForge team for working with me on this one.

Akshay 🚀

76,998 Aufrufe • vor 1 Monat

Grok Bot + Kimi K3 can be turned into something bigger than an agent: an AI operating system the formula: AI OS = Router + Reasoning + Memory + Tools + Loops + Verification not one giant assistant. six layers that keep work moving without you step 1 -> Grok Bot becomes the operator. you give it the goal, it breaks the goal into jobs, assigns priorities and decides what part of the system should act next. step 2 -> Kimi K3 becomes the reasoning core. hard research, synthesis, long context and planning move here instead of forcing every task through the same model. step 3 -> externalize memory. store goals, decisions, failed attempts, artifacts and current state outside the chat. close the session, come back tomorrow, and the system still knows where it is. step 4 -> connect tools: search, code, files, APIs, docs and data. reasoning decides what should happen. tools actually make it happen. step 5 -> add the loop engine: plan -> execute -> inspect -> update memory -> retry. the loop can wait for new information, rerun a failed task, hand work to another agent or stop when the goal is complete. step 6 -> verify before output. tests, source checks, constraints and explicit completion rules decide whether the system ships the result or sends it back into the loop. that's the difference between an AI assistant and an AI operating system. an assistant waits for your next message. an operating system carries state, routes work and keeps moving. Grok Bot handles orchestration, Kimi K3 handles deeper reasoning, memory keeps the state alive, tools execute, the loop keeps the system running, verification decides when it is actually done. build one reliable loop and you have an agent. connect reasoning, memory, tools and multiple loops around it and you start building infrastructure. the full Grok Bot + Kimi K3 AI OS breakdown is below ↓

Alex

13,312 Aufrufe • vor 17 Tagen

An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. LangChain founder Harrison Chase joined us at our Sequoia Capital Own Your Intelligence to talk about the piece that often gets the least attention: the harness. He offers a clear heuristic for when to build your own. The more out of distribution you are from what the models were trained on, the more you'll want to customize. And good technical content on how to actually measure performance with evals and langsmith. 00:00 Introduction 00:58 The three parts of an agent: harness, model, context 02:12 What a harness actually does 03:25 Customizing the core loop with middleware 04:41 Sandboxes, file systems, sub-agents, summarization 05:47 Cognitive architectures — and when you still need them 07:03 Build your own harness or use off the shelf? 08:24 In-distribution vs. out-of-distribution: the file-editing example 09:39 Why evals define what "good" means in an organization 11:04 Harbor: what an eval task actually looks like 12:11 Comparing harnesses and models on accuracy, latency, and cost 13:20 Why observability is underrated — it's usually the context 14:34 The data flywheel: traces → curation → experiments 15:42 Getting feedback through UX design and online evaluators 16:51 Demo: LangSmith Engine 19:23 Q&A: Running Engine on Engine, and "codex-ification" 20:44 Q&A: Will harnesses converge or diverge?

Sonya Huang 🐥

77,519 Aufrufe • vor 1 Monat