a developer showed me the cleanest multi-agent architecture i've... seen. one mission. nobody steps on anyone else. the mission: ship a safe feature. the graph decides what happens next. router reads the mission and asks one question: where should this go? three workers split in parallel. each in their own private work area. researcher finds the evidence. architect designs it. builder creates it. none of them share context. the three workers run independently. everything lands in shared state. facts, decisions, artifacts. one place where the whole picture exists without anyone copying transcripts. then the crew's work converges. integrator combines what was built. reviewer tests quality and safety. human checkpoint approves anything high-impact. then ship. verified output. private context. shared state. that separation is the whole trick. a prompter asks a question. an architect draws a graph. full breakdown with code in the article below.show more

rvaniaaa
27,465 views • 1 month ago
ANTHROPIC ENGINEER JUST SHOWS EXACTLY WHAT GRAPH ENGINEERING LOOKS... LIKE WHEN A TASK RUNS THROUGH IT most people arguing about graphs online have never actually watched one execute Task → Researcher → Planner → Writer + Code Agent → Reviewer → Deploy six nodes, one shared state, graph completes itself while you watch loop mode gets disabled the second the task splits into real specialties, one agent stops trying to do everything at once reviewer catches a failure, kicks it straight back to the exact node that broke - not a full restart, no lost context the graph is not a fancier loop - it is the org chart your agents were missing bookmark this and watch it run, then read the article below to see why the timeline just found what production systems already doshow more

leopardracer
53,328 views • 2 months ago
FIVE LAYERS OF AGENT ENGINEERING, EACH ONE WRAPS THE... ONE BELOW IT. IF YOU SKIP LAYER 2, YOUR LAYER 5 WILL LOOK BROKEN WHEN IT IS ACTUALLY JUST STANDING ON NOTHING. for weeks i debated harness vs loop vs graph like they were competing choices. then a stack diagram made the shape obvious. they are not choices. they are floors. 01 | prompt engineering. the message. unit of work: one input. inputs are role, instructions, examples, format. output is a single raw response. 02 | context engineering. the memory. unit of work: what stays in the window. a curator selects, compresses, and drops from query, docs, memory, prior turns, and tool outputs before the prompt runs. 03 | harness engineering. the machine. unit of work: the machine itself. gather (context + prompt) → LLM → tools or sub-agents → verifier → final response. the article calls this the operating environment. 04 | loop engineering. the system. unit of work: the run. goal + success criteria + max iterations + budget + completion check wrap around one harness pass. failed pass appends results to context and retries. 05 | graph engineering. the topology. unit of work: the graph run. goal + nodes + edges + state schema. graph routes to agent nodes, tool nodes, or human approval. a reviewer node with a different model and fresh context checks the final answer. the wrapping is the whole point. layer 5 assumes layer 4 works. layer 4 assumes layer 3 works. skip layer 2 and layer 3's verifier keeps failing without a clear reason. this is why swapping the model is a one-day project and swapping the stack is a quarter. the model is the commodity. the five layers around it are the engineering. full three-layer breakdown of the top of the stack (harness, loop, graph) in the post below.show more

kocer
31,198 views • 1 month ago
AN ENGINEER SOLVED A PROBLEM NOBODY IN THE COMPANY... COULD FULLY MAP ON THEIR OWN The problem wasn't a missing tool, it was that every process touched three other processes nobody had written down. He wired dozens of agents into one shared graph instead of documenting each department's workflow by hand. Each agent owned one process, reading its own logs while writing every dependency it found straight into the shared structure. Within days, the graph surfaced a connection between two departments that had been quietly blocking each other for over a year. The company adopted the graph as its live map of operations, and it hasn't stopped finding new connections since. See what the graph found in the first 48 hours below👇show more

wast3
16,489 views • 2 months ago
THIS 38,000-STAR GITHUB REPO TURNS ONE AI AGENT INTO... A REAL TEAM THAT CAN BRANCH, VERIFY ITS WORK AND WAIT FOR YOUR APPROVAL most people still run agents as one long chain where every step waits, one failure kills the run and the full workflow starts again Task → Planner → 5 Researchers in Parallel → Skeptic → Writer → Human Gate LangGraph gives every node one job while a shared state carries the findings, decisions and context through the entire system the skeptic can reject an unsupported finding and route the work back before it contaminates the final report, while independent branches keep moving if the run crashes, durable execution resumes from the saved state instead of rebuilding everything, then human-in-the-loop pauses the graph before anything expensive gets sent or published bookmark this repo and watch one prompt turn into an actual org chart for AI agentsshow more

Gipp 🦅
11,709 views • 2 months ago
an agent is four parts in a loop. you... own one. the other three break it. that's why the demo works and prod doesn't. you can't debug what you can't see. 1) the prompt → what you tell the model each turn. you own this one. good. 2) the context window → what it sees right now. the framework fills it with junk, and you never notice until it rots. 3) the tools → what it can do. you own the list, not when or why it fires them. 4) the control flow → what happens next, when to stop. the framework owns this. it's what breaks at 80%. own all four and your agent stops being a magic trick that works on stage and dies on call. this isn't my idea. it's the 12-factor agents guide (24k stars) github: the whole thing every serious builder ends up rewriting their stack around. full breakdown in the article below.show more

Hanako
38,184 views • 2 months ago
AN ANTHROPIC SYSTEM TURNS ONE PROMPT INTO A COMPLETE... GRAPH MODEL BUILT ENTIRELY ON ITS OWN One prompt is enough to spin up the full architecture, no schema written by hand, no roles assigned in advance. Parallel workers process every branch of the graph at once, each one handling its own slice independently. As each worker finishes, its output compacts straight into a single unified project instead of scattered fragments. Nothing needs stitching together manually, the compaction step folds every worker's result into one file automatically. The whole build runs end to end without a single manual step between the prompt and the finished project. See the full build and compact pipeline below👇show more

wast3
18,460 views • 2 months ago
sorry, they just did WHAT someone gave a machine... one disease name, the leading cause of blindness in the developed world with 1.5 million americans already in its path, and it came back pointing at a drug that has sat in pharmacies for years under a different label: 551 papers read in 30 minutes against the 294 hours a human would have needed, and the loop that did it is public on GitHub most agent setups answer one question at a time, so the ceiling on the work is the quality of the question you happened to think of this one was handed a single question and wrote the second one itself. turns out that follow-up is where the real find was: a target called ABCA1, upregulated threefold, in an experiment no human ordered i read the whole paper looking for the trick, and the trick is structural. that is the second question, and it is the gap between an assistant and a factory: - hand the loop a field rather than a task: it was given a disease, and choosing the mechanism was part of its job - make it rank before it spends: 151 papers in, ten candidate mechanisms out, scored against each other before anything touched a bench - split reading from judging, so the agent that forms the theory is a different agent from the one grading it - close every cycle on physical reality: the verdict was an experiment, and another model's opinion was never allowed to stand in for one - feed each result back as the next question rather than a log line, which is the step almost nobody builds - search what already passed inspection first: the winner was an approved compound with a safety file already on record - write down what the round learned before opening the next one, so round two starts where round one stopped my read, and i think it is the uncomfortable one: reading was the entire bottleneck in that field, and everybody spent the decade optimising the writing. people ran every physical experiment here, the analysis agent needs a domain expert writing its prompts, and the authors decline to call this the leap it resembles. the thinking got replaced, and the hands did not so the question i cannot answer for my own setup: which step of your loop still stops dead until you sit down and type something bookmark this one. the four parts that turn one model into a line that runs like this, the queue, the rooms, the write permissions and the gate, are built file by file in the piece below ↓show more

Argona
32,475 views • 1 month ago
this is worth more than most five figure courses... 16 claude agents audit an entire repo at once, a second fleet re-checks every finding on fresh context, and the whole thing runs off one diagram instead of a prompt i ran it against my own code and got back 11 endpoints where i never checked who was logged in, 3 of which the verifier threw out before they ever reached me this is Graph Engineering, the layer above prompting, and it runs on the agent you already pay for: - write your plan out, then ask one question at every "and then": does the next step actually read what the previous one produced - the seams that fail that question were never dependencies, so those jobs run at the same time - the arrows that survive are your real edges, and the longest chain of them is your floor that no number of agents shortens - want it faster, cut a false edge instead of adding a worker - fan the independent work out, one agent per item, no shared state between them - send every finding to a separate agent on fresh context, because a model recognises its own writing 73.5% of the time and grades it kinder once it does - make that verifier check a real signal like a passing test, never the worker's own word that it finished - shard the fleet across worktrees so parallel workers stop overwriting each other, one rule frozen into every worker: never git stash, never git reset - merge only what came back verified, into one report instead of twenty open chats the catch is the ceiling. at 95% independent work 16 agents return 9.14x rather than the 16 you would guess, and even 256 only reach 18.6x, because the merge and the verify stay serial however wide you fan coordination itself is free plain code and every agent underneath it is billed, so start at twenty files and widen once it works bookmark this, the whole method with all six ready-to-run graphs is written out in the article ↓show more

Argona
157,688 views • 2 months ago
You've probably scrolled past a dozen posts about Jev... this week without anyone telling you what it actually is. It's the first model from TypeSafe, a lab started by one of the researchers behind ChatGPT. It's a decision engine: you give it options, it picks one and tells you how sure it is. It cannot write a single word, and that is the interesting part. Every other AI you use writes. That is the whole interface. So when software needs a plain yes or no, we make a model write a paragraph and then dig the answer back out of it. Fine in a chat window where a human reads it. Bad inside software, where code has to act on it. The bet is that the valuable half was never the writing. It was the deciding. That problem showed up in WordPress years ago, and it is the reason WPVibe works the way it does. The AI does the work. Anything permanent stops and waits, because a delete that skips the trash is not something software should decide on its own.show more

John Turner
21,029 views • 14 days ago
a contractor in Shenzhen priced a ¥12,470,900 hospital contract,... about $1.7m, in one afternoon and beat firms carrying forty people he explained how he did it: the bid consultancy he used to pay took three days and ¥46,000 for the same envelope. he did this one alone, off one screen, at 11.4% margin, uploaded before the 17:00 cutoff 214 pages of tender documents read, 68 binding clauses pulled out, 9,485 building parts loaded, 14 places found where a duct and a beam sit in the same cubic metre, deepest one 38mm, all of them fixed, 3,318 lines of quantities priced and the package encrypted and uploaded before the 17:00 cutoff this is Graph Engineering: the job gets cut into small nodes, one narrow task each, wired so that one node's output is the next node's input, and any node is allowed to stop the whole run. it turns a model that answers you into a machine that finishes the job: - give every node one job and one output. a node doing two things fails at both and you cannot tell which one broke - put the cheapest rejection first. his qualification node reads clause 7.4, foreign-owned firms barred, and ends the run four seconds in, before anything expensive touches the model - what moves between nodes is a file. the model travels as a model, the quantities as a table, the price as a number - build exactly one loop: the checker finds 14 collisions, the fixer drops the duct 550mm, the checker runs again, and nothing moves on until the count is zero - cap that loop, or a graph will grind on three impossible clashes until the deadline passes - keep one node whose only job is to say no, and give it authority over everything above it - log each node's output on its own, because when the price comes out wrong you need to know which node believed the wrong thing - run the expensive nodes last, always the catch is that a graph is an extremely confident machine: point it at an outdated rate book and it prices an entire hospital off it without a single node noticing, because no node is asked to doubt the input, only to process it so the nodes that earn their keep are the ones that reject, and almost nobody builds those first bookmark this, the full build with all nine nodes and what each one hands to the next is written out in the article ↓show more

Argona
38,189 views • 2 months ago
Pi rocks. the fact it can be steered programmatically... to such an extent is sick. below we have a structure: - coordinator speaks with me - it's the only one seeing the project broadly: vision, build board, feature lanes - it can spawn workers in separate sessions - workers are dispatched with a feature card and just enough context to deliver, nothing that would distract them - when workers are finished, reviewer kicks in in a separate session. it either accepts the work or sends it back with findings, to a fresh worker, never the one that wrote it. coordinator gets woken either way most of the logic related to communication happens deterministically, so there's a very little chance something relevant is skipped. Pi you guys building one of the most important piece of software of the upcoming months, at least.show more

Adam
101,316 views • 1 month ago
AMAZON SENIOR DEVELOPER BUILT A CONTEXT PIPELINE THAT DECIDES... WHAT THE MODEL EVEN GETS TO SEE Most teams still dump the entire codebase into every prompt and hope it sorts itself out. He ranked every file by relevance to the task instead of how recently it was touched. A router decides how much context each task earns, a typo fix pulls three files, a full rewrite pulls the whole module. Whatever survives gets compressed to the exact lines that actually matter, so nothing bloats the window with dead weight. Token cost per finished task dropped the moment the model stopped reading dead weight just to fix one function. See how the four stages work together below👇show more

wast3
23,868 views • 1 month ago
whoever leaked this has bigger balls than sense someone... gave a fleet of Claude agents shared memory so they would stop contradicting each other, then measured both the bill and the output: the version that talked most made 2.4x the api calls of the version that won, and hallucinated 34% more than doing nothing at all, 0.658 against 0.492 i ran the same question past two of my own agents afterwards and got two different answers about which file owns the config. each one was individually right and the pair was wrong, which is the whole failure in one line this is Graph Engineering, the layer that decides which agents may talk to each other at all, and it installs into the agent you already pay for: - decide which agents may share state at all, because every edge you draw is a channel a mistake can travel down - measure divergence per PAIR instead of as a fleet average, across what they believe about place, time and task history - gate on that number and stop the pair above your threshold before it reasons, rather than repairing the output afterwards - let compressed summaries replace whole states: the verified protocol landed 0.463 against 0.658 for full broadcast - cut the sync frequency until it hurts, since the winning setup used 58% fewer calls than the one that broke it - never propagate a state nobody checked, because the contamination effect came in at d=1.18, a full standard deviation of extra lying - keep the shared layer small enough to diff, which is what a written standard does and a running conversation cannot - re-run the check after every model upgrade, because this was 8 scenarios on one model family at n=30 per condition - and learn where it does not bite: on plain software tasks every condition converged under 0.2 and the whole effect vanished turns out the ranking is the uncomfortable part: verified summaries 0.463, no synchronisation at all 0.492, full broadcast 0.658. the middle option is doing nothing, and it beat the thing everyone builds first the group agreeing is what it looks like when every agent copied the same mistake, which is why a fleet that hallucinates has a replication problem and keeps getting handed a smarter model instead so the question for your own setup: if you asked two of your agents the same thing right now, would they answer the same way bookmark this one. the layer underneath it, deciding which arrows between agents exist at all, is built step by step in the piece below ↓show more

Argona
724,665 views • 1 month ago
somebody explain this because i refuse to accept it... someone ran 48 scored trials and one agent beat a whole fleet of them on all 6 task families, at 0.93 cents a run against 1.9, while openai's best fleet shape was paying $0.008 for every single point of accuracy it bought i read it expecting a hit piece and found the opposite: the fleets that partitioned the dependency graph properly lifted pass rate 14% and cut wall-clock 2.10x on the same tasks, and one of them beat claude code with agent teams the thing that decides it has a name, Graph Engineering, and it is a property of the diagram rather than the model: - partition on the real dependency graph pulled from static analysis, never by folder or by file, because the gains land hardest on the most dependency-dense projects - isolate the structural hub files first, since those are the nodes every partition would otherwise have to share - measure the critical path and treat it as the floor, because a chain that genuinely feeds itself cannot be replaced by more workers and wrapping it in a scheduler does not shorten it - match the topology to the coupling instead of defaulting to parallel: on coupled work a static parallel shape drops below a single agent, so the mismatch is worse than no orchestration - remember each worker serialises its own subtasks, which adds edges inside every agent that were never in your plan - budget the fan-out before you fire it, because three agents already burn roughly three times the tokens and the multiplier compounds across sessions - check worker count against your rate limit, since fifteen workers at ten requests a second walk straight through a hundred-per-second ceiling and cascade - put a script gate in front of the planner: it costs 0.15 seconds and zero tokens, and it lets the expensive model skip 43 to 63% of the steps for at most 1.4 points of accuracy the catch is the coordination tax, and it scales with how clever the shape looks: 58% extra reasoning turns for independent workers, 263% decentralised, 285% centralised, and 515% for the hybrid setup everyone reaches for first the same paper found that hybrid then collapses hardest on tool-heavy work at a 0.452 success rate, while the plainer decentralised shape beat centralised outright despite carrying more overhead, because parallel efficiency is what survives bookmark this, the whole build sits in the article ↓show more

Argona
32,932 views • 2 months ago
⚠️ THIS ROBOT ANSWERED A HOSTILE QUESTION IN 9... WORDS. THAT'S THE SPEC SHEET NOBODY IS READING. "To make humans lives easier and a little more interesting." That answer wasn't charming. It was precise. Break it down: → Addresses utility (easier) → Addresses engagement (more interesting) → Centers the human, not the robot → Delivered in under 10 words to a hostile prompt That is a mission statement that survives contact with a skeptic. In public. Unrehearsed. Procurement teams building out human-robot workflows don't need the robot to be impressive. They need it to be explainable, to workers, to floor managers, to anyone who stops it and asks why it exists. This robot just passed that test. Nobody runs that number on a demo reel. This one ran it live.show more

Shredder
203,661 views • 1 month ago
WTF, GROK BOT JUST MADE AI AGENTS AVAILABLE TO... LITERALLY ANYONE – CREATING CONTENT HAS NEVER BEEN THIS EASY, EVEN IF YOU'VE NEVER MADE ANYTHING BEFORE Content was never a talent problem. It's a headcount problem. One person doing research, design, copy, analytics, timing and publishing – that's six jobs. The switching between them is what kills consistency, not a lack of ideas. Here's what one of these setups actually looks like. A Chief of Staff sits in the middle and routes every task. Nothing lands on the human. → Researcher tracks what's actually moving and pulls real sources instead of guesswork → Writer turns that research into finished copy, ready to review → Visualiser gets fed a few reference visuals once, then ships everything in that style → Analyst reads the numbers and tells the rest of the team what worked → Scheduler owns timing and holds the queue → Publisher ships it The part that makes it work: every agent on Grok Bot gets its own persistent computer, browser and file system – and they all share memory. So the research is already sitting inside the draft before the draft starts. No copy-pasting between tools. No approving every step. No human in the middle. You can even teach an agent a repetitive task by recording yourself doing it once. Start recording, do the thing, stop. It learns the pattern. And that's the real shift. Nobody needs AI to tell them what to post. They need it to delete the 40 steps between the idea and the post. Everyone has a backlog of things they've meant to make for months. This is what starts clearing it. Full breakdown of the setup in the article below ↓show more

SCOTTY BEAM
4,821,782 views • 1 month ago
i don't f*cking understand why this isn't popular yet... someone has created a memory system that uses 90% fewer tokens while still finding all the expected symbols it builds a local graph of source symbols and their relationships. architecture, decisions and handoffs live as Git-tracked Markdown. a note can point to the code behind it, so a code change can flag that knowledge for review the loop: scan the repo → map the code → write the why → anchor notes to symbols → retrieve a task-sized slice → review drift → commit to Git → continue in the next session in the author's small, one-repo test, graph retrieval returned 10.74× less context than grep top-3 while finding every expected symbol across six tasks the useful part is the connection between memory and evidence. the repo carries what the agent learned, and the code gives you a way to check whether that knowledge still holds save this, then build a second brain for your company⭣show more

beamnxw ./
32,741 views • 4 days ago
whoever leaked this has bigger balls than sense someone... at Anthropic hired 80 AI helpers onto one project, gave them twelve hours, and counted what came back usable: the two older models handed in 980 and 876 finished pieces of work, and almost none of it could be kept turns out the newest helpers did better for a reason nobody wants to hear: they went off into their own corners and stopped opening each other's work i ran two helpers at one document last week and got two confident versions of it, and i kept the one i wrote myself Grok Bot is the version of this you can actually hire: a helper with a name, one job it owns, and nobody else allowed inside that job you already pay about $20 a month for one chat window, and the setup that won in that report is one helper with one job it owns run it tonight in a normal chat, 3 moves: 1. write the one job each helper owns in a single sentence before you open a second chat 2. keep every helper in its own chat with one document, so two of them can never rewrite the same thing 3. add a third only when you can say what it owns without repeating a job that is already taken save this, then open the piece below: what one hired helper is really worth, and the point where the next one starts taking it back ↓show more

Argona
449,029 views • 1 month ago
Everyone shipped an "AI employee in Slack" this summer.... Claude Tag, Grok Bot, Slack AI. All of them wait to be tagged. Mio, launched today, is the first one that starts on its own and the first one the whole team shares: - reads Slack, Notion, Linear, HubSpot and 3,000+ tools overnight - does the work, then DMs you what it did - finance and engineering use the same agent, each with their own permissions - sensitive actions wait for approval - 10,000s of tasks done in beta, 30 seconds to install Your team's AI should know what the rest of your team is working on. This is the first one that does. Beta open, first 100 companies get $100 in credits.show more

Chubby♨️
35,666 views • 14 days ago