Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Vercel just solved the biggest problem with trusting agents to write your code. yesterday they shipped Foreman - an open-source software factory. 58 stars. Nobody's seen it yet. The idea? Instead of one agent writing, reviewing and merging its own work - rubber-stamping its own bugs… Foreman splits the...

126,235 Aufrufe • vor 1 Monat •via X (Twitter)

25 Kommentare

Profilbild von Jurly
Jurlyvor 1 Monat

The separate-vendor reviewer reduces shared blind spots without pretending agents are trustworthy.

Profilbild von Granite
Granitevor 1 Monat

you got the point right, congrats

Profilbild von Argona
Argonavor 1 Monat

trying this before it blows up

Profilbild von Granite
Granitevor 1 Monat

def do it! let me know what you think after, your take

Profilbild von Fusion Software Factory
Fusion Software Factoryvor 1 Monat

Fusion.

Profilbild von Trung Le
Trung Levor 1 Monat

I automate this Loop with RealtimeX Loop I shared it with @Scobleizer here

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

different vendor reviewing is smart

Profilbild von Granite
Granitevor 1 Monat

yeah, thats the part that actually matters. a reviewer that never sees the implementers reasoning cant just agree with it -_-

Profilbild von AI Mastery Guide
AI Mastery Guidevor 1 Monat

Fresh eyes catch real flaws

Profilbild von Gal Vered
Gal Veredvor 1 Monat

We tested it at Checksum and it did not go well. Reviewers without context often do worse

Profilbild von Ben Sabic
Ben Sabicvor 1 Monat

Thank you so much for sharing the template! Let me know if you run into any issues or have feature requests.

Profilbild von Granite
Granitevor 1 Monat

thanks for making this! alright, if i notice an issue beyond just my end, i'll message you

Profilbild von pms190
pms190vor 1 Monat

This story is months old. Check out

Profilbild von Bally_AgenticAI
Bally_AgenticAIvor 1 Monat

Different vendor only buys independence if failures are uncorrelated - models trained on the same public code share blind spots. The real gate is a signal the implementer never had: tests, runtime traces, prod telemetry. #AgenticAI #MultiAgent @vercel

Profilbild von Kai Lennox
Kai Lennoxvor 1 Monat

The different-model reviewer is the part that really matters.

Profilbild von Miles AI Wizard
Miles AI Wizardvor 1 Monat

Wait, the reviewer never sees the implementer's reasoning? Only the pushed branch. That's actual independence.

Profilbild von Renzo
Renzovor 1 Monat

Mmm I’d rather use

Profilbild von Emanuel Farauanu
Emanuel Farauanuvor 1 Monat

I feel like the "Analyst" role is the weak link here. I can see how most of the steps can be fully automated, but at the moment I still feel like I need to spend ages planning correct implementations with the grill-me skill.

Profilbild von Dani Acosta
Dani Acostavor 1 Monat

For me this don’t replace a Dynamic Workflow on Claude I think is interesting to see more of these harnessing in the open and probably we will never see (again) the code of Claude Code that is their only moat as models get commodified

Profilbild von Emad Ghorbaninia
Emad Ghorbaniniavor 1 Monat

Splitting reviewer and implementer across different model vendors so neither can rationalize the other's shortcuts is the first agent-pipeline design I've seen that assumes the agents will collude if you let them. Correct assumption, underrated fix.

Profilbild von AI Alpha By Mario
AI Alpha By Mariovor 1 Monat

The separation between builder and reviewer agents feels like the key. Curious how Foreman handles conflicting review signals

Profilbild von Max Bevza
Max Bevzavor 1 Monat

vercel cooking up literal magic here

Profilbild von Kristoph
Kristophvor 1 Monat

This is just not that revolutionary.

Profilbild von neuralforge
neuralforgevor 1 Monat

how about its memory?

Profilbild von oscar
oscarvor 1 Monat

No one's using this because it would be wildly expensive and generally impractical for most organizations and developers. Most people need something like instead.

Ähnliche Videos

WHAT IS AN AI "SOFTWARE FACTORY" AND IS IT HYPE (31 MINUTE BREAKDOWN) I think it's a silly name for a genuinely USEFUL idea! A software factory is 5-6 markdown files that sit next to your code and tell your agents how you like to work, so you can build high quality apps 24/7. It's going viral because AI coding has a trust problem. The model can build the feature, but with no structure around it you end up babysitting the agent, wondering what changed and hoping it didn't break something important. So you build with agents the same way a factory builds physical products! 1. Each feature gets its own station, which in software means its own branch, so multiple agents can work at the same time without stepping on each other. 2. The build station gives the agent rules for how to write the code, because "it works" is very different from "a developer could open this repo next month and understand what happened." 3. The proof station makes the agent show evidence. Screenshots, videos, speed numbers, before-and-after states. It has to prove the thing works instead of saying it works. 4. The review station runs the work through a code review agent, and if it doesn't clear the bar, it goes back through the line. 5. Then you show up at the end to merge. For a 100+ years people have run production this way, and it worked because the structure is good. The full episode on what’s a software factory is NOW live on The Startup Ideas Podcast (SIP) 🧃 with the wonderful Micky Watch: So is it hype?!? I don't think it is, because of what it does to your output! WITHOUT a factory, you build ONE feature at a time and you're the bottleneck at every step, prompting, checking the diff, testing it yourself, hoping nothing else broke (spoiler alert it often does). WITH a factory, EACH feature runs in its own isolated copy of the app, so you can have 10+ of them going at once, and each agent has to prove its own work and pass a code review before it ever reaches you. Instead of supervising the work, you're APPROVING finished work that already has evidence attached. REALLY interesting to see how work with agents is evolving to be….well, similar to working with people!

GREG ISENBERG

30,787 Aufrufe • vor 20 Tagen

THIS GUY CONNECTED HIS AI AGENTS TO HIS OBSIDIAN AND BUILT A BRAIN THAT LEARNS ON ITS OWN. HERE'S HOW TO BUILD IT Obsidian is just markdown files sitting in a folder. That turns out to be the perfect memory for an AI agent, because an agent can read and write those files directly. He wired his agents into the vault so they pull context from it, do the work, and write what they learned back. The notes aren't the point. The loop is, and it gets sharper every cycle How to build it: 1. Point an agent at your vault. The fastest way, no plugins, no API keys: open a terminal and run npx obsidian-mcp /path/to/your/vault. That exposes your Obsidian folder to Claude as a tool it can read, search, and write to. Add it to your Claude Code or Cowork config and restart 2. Confirm it can see the brain. Ask it: "list the notes in my vault and summarize what's in them." If it reads them back, the connection is live. Now it starts every task with everything the vault already holds instead of from zero 3. Give each agent one job and a write-back rule. Tell it: "research this, then save what you found as a new note in /brain with links to related notes." One agent researches, one summarizes, one plans. Each writes its output back into the vault 4. Close the loop. Add one line to every agent's instructions: "read /brain before starting, write your result back when done." Now each task leaves the vault richer, and the next run reads that before it works. It compounds instead of resetting 5. You only steer. Review what the brain produces, point it at the next thing. The agents handle the reading, writing, and connecting The edge isn't better notes. It's a brain that feeds itself, so the work gets sharper every cycle instead of starting over Bookmark this

Yarchi

58,591 Aufrufe • vor 3 Monaten

BlackRock runs on 20,000 people. Elon's Grok Bot runs the same shape for $300 a month, and it hires its own staff. You do not get an assistant. You get a company that hires. It does not throw ten agents at your problem and hand you the pile. It makes one agent that makes 10, and those ten make a 100. > LAYER ONE is one agent, the chief of staff, and it never touches the market > LAYER TWO is six desk heads, one job each, every one on its own computer with its own logins > LAYER THREE is whatever those six decide they need, spun up on the spot and shut down when the work is done Nobody writes a task list. You hand out job titles and the org fills itself in underneath. The swarm is never the same twice. Agents get spun up for one job, finish it, and are gone before I ever read their names. Not one of them sees the whole picture. The answer only exists after they hand off to each other. Wall Street cannot copy that. You cannot hire a hundred people for eleven minutes. BlackRock holds that shape together with a risk system called Aladdin. Mine holds it together with one agent that is only allowed to say no. I gave it $1,000 and told it to grow the money or get deleted. 15 hours later it was holding $3,900, on an address anyone can open and read. I was asleep for most of it, and I have still not written a line of code. The whole thing runs with my laptop shut, because none of it lives on my laptop. Setup is one evening. Create the chief, hand out the titles, run one trade on your screen while they watch, connect Telegram. Ten years ago a machine this shape had its name on a tower. Mine has a name I typed into a box. Save this while the whole thing still fits on one screen.

cvxv666

45,488 Aufrufe • vor 1 Monat

Every AI agent you've tried has amnesia. It does one task, forgets everything, and tomorrow you start from zero. That's not an employee. That's a temp you have to retrain every single morning. Hyperagent by Airtable is the first platform I've used that actually fixes this. Here's what got me: 1. Agents that compound. Each agent has memory. The one running today is smarter than the one you shipped three weeks ago. Same prompt, same integrations, but weeks of your judgment baked in. 2. Real deliverables, real receipts. You don't get a chat transcript. You get finished work with the cost and runtime printed right on it. A full research report for under ten bucks. Try getting that invoice from an agency. 3. A fleet, not a chatbot. Build a specialist for outreach, another for research, another for reporting. Give each one its own tools, its own memory, and its own budget cap so nothing runs away with your credits. 4. Deploy to Slack and your whole team uses the agent you built. One competitive intel agent, @ mentioned by everyone. Airtable runs its own data team this way. 5. Each agent gets its own cloud machine with a real browser and code execution. It works while you sleep. No babysitting, no local setup, no laptop that has to stay open. I put it to work in the video below. Watch what it builds. The teams treating agents as durable assets instead of one-off prompts are going to lap everyone else. This is the first tool that actually treats them that way. #ad Hyperagent

Leonard Rodman

94,961 Aufrufe • vor 2 Monaten

everyone in iOS development should watch this. seriously, it might change the whole industry. i pointed claude code at a live ios device running on revyl, typed "test everything," and walked away. here's what's actually happening: ① you don't write the tests. no scripts, no selectors, no test plan. i never told it which screens to open or what to check. it read the app, decided what mattered, and tested it. the entire instruction was "test everything." ② it built its own test team. it looked at the app, clocked that it's basically four mini apps (rides, delivery, services, account), and split itself into 4 agents, one per surface. scoping coverage like that is usually a person's whole afternoon. it did it in seconds, unprompted. ③ all four ran at the same time, each on its own live device. this is where revyl comes in. every agent gets its own live ios session in the cloud, so four running apps get tested in parallel instead of taking turns on one simulator. serial testing turns coverage into a time tax. running all of it at once removes the tax. ④ it tests like a person, not like a script. each agent drives the app the way a user would, taps through the flows, and visually checks each screen against what it expected to see. nothing is pinned to a brittle element id, so renaming a button doesn't take down half your suite. that one detail is the most annoying thing about how we test today, and it just quietly goes away. ⑤ no xcuitest, no sims melting your laptop. i didn't write a single xcuitest script, and there were no simulators booting on my machine. the agents run on cloud devices, so coverage stops being capped by what your laptop can handle. the part that got me isn't that an agent tested an app. it's that i never told it how. i handed it a device and an intent, and it figured out the scoping, the parallelizing, and the driving on its own. if you still write and maintain mobile ui tests by hand, i'm not sure that lasts the year.

Landseer Enga

23,963 Aufrufe • vor 4 Monaten

Introducing Headlong, an open source microharness for persistent agents: self-guided agents that think continuously. Most agent harnesses are reactive: you send a task, the agent completes it, and then it sits frozen until the next request. Cron jobs and heartbeats wake it up to run a checklist and put it back to sleep. A Headlong agent is never asleep. It keeps generating thoughts about whatever it decides is interesting, in a self-guided loop inspired by human inner monologue. Your message doesn't start a session. It's one more observation that lands in the agent's thought stream, and the agent decides if and when to reply. Headlong is built on the idea of persistent agency: continuous inner thought generation between external interactions. The agent sets its own interests and priorities, comes up with its own projects, and sometimes pings you unprompted with progress. To keep our prototype as simple and small as possible, we implemented Headlong as a microharness: a complete agent harness in under 10K lines of Bash, organized as a handful of small executables. It includes a loop that generates the next thought, shellm (a recursive language model written in Bash), a trajectory stored as a DAG of jsonl files, and context as a projection of that trajectory. We've been running one Headlong agent internally at Laude for several weeks. The whole team talks to it over Slack and Telegram, and every conversation lands in its single stream of thought. It works in its own fork of Headlong and we've pulled over 50 of its commits into main. One night, with nobody talking to it, it went back to check whether a recall process it had built was actually wired into its mind, found that it wasn't, diagnosed and fixed the bug, and verified the fix end to end. 48 minutes, no human asked for the fix or was in the loop at any point. Every step is a timestamped line in its log. Things broke too, and we wrote those up. Background thinking costs us $1 to $2 an hour, our agent stopped its own service three times by accident, and self-delegation died on day one. Details in the post. One line installs everything and starts an agent. Use a dedicated sandbox and spend-capped API key; it runs real shell commands and thinks around the clock. Headlong is research software, be careful! curl -fsSL | bash Launch post: Repo: Headlong is a Laude Institute / MIT collaboration.

Andy Konwinski

361,750 Aufrufe • vor 1 Monat