正在加载视频...
视频加载失败
“I don't believe Claude Code will exist in its current form in six months” Benۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ☁️, CEO & Co-Founder Freestyle, thinks coding agents are moving from local machines to the cloud, where a single task could get the attention of 20+ agents at once. Each gets a complete copy of... show more
14,543 次观看 • 5 天前 •via X (Twitter)
14 条评论

@benswerd @freestyle_dev local agent on one machine was always a temporary shape. once the task can fork twenty sandboxes the product has to be the control plane.

@benswerd @freestyle_dev local agents have one fatal flaw and it is called closing your laptop

@benswerd @freestyle_dev the freestyle detail that stuck with me: 700 tests and 90 metrics before they call a branch the winner. without a scoreboard like that, twenty cloud agents just spend money in parallel.

@benswerd @freestyle_dev if those twenty can run the 700 tests, a week mostly rewards whoever games the suite fastest. hold out a slice none of them can see, score on that, and the bake-off measures fixes instead of tests.

@benswerd @freestyle_dev local agent to cloud swarm in six months. that timeline is wild

@benswerd @freestyle_dev cloud agents make sense to me. mine kept dying when my laptop slept mid run, so i moved them off my machine and stopped babysitting them.

@benswerd @freestyle_dev This turns coding into a bake-off: give isolated agents the same starting machine and measurable goal, then promote the approach that actually improves the tests.

@benswerd @freestyle_dev Great interview. Constantly putting insights from papers into production sounds fun!

@benswerd @freestyle_dev 20 agents on one task means 19 losing branches you still pay for. the 700 tests and 90 metrics are the real story here. without a scoreboard like that, parallel agents just multiply the bill. the number that has to fall 99% is cost per accepted result, not cost per run.

@benswerd @freestyle_dev 20 agents on one task makes the bill question sharper, not smaller. if every task fans out to 20 copies, someone has to see cost per merged change, or parallel just means paying 20x for the one that won

@benswerd @freestyle_dev Twenty agents can explore more paths, but twenty copies of the same context can also repeat the same blind spot. The useful control is independent verification and one clear stop condition before any agent acts outside the sandbox.

@benswerd @freestyle_dev Giving 20 agents a full week to disagree sounds like either the future of software or the world’s most expensive group chat. The VM + 700 tests detail makes the idea feel a lot less sci-fi.

@benswerd @freestyle_dev six months is a bold call. did he say what he thinks replaces it?

@benswerd @freestyle_dev spinning up 20 copies of production is the part nobody has solved. containers are fine, it's the data that won't clone, and a stack you can't rebuild from scratch caps the whole thing.
