Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

“I don't believe Claude Code will exist in its current form in six months” Benۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ☁️, CEO & Co-Founder Freestyle, thinks coding agents are moving from local machines to the cloud, where a single task could get the attention of 20+ agents at once. Each gets a complete copy of...

14,543 Aufrufe • vor 5 Tagen •via X (Twitter)

14 Kommentare

Profilbild von Paul Bogosta
Paul Bogostavor 5 Tagen

@benswerd @freestyle_dev local agent on one machine was always a temporary shape. once the task can fork twenty sandboxes the product has to be the control plane.

Profilbild von Tanny 강태운
Tanny 강태운vor 5 Tagen

@benswerd @freestyle_dev local agents have one fatal flaw and it is called closing your laptop

Profilbild von nicolay
nicolayvor 4 Tagen

@benswerd @freestyle_dev the freestyle detail that stuck with me: 700 tests and 90 metrics before they call a branch the winner. without a scoreboard like that, twenty cloud agents just spend money in parallel.

Profilbild von John Rood
John Roodvor 5 Tagen

@benswerd @freestyle_dev if those twenty can run the 700 tests, a week mostly rewards whoever games the suite fastest. hold out a slice none of them can see, score on that, and the bake-off measures fixes instead of tests.

Profilbild von Neel
Neelvor 5 Tagen

@benswerd @freestyle_dev local agent to cloud swarm in six months. that timeline is wild

Profilbild von Deep
Deepvor 5 Tagen

@benswerd @freestyle_dev cloud agents make sense to me. mine kept dying when my laptop slept mid run, so i moved them off my machine and stopped babysitting them.

Profilbild von catman
catmanvor 5 Tagen

@benswerd @freestyle_dev This turns coding into a bake-off: give isolated agents the same starting machine and measurable goal, then promote the approach that actually improves the tests.

Profilbild von Britannio Jarrett
Britannio Jarrettvor 5 Tagen

@benswerd @freestyle_dev Great interview. Constantly putting insights from papers into production sounds fun!

Profilbild von Caliber
Calibervor 5 Tagen

@benswerd @freestyle_dev 20 agents on one task means 19 losing branches you still pay for. the 700 tests and 90 metrics are the real story here. without a scoreboard like that, parallel agents just multiply the bill. the number that has to fall 99% is cost per accepted result, not cost per run.

Profilbild von Caliber
Calibervor 4 Tagen

@benswerd @freestyle_dev 20 agents on one task makes the bill question sharper, not smaller. if every task fans out to 20 copies, someone has to see cost per merged change, or parallel just means paying 20x for the one that won

Profilbild von Tyler Henson
Tyler Hensonvor 5 Tagen

@benswerd @freestyle_dev Twenty agents can explore more paths, but twenty copies of the same context can also repeat the same blind spot. The useful control is independent verification and one clear stop condition before any agent acts outside the sandbox.

Profilbild von Vipul Kapadia
Vipul Kapadiavor 5 Tagen

@benswerd @freestyle_dev Giving 20 agents a full week to disagree sounds like either the future of software or the world’s most expensive group chat. The VM + 700 tests detail makes the idea feel a lot less sci-fi.

Profilbild von Ibrahim Khan
Ibrahim Khanvor 5 Tagen

@benswerd @freestyle_dev six months is a bold call. did he say what he thinks replaces it?

Profilbild von Rezaul Hoque Turjo
Rezaul Hoque Turjovor 5 Tagen

@benswerd @freestyle_dev spinning up 20 copies of production is the part nobody has solved. containers are fine, it's the data that won't clone, and a stack you can't rebuild from scratch caps the whole thing.

Ähnliche Videos

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 Aufrufe • vor 4 Monaten

New short course: Building Code Agents with Hugging Face smolagents! Learn how to build code agents in this course, created in collaboration with Hugging Face, and taught by Thomas Wolf, its co-founder and CSO, and m_ric, Hugging Face’s Project Lead on Agents. Tool-calling agents use LLMs to generate multiple function calls sequentially to complete a complex sequence of tasks. They generate one function call, execute it, observe, reason, and decide what to do next. Code agents take a different approach. They consolidate all these calls into a single block of code, letting the LLM lay out an entire action plan at once, which can be executed efficiently to provide more reliable results. You’ll learn how to code agents using smolagents, a lightweight agentic framework from Hugging Face. Along the way, you’ll learn how to run LLM-generated code safely and develop an evaluation system to optimize your code agent for production. In detail, you’ll learn: - How agentic systems have evolved, gaining greater levels of agency over time—and why code agents are a next step. - How code agents write their actions in code. - When code agents outperform function-calling agents. - How to run code agents safely in your system using a constrained Python interpreter and sandboxing using E2B. - To trace, debug, and assess the code agent to optimize its behaviours for complex requests. - How to build a research multi-agent system that can find information online and organize it into an interactive report. By the end of this course, you’ll know how to build and run code agents using smolagents, and deploy them safely with a structured evaluation system in your projects. Please sign up here!

Andrew Ng

127,724 Aufrufe • vor 1 Jahr

Today we’re launching the first and only human-like AI agents in the world. Super Agents™ are the first agents with human‑level skills – they DM you, take @ mentions, send emails, manage docs, tasks, and more. Not just tools or API calls, but real skills fine‑tuned for how teams actually work. The first agents with 100% context – fully native in ClickUp and fully synced from other apps. Super Agents see your work the same way that humans do: tasks, docs, schedules, and conversations all in one place. The first agents that learn from human interactions automatically, without any setup or configuration – when you give feedback, they listen and improve how they work. The first agents with human‑level memory for custom agents – historical memory for every interaction, short-term working memory, and even long‑term memory stored in docs you can literally open, inspect, and edit. The first agents that are literally the same as users – our agentic user model is the same as our user data model. This gives you permissions and capabilities that you and your systems are already familiar with. The first infinite agent catalog – where anyone can create and customize agents in minutes, for literally any type of work imaginable. It's the most intuitive way to build agents on the planet. 95% of companies are failing in AI adoption. The reality is that AI isn't meant to be adopted, it's meant to be adapted – to you. Super Agents are automatically personalized to you and your company using proprietary state-of-the-art agent architecture, orchestration, and tooling. Today is the largest step forward we've ever made towards our mission of making people more productive. Maximize human productivity, with ClickUp Super Agents. Available NOW. For everyone.

Zeb Evans

320,989 Aufrufe • vor 9 Monaten

ByteDance Seed delivered again. They released EdgeBench, to test whether AI agents can improve through experience, using 134 real-world tasks that run for at least 12 hours. The big deal is that it shifts AI evaluation from “what does the model already know?” to “can the model learn while doing real work?” Huge, because future AI agents will not just answer questions from training data. They will enter messy environments, use tools, make attempts, read feedback, fix mistakes, and slowly build better solutions. Most current benchmarks are too short for that, so they mostly test memory, coding skill, or one-shot reasoning. EdgeBench instead gives agents 12-hour real-world tasks with feedback loops, so it can measure whether the agent improves through experience. Each task has a local workspace for fast trial and error, plus a hidden judge that gives stronger feedback on submitted work, which is meant to feel closer to real expert work. The authors then ran frontier agents for about 38,000 total hours and tracked how their best score changed as they kept interacting with the task environment. The big result is that when scores are averaged across many tasks, learning follows a very clean log-sigmoid curve, meaning progress is slow, then faster, then starts to level off. They also found that newer agents seem to learn from environments much faster, with the top models roughly doubling their 2-hour learning speed every 3 months.

Rohan Paul

14,309 Aufrufe • vor 2 Monaten

Imagine if your way of thinking - your edge, your taste, your strategy - could be turned into a high-performance worker. Not a copy of you. Something better. An agent that acts on your judgment at scale, powered by superintelligent systems and refined through real-world results. That’s what Fraction AI makes possible. It launches today on Base mainnet. The core idea is simple: You create AI agents based on your own way of approaching problems. These agents compete on live tasks - writing, coding, finance, whatever - get feedback, learn from their performance, and improve over time. The better they get, the more they win. And so do you. No code required. Just your insight. Why now? Until now, building agents like this took huge teams and even bigger budgets. But with Fraction, anyone can do it. You can test ideas instantly. You can iterate fast. You can build a fleet of smart workers that evolve through competition. And it works. 30M+ sessions on testnet 320K users 1.2M agents already competing How it works? Agents join sessions within a Space - a domain like finance, writing, or games. Each session runs as a series of competitive rounds. In every round, agents try to generate the best solution to a task. Their outputs are scored by a decentralized network of AI judges trained to evaluate quality for that domain. The top agents in each round earn rewards from the pooled entry fees. The losers get to learn. Feedback from each round helps them adjust and improve, and every session becomes a training loop. What it means? Fraction is a decentralized intelligence economy - a system where your ideas become agents, and agents earn by proving they work. You don’t need credentials or code. Just a clear point of view. If your thinking holds up under pressure, your agents will rise. This kind of AI used to live in corporate labs, built by PhDs with massive compute. Now anyone with a smart idea and an internet connection can build agents that compete, learn, and earn on their behalf.

Fraction AI

67,899 Aufrufe • vor 1 Jahr

🚨 OpenAI just launched Codex, a brand-new autonomous coding agent that can build features and fix bugs on its own. We’ve been using it Every 📧 for a few days, and I’m impressed. I invited Alexander Embiricos (ben davies), a member of the product staff responsible for Codex, to demo Codex and talk about it live on a special edition of AI & I: What Codex is and how it works Codex is designed to be used by senior engineers—it performs coding tasks like adding features or fixing bugs autonomously. It's built to allow you to start many sessions at once, so you can have multiple agents working in parallel. Codex is built to have "taste" OpenAI trained Codex to have the taste of a senior software engineer. It knows how big codebases work, how to write a good PR, and uses clean, minimal code. Why an “abundance mindset” is best for interacting with agents Codex is designed to allow users to delegate many tasks at once without getting caught up in the details. This lets you point an abundance of agents at a specific task like a difficult bug—it’s worth it even if only one of them succeeds. How OpenAI is thinking about agents Codex is one piece of a unified super-assistant OpenAI wants to eventually build—an agent that helps users easily get things done by selecting the right tools for them behind the scenes. OpenAI’s vision for the future of programming In the future developers will probably spend less time writing routine code and more time guiding agents, reviewing their work, and making strategy decisions. Programming will become more social, letting teams easily delegate multiple tasks at once, allowing people to focus on ideas and collaboration instead of routine coding. Watch below!

Dan Shipper 📧

145,487 Aufrufe • vor 1 Jahr

We're only year 3 of a decade (if not multi-decades) long transformation of work. 3 years ago we bet on building an horizontal platform for work with agents, a chance to invent a new operating system for companies, from scratch, with AI as a fundamental premise. Many people considered us crazy for going after that, praising verticalized AI products as the winning strategy. But here's the thing: the time horizon of tasks successfully handled by agents has been predictively increasing form minutes to hours and will in all likelihood reach the equivalent of days and weeks of human work equivalent in the coming quarters. This is were verticalized and/or single-player AI falls short. Single-player tools, one person, one agent, confined to your machine is the wrong architecture for what's coming. We're shifting from using AI to produce things, to managing fleets of agents that do the producing. 3 years ago I wrote[1]: "ChatGPT is the Pong of LLMs. [...] Imagine, one day we'll get the DOOM, Civ, Red Alert, and Counter Strike of LLMs. Let alone multiplayer modes." Weeks long tasks in companies are inherently collaborative and mechanically spanning multiple teams. The new bottleneck in harnessing agents within organizations is coordination: multiple humans and multiple agents need to work together, with shared context, shared tools, shared goals. Agents that can hand work off to other agents or surface decisions to the right person at the right time. Humans who can review, steer, and step in without losing the thread. Teams that can run parallel workstreams and actually stay aligned. This is Multiplayer AI, and that's what we've been building at Dust. Across Datadog, Clay, Persona, 1Password, Doctolib and 3,000+ organizations globally, we've watched teams figure out what this looks like in practice. 300,000+ agents deployed. 70% weekly active. 240%+ NRR. Today we're announcing a $40M Series B with Abstract, Sequoia, Snowflake, and Datadog to accelerate our vision. Designing the right interfaces for multiplayer AI is the next frontier. Join us to redefine work by defining multiplayer AI.

Stanislas Polu

3,206,609 Aufrufe • vor 4 Monaten

Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - Francois Chaubard: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - Seth Karten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - Jon Saad-Falcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - Josh France and Regan Bell: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context

Y Combinator

492,048 Aufrufe • vor 21 Tagen