Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. LangChain founder Harrison Chase joined us at our Sequoia Capital Own Your Intelligence to talk about the piece that often gets the least attention:...

77,519 görüntüleme • 1 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the frontier to do a deep dive into the state of harnesses. We cover how we got to this point, the case for making your harness as expressive as possible, and what YC learned building an agent for every employee in the company. 00:00 - Francois Chaubard: Why harnesses matter 04:27 - Building an auto-researcher by accident 07:13 - A five minute history of harnesses 13:56 - Self-improving harnesses 18:35 - Seth Karten: Prime Agent, a self-improving RLM harness 21:50 - Context as an L1, L2, L3 cache 24:51 - From Turing machine to von Neumann computer 28:33 - Messaging between agents 30:04 - ARC-AGI results 33:09 - Emulator Bench and GPU kernels 37:30 - Jon Saad-Falcon: OpenJarvis, personal AI on personal devices 38:26 - How far behind are local models 39:21 - The five primitives of a personal AI stack 42:47 - Letting cloud models optimize your local stack 43:53 - 800x cheaper than the cloud 45:58 - Josh France and Regan Bell: QM, YC's agent harness for work 47:29 - A history of YC's internal agents 49:24 - OpenClaw and a fleet of 50 agents 51:04 - Pulling the brain out of the sandbox 54:43 - Letting the agent choose its own sandbox and model 57:16 - The grind tool: budgets on goals 58:50 - Agents don't understand social context

Y Combinator

492,800 görüntüleme • 25 gün önce

OWN YOUR INTELLIGENCE Last year, building on open-weight models was primarily a cost rationalization exercise. Slightly worse performance for a much cheaper price. Now, it is increasingly an existential and strategic topic for our portfolio. Intelligence is the product. Companies want to shape it and own it and let it compound within their own walls. Not your weights, not your product. Now, with frontier open-weight models and fantastic tooling/infrastructure, owning your intelligence at the frontier is finally becoming possible. The result: every application company we work with is embarking on the journey of doing their own research on post-training, evals, harnesses, etc. The hottest neolabs may just be Harvey, Factory, RamPrasad "RamP!" Moudgalya, etc. The list goes on. We held a summit Sequoia Capital to convene our portfolio on this topic, together with Gabe Pereyra (Harvey) on building Harvey Labs, Lin Qiao (Fireworks) on post-training, Harrison Chase (LangChain) on harnesses + evals, Brendan (can/do) () on RL environments and synthetic data, Arjun Karanam (Trajectory) on online continual learning. Opening talk below; rest to come this week! 00:00 What is sovereign AI (and what it isn't) 01:24 Centralized vs. decentralized intelligence 02:54 Four reasons companies own their models: cost, speed, performance, destiny 04:22 "Not your weights, not your product" 05:32 The application companies are the newest neo labs 07:05 Step 1: Deciding what to own vs. rent 09:51 Step 2: Build the team (and don't shoehorn your platform team) 11:17 Step 3: Legibility – why your research has to be visible 12:33 Step 4: The technical roadmap 13:56 The stack: production vs. development 15:16 Opening Pandora's box – base models, harnesses, context

Sonya Huang 🐥

128,779 görüntüleme • 1 ay önce

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 görüntüleme • 4 ay önce

Finally, an open-source runtime security layer for your agent harnesses. AI agents can call tools, run shell commands, modify files, and interact with production systems. But when something goes wrong, teams often have to reconstruct what the agent actually did from scattered logs after the fact. Agent Beacon is built to change that. It runs locally and records agent activity as it happens, including: → Tool calls → Shell commands → File changes → Approval decisions → Session and harness context The useful part is that it normalizes all of this into the same event format across 23+ agent harnesses. So instead of writing different detection logic for Claude Code, Codex, or another harness, security teams can reason about the underlying action itself. A command executed is a command executed, regardless of which harness produced it. Beacon also records how confidently an event was captured. An action can be directly observed from the runtime or inferred from indirect evidence. That matters when you start writing security rules against this data. For example, you can detect when an agent reads an environment variable, modifies a sensitive file, executes a suspicious command, or performs an action without the expected approval. And the detection happens while the session is still unfolding, not days later during incident reconstruction. Everything runs locally by default. You can inspect sessions through the dashboard, write your own detection rules, or forward the same normalized events to tools like Splunk, Datadog, Elastic, Sentinel, or CrowdStrike. Beacon GitHub: (don't forget to star 🌟) I also wrote a detailed article on the same, covering how important runtime security is and how to build it for your agent harnesses. The article is quoted below.

Akshay 🚀

27,859 görüntüleme • 22 gün önce

Watch what Nadella did on the Microsoft earnings call tonight. An analyst asked how Microsoft benefits from enterprises adopting open models when it carries all that frontier lab exposure. Instead of defending the lab relationship, Nadella laid out an architecture: "You've got to keep your harness separate from the model. The harness will ensure that your memory, your context, all of that is external. That means any given model at any given time is swappable." And then the part that should worry anyone underwriting model moats: use frontier models where they earn it, low-cost models where they don't, "and in fact, train your own model when you don't want to use any external model itself because after all, you have all the outputs, you have all the traces, you have all the context." The firm keeps the harness. The models compete for slots inside it. Jensen Huang said most companies will be built on harnesses at the LangChain fireside on July 8. I published the full framework on July 12, five launches in three days, all converging on the same architecture. Tonight the largest enterprise software company on earth made it the official pitch on an earnings call. The moat question in enterprise AI just moved from who has the best weights to who owns the loop around them. And if the completed task is the unit everyone now competes on, someone has to price it for the buyer. That is the next thing we are building at BEP Research: a cost per task tool for enterprises. More on that soon. I also took the paywall off the full framework piece tonight, so the whole thing is free to read:

Ben Pouladian

53,433 görüntüleme • 2 ay önce