Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the...

491,785 görüntüleme • 21 gün önce •via X (Twitter)

35 Yorum

Y Combinator profil fotoğrafı
Y Combinator21 gün önce

Tune in:

Francois Chaubard profil fotoğrafı
Francois Chaubard21 gün önce

autoresearch repo discussed in intro:

Peter Bordes profil fotoğrafı
Peter Bordes21 gün önce

@truefoundry just launched a fantastic next-generation open source agent harness. Paired with their AI Gateway lets enterprise teams build, deploy, and govern production agents on any model or MCP server. As an alternative to Claude Managed Agents at 50% lower cost.

Tereza Tizkova profil fotoğrafı
Tereza Tizkova21 gün önce

Harness doesn’t matter until @FactoryAI completes task faster than Claude for cheaper

Florian Brand profil fotoğrafı
Florian Brand21 gün önce

@sethkarten Seth 🐐

Raven profil fotoğrafı
Raven21 gün önce

30% weights, 95% harness, one extremely employed horse

Francois Chaubard profil fotoğrafı
Francois Chaubard21 gün önce

Next YC Paper Club Signup: Alternative Compute Paradigms 9.23 5pm Mountain View

Scott N profil fotoğrafı
Scott N21 gün önce

Same weights, triple the score. The wrapper is the weapon system. A round without fire control is just an expensive rock.

Flo profil fotoğrafı
Flo21 gün önce

Most understand how important the harness is, once the project grows beyond a certain size. It’s impossible to manage a large project without a harness built around it. Some think a harness shouldn’t even do context management and let the model figure it out each time, VERY WRONG

StartupHakk profil fotoğrafı
StartupHakk21 gün önce

We have been preaching this same thing for months!! What we have pulled off with an Open Source, freez harness ( has blown our minds. We are building real platforms around it which are competitive to frontier level with hardware you run under your desk!

Fitih "Fitz" Cinnor profil fotoğrafı
Fitih "Fitz" Cinnor21 gün önce

Harnesses were the missing layer to unlock deeper coordination and execution across tasks and domains. Model advancement will likely obviate the need over time. For now, structured problem solving requires imposed structure and/or guidance from outside.

Sebastian the Euro VC profil fotoğrafı
Sebastian the Euro VC21 gün önce

In Europe we have used harnesses for 400 years. We call them lederhosen. Incredible retention, zero scale.

Laurent Bindschaedler profil fotoğrafı
Laurent Bindschaedler20 gün önce

Some of us academics have been singing this song for over a year, and we have prototypes with cool ideas that go beyond some of what you are showing here. We would welcome more academic engagement and comparing notes if that is ever of interest 🙂

cackles (jeff weisbein) profil fotoğrafı
cackles (jeff weisbein)21 gün önce

using agents every day has made this pretty obvious to me. a lot of the progress isn’t switching models. it’s getting them to remember what we already figured out and finish things without me having to keep checking in.

dFusion AI Protocol profil fotoğrafı
dFusion AI Protocol21 gün önce

People underestimate the harness layer.

Varun Mundra profil fotoğrafı
Varun Mundra21 gün önce

@garrytan Yes they are powerful. Open sourced this one- Build a @PalantirTech of your own!

andoni profil fotoğrafı
andoni21 gün önce

Todo el mundo debería verlo

ilyas profil fotoğrafı
ilyas21 gün önce

interesting but having thousands of harnesses doing the same thing is what the problem is. if someone makes groundbreaking harness thats very cool

atharva ☆ profil fotoğrafı
atharva ☆21 gün önce

@PrajwalAvhad8

Jim Monge profil fotoğrafı
Jim Monge21 gün önce

Very educational. I've built a couple of harnesses myself, and the process is honestly not that easy. Building complexity must be addressed in the same way that the creation of AI agents was simplified

Marko Tasic profil fotoğrafı
Marko Tasic21 gün önce

This doesn't make any sense. If model improves in general, it should work across different harnesses in general. We are so far from AGI as we were back in 2022.

Pitch profil fotoğrafı
Pitch21 gün önce

Scaffolding is compute.

Yashas profil fotoğrafı
Yashas21 gün önce

@harjtaggar Oh cool, I'd love seeing this video tonight looks fun, would be cool to explore stuff as I'm building my harness too!

Fajar M Reza profil fotoğrafı
Fajar M Reza21 gün önce

Harness quality can change benchmark scores, making evaluation a systems design problem.

Reyaz profil fotoğrafı
Reyaz21 gün önce

Same weights, different harness, 30 to 95. That gap is basically the distance between 'can't' and 'can if you structure the problem well. ' How many capabilities are dormant just because the evaluation wrapper is lazy?

Atomic Strata profil fotoğrafı
Atomic Strata20 gün önce

Persistent memory and feedback loops are what turn model output into something teams can actually rely on.

Lorenzo Price profil fotoğrafı
Lorenzo Price21 gün önce

The ARC-AGI benchmark illustrates the shift perfectly. Look at the performance gap when companies like OpenAI switched from raw prompting to agentic loops; the "o1" series effectively functioned as its own harness to hit those 90%+ reasoning scores.

Adam Elkassas profil fotoğrafı
Adam Elkassas21 gün önce

from back in May

Arnold profil fotoğrafı
Arnold20 gün önce

A 30% to 95% jump makes the harness part of the model system, not mere scaffolding. How do you test that the gain is not benchmark-specific overfitting?

Aniket Kadam profil fotoğrafı
Aniket Kadam21 gün önce

This is a fascinating shift in how we think about AI agents. The model itself may not be the whole story the harness around it can unlock a huge amount of capability

Visiting Fellow, Ph.D. profil fotoğrafı
Visiting Fellow, Ph.D.21 gün önce

This was a great overview.

Scott 👀 base.eth profil fotoğrafı
Scott 👀 base.eth21 gün önce

@garrytan true, the infrastructure work is honestly the most underrated part.

Hardik Hindocha profil fotoğrafı
Hardik Hindocha21 gün önce

the harness IS the product. same thing in outbound - same list, same copy, better sequencing infra and replies go 3-4x. nobody wants to admit the wrapper is where the value lives

Keep Your Head profil fotoğrafı
Keep Your Head21 gün önce

If the same weights go 30 to 95 on harness alone, the benchmark is measuring the system, not the model. Most model comparisons still ignore that.

Hezo profil fotoğrafı
Hezo20 gün önce

💯 The harness makes a huge difference which is why we've spent a lot of time iterating towards the perfect one in Hezo.

Benzer Videolar

In the future, you’ll be able to accomplish a goal by just giving Claude an outcome and a budget. That’s the direction Anthropic is building in with its new Managed Agents features, announced at this week’s Code with Claude developer event. The basic idea: Claude, wrapped in a computer in the cloud, that you can spin up, scale, and manage as needed. Anthropic is taking on the infrastructure that kills most agent products, and making sure that it scales to meet the needs of agents running 24/7. On this week’s AI & I from Every 📧, I talk with Angela Jiang (Angela Jiang), head of product for the Claude platform, and Katelyn Lesse (Katelyn Lesse), head of engineering for the Claude platform, about what Anthropic is building and what it takes to make agents reliable in production. We get into: - Why the "build a generic harness, hot-swap any model behind it" playbook is already outdated. Angela points to eval data on Memory where the same task across different harnesses performed drastically differently. - The infrastructure wall every team hits in production—and why Katelyn thinks “my sandbox died and took the agent with it” is the real reason internal agents don't ship. - Why Anthropic is so bullish on using file systems and skills within Claude, including Angela's argument that those early design choices can compound for years. This is a must-watch for anyone trying to take an agent past the demo and into production. Watch below! Timestamps: How the Claude platform evolved from API to agents: 00:01:48 The primitives that make up Claude Managed Agents: 00:04:09 Why the harness and the model are becoming a single unit: 00:10:37 The infrastructure wall that kills most agent projects in production: 00:18:49 Why team agents need a different shape than individual productivity tools: 00:24:49 How Anthropic's legal team uses an agent to review marketing copy: 00:26:36 Using multi-agent orchestration for advisor strategies, adversarial pairs, and swarms: 00:34:24 How to measure agent success with outcome and budget as the end state: 00:35:50 What the platform looks like a year from now, when Claude writes its own harness: 00:39:11

Dan Shipper

66,871 görüntüleme • 4 ay önce

An agent is three things: a harness, a model, and context. If you're serious about owning your intelligence, you probably want to own all three. LangChain founder Harrison Chase joined us at our Sequoia Capital Own Your Intelligence to talk about the piece that often gets the least attention: the harness. He offers a clear heuristic for when to build your own. The more out of distribution you are from what the models were trained on, the more you'll want to customize. And good technical content on how to actually measure performance with evals and langsmith. 00:00 Introduction 00:58 The three parts of an agent: harness, model, context 02:12 What a harness actually does 03:25 Customizing the core loop with middleware 04:41 Sandboxes, file systems, sub-agents, summarization 05:47 Cognitive architectures — and when you still need them 07:03 Build your own harness or use off the shelf? 08:24 In-distribution vs. out-of-distribution: the file-editing example 09:39 Why evals define what "good" means in an organization 11:04 Harbor: what an eval task actually looks like 12:11 Comparing harnesses and models on accuracy, latency, and cost 13:20 Why observability is underrated — it's usually the context 14:34 The data flywheel: traces → curation → experiments 15:42 Getting feedback through UX design and online evaluators 16:51 Demo: LangSmith Engine 19:23 Q&A: Running Engine on Engine, and "codex-ification" 20:44 Q&A: Will harnesses converge or diverge?

Sonya Huang 🐥

77,519 görüntüleme • 1 ay önce

“I don't believe Claude Code will exist in its current form in six months” Benۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ☁️, CEO & Co-Founder Freestyle, thinks coding agents are moving from local machines to the cloud, where a single task could get the attention of 20+ agents at once. Each gets a complete copy of your stack, production environment included. Each can spend a week testing and refining its approach. You compare the results and take the best one forward. The economics aren’t there yet. Ben expects cost per task to fall another 99% over the next four years, making that level of parallel work practical. Freestyle builds the computers for it: full Linux VMs for tasks that run for hours, days, or weeks. Clone a running machine, memory included, and let agents pursue different approaches from the same starting point. Pause and resume with their state intact. Inside Freestyle, Ben already gives agents a week to find improvements to its VM technology. Roughly 700 tests and 90 metrics help the team judge whether the work made things better. He calls this goal engineering: define the outcome clearly, give agents time to work toward it, and measure whether they’re making progress. "Full Linux VMs for AI Agents" Built for complex tasks that run for hours, days, or weeks 🎙️ Benۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗۗ☁️, CEO & Co-Founder, Freestyle on Fondo.com START 00:57 Why coding agents could move from local to the cloud 02:22 From harness engineering to goal engineering 03:08 Giving agents a week to improve measurable results 04:13 Why defining the problem becomes the bottleneck 04:48 How early access to o1 changed Freestyle’s direction 05:51 Why frustrated sandbox users revealed a bigger opportunity 07:28 Giving agents a computer instead of building custom tools 08:12 Getting into YC 11:17 Snapshotting VMs and running agents in parallel 12:24 The economics of letting multiple agents attempt every task 13:26 How GPU supply could drive down agent costs 15:09 Try Freestyle and follow Ben

David J Phillips

14,543 görüntüleme • 5 gün önce