Video yükleniyor...
Video Yüklenemedi
Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be farther from the truth. The same model weights that score 30% on ARC-AGI score 95% with a better harness. So we gathered a group of researchers and founders working at the... show more
491,785 görüntüleme • 21 gün önce •via X (Twitter)
35 Yorum

Tune in:

autoresearch repo discussed in intro:

@truefoundry just launched a fantastic next-generation open source agent harness. Paired with their AI Gateway lets enterprise teams build, deploy, and govern production agents on any model or MCP server. As an alternative to Claude Managed Agents at 50% lower cost.

Harness doesn’t matter until @FactoryAI completes task faster than Claude for cheaper

@sethkarten Seth 🐐

30% weights, 95% harness, one extremely employed horse

Next YC Paper Club Signup: Alternative Compute Paradigms 9.23 5pm Mountain View

Same weights, triple the score. The wrapper is the weapon system. A round without fire control is just an expensive rock.

Most understand how important the harness is, once the project grows beyond a certain size. It’s impossible to manage a large project without a harness built around it. Some think a harness shouldn’t even do context management and let the model figure it out each time, VERY WRONG

We have been preaching this same thing for months!! What we have pulled off with an Open Source, freez harness ( has blown our minds. We are building real platforms around it which are competitive to frontier level with hardware you run under your desk!

Harnesses were the missing layer to unlock deeper coordination and execution across tasks and domains. Model advancement will likely obviate the need over time. For now, structured problem solving requires imposed structure and/or guidance from outside.

In Europe we have used harnesses for 400 years. We call them lederhosen. Incredible retention, zero scale.

Some of us academics have been singing this song for over a year, and we have prototypes with cool ideas that go beyond some of what you are showing here. We would welcome more academic engagement and comparing notes if that is ever of interest 🙂

using agents every day has made this pretty obvious to me. a lot of the progress isn’t switching models. it’s getting them to remember what we already figured out and finish things without me having to keep checking in.

People underestimate the harness layer.

@garrytan Yes they are powerful. Open sourced this one- Build a @PalantirTech of your own!

Todo el mundo debería verlo

interesting but having thousands of harnesses doing the same thing is what the problem is. if someone makes groundbreaking harness thats very cool

@PrajwalAvhad8

Very educational. I've built a couple of harnesses myself, and the process is honestly not that easy. Building complexity must be addressed in the same way that the creation of AI agents was simplified

This doesn't make any sense. If model improves in general, it should work across different harnesses in general. We are so far from AGI as we were back in 2022.

Scaffolding is compute.

@harjtaggar Oh cool, I'd love seeing this video tonight looks fun, would be cool to explore stuff as I'm building my harness too!

Harness quality can change benchmark scores, making evaluation a systems design problem.

Same weights, different harness, 30 to 95. That gap is basically the distance between 'can't' and 'can if you structure the problem well. ' How many capabilities are dormant just because the evaluation wrapper is lazy?

Persistent memory and feedback loops are what turn model output into something teams can actually rely on.

The ARC-AGI benchmark illustrates the shift perfectly. Look at the performance gap when companies like OpenAI switched from raw prompting to agentic loops; the "o1" series effectively functioned as its own harness to hit those 90%+ reasoning scores.

from back in May

A 30% to 95% jump makes the harness part of the model system, not mere scaffolding. How do you test that the gain is not benchmark-specific overfitting?

This is a fascinating shift in how we think about AI agents. The model itself may not be the whole story the harness around it can unlock a huge amount of capability

This was a great overview.

@garrytan true, the infrastructure work is honestly the most underrated part.

the harness IS the product. same thing in outbound - same list, same copy, better sequencing infra and replies go 3-4x. nobody wants to admit the wrapper is where the value lives

If the same weights go 30 to 95 on harness alone, the benchmark is measuring the system, not the model. Most model comparisons still ignore that.

💯 The harness makes a huge difference which is why we've spent a lot of time iterating towards the perfect one in Hezo.

