正在加载视频...
视频加载失败
What is the best LLM for agentic software engineering? Today, we're releasing The OpenHands Index to answer this question. It's the first broad-coverage benchmark for AI coding agents, comparing them on accuracy, cost, and runtime across 5 task domains.
31,924 次观看 • 8 个月前 •via X (Twitter)
19 条评论

One well-known benchmark for coding agents is SWE-Bench. It's a great benchmark, but only covers one sub-task of software engineering: resolving github issues on open-source Python repositories.

There's a few issues with this: 1. Of course there's more to software engineering than this! 2. Many of the models are optimized to do well on this benchmark, and there's only a few points of difference between them, even though the models feel quite different.

To fix this, we picked 5 tasks that we have observed are important to users of coding agents: 1. Issue Resolution 2. Frontend Development 3. Greenfield Development 4. Software Testing 5. Information Gathering We prepared a benchmark to evaluate each ability.

So what do the results look like? Here's an overview from the cost-accuracy perspective across the entire benchmark.

Claude 4.5 Opus stands out from the pack with the overall highest scores, and #1 in 3 categories. This shouldn't be a surprise for those paying attention - social media has been abuzz about it for a couple months now. But it's also the most expensive by a good margin.

GPT-5.2-Codex is also a strong contender, with the highest scores in categories such as greenfield app development, where it needs to create an app from start to finish.

Gemini Flash and Deepseek-v3.2 Reasoner are also strong contenders with respect to price-performance. If you'd like to see other models added we'd love to hear your suggestions as well!

There's much much more on the OpenHands Index site: Read more on the blog: We'll continue updating this, so comments+requests are welcome!

Love it! I wish there was a spot for a 6th benchmark in there, swe-fficiency, algotune or codeclash would be great candidates

These are good benchmarks! Maybe 1 in the next version.

You should add Kimi K2.5 and GLM 4.7

clean benchmark for agentic coding agents. accuracy and runtime are start, but the killer metric is state fidelity over 20+ loops. how do you penalize drift in long sessions? 🇳🇴

I recommend it to all my followers

The cost-per-task breakdown is what I've been waiting for. Raw accuracy leaderboards don't mean much when you're burning $50 on a simple refactor. Curious how the pareto frontier shifts once Sonnet 5 drops. Has anyone run these domains on local models like Qwen or DeepSeek for comparison?

This is exactly what people needed.

the cost vs accuracy framing is the most underrated part of this. everyone benchmarks accuracy but nobody talks about how much it costs to get there. an agent that scores 5% lower but costs 10x less is the one people actually ship with

many say kimi 2.5 is goated :) !

Could you guys actually address your users issues, I was wanting to try OH for the first time and what a disappointment. I tried your UV install method and the docker pull and both have the same issues that others are reporting. Running linux, Debian!

Hi @OG_Pinochet, apologies that this is not working. mamoodi just raised this on slack today and we will take a look.
