正在加载视频...

视频加载失败

What is the best LLM for agentic software engineering? Today, we're releasing The OpenHands Index to answer this question. It's the first broad-coverage benchmark for AI coding agents, comparing them on accuracy, cost, and runtime across 5 task domains.

31,924 次观看 • 8 个月前 •via X (Twitter)

19 条评论

OpenHands 的头像
OpenHands8 个月前

One well-known benchmark for coding agents is SWE-Bench. It's a great benchmark, but only covers one sub-task of software engineering: resolving github issues on open-source Python repositories.

OpenHands 的头像
OpenHands8 个月前

There's a few issues with this: 1. Of course there's more to software engineering than this! 2. Many of the models are optimized to do well on this benchmark, and there's only a few points of difference between them, even though the models feel quite different.

OpenHands 的头像
OpenHands8 个月前

To fix this, we picked 5 tasks that we have observed are important to users of coding agents: 1. Issue Resolution 2. Frontend Development 3. Greenfield Development 4. Software Testing 5. Information Gathering We prepared a benchmark to evaluate each ability.

OpenHands 的头像
OpenHands8 个月前

So what do the results look like? Here's an overview from the cost-accuracy perspective across the entire benchmark.

OpenHands 的头像
OpenHands8 个月前

Claude 4.5 Opus stands out from the pack with the overall highest scores, and #1 in 3 categories. This shouldn't be a surprise for those paying attention - social media has been abuzz about it for a couple months now. But it's also the most expensive by a good margin.

OpenHands 的头像
OpenHands8 个月前

GPT-5.2-Codex is also a strong contender, with the highest scores in categories such as greenfield app development, where it needs to create an app from start to finish.

OpenHands 的头像
OpenHands8 个月前

Gemini Flash and Deepseek-v3.2 Reasoner are also strong contenders with respect to price-performance. If you'd like to see other models added we'd love to hear your suggestions as well!

OpenHands 的头像
OpenHands8 个月前

There's much much more on the OpenHands Index site: Read more on the blog: We'll continue updating this, so comments+requests are welcome!

Noema 的头像
Noema8 个月前

Love it! I wish there was a spot for a 6th benchmark in there, swe-fficiency, algotune or codeclash would be great candidates

OpenHands 的头像
OpenHands8 个月前

These are good benchmarks! Maybe 1 in the next version.

Alex Chapin 的头像
Alex Chapin8 个月前

You should add Kimi K2.5 and GLM 4.7

SynthesisLedger 的头像
SynthesisLedger8 个月前

clean benchmark for agentic coding agents. accuracy and runtime are start, but the killer metric is state fidelity over 20+ loops. how do you penalize drift in long sessions? 🇳🇴

Kshitij Mishra | AI & Tech 的头像
Kshitij Mishra | AI & Tech8 个月前

I recommend it to all my followers

Sol Nyx 的头像
Sol Nyx8 个月前

The cost-per-task breakdown is what I've been waiting for. Raw accuracy leaderboards don't mean much when you're burning $50 on a simple refactor. Curious how the pareto frontier shifts once Sonnet 5 drops. Has anyone run these domains on local models like Qwen or DeepSeek for comparison?

Aaliya 的头像
Aaliya8 个月前

This is exactly what people needed.

synthline studio 的头像
synthline studio8 个月前

the cost vs accuracy framing is the most underrated part of this. everyone benchmarks accuracy but nobody talks about how much it costs to get there. an agent that scores 5% lower but costs 10x less is the one people actually ship with

Saïd Aitmbarek 的头像
Saïd Aitmbarek8 个月前

many say kimi 2.5 is goated :) !

Real Pinochet 的头像
Real Pinochet7 个月前

Could you guys actually address your users issues, I was wanting to try OH for the first time and what a disappointment. I tried your UV install method and the docker pull and both have the same issues that others are reporting. Running linux, Debian!

OpenHands 的头像
OpenHands7 个月前

Hi @OG_Pinochet, apologies that this is not working. mamoodi just raised this on slack today and we will take a look.

相关视频