Loading video...
Video Failed to Load
Claude Code vs Codex vs Pi: which coding agent wins? Melissa Pan, PhD candidate at UC Berkeley’s Sky Computing Lab and previous Arena.ai intern, explored the hidden “harness tax”: how the system surrounding an AI model affects its cost and performance. She reports three surprising findings. One: harness choice... show more
85,347 views • 11 days ago •via X (Twitter)
22 Comments

@melissapan Swap the model and barely anything changes. Rebuild the scaffolding around it and suddenly it's a different tool.

@melissapan Well said.

@melissapan Arena 对照 Claude Code、Codex 与 Pi,指出外围 harness 带来的成本差往往大于准确率差。代理任务的费用不只取决于模型单价,工具编排、重试与上下文装填会显著改写单次账单;选栈时应把 harness 与模型一并计入。

@melissapan 同模型换壳,账单和成功率一起飘——这句比榜单管用。一人公司选型别只盯 benchmark:环路怎么验、失败怎么停、状态写在哪。壳可以换;「过/不过」别寄存在某一个壳的默认里。

@melissapan Harness mattering more for cost than accuracy makes sense when so much of the bill is context re-sent every turn. Would love to see cost per solved task broken out by harness.

@melissapan @trycodegraff! as well!! it’s a self evolving harness haha

@melissapan 原来测试套件才是偷偷抬成本的那只手

@melissapan harness tax is the real scoreboard. same model, totally different bill depending on how messy the loop around it is

@melissapan The harness tax is a fascinating angle. Really useful analysis for building efficient coding agents.

@melissapan can someone give me the tldr 🤓

@melissapan The 'which coding agent wins' answer is usually whichever harness stopped fighting the model. Measure tool-loop stability under load, not demo speed. #ClaudeCode

@melissapan The harness tax is an often overlooked performance killer

@melissapan if harness moves cost more than accuracy, you're ranking the wrapper as much as the model. was that Claude Code vs Codex vs Pi on the same model, or did each keep its default?

@melissapan I think pi cost effective and powerful coding agent right now

@melissapan the model gets graded, the harness quietly invoices everyone

@melissapan This is awesome/helpful Melissa & Arena!

@melissapan would love to see how codewhale performs!

Harness tax decides the winner more than the model card. Tonight: time-to-first-kept-diff on the same bug across Claude Code / Codex / your third agent, with identical AGENTS.md. Whichever burns fewer tokens per kept line for a week becomes the default lane — swap the rest to overflow only.

@melissapan Agent benchmarks need matched tasks and budgets; harnesses can dominate outcomes.

@melissapan Female woman PagMan

@melissapan Benchmark the whole harness on representative tasks, including tool calls and retries, before picking a model. If harness choice drives cost more than accuracy, optimize routing and stop conditions first.

@melissapan did she break down prompt cache hit rate per harness? I'd guess that's where most of the cost gap comes from, more than anything the harness does to accuracy
