Loading video...
Video Failed to Load
We beat Claude Code in our evals: 61% to 53%. Yes, evals are actually important haha And now we’re open-sourcing our entire codebase and launching an agent framework! Now you can build custom agents for your project and get better results 🔥
276,953 views • 1 year ago •via X (Twitter)
46 Comments

We’re launching so many things that we need the whole week to fill you in on it all. Today is just Day 1. Codebuff first launched 1 year ago in YC F24, raised $1.6M, and has since been working to build the best coding agent.

2-3 months ago, we refactored our multi-agent system into an internal framework of composable agents. Each agent can have context specialized to its exact task, which we found to produce better results. The increase in our eval scores speak for themselves.

The BuffBench eval is different: It imitates the actual experience of using a coding agent. We have an agent pretend to be a human that sends prompts over multiple turns to a coding agent. Then a judge compares the output vs a real git commit diff! Read more here:

So why use Codebuff? - Best eval results out of the box - Powerful, customizable agents with our new framework - Flexible: compose and use published agents, run them with our new SDK

Try it now: > npm install -g codebuff > codebuff init-agents See our github page for more info! (And give us a star haha). Stay tuned for Launch Days 2-5 where we will elaborate on the vision and capabilities of Codebuff.

Usage going up, up, up!

Get more detail on this launch in our writeup!

👀

we were so young and innocent a year ago 🥹

If it's that good, I hope to see it at the top of the SWE-Bench in the coming days.

Nice! We'll see. But SWE bench doesn't measure code quality! Only pass/fail.

Absolutely amazing! I’ve been waiting forever for the agent framework! So excited 💪💪💪

Yup, pretty excited to see how people use it!

Let’s goooo🔥🔥🔥

Only up from here!

Let’s gooo

Insane!! congrats guys, hyped to try this

Very cool, congrats!

Thanks for open sourcing this. Can’t wait to dig into it.

Let me know if you find anything confusing! Haven't had fresh eyes on the codebase since our last hire haha

looks cool, excited to try!

sounds like a w move

This is awesome!

crazy stuff!!

congrats guys!

Sounds exciting! Can't wait to see it!

hard choice, eval performance is great but pricing on max plan is hard to beat, someone should make a true pice based on diff error rate, or add $ savings to buffbench

i remember overhearing this at builder sundays in shopify torontos HQ last weekend

Awesome launch! 🚀

Thanks Akio!

@rs545837 do your thing

Is it available in python?

this is awesome! cant wait to try. Does it support ollama?

Currently we support going through any model on OpenRouter! But since it's open source, you could try running it locally and swapping out to a local model. Probably not that easy to do, yet!

nice - Ill give it a shot.

So you're providing context engine to claude code so it can better solve an issue?

We are using Claude Sonnet (among other models), but not Claude Code. We also happen to have one of the best context engines ever IMO:

Really nice, you're solve an issue that i've been missing since i stopped using @augmentcode and @claudeai code doesn't really have a memory so it repeats itself and burn token every new task 😅

Thank you for contributing to open source! We'll check out your cool framework

Not. Another. Agent. Framework

Open source & agents? Alpha move, builders only. WAGMI.

This is very cool The SDK is only in Typescript A Python version would great as well.

Glad to have a new agentic coding bench, but crowing over beating competitors on evals you made up feels a bit disingenuous. A bit like inventing an intelligence test based on things you know off the top of your head and proclaiming yourself the smartest person in the world...

Open-sourcing an agent framework right after outperforming Claude Code is a bold move. It lowers the barrier for teams to build tailored agents and push results even further.

how does it score on the common evals?

I'd love to see other tools evaluated on that benchmark. Mainly codex but also codebuff with other models like Kimi and glm



