Loading video...

Video Failed to Load

Go Home

We beat Claude Code in our evals: 61% to 53%. Yes, evals are actually important haha And now we’re open-sourcing our entire codebase and launching an agent framework! Now you can build custom agents for your project and get better results 🔥

276,953 views • 1 year ago •via X (Twitter)

46 Comments

James Grugett's profile picture
James Grugett1 year ago

We’re launching so many things that we need the whole week to fill you in on it all. Today is just Day 1. Codebuff first launched 1 year ago in YC F24, raised $1.6M, and has since been working to build the best coding agent.

James Grugett's profile picture
James Grugett1 year ago

2-3 months ago, we refactored our multi-agent system into an internal framework of composable agents. Each agent can have context specialized to its exact task, which we found to produce better results. The increase in our eval scores speak for themselves.

James Grugett's profile picture
James Grugett1 year ago

The BuffBench eval is different: It imitates the actual experience of using a coding agent. We have an agent pretend to be a human that sends prompts over multiple turns to a coding agent. Then a judge compares the output vs a real git commit diff! Read more here:

James Grugett's profile picture
James Grugett1 year ago

So why use Codebuff? - Best eval results out of the box - Powerful, customizable agents with our new framework - Flexible: compose and use published agents, run them with our new SDK

James Grugett's profile picture
James Grugett1 year ago

Try it now: > npm install -g codebuff > codebuff init-agents See our github page for more info! (And give us a star haha). Stay tuned for Launch Days 2-5 where we will elaborate on the vision and capabilities of Codebuff.

James Grugett's profile picture
James Grugett1 year ago

Usage going up, up, up!

James Grugett's profile picture
James Grugett1 year ago

Get more detail on this launch in our writeup!

OpenRouter's profile picture
OpenRouter1 year ago

👀

Brandon Chen's profile picture
Brandon Chen1 year ago

we were so young and innocent a year ago 🥹

RootFTW's profile picture
RootFTW1 year ago

If it's that good, I hope to see it at the top of the SWE-Bench in the coming days.

James Grugett's profile picture
James Grugett1 year ago

Nice! We'll see. But SWE bench doesn't measure code quality! Only pass/fail.

Tiger's profile picture
Tiger1 year ago

Absolutely amazing! I’ve been waiting forever for the agent framework! So excited 💪💪💪

James Grugett's profile picture
James Grugett1 year ago

Yup, pretty excited to see how people use it!

Akshay Iyer's profile picture
Akshay Iyer1 year ago

Let’s goooo🔥🔥🔥

Alex Aridgides's profile picture
Alex Aridgides1 year ago

Only up from here!

Felipe's profile picture
Felipe1 year ago

Let’s gooo

John Yeo's profile picture
John Yeo1 year ago

Insane!! congrats guys, hyped to try this

Essam Sleiman's profile picture
Essam Sleiman1 year ago

Very cool, congrats!

Akbar Ahmed's profile picture
Akbar Ahmed1 year ago

Thanks for open sourcing this. Can’t wait to dig into it.

James Grugett's profile picture
James Grugett1 year ago

Let me know if you find anything confusing! Haven't had fresh eyes on the codebase since our last hire haha

Raymond Weitekamp's profile picture
Raymond Weitekamp1 year ago

looks cool, excited to try!

Dodo's profile picture
Dodo1 year ago

sounds like a w move

Abby Grills's profile picture
Abby Grills1 year ago

This is awesome!

Sam Park's profile picture
Sam Park1 year ago

crazy stuff!!

Gautam Paranjape's profile picture
Gautam Paranjape1 year ago

congrats guys!

Md Fahim's profile picture
Md Fahim1 year ago

Sounds exciting! Can't wait to see it!

Swissy's profile picture
Swissy1 year ago

hard choice, eval performance is great but pricing on max plan is hard to beat, someone should make a true pice based on diff error rate, or add $ savings to buffbench

danialhasan's profile picture
danialhasan1 year ago

i remember overhearing this at builder sundays in shopify torontos HQ last weekend

Akio's profile picture
Akio1 year ago

Awesome launch! 🚀

James Grugett's profile picture
James Grugett1 year ago

Thanks Akio!

Elliot Arledge's profile picture
Elliot Arledge1 year ago

@rs545837 do your thing

King's profile picture
King1 year ago

Is it available in python?

vish's profile picture
vish1 year ago

this is awesome! cant wait to try. Does it support ollama?

James Grugett's profile picture
James Grugett1 year ago

Currently we support going through any model on OpenRouter! But since it's open source, you could try running it locally and swapping out to a local model. Probably not that easy to do, yet!

vish's profile picture
vish1 year ago

nice - Ill give it a shot.

Sean's profile picture
Sean1 year ago

So you're providing context engine to claude code so it can better solve an issue?

James Grugett's profile picture
James Grugett1 year ago

We are using Claude Sonnet (among other models), but not Claude Code. We also happen to have one of the best context engines ever IMO:

Sean's profile picture
Sean1 year ago

Really nice, you're solve an issue that i've been missing since i stopped using @augmentcode and @claudeai code doesn't really have a memory so it repeats itself and burn token every new task 😅

Team Reagent's profile picture
Team Reagent1 year ago

Thank you for contributing to open source! We'll check out your cool framework

Max Schulze-Melander's profile picture
Max Schulze-Melander1 year ago

Not. Another. Agent. Framework

Chain Alpha's profile picture
Chain Alpha1 year ago

Open source & agents? Alpha move, builders only. WAGMI.

Ebuka Gabriel 🥋's profile picture
Ebuka Gabriel 🥋1 year ago

This is very cool The SDK is only in Typescript A Python version would great as well.

Michael Bleigh's profile picture
Michael Bleigh1 year ago

Glad to have a new agentic coding bench, but crowing over beating competitors on evals you made up feels a bit disingenuous. A bit like inventing an intelligence test based on things you know off the top of your head and proclaiming yourself the smartest person in the world...

Suhrab Khan⚡️'s profile picture
Suhrab Khan⚡️1 year ago

Open-sourcing an agent framework right after outperforming Claude Code is a bold move. It lowers the barrier for teams to build tailored agents and push results even further.

George Pickett's profile picture
George Pickett1 year ago

how does it score on the common evals?

Leandro Narosky's profile picture
Leandro Narosky1 year ago

I'd love to see other tools evaluated on that benchmark. Mainly codex but also codebuff with other models like Kimi and glm

Related Videos