Video wird geladen...
Video konnte nicht geladen werden
Introducing Arena Mode in Windsurf: One prompt. Two models. Your vote. Benchmarks don't reflect real-world coding quality. The best model for you depends on your codebase and stack. So we made real-world coding the benchmark. Free for the next week. May the best model win.
1,069,578 Aufrufe • vor 7 Monaten •via X (Twitter)
39 Kommentare

You can use Arena Mode by choosing up to five models or through Battle Groups. With Battle Groups, two models will be randomly selected each turn. To celebrate the launch, Battle Groups will consume 0x credits for the next week (for trial and paid users)! To start using Arena Mode, install or update Windsurf: Leaderboard: Blog post:

Why we built Arena Mode:

Arena Mode is part of Wave 14! We've also added Plan Mode to Cascade. Check out the blog post for more details:

May the best model win!

There's a lot of buzz around Kimi K2.5 and Opus 4.5. Kimi K2.5 is now in Windsurf and Frontier Arena!

How to use Plan Mode:

This is fantastic and looks like a direct follow-up to our work on Copilot Arena (an in IDE Arena) from ICML 2025. Really excited to see this line of work with professional support. Check it out and hope it helps!

whoa cool!

What a great way to burn through even more tokens

Finally, a benchmark that looks like real work instead of leaderboard gaming. Love this. benchmarks rarely capture refactors, edge cases, or messy real code.Preference based evals matter more than scores.

This just feels like you’re asking users to train models for you for free

I don’t know which dev would use this but def not me. I even hate it when chatbot ask me to compare and choose between 2 responses: it takes longer to generate and it hammers my cognitive ability by asking me to read something twice. I can only imagine the problem is exacerbated in coding task where the responses are longer and harder to compare.

Just tested it. It's amazing. I know the clear winner for me

Let's go

Okay I will cancel my cancelation of windsurf which I did yesterday.

hey i did it first

Let’s go 🥊🥊🥊🥊

“Your codebase is the benchmark.” Spicy.

Can we use agents natively inside of Windsurf yet with their respective models? I want to do this with Claude Code and Codex, not Opus & GPT-5.2, if that makes sense.

some background for those intersted in details

Benchmarks are vanity metrics for model companies. The only benchmark that matters is "did my code ship without me babysitting it." Curious if Arena Mode tracks that.

Great benchmark approach! I believe it won't be more expensive and will allow us to pick the best result that way.

this is the right call. every synthetic benchmark gets gamed within a week. testing on your actual codebase with your actual stack is the only eval that matters. curious how the vote data looks after a month of real usage

But how is it useful? Many free arenas already exist where people can check model efficiency. Why would someone want to test a model by burning their own credits?

42 min in and SWE is now my go to.

It can be helpful - You'll be able to compare which model will be best for your project

This is genius. Benchmarks lie, but your own codebase doesn't.

Error

@windsurf, check out nordlys models,we build context specific routing, we took a bunch of swe data we clustered it to find real world tasks on a large scale, and built Nordlys Hypernova 75.6% on the SWE-Bench

This is actually smart. benchmarks mean nothing when you're 3 hours into debugging a specific codebase

But where s kimi ?

guys theres no switch or toggle for plan mode , check please

Who will win?

❤️

Will you train on this data?

@grok I thought windsurf was sunsetted and turn into antigravity?

I thought windsurf was dead

Just for paid member free

Add a damn confirmation on selecting a winner! 3 times now I've accidentally it whilst other models were still running (as normally the poorer model - sonnet 4.5?!? finishes early)


