Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Arena Mode in Windsurf: One prompt. Two models. Your vote. Benchmarks don't reflect real-world coding quality. The best model for you depends on your codebase and stack. So we made real-world coding the benchmark. Free for the next week. May the best model win.

1,069,578 Aufrufe • vor 7 Monaten •via X (Twitter)

39 Kommentare

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

You can use Arena Mode by choosing up to five models or through Battle Groups. With Battle Groups, two models will be randomly selected each turn. To celebrate the launch, Battle Groups will consume 0x credits for the next week (for trial and paid users)! To start using Arena Mode, install or update Windsurf: Leaderboard: Blog post:

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

Why we built Arena Mode:

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

Arena Mode is part of Wave 14! We've also added Plan Mode to Cascade. Check out the blog post for more details:

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

May the best model win!

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

There's a lot of buzz around Kimi K2.5 and Opus 4.5. Kimi K2.5 is now in Windsurf and Frontier Arena!

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

How to use Plan Mode:

Profilbild von Wayne Chi
Wayne Chivor 7 Monaten

This is fantastic and looks like a direct follow-up to our work on Copilot Arena (an in IDE Arena) from ICML 2025. Really excited to see this line of work with professional support. Check it out and hope it helps!

Profilbild von Devin Desktop
Devin Desktopvor 7 Monaten

whoa cool!

Profilbild von Big WUM
Big WUMvor 7 Monaten

What a great way to burn through even more tokens

Profilbild von Abhishek Yadav
Abhishek Yadavvor 7 Monaten

Finally, a benchmark that looks like real work instead of leaderboard gaming. Love this. benchmarks rarely capture refactors, edge cases, or messy real code.Preference based evals matter more than scores.

Profilbild von Carlos
Carlosvor 7 Monaten

This just feels like you’re asking users to train models for you for free

Profilbild von Thanh Nguyen
Thanh Nguyenvor 7 Monaten

I don’t know which dev would use this but def not me. I even hate it when chatbot ask me to compare and choose between 2 responses: it takes longer to generate and it hammers my cognitive ability by asking me to read something twice. I can only imagine the problem is exacerbated in coding task where the responses are longer and harder to compare.

Profilbild von Abhimanyu ARYAN
Abhimanyu ARYANvor 7 Monaten

Just tested it. It's amazing. I know the clear winner for me

Profilbild von Eashan Sinha
Eashan Sinhavor 7 Monaten

Let's go

Profilbild von Tony Lee
Tony Leevor 7 Monaten

Okay I will cancel my cancelation of windsurf which I did yesterday.

Profilbild von Nate Chan
Nate Chanvor 7 Monaten

hey i did it first

Profilbild von Alex Kaplan
Alex Kaplanvor 7 Monaten

Let’s go 🥊🥊🥊🥊

Profilbild von Nnenna 👩🏽‍💻✨
Nnenna 👩🏽‍💻✨vor 7 Monaten

“Your codebase is the benchmark.” Spicy.

Profilbild von abuah
abuahvor 7 Monaten

Can we use agents natively inside of Windsurf yet with their respective models? I want to do this with Claude Code and Codex, not Opus & GPT-5.2, if that makes sense.

Profilbild von swyx
swyxvor 7 Monaten

some background for those intersted in details

Profilbild von Uncle J
Uncle Jvor 7 Monaten

Benchmarks are vanity metrics for model companies. The only benchmark that matters is "did my code ship without me babysitting it." Curious if Arena Mode tracks that.

Profilbild von Huseyn Hajiyev
Huseyn Hajiyevvor 7 Monaten

Great benchmark approach! I believe it won't be more expensive and will allow us to pick the best result that way.

Profilbild von Charles Lazaroni
Charles Lazaronivor 7 Monaten

this is the right call. every synthetic benchmark gets gamed within a week. testing on your actual codebase with your actual stack is the only eval that matters. curious how the vote data looks after a month of real usage

Profilbild von Siddik
Siddikvor 7 Monaten

But how is it useful? Many free arenas already exist where people can check model efficiency. Why would someone want to test a model by burning their own credits?

Profilbild von Micah Pierce
Micah Piercevor 7 Monaten

42 min in and SWE is now my go to.

Profilbild von Oskar Minor
Oskar Minorvor 7 Monaten

It can be helpful - You'll be able to compare which model will be best for your project

Profilbild von Pratik
Pratikvor 7 Monaten

This is genius. Benchmarks lie, but your own codebase doesn't.

Profilbild von bams
bamsvor 7 Monaten

Error

Profilbild von Botir Khaltaev
Botir Khaltaevvor 7 Monaten

@windsurf, check out nordlys models,we build context specific routing, we took a bunch of swe data we clustered it to find real world tasks on a large scale, and built Nordlys Hypernova 75.6% on the SWE-Bench

Profilbild von Pratik
Pratikvor 7 Monaten

This is actually smart. benchmarks mean nothing when you're 3 hours into debugging a specific codebase

Profilbild von Razex
Razexvor 7 Monaten

But where s kimi ?

Profilbild von brgma
brgmavor 7 Monaten

guys theres no switch or toggle for plan mode , check please

Profilbild von Vinícius Ragazzi
Vinícius Ragazzivor 7 Monaten

Who will win?

Profilbild von Alex Atallah
Alex Atallahvor 7 Monaten

❤️

Profilbild von Apoorv Sharma
Apoorv Sharmavor 7 Monaten

Will you train on this data?

Profilbild von 8 Bit Elon
8 Bit Elonvor 7 Monaten

@grok I thought windsurf was sunsetted and turn into antigravity?

Profilbild von Akshat Bahety
Akshat Bahetyvor 7 Monaten

I thought windsurf was dead

Profilbild von Echo
Echovor 7 Monaten

Just for paid member free

Profilbild von Sky
Skyvor 7 Monaten

Add a damn confirmation on selecting a winner! 3 times now I've accidentally it whilst other models were still running (as normally the poorer model - sonnet 4.5?!? finishes early)

Ähnliche Videos

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

132,623 Aufrufe • vor 1 Monat