Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Arena Mode in Windsurf: One prompt. Two models. Your vote. Benchmarks don't reflect real-world coding quality. The best model for you depends on your codebase and stack. So we made real-world coding the benchmark. Free for the next week. May the best model win.

1,069,578 görüntüleme • 7 ay önce •via X (Twitter)

39 Yorum

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

You can use Arena Mode by choosing up to five models or through Battle Groups. With Battle Groups, two models will be randomly selected each turn. To celebrate the launch, Battle Groups will consume 0x credits for the next week (for trial and paid users)! To start using Arena Mode, install or update Windsurf: Leaderboard: Blog post:

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

Why we built Arena Mode:

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

Arena Mode is part of Wave 14! We've also added Plan Mode to Cascade. Check out the blog post for more details:

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

May the best model win!

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

There's a lot of buzz around Kimi K2.5 and Opus 4.5. Kimi K2.5 is now in Windsurf and Frontier Arena!

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

How to use Plan Mode:

Wayne Chi profil fotoğrafı
Wayne Chi7 ay önce

This is fantastic and looks like a direct follow-up to our work on Copilot Arena (an in IDE Arena) from ICML 2025. Really excited to see this line of work with professional support. Check it out and hope it helps!

Devin Desktop profil fotoğrafı
Devin Desktop7 ay önce

whoa cool!

Big WUM profil fotoğrafı
Big WUM7 ay önce

What a great way to burn through even more tokens

Abhishek Yadav profil fotoğrafı
Abhishek Yadav7 ay önce

Finally, a benchmark that looks like real work instead of leaderboard gaming. Love this. benchmarks rarely capture refactors, edge cases, or messy real code.Preference based evals matter more than scores.

Carlos profil fotoğrafı
Carlos7 ay önce

This just feels like you’re asking users to train models for you for free

Thanh Nguyen profil fotoğrafı
Thanh Nguyen7 ay önce

I don’t know which dev would use this but def not me. I even hate it when chatbot ask me to compare and choose between 2 responses: it takes longer to generate and it hammers my cognitive ability by asking me to read something twice. I can only imagine the problem is exacerbated in coding task where the responses are longer and harder to compare.

Abhimanyu ARYAN profil fotoğrafı
Abhimanyu ARYAN7 ay önce

Just tested it. It's amazing. I know the clear winner for me

Eashan Sinha profil fotoğrafı
Eashan Sinha7 ay önce

Let's go

Tony Lee profil fotoğrafı
Tony Lee7 ay önce

Okay I will cancel my cancelation of windsurf which I did yesterday.

Nate Chan profil fotoğrafı
Nate Chan7 ay önce

hey i did it first

Alex Kaplan profil fotoğrafı
Alex Kaplan7 ay önce

Let’s go 🥊🥊🥊🥊

Nnenna 👩🏽‍💻✨ profil fotoğrafı
Nnenna 👩🏽‍💻✨7 ay önce

“Your codebase is the benchmark.” Spicy.

abuah profil fotoğrafı
abuah7 ay önce

Can we use agents natively inside of Windsurf yet with their respective models? I want to do this with Claude Code and Codex, not Opus & GPT-5.2, if that makes sense.

swyx profil fotoğrafı
swyx7 ay önce

some background for those intersted in details

Uncle J profil fotoğrafı
Uncle J7 ay önce

Benchmarks are vanity metrics for model companies. The only benchmark that matters is "did my code ship without me babysitting it." Curious if Arena Mode tracks that.

Huseyn Hajiyev profil fotoğrafı
Huseyn Hajiyev7 ay önce

Great benchmark approach! I believe it won't be more expensive and will allow us to pick the best result that way.

Charles Lazaroni profil fotoğrafı
Charles Lazaroni7 ay önce

this is the right call. every synthetic benchmark gets gamed within a week. testing on your actual codebase with your actual stack is the only eval that matters. curious how the vote data looks after a month of real usage

Siddik profil fotoğrafı
Siddik7 ay önce

But how is it useful? Many free arenas already exist where people can check model efficiency. Why would someone want to test a model by burning their own credits?

Micah Pierce profil fotoğrafı
Micah Pierce7 ay önce

42 min in and SWE is now my go to.

Oskar Minor profil fotoğrafı
Oskar Minor7 ay önce

It can be helpful - You'll be able to compare which model will be best for your project

Pratik profil fotoğrafı
Pratik7 ay önce

This is genius. Benchmarks lie, but your own codebase doesn't.

bams profil fotoğrafı
bams7 ay önce

Error

Botir Khaltaev profil fotoğrafı
Botir Khaltaev7 ay önce

@windsurf, check out nordlys models,we build context specific routing, we took a bunch of swe data we clustered it to find real world tasks on a large scale, and built Nordlys Hypernova 75.6% on the SWE-Bench

Pratik profil fotoğrafı
Pratik7 ay önce

This is actually smart. benchmarks mean nothing when you're 3 hours into debugging a specific codebase

Razex profil fotoğrafı
Razex7 ay önce

But where s kimi ?

brgma profil fotoğrafı
brgma7 ay önce

guys theres no switch or toggle for plan mode , check please

Vinícius Ragazzi profil fotoğrafı
Vinícius Ragazzi7 ay önce

Who will win?

Alex Atallah profil fotoğrafı
Alex Atallah7 ay önce

❤️

Apoorv Sharma profil fotoğrafı
Apoorv Sharma7 ay önce

Will you train on this data?

8 Bit Elon profil fotoğrafı
8 Bit Elon7 ay önce

@grok I thought windsurf was sunsetted and turn into antigravity?

Akshat Bahety profil fotoğrafı
Akshat Bahety7 ay önce

I thought windsurf was dead

Echo profil fotoğrafı
Echo7 ay önce

Just for paid member free

Sky profil fotoğrafı
Sky7 ay önce

Add a damn confirmation on selecting a winner! 3 times now I've accidentally it whilst other models were still running (as normally the poorer model - sonnet 4.5?!? finishes early)

Benzer Videolar

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

132,623 görüntüleme • 1 ay önce