Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Arena Mode in Windsurf: One prompt. Two models. Your vote. Benchmarks don't reflect real-world coding quality. The best model for you depends on your codebase and stack. So we made real-world coding the benchmark. Free for the next week. May the best model win.

1,069,578 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 39

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

You can use Arena Mode by choosing up to five models or through Battle Groups. With Battle Groups, two models will be randomly selected each turn. To celebrate the launch, Battle Groups will consume 0x credits for the next week (for trial and paid users)! To start using Arena Mode, install or update Windsurf: Leaderboard: Blog post:

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

Why we built Arena Mode:

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

Arena Mode is part of Wave 14! We've also added Plan Mode to Cascade. Check out the blog post for more details:

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

May the best model win!

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

There's a lot of buzz around Kimi K2.5 and Opus 4.5. Kimi K2.5 is now in Windsurf and Frontier Arena!

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

How to use Plan Mode:

Фото профиля Wayne Chi
Wayne Chi7 месяцев назад

This is fantastic and looks like a direct follow-up to our work on Copilot Arena (an in IDE Arena) from ICML 2025. Really excited to see this line of work with professional support. Check it out and hope it helps!

Фото профиля Devin Desktop
Devin Desktop7 месяцев назад

whoa cool!

Фото профиля Big WUM
Big WUM7 месяцев назад

What a great way to burn through even more tokens

Фото профиля Abhishek Yadav
Abhishek Yadav7 месяцев назад

Finally, a benchmark that looks like real work instead of leaderboard gaming. Love this. benchmarks rarely capture refactors, edge cases, or messy real code.Preference based evals matter more than scores.

Фото профиля Carlos
Carlos7 месяцев назад

This just feels like you’re asking users to train models for you for free

Фото профиля Thanh Nguyen
Thanh Nguyen7 месяцев назад

I don’t know which dev would use this but def not me. I even hate it when chatbot ask me to compare and choose between 2 responses: it takes longer to generate and it hammers my cognitive ability by asking me to read something twice. I can only imagine the problem is exacerbated in coding task where the responses are longer and harder to compare.

Фото профиля Abhimanyu ARYAN
Abhimanyu ARYAN7 месяцев назад

Just tested it. It's amazing. I know the clear winner for me

Фото профиля Eashan Sinha
Eashan Sinha7 месяцев назад

Let's go

Фото профиля Tony Lee
Tony Lee7 месяцев назад

Okay I will cancel my cancelation of windsurf which I did yesterday.

Фото профиля Nate Chan
Nate Chan7 месяцев назад

hey i did it first

Фото профиля Alex Kaplan
Alex Kaplan7 месяцев назад

Let’s go 🥊🥊🥊🥊

Фото профиля Nnenna 👩🏽‍💻✨
Nnenna 👩🏽‍💻✨7 месяцев назад

“Your codebase is the benchmark.” Spicy.

Фото профиля abuah
abuah7 месяцев назад

Can we use agents natively inside of Windsurf yet with their respective models? I want to do this with Claude Code and Codex, not Opus & GPT-5.2, if that makes sense.

Фото профиля swyx
swyx7 месяцев назад

some background for those intersted in details

Фото профиля Uncle J
Uncle J7 месяцев назад

Benchmarks are vanity metrics for model companies. The only benchmark that matters is "did my code ship without me babysitting it." Curious if Arena Mode tracks that.

Фото профиля Huseyn Hajiyev
Huseyn Hajiyev7 месяцев назад

Great benchmark approach! I believe it won't be more expensive and will allow us to pick the best result that way.

Фото профиля Charles Lazaroni
Charles Lazaroni7 месяцев назад

this is the right call. every synthetic benchmark gets gamed within a week. testing on your actual codebase with your actual stack is the only eval that matters. curious how the vote data looks after a month of real usage

Фото профиля Siddik
Siddik7 месяцев назад

But how is it useful? Many free arenas already exist where people can check model efficiency. Why would someone want to test a model by burning their own credits?

Фото профиля Micah Pierce
Micah Pierce7 месяцев назад

42 min in and SWE is now my go to.

Фото профиля Oskar Minor
Oskar Minor7 месяцев назад

It can be helpful - You'll be able to compare which model will be best for your project

Фото профиля Pratik
Pratik7 месяцев назад

This is genius. Benchmarks lie, but your own codebase doesn't.

Фото профиля bams
bams7 месяцев назад

Error

Фото профиля Botir Khaltaev
Botir Khaltaev7 месяцев назад

@windsurf, check out nordlys models,we build context specific routing, we took a bunch of swe data we clustered it to find real world tasks on a large scale, and built Nordlys Hypernova 75.6% on the SWE-Bench

Фото профиля Pratik
Pratik7 месяцев назад

This is actually smart. benchmarks mean nothing when you're 3 hours into debugging a specific codebase

Фото профиля Razex
Razex7 месяцев назад

But where s kimi ?

Фото профиля brgma
brgma7 месяцев назад

guys theres no switch or toggle for plan mode , check please

Фото профиля Vinícius Ragazzi
Vinícius Ragazzi7 месяцев назад

Who will win?

Фото профиля Alex Atallah
Alex Atallah7 месяцев назад

❤️

Фото профиля Apoorv Sharma
Apoorv Sharma7 месяцев назад

Will you train on this data?

Фото профиля 8 Bit Elon
8 Bit Elon7 месяцев назад

@grok I thought windsurf was sunsetted and turn into antigravity?

Фото профиля Akshat Bahety
Akshat Bahety7 месяцев назад

I thought windsurf was dead

Фото профиля Echo
Echo7 месяцев назад

Just for paid member free

Фото профиля Sky
Sky7 месяцев назад

Add a damn confirmation on selecting a winner! 3 times now I've accidentally it whilst other models were still running (as normally the poorer model - sonnet 4.5?!? finishes early)

Похожие видео

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

132,623 просмотров • 1 месяц назад