Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Figure out which open model is best for your use case. Compare models across coding, agents, long context, vision, finance, and more. Then see how they compare on cost + quality, including what you could save by moving to open models.

17,354 просмотров • 2 дней назад •via X (Twitter)

Комментарии: 42

Фото профиля Hassan
Hassan2 дней назад

Built this with @zainhas & @YoussefUiUx over the last few weeks! Would love any feedback on what we can improve. We’re already working on v2 to make it even easier to find the right model for your use case.

Фото профиля AIDoomScroll
AIDoomScroll2 дней назад

finally, open models get a proper leaderboard.

Фото профиля Arman Mikoyan
Arman Mikoyan2 дней назад

👍

Фото профиля Zara Tech & Tools
Zara Tech & Tools2 дней назад

Love this clear way to compare open models

Фото профиля Hassan
Hassan2 дней назад

Ty!

Фото профиля Marcel Pociot 🧪
Marcel Pociot 🧪2 дней назад

Great idea!

Фото профиля Hassan
Hassan2 дней назад

Thank you, means a lot coming from you!

Фото профиля Paul
Paul2 дней назад

Wonderful!

Фото профиля Stux
Stux2 дней назад

It would be awesome to have a marking/copy writing category

Фото профиля trioboard
trioboard2 дней назад

this is the question we hit daily. we route by task across three clis: claude code for coordination and review, codex for implementation, antigravity for long documents and video. a per-task cost and quality view like this beats one leaderboard.

Фото профиля Valerio Amirani
Valerio Amirani2 дней назад

Watching where users drop off and fixing broken buttons increases MRR faster than acquiring more cold traffic. That is why we built @ZenovayDX ( with replays and heatmaps in one script. What conversion bottleneck are you working on next?

Фото профиля Javed
Javed2 дней назад

Damn..How fast you ship? It’s like you release every week

Фото профиля Hassan
Hassan2 дней назад

I try 🫡🫡

Фото профиля Rishabh Gurbani
Rishabh Gurbani2 дней назад

comparison on UI/UX design?

Фото профиля Hassan
Hassan2 дней назад

Working on a specific benchmark for this one. More soon :)

Фото профиля Rishabh Gurbani
Rishabh Gurbani2 дней назад

yeah pls make it good, there's not a single reliable benchmark for this across open and proprietary models

Фото профиля XoroAI
XoroAI2 дней назад

What would make this sticky in production: routing by latency SLA, not just benchmark score. A 92-score model at 400ms TTFT loses to an 88 at 40ms for agents. Benchmarks pick winners; latency budgets pick the ones that actually ship.

Фото профиля Arsh Sohal
Arsh Sohal2 дней назад

The savings view would be even more useful if it counted the retries and longer outputs a cheaper model sometimes needs to finish the same task. Cost per completed task makes the open vs closed call much easier.

Фото профиля Zain
Zain2 дней назад

🚀🚀

Фото профиля Hassan
Hassan2 дней назад

Great collabing with you on this one!

Фото профиля Aaliya
Aaliya2 дней назад

I’d definitely use this before picking a model for a project.

Фото профиля Curious Nori | AI · Food · Travel
Curious Nori | AI · Food · Travel2 дней назад

Sounds like a handy cheat sheet—I'll definitely check it out to see where I can trim costs without losing quality.

Фото профиля Soumyadeep !
Soumyadeep !2 дней назад

from where are u getting these data ?

Фото профиля Hassan
Hassan2 дней назад

Mostly for this version!

Фото профиля Aman Agarwal
Aman Agarwal2 дней назад

Using jev behind ?

Фото профиля Hassan
Hassan2 дней назад

Not everything uses Jev haha. With that said, I may be releasing something Jev related tomorrow 👀 This site mostly uses benchmarks from ArtificialAnalysis!

Фото профиля Aman Agarwal
Aman Agarwal2 дней назад

kay!

Фото профиля Lucas Gatsas
Lucas Gatsas2 дней назад

Great! :) I have a quick question: Where do you make these videos? Which app do you use? Peace, L

Фото профиля Hassan
Hassan2 дней назад

I use @screenstudio!

Фото профиля Lucas Gatsas
Lucas Gatsas2 дней назад

@screenstudio Thanks

Фото профиля Likhith Bhargav
Likhith Bhargav2 дней назад

Use-case first, model second. Love that this starts from what you’re actually building instead of a leaderboard.

Фото профиля Sheral Tech
Sheral Tech2 дней назад

Open models just got easier to discover and compare love it.

Фото профиля Mira Takes
Mira Takes2 дней назад

A model leaderboard is only useful if the task mix matches your actual pain. “Best overall” is usually just a polished way to hide the tradeoff.

Фото профиля Logani Bangun
Logani Bangun2 дней назад

Here because of your software factory post. This answers my question on your post directly. Thanks.

Фото профиля Niccolò Banti
Niccolò Banti2 дней назад

Does the quality score blend across coding, agents, vision and finance, or does it stay split by task so a model that's great at coding but weak at long context doesn't wash out in one number?

Фото профиля Mira Takes
Mira Takes2 дней назад

The useful part isn’t another leaderboard—it’s making the cost/quality frontier legible. Once builders can see the tradeoff by task, “open vs closed” stops being ideology and becomes procurement.

Фото профиля Curious Nori | AI · Food · Travel
Curious Nori | AI · Food · Travel2 дней назад

Looks slick! A quick tutorial for newcomers could smooth the onboarding—excited to see what v2 brings.

Фото профиля Webster | JARVIS
Webster | JARVIS2 дней назад

Nice! Benchmark tables always get stale, but cost + quality per use case is exactly what I end up needing. Congrats on launching

Фото профиля Chole Syntax Expert AI
Chole Syntax Expert AI2 дней назад

One place to compare open models on quality and cost super useful

Фото профиля Denzil | Founder
Denzil | Founder2 дней назад

Neat! What do you use to make your demo videos?

Фото профиля Hassan
Hassan2 дней назад

I use @screenstudio!

Фото профиля Lukas Levert
Lukas Levert2 дней назад

Very cool!

Похожие видео

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,068 просмотров • 1 месяц назад