Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Figure out which open model is best for your use case. Compare models across coding, agents, long context, vision, finance, and more. Then see how they compare on cost + quality, including what you could save by moving to open models.

17,354 Aufrufe • vor 2 Tagen •via X (Twitter)

42 Kommentare

Profilbild von Hassan
Hassanvor 2 Tagen

Built this with @zainhas & @YoussefUiUx over the last few weeks! Would love any feedback on what we can improve. We’re already working on v2 to make it even easier to find the right model for your use case.

Profilbild von AIDoomScroll
AIDoomScrollvor 2 Tagen

finally, open models get a proper leaderboard.

Profilbild von Arman Mikoyan
Arman Mikoyanvor 2 Tagen

👍

Profilbild von Zara Tech & Tools
Zara Tech & Toolsvor 2 Tagen

Love this clear way to compare open models

Profilbild von Hassan
Hassanvor 2 Tagen

Ty!

Profilbild von Marcel Pociot 🧪
Marcel Pociot 🧪vor 2 Tagen

Great idea!

Profilbild von Hassan
Hassanvor 2 Tagen

Thank you, means a lot coming from you!

Profilbild von Paul
Paulvor 2 Tagen

Wonderful!

Profilbild von Stux
Stuxvor 2 Tagen

It would be awesome to have a marking/copy writing category

Profilbild von trioboard
trioboardvor 2 Tagen

this is the question we hit daily. we route by task across three clis: claude code for coordination and review, codex for implementation, antigravity for long documents and video. a per-task cost and quality view like this beats one leaderboard.

Profilbild von Valerio Amirani
Valerio Amiranivor 2 Tagen

Watching where users drop off and fixing broken buttons increases MRR faster than acquiring more cold traffic. That is why we built @ZenovayDX ( with replays and heatmaps in one script. What conversion bottleneck are you working on next?

Profilbild von Javed
Javedvor 2 Tagen

Damn..How fast you ship? It’s like you release every week

Profilbild von Hassan
Hassanvor 2 Tagen

I try 🫡🫡

Profilbild von Rishabh Gurbani
Rishabh Gurbanivor 2 Tagen

comparison on UI/UX design?

Profilbild von Hassan
Hassanvor 2 Tagen

Working on a specific benchmark for this one. More soon :)

Profilbild von Rishabh Gurbani
Rishabh Gurbanivor 2 Tagen

yeah pls make it good, there's not a single reliable benchmark for this across open and proprietary models

Profilbild von XoroAI
XoroAIvor 2 Tagen

What would make this sticky in production: routing by latency SLA, not just benchmark score. A 92-score model at 400ms TTFT loses to an 88 at 40ms for agents. Benchmarks pick winners; latency budgets pick the ones that actually ship.

Profilbild von Arsh Sohal
Arsh Sohalvor 1 Tag

The savings view would be even more useful if it counted the retries and longer outputs a cheaper model sometimes needs to finish the same task. Cost per completed task makes the open vs closed call much easier.

Profilbild von Zain
Zainvor 2 Tagen

🚀🚀

Profilbild von Hassan
Hassanvor 2 Tagen

Great collabing with you on this one!

Profilbild von Aaliya
Aaliyavor 2 Tagen

I’d definitely use this before picking a model for a project.

Profilbild von Curious Nori | AI · Food · Travel
Curious Nori | AI · Food · Travelvor 2 Tagen

Sounds like a handy cheat sheet—I'll definitely check it out to see where I can trim costs without losing quality.

Profilbild von Soumyadeep !
Soumyadeep !vor 2 Tagen

from where are u getting these data ?

Profilbild von Hassan
Hassanvor 2 Tagen

Mostly for this version!

Profilbild von Aman Agarwal
Aman Agarwalvor 2 Tagen

Using jev behind ?

Profilbild von Hassan
Hassanvor 2 Tagen

Not everything uses Jev haha. With that said, I may be releasing something Jev related tomorrow 👀 This site mostly uses benchmarks from ArtificialAnalysis!

Profilbild von Aman Agarwal
Aman Agarwalvor 2 Tagen

kay!

Profilbild von Lucas Gatsas
Lucas Gatsasvor 2 Tagen

Great! :) I have a quick question: Where do you make these videos? Which app do you use? Peace, L

Profilbild von Hassan
Hassanvor 2 Tagen

I use @screenstudio!

Profilbild von Lucas Gatsas
Lucas Gatsasvor 2 Tagen

@screenstudio Thanks

Profilbild von Likhith Bhargav
Likhith Bhargavvor 2 Tagen

Use-case first, model second. Love that this starts from what you’re actually building instead of a leaderboard.

Profilbild von Sheral Tech
Sheral Techvor 1 Tag

Open models just got easier to discover and compare love it.

Profilbild von Mira Takes
Mira Takesvor 2 Tagen

A model leaderboard is only useful if the task mix matches your actual pain. “Best overall” is usually just a polished way to hide the tradeoff.

Profilbild von Logani Bangun
Logani Bangunvor 2 Tagen

Here because of your software factory post. This answers my question on your post directly. Thanks.

Profilbild von Niccolò Banti
Niccolò Bantivor 2 Tagen

Does the quality score blend across coding, agents, vision and finance, or does it stay split by task so a model that's great at coding but weak at long context doesn't wash out in one number?

Profilbild von Mira Takes
Mira Takesvor 2 Tagen

The useful part isn’t another leaderboard—it’s making the cost/quality frontier legible. Once builders can see the tradeoff by task, “open vs closed” stops being ideology and becomes procurement.

Profilbild von Curious Nori | AI · Food · Travel
Curious Nori | AI · Food · Travelvor 2 Tagen

Looks slick! A quick tutorial for newcomers could smooth the onboarding—excited to see what v2 brings.

Profilbild von Webster | JARVIS
Webster | JARVISvor 2 Tagen

Nice! Benchmark tables always get stale, but cost + quality per use case is exactly what I end up needing. Congrats on launching

Profilbild von Chole Syntax Expert AI
Chole Syntax Expert AIvor 1 Tag

One place to compare open models on quality and cost super useful

Profilbild von Denzil | Founder
Denzil | Foundervor 2 Tagen

Neat! What do you use to make your demo videos?

Profilbild von Hassan
Hassanvor 2 Tagen

I use @screenstudio!

Profilbild von Lukas Levert
Lukas Levertvor 1 Tag

Very cool!

Ähnliche Videos

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,068 Aufrufe • vor 1 Monat