Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Figure out which open model is best for your use case. Compare models across coding, agents, long context, vision, finance, and more. Then see how they compare on cost + quality, including what you could save by moving to open models.

17,354 görüntüleme • 2 gün önce •via X (Twitter)

42 Yorum

Hassan profil fotoğrafı
Hassan2 gün önce

Built this with @zainhas & @YoussefUiUx over the last few weeks! Would love any feedback on what we can improve. We’re already working on v2 to make it even easier to find the right model for your use case.

AIDoomScroll profil fotoğrafı
AIDoomScroll2 gün önce

finally, open models get a proper leaderboard.

Arman Mikoyan profil fotoğrafı
Arman Mikoyan2 gün önce

👍

Zara Tech & Tools profil fotoğrafı
Zara Tech & Tools2 gün önce

Love this clear way to compare open models

Hassan profil fotoğrafı
Hassan2 gün önce

Ty!

Marcel Pociot 🧪 profil fotoğrafı
Marcel Pociot 🧪2 gün önce

Great idea!

Hassan profil fotoğrafı
Hassan2 gün önce

Thank you, means a lot coming from you!

Paul profil fotoğrafı
Paul2 gün önce

Wonderful!

Stux profil fotoğrafı
Stux2 gün önce

It would be awesome to have a marking/copy writing category

trioboard profil fotoğrafı
trioboard2 gün önce

this is the question we hit daily. we route by task across three clis: claude code for coordination and review, codex for implementation, antigravity for long documents and video. a per-task cost and quality view like this beats one leaderboard.

Valerio Amirani profil fotoğrafı
Valerio Amirani2 gün önce

Watching where users drop off and fixing broken buttons increases MRR faster than acquiring more cold traffic. That is why we built @ZenovayDX ( with replays and heatmaps in one script. What conversion bottleneck are you working on next?

Javed profil fotoğrafı
Javed2 gün önce

Damn..How fast you ship? It’s like you release every week

Hassan profil fotoğrafı
Hassan2 gün önce

I try 🫡🫡

Rishabh Gurbani profil fotoğrafı
Rishabh Gurbani2 gün önce

comparison on UI/UX design?

Hassan profil fotoğrafı
Hassan2 gün önce

Working on a specific benchmark for this one. More soon :)

Rishabh Gurbani profil fotoğrafı
Rishabh Gurbani2 gün önce

yeah pls make it good, there's not a single reliable benchmark for this across open and proprietary models

XoroAI profil fotoğrafı
XoroAI2 gün önce

What would make this sticky in production: routing by latency SLA, not just benchmark score. A 92-score model at 400ms TTFT loses to an 88 at 40ms for agents. Benchmarks pick winners; latency budgets pick the ones that actually ship.

Arsh Sohal profil fotoğrafı
Arsh Sohal1 gün önce

The savings view would be even more useful if it counted the retries and longer outputs a cheaper model sometimes needs to finish the same task. Cost per completed task makes the open vs closed call much easier.

Zain profil fotoğrafı
Zain2 gün önce

🚀🚀

Hassan profil fotoğrafı
Hassan2 gün önce

Great collabing with you on this one!

Aaliya profil fotoğrafı
Aaliya2 gün önce

I’d definitely use this before picking a model for a project.

Curious Nori | AI · Food · Travel profil fotoğrafı
Curious Nori | AI · Food · Travel2 gün önce

Sounds like a handy cheat sheet—I'll definitely check it out to see where I can trim costs without losing quality.

Soumyadeep ! profil fotoğrafı
Soumyadeep !2 gün önce

from where are u getting these data ?

Hassan profil fotoğrafı
Hassan2 gün önce

Mostly for this version!

Aman Agarwal profil fotoğrafı
Aman Agarwal2 gün önce

Using jev behind ?

Hassan profil fotoğrafı
Hassan2 gün önce

Not everything uses Jev haha. With that said, I may be releasing something Jev related tomorrow 👀 This site mostly uses benchmarks from ArtificialAnalysis!

Aman Agarwal profil fotoğrafı
Aman Agarwal2 gün önce

kay!

Lucas Gatsas profil fotoğrafı
Lucas Gatsas2 gün önce

Great! :) I have a quick question: Where do you make these videos? Which app do you use? Peace, L

Hassan profil fotoğrafı
Hassan2 gün önce

I use @screenstudio!

Lucas Gatsas profil fotoğrafı
Lucas Gatsas2 gün önce

@screenstudio Thanks

Likhith Bhargav profil fotoğrafı
Likhith Bhargav2 gün önce

Use-case first, model second. Love that this starts from what you’re actually building instead of a leaderboard.

Sheral Tech profil fotoğrafı
Sheral Tech2 gün önce

Open models just got easier to discover and compare love it.

Mira Takes profil fotoğrafı
Mira Takes2 gün önce

A model leaderboard is only useful if the task mix matches your actual pain. “Best overall” is usually just a polished way to hide the tradeoff.

Logani Bangun profil fotoğrafı
Logani Bangun2 gün önce

Here because of your software factory post. This answers my question on your post directly. Thanks.

Niccolò Banti profil fotoğrafı
Niccolò Banti2 gün önce

Does the quality score blend across coding, agents, vision and finance, or does it stay split by task so a model that's great at coding but weak at long context doesn't wash out in one number?

Mira Takes profil fotoğrafı
Mira Takes2 gün önce

The useful part isn’t another leaderboard—it’s making the cost/quality frontier legible. Once builders can see the tradeoff by task, “open vs closed” stops being ideology and becomes procurement.

Curious Nori | AI · Food · Travel profil fotoğrafı
Curious Nori | AI · Food · Travel2 gün önce

Looks slick! A quick tutorial for newcomers could smooth the onboarding—excited to see what v2 brings.

Webster | JARVIS profil fotoğrafı
Webster | JARVIS2 gün önce

Nice! Benchmark tables always get stale, but cost + quality per use case is exactly what I end up needing. Congrats on launching

Chole Syntax Expert AI profil fotoğrafı
Chole Syntax Expert AI1 gün önce

One place to compare open models on quality and cost super useful

Denzil | Founder profil fotoğrafı
Denzil | Founder2 gün önce

Neat! What do you use to make your demo videos?

Hassan profil fotoğrafı
Hassan2 gün önce

I use @screenstudio!

Lukas Levert profil fotoğrafı
Lukas Levert2 gün önce

Very cool!

Benzer Videolar

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,068 görüntüleme • 1 ay önce