Video yükleniyor...
Video Yüklenemedi
Introducing Figure out which open model is best for your use case. Compare models across coding, agents, long context, vision, finance, and more. Then see how they compare on cost + quality, including what you could save by moving to open models.
17,354 görüntüleme • 2 gün önce •via X (Twitter)
42 Yorum

Built this with @zainhas & @YoussefUiUx over the last few weeks! Would love any feedback on what we can improve. We’re already working on v2 to make it even easier to find the right model for your use case.

finally, open models get a proper leaderboard.

👍

Love this clear way to compare open models

Ty!

Great idea!

Thank you, means a lot coming from you!

Wonderful!

It would be awesome to have a marking/copy writing category

this is the question we hit daily. we route by task across three clis: claude code for coordination and review, codex for implementation, antigravity for long documents and video. a per-task cost and quality view like this beats one leaderboard.

Watching where users drop off and fixing broken buttons increases MRR faster than acquiring more cold traffic. That is why we built @ZenovayDX ( with replays and heatmaps in one script. What conversion bottleneck are you working on next?

Damn..How fast you ship? It’s like you release every week

I try 🫡🫡

comparison on UI/UX design?

Working on a specific benchmark for this one. More soon :)

yeah pls make it good, there's not a single reliable benchmark for this across open and proprietary models

What would make this sticky in production: routing by latency SLA, not just benchmark score. A 92-score model at 400ms TTFT loses to an 88 at 40ms for agents. Benchmarks pick winners; latency budgets pick the ones that actually ship.

The savings view would be even more useful if it counted the retries and longer outputs a cheaper model sometimes needs to finish the same task. Cost per completed task makes the open vs closed call much easier.

🚀🚀

Great collabing with you on this one!

I’d definitely use this before picking a model for a project.

Sounds like a handy cheat sheet—I'll definitely check it out to see where I can trim costs without losing quality.

from where are u getting these data ?

Mostly for this version!

Using jev behind ?

Not everything uses Jev haha. With that said, I may be releasing something Jev related tomorrow 👀 This site mostly uses benchmarks from ArtificialAnalysis!

kay!

Great! :) I have a quick question: Where do you make these videos? Which app do you use? Peace, L

I use @screenstudio!

@screenstudio Thanks

Use-case first, model second. Love that this starts from what you’re actually building instead of a leaderboard.

Open models just got easier to discover and compare love it.

A model leaderboard is only useful if the task mix matches your actual pain. “Best overall” is usually just a polished way to hide the tradeoff.

Here because of your software factory post. This answers my question on your post directly. Thanks.

Does the quality score blend across coding, agents, vision and finance, or does it stay split by task so a model that's great at coding but weak at long context doesn't wash out in one number?

The useful part isn’t another leaderboard—it’s making the cost/quality frontier legible. Once builders can see the tradeoff by task, “open vs closed” stops being ideology and becomes procurement.

Looks slick! A quick tutorial for newcomers could smooth the onboarding—excited to see what v2 brings.

Nice! Benchmark tables always get stale, but cost + quality per use case is exactly what I end up needing. Congrats on launching

One place to compare open models on quality and cost super useful

Neat! What do you use to make your demo videos?

I use @screenstudio!

Very cool!



