Loading video...

Video Failed to Load

Go Home

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads...

133,068 views • 1 month ago •via X (Twitter)

32 Comments

Artificial Analysis's profile picture
Artificial Analysis1 month ago

Build and run your own benchmark today:

Chris 🇨🇦's profile picture
Chris 🇨🇦1 month ago

🐐🐐🐐

Command Code's profile picture
Command Code1 month ago

Congrats y'all!!

Jeet Jain's profile picture
Jeet Jain1 month ago

Woah, this is actually useful, though a very niche case, but can be a better reporting of actual results, on someone's personal use case

Kol Tregaskes's profile picture
Kol Tregaskes1 month ago

Well this could be good timing. 👀

Perceptron.cloud's profile picture
Perceptron.cloud1 month ago

Way to go 🚀

Hermann Saah's profile picture
Hermann Saah1 month ago

That's great 👍🏽

Inflectiv AI ⧉'s profile picture
Inflectiv AI ⧉1 month ago

Adding cost and latency makes model evaluation far more practical for production teams. The best model is not always the highest-scoring one.

Asharib Ali's profile picture
Asharib Ali1 month ago

one of the best product for enterprise to find-out the right ai system for their use-cases, cool stuff ❤️

Jacob Beers's profile picture
Jacob Beers1 month ago

Another fantastic idea. You guys are absolutely killing adding practical and useful features to AA.

agenttrader's profile picture
agenttrader1 month ago

@_micah_h Smart. A lot of companies are missing the bigger picture

Branding Waves's profile picture
Branding Waves1 month ago

Benchmarking your own workload is far more useful than generic scores.

Sebastian Buzdugan's profile picture
Sebastian Buzdugan1 month ago

i've seen custom benchmarks become training sets within weeks, how does optima detect leakage

Josh's profile picture
Josh1 month ago

public leaderboards never match my actual workload, so benchmarking on my own traces is the only number I trust. throwing a messy agent trace at this tonight.

Jaime Enrique Lincovil Curivil's profile picture
Jaime Enrique Lincovil Curivil1 month ago

Nice

Whisper Create's profile picture
Whisper Create1 month ago

is this free?

Nitish Pandey's profile picture
Nitish Pandey1 month ago

@grok how is pairwise judging approach (as mentioned in the post) automated ? or is it manual ? how is it done ? tell me in detail

Deyu Wu's profile picture
Deyu Wu1 month ago

This is very useful for companies struggling to catch up with the model race. The difficulty is not selecting the model, it is building your own eval. Great launch!

zhod's profile picture
zhod1 month ago

Good stuff AA

Ella Tech & Tool's profile picture
Ella Tech & Tool1 month ago

This is huge Custom benchmarks cost/time metrics is exactly what teams need

Aiden Milan's profile picture
Aiden Milan1 month ago

Finding a model that’s 10x cheaper or faster for your specific task is more useful than another public leaderboard flex.

Aviv Steiner's profile picture
Aviv Steiner1 month ago

Can I optimize cached input tokens where applicable and take it into account in the benchmark?

GG WP Brawl Stars's profile picture
GG WP Brawl Stars1 month ago

You're on the payroll! You're hiding GLM 5.3.

Ricci Research's profile picture
Ricci Research1 month ago

The real benchmark isn’t who tops a leaderboard—it’s who survives your workload without setting your token budget on fire.

Ian Lucas's profile picture
Ian Lucas1 month ago

Cool idea! When it says "bring your own agent," can that include our own harness? Could we run the same model through multiple harnesses?

Artificial Analysis's profile picture
Artificial Analysis1 month ago

Yes, exactly - you can run a eval on model(s) running in your own agent harness. We expose a HTTP endpoint for you to send outputs to Optima that can then be evaluated

Ian Lucas's profile picture
Ian Lucas1 month ago

nice!

Mr. Walker's profile picture
Mr. Walker1 month ago

The cost-per-task and latency tracking is the part that actually matters. Once people start ranking models on their own agent traces instead of public benches, a lot of the current “top model” rankings are going to look very different.

Code is a Battlefield's profile picture
Code is a Battlefield1 month ago

Optima von Artificial Analysis ist unspektakulärer als „NEUES MODELL!!!“ – und wahrscheinlich wichtiger. Denn öffentliche Benchmarks fragen: Wer gewinnt unseren Test? Eigene Benchmarks fragen: Wer gewinnt meinen tatsächlichen Anwendungsfall? Das klingt nach einer kleinen methodischen Änderung. Für Modellrankings ist es eher der Moment, in dem der Teppich weggezogen wird. Plötzlich zählt nicht mehr, wer auf der Bühne die schönste Zahl hochhält, sondern wer bei der echten Arbeit liefert. Marketingabteilungen entdecken gerade vermutlich eine neue Benchmark: Fluchtgeschwindigkeit zum Notausgang.

ThirdWorldAmerican's profile picture
ThirdWorldAmerican1 month ago

Cool

Maziyar PANAHI's profile picture
Maziyar PANAHI1 month ago

very interesting! i have edge evals for clinical and biomedical tasks, ranging across safety/classification/longevity/medical reasoning. i have kept it private because i don't want this to be contaminated. how private can the evals be on Optima?

Kevin Morales's profile picture
Kevin Morales1 month ago

The use case I'd benchmark first is boring and unglamorous: which model reads a photographed restaurant menu in Spanish and returns clean dish/price JSON. Public leaderboards never tell me that, and it's the only number that decides my model spend.

Related Videos

Today, we’re pushing a major update to Edison Analysis, our data analysis agent, which is tuned for scientific research and SOTA across data analysis benchmarks. In contrast to Kosmos, which runs for 6-12 hours and produces tens of thousands of lines of code, Edison Analysis runs for seconds to minutes and is best for specific, well-defined computational tasks. It is available both on our platform under the Analysis tab, and via API, and costs only one credit per run, so it is available to users on both free and paid tiers. Edison Analysis is a modified version of the data analysis agent Kosmos uses in its trajectories. Try it out! One of the most important improvements over our previous data analysis agents has been the addition of a specialized data retrieval tool. Edison Analysis can either use this tool to access data, or can pull data down directly via API. To evaluate this tool, we ranked the most commonly used public data repositories across recent papers from BioRxiv, and created a new benchmark that measures the ability of a language agent system to retrieve raw data from those sources. Edison Analysis gets 71% on this benchmark, and we’ll be working to increase this over time. You can read more about our benchmarks in the our blog post, link below. Some features worth highlighting: 1. Edison Analysis produces a report on the analysis it runs, along with a Jupyter notebook that you can download to reproduce the analysis yourself. Every figure it produces is linked back to the specific lines of code used to produce the figure, to make it easy to reproduce. 2. It works well with both Python and R. 3. One of the best uses for Edison Analysis is to use it to retrieve datasets that you can then analyze with Kosmos. We have a bunch of major improvements to Edison Analysis coming in the next few months that we’re excited to share. In the meantime, congratulations to the team, especially Ludovico Mitchener, Jon Laurent, Conor Igoe , Alex Andonian, and many more.

Sam Rodriques

62,017 views • 10 months ago

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 views • 2 years ago

At DataInsta , we just rebranded InstaAgents to Daiko. In Japanese, Daiko means acting on your behalf. We chose this name because our AI builder creates autonomous agents that take ownership of complex tasks and do the heavy lifting for your business. Most AI projects fail because teams lack the tools, the talent, and a reliable way to evaluate agents. Testing is hard, especially for multiple agent setups and long workflows. We built something different. Daiko is battle tested and evaluated across many industries. With Daiko you can: ◾ Build and deploy agents that automate real workflows ◾ Describe your task and goals so our AI generator can build full multiple agent flows ◾ Choose from over 60 templates and use cases or create your own ◾ Connect your data and APIs, deploy anywhere, and set up evaluations and guardrails ◾ Host the platform on your own servers with white label options We designed this for real use cases in healthcare, deep tech, manufacturing, logistics, retail, legal, finance, and operations. We can handle everything for you. We manage the entire process from strategy to full scale deployment in just a few days. We charge nothing until you are happy with the value we create, and there are no upfront fees. Or, you can do it yourself. You get immediate access to our AI builder and the DataInsta marketplace. Just sign up, post your project, and hire from thousands of vetted AI experts by the hour or per project to lead the build on your terms. More Details :

Abu

28,610 views • 6 months ago