Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads...

133,068 Aufrufe • vor 1 Monat •via X (Twitter)

32 Kommentare

Profilbild von Artificial Analysis
Artificial Analysisvor 1 Monat

Build and run your own benchmark today:

Profilbild von Chris 🇨🇦
Chris 🇨🇦vor 1 Monat

🐐🐐🐐

Profilbild von Command Code
Command Codevor 1 Monat

Congrats y'all!!

Profilbild von Jeet Jain
Jeet Jainvor 1 Monat

Woah, this is actually useful, though a very niche case, but can be a better reporting of actual results, on someone's personal use case

Profilbild von Kol Tregaskes
Kol Tregaskesvor 1 Monat

Well this could be good timing. 👀

Profilbild von Perceptron.cloud
Perceptron.cloudvor 1 Monat

Way to go 🚀

Profilbild von Hermann Saah
Hermann Saahvor 1 Monat

That's great 👍🏽

Profilbild von Inflectiv AI ⧉
Inflectiv AI ⧉vor 1 Monat

Adding cost and latency makes model evaluation far more practical for production teams. The best model is not always the highest-scoring one.

Profilbild von Asharib Ali
Asharib Alivor 1 Monat

one of the best product for enterprise to find-out the right ai system for their use-cases, cool stuff ❤️

Profilbild von Jacob Beers
Jacob Beersvor 1 Monat

Another fantastic idea. You guys are absolutely killing adding practical and useful features to AA.

Profilbild von agenttrader
agenttradervor 1 Monat

@_micah_h Smart. A lot of companies are missing the bigger picture

Profilbild von Branding Waves
Branding Wavesvor 1 Monat

Benchmarking your own workload is far more useful than generic scores.

Profilbild von Sebastian Buzdugan
Sebastian Buzduganvor 1 Monat

i've seen custom benchmarks become training sets within weeks, how does optima detect leakage

Profilbild von Josh
Joshvor 1 Monat

public leaderboards never match my actual workload, so benchmarking on my own traces is the only number I trust. throwing a messy agent trace at this tonight.

Profilbild von Jaime Enrique Lincovil Curivil
Jaime Enrique Lincovil Curivilvor 1 Monat

Nice

Profilbild von Whisper Create
Whisper Createvor 1 Monat

is this free?

Profilbild von Nitish Pandey
Nitish Pandeyvor 1 Monat

@grok how is pairwise judging approach (as mentioned in the post) automated ? or is it manual ? how is it done ? tell me in detail

Profilbild von Deyu Wu
Deyu Wuvor 1 Monat

This is very useful for companies struggling to catch up with the model race. The difficulty is not selecting the model, it is building your own eval. Great launch!

Profilbild von zhod
zhodvor 1 Monat

Good stuff AA

Profilbild von Ella Tech & Tool
Ella Tech & Toolvor 1 Monat

This is huge Custom benchmarks cost/time metrics is exactly what teams need

Profilbild von Aiden Milan
Aiden Milanvor 1 Monat

Finding a model that’s 10x cheaper or faster for your specific task is more useful than another public leaderboard flex.

Profilbild von Aviv Steiner
Aviv Steinervor 1 Monat

Can I optimize cached input tokens where applicable and take it into account in the benchmark?

Profilbild von GG WP Brawl Stars
GG WP Brawl Starsvor 1 Monat

You're on the payroll! You're hiding GLM 5.3.

Profilbild von Ricci Research
Ricci Researchvor 1 Monat

The real benchmark isn’t who tops a leaderboard—it’s who survives your workload without setting your token budget on fire.

Profilbild von Ian Lucas
Ian Lucasvor 1 Monat

Cool idea! When it says "bring your own agent," can that include our own harness? Could we run the same model through multiple harnesses?

Profilbild von Artificial Analysis
Artificial Analysisvor 1 Monat

Yes, exactly - you can run a eval on model(s) running in your own agent harness. We expose a HTTP endpoint for you to send outputs to Optima that can then be evaluated

Profilbild von Ian Lucas
Ian Lucasvor 1 Monat

nice!

Profilbild von Mr. Walker
Mr. Walkervor 1 Monat

The cost-per-task and latency tracking is the part that actually matters. Once people start ranking models on their own agent traces instead of public benches, a lot of the current “top model” rankings are going to look very different.

Profilbild von Code is a Battlefield
Code is a Battlefieldvor 1 Monat

Optima von Artificial Analysis ist unspektakulärer als „NEUES MODELL!!!“ – und wahrscheinlich wichtiger. Denn öffentliche Benchmarks fragen: Wer gewinnt unseren Test? Eigene Benchmarks fragen: Wer gewinnt meinen tatsächlichen Anwendungsfall? Das klingt nach einer kleinen methodischen Änderung. Für Modellrankings ist es eher der Moment, in dem der Teppich weggezogen wird. Plötzlich zählt nicht mehr, wer auf der Bühne die schönste Zahl hochhält, sondern wer bei der echten Arbeit liefert. Marketingabteilungen entdecken gerade vermutlich eine neue Benchmark: Fluchtgeschwindigkeit zum Notausgang.

Profilbild von ThirdWorldAmerican
ThirdWorldAmericanvor 1 Monat

Cool

Profilbild von Maziyar PANAHI
Maziyar PANAHIvor 1 Monat

very interesting! i have edge evals for clinical and biomedical tasks, ranging across safety/classification/longevity/medical reasoning. i have kept it private because i don't want this to be contaminated. how private can the evals be on Optima?

Profilbild von Kevin Morales
Kevin Moralesvor 1 Monat

The use case I'd benchmark first is boring and unglamorous: which model reads a photographed restaurant menu in Spanish and returns clean dish/price JSON. Public leaderboards never tell me that, and it's the only number that decides my model spend.

Ähnliche Videos

Today, we’re pushing a major update to Edison Analysis, our data analysis agent, which is tuned for scientific research and SOTA across data analysis benchmarks. In contrast to Kosmos, which runs for 6-12 hours and produces tens of thousands of lines of code, Edison Analysis runs for seconds to minutes and is best for specific, well-defined computational tasks. It is available both on our platform under the Analysis tab, and via API, and costs only one credit per run, so it is available to users on both free and paid tiers. Edison Analysis is a modified version of the data analysis agent Kosmos uses in its trajectories. Try it out! One of the most important improvements over our previous data analysis agents has been the addition of a specialized data retrieval tool. Edison Analysis can either use this tool to access data, or can pull data down directly via API. To evaluate this tool, we ranked the most commonly used public data repositories across recent papers from BioRxiv, and created a new benchmark that measures the ability of a language agent system to retrieve raw data from those sources. Edison Analysis gets 71% on this benchmark, and we’ll be working to increase this over time. You can read more about our benchmarks in the our blog post, link below. Some features worth highlighting: 1. Edison Analysis produces a report on the analysis it runs, along with a Jupyter notebook that you can download to reproduce the analysis yourself. Every figure it produces is linked back to the specific lines of code used to produce the figure, to make it easy to reproduce. 2. It works well with both Python and R. 3. One of the best uses for Edison Analysis is to use it to retrieve datasets that you can then analyze with Kosmos. We have a bunch of major improvements to Edison Analysis coming in the next few months that we’re excited to share. In the meantime, congratulations to the team, especially Ludovico Mitchener, Jon Laurent, Conor Igoe , Alex Andonian, and many more.

Sam Rodriques

62,017 Aufrufe • vor 10 Monaten

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 Aufrufe • vor 2 Jahren

At DataInsta , we just rebranded InstaAgents to Daiko. In Japanese, Daiko means acting on your behalf. We chose this name because our AI builder creates autonomous agents that take ownership of complex tasks and do the heavy lifting for your business. Most AI projects fail because teams lack the tools, the talent, and a reliable way to evaluate agents. Testing is hard, especially for multiple agent setups and long workflows. We built something different. Daiko is battle tested and evaluated across many industries. With Daiko you can: ◾ Build and deploy agents that automate real workflows ◾ Describe your task and goals so our AI generator can build full multiple agent flows ◾ Choose from over 60 templates and use cases or create your own ◾ Connect your data and APIs, deploy anywhere, and set up evaluations and guardrails ◾ Host the platform on your own servers with white label options We designed this for real use cases in healthcare, deep tech, manufacturing, logistics, retail, legal, finance, and operations. We can handle everything for you. We manage the entire process from strategy to full scale deployment in just a few days. We charge nothing until you are happy with the value we create, and there are no upfront fees. Or, you can do it yourself. You get immediate access to our AI builder and the DataInsta marketplace. Just sign up, post your project, and hire from thousands of vetted AI experts by the hour or per project to lead the build on your terms. More Details :

Abu

28,610 Aufrufe • vor 6 Monaten