Загрузка видео...

Не удалось загрузить видео

На главную

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads...

133,068 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 32

Фото профиля Artificial Analysis
Artificial Analysis1 месяц назад

Build and run your own benchmark today:

Фото профиля Chris 🇨🇦
Chris 🇨🇦1 месяц назад

🐐🐐🐐

Фото профиля Command Code
Command Code1 месяц назад

Congrats y'all!!

Фото профиля Jeet Jain
Jeet Jain1 месяц назад

Woah, this is actually useful, though a very niche case, but can be a better reporting of actual results, on someone's personal use case

Фото профиля Kol Tregaskes
Kol Tregaskes1 месяц назад

Well this could be good timing. 👀

Фото профиля Perceptron.cloud
Perceptron.cloud1 месяц назад

Way to go 🚀

Фото профиля Hermann Saah
Hermann Saah1 месяц назад

That's great 👍🏽

Фото профиля Inflectiv AI ⧉
Inflectiv AI ⧉1 месяц назад

Adding cost and latency makes model evaluation far more practical for production teams. The best model is not always the highest-scoring one.

Фото профиля Asharib Ali
Asharib Ali1 месяц назад

one of the best product for enterprise to find-out the right ai system for their use-cases, cool stuff ❤️

Фото профиля Jacob Beers
Jacob Beers1 месяц назад

Another fantastic idea. You guys are absolutely killing adding practical and useful features to AA.

Фото профиля agenttrader
agenttrader1 месяц назад

@_micah_h Smart. A lot of companies are missing the bigger picture

Фото профиля Branding Waves
Branding Waves1 месяц назад

Benchmarking your own workload is far more useful than generic scores.

Фото профиля Sebastian Buzdugan
Sebastian Buzdugan1 месяц назад

i've seen custom benchmarks become training sets within weeks, how does optima detect leakage

Фото профиля Josh
Josh1 месяц назад

public leaderboards never match my actual workload, so benchmarking on my own traces is the only number I trust. throwing a messy agent trace at this tonight.

Фото профиля Jaime Enrique Lincovil Curivil
Jaime Enrique Lincovil Curivil1 месяц назад

Nice

Фото профиля Whisper Create
Whisper Create1 месяц назад

is this free?

Фото профиля Nitish Pandey
Nitish Pandey1 месяц назад

@grok how is pairwise judging approach (as mentioned in the post) automated ? or is it manual ? how is it done ? tell me in detail

Фото профиля Deyu Wu
Deyu Wu1 месяц назад

This is very useful for companies struggling to catch up with the model race. The difficulty is not selecting the model, it is building your own eval. Great launch!

Фото профиля zhod
zhod1 месяц назад

Good stuff AA

Фото профиля Ella Tech & Tool
Ella Tech & Tool1 месяц назад

This is huge Custom benchmarks cost/time metrics is exactly what teams need

Фото профиля Aiden Milan
Aiden Milan1 месяц назад

Finding a model that’s 10x cheaper or faster for your specific task is more useful than another public leaderboard flex.

Фото профиля Aviv Steiner
Aviv Steiner1 месяц назад

Can I optimize cached input tokens where applicable and take it into account in the benchmark?

Фото профиля GG WP Brawl Stars
GG WP Brawl Stars1 месяц назад

You're on the payroll! You're hiding GLM 5.3.

Фото профиля Ricci Research
Ricci Research1 месяц назад

The real benchmark isn’t who tops a leaderboard—it’s who survives your workload without setting your token budget on fire.

Фото профиля Ian Lucas
Ian Lucas1 месяц назад

Cool idea! When it says "bring your own agent," can that include our own harness? Could we run the same model through multiple harnesses?

Фото профиля Artificial Analysis
Artificial Analysis1 месяц назад

Yes, exactly - you can run a eval on model(s) running in your own agent harness. We expose a HTTP endpoint for you to send outputs to Optima that can then be evaluated

Фото профиля Ian Lucas
Ian Lucas1 месяц назад

nice!

Фото профиля Mr. Walker
Mr. Walker1 месяц назад

The cost-per-task and latency tracking is the part that actually matters. Once people start ranking models on their own agent traces instead of public benches, a lot of the current “top model” rankings are going to look very different.

Фото профиля Code is a Battlefield
Code is a Battlefield1 месяц назад

Optima von Artificial Analysis ist unspektakulärer als „NEUES MODELL!!!“ – und wahrscheinlich wichtiger. Denn öffentliche Benchmarks fragen: Wer gewinnt unseren Test? Eigene Benchmarks fragen: Wer gewinnt meinen tatsächlichen Anwendungsfall? Das klingt nach einer kleinen methodischen Änderung. Für Modellrankings ist es eher der Moment, in dem der Teppich weggezogen wird. Plötzlich zählt nicht mehr, wer auf der Bühne die schönste Zahl hochhält, sondern wer bei der echten Arbeit liefert. Marketingabteilungen entdecken gerade vermutlich eine neue Benchmark: Fluchtgeschwindigkeit zum Notausgang.

Фото профиля ThirdWorldAmerican
ThirdWorldAmerican1 месяц назад

Cool

Фото профиля Maziyar PANAHI
Maziyar PANAHI1 месяц назад

very interesting! i have edge evals for clinical and biomedical tasks, ranging across safety/classification/longevity/medical reasoning. i have kept it private because i don't want this to be contaminated. how private can the evals be on Optima?

Фото профиля Kevin Morales
Kevin Morales1 месяц назад

The use case I'd benchmark first is boring and unglamorous: which model reads a photographed restaurant menu in Spanish and returns clean dish/price JSON. Public leaderboards never tell me that, and it's the only number that decides my model spend.

Похожие видео

Today, we’re pushing a major update to Edison Analysis, our data analysis agent, which is tuned for scientific research and SOTA across data analysis benchmarks. In contrast to Kosmos, which runs for 6-12 hours and produces tens of thousands of lines of code, Edison Analysis runs for seconds to minutes and is best for specific, well-defined computational tasks. It is available both on our platform under the Analysis tab, and via API, and costs only one credit per run, so it is available to users on both free and paid tiers. Edison Analysis is a modified version of the data analysis agent Kosmos uses in its trajectories. Try it out! One of the most important improvements over our previous data analysis agents has been the addition of a specialized data retrieval tool. Edison Analysis can either use this tool to access data, or can pull data down directly via API. To evaluate this tool, we ranked the most commonly used public data repositories across recent papers from BioRxiv, and created a new benchmark that measures the ability of a language agent system to retrieve raw data from those sources. Edison Analysis gets 71% on this benchmark, and we’ll be working to increase this over time. You can read more about our benchmarks in the our blog post, link below. Some features worth highlighting: 1. Edison Analysis produces a report on the analysis it runs, along with a Jupyter notebook that you can download to reproduce the analysis yourself. Every figure it produces is linked back to the specific lines of code used to produce the figure, to make it easy to reproduce. 2. It works well with both Python and R. 3. One of the best uses for Edison Analysis is to use it to retrieve datasets that you can then analyze with Kosmos. We have a bunch of major improvements to Edison Analysis coming in the next few months that we’re excited to share. In the meantime, congratulations to the team, especially Ludovico Mitchener, Jon Laurent, Conor Igoe , Alex Andonian, and many more.

Sam Rodriques

62,017 просмотров • 10 месяцев назад

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 просмотров • 2 лет назад

At DataInsta , we just rebranded InstaAgents to Daiko. In Japanese, Daiko means acting on your behalf. We chose this name because our AI builder creates autonomous agents that take ownership of complex tasks and do the heavy lifting for your business. Most AI projects fail because teams lack the tools, the talent, and a reliable way to evaluate agents. Testing is hard, especially for multiple agent setups and long workflows. We built something different. Daiko is battle tested and evaluated across many industries. With Daiko you can: ◾ Build and deploy agents that automate real workflows ◾ Describe your task and goals so our AI generator can build full multiple agent flows ◾ Choose from over 60 templates and use cases or create your own ◾ Connect your data and APIs, deploy anywhere, and set up evaluations and guardrails ◾ Host the platform on your own servers with white label options We designed this for real use cases in healthcare, deep tech, manufacturing, logistics, retail, legal, finance, and operations. We can handle everything for you. We manage the entire process from strategy to full scale deployment in just a few days. We charge nothing until you are happy with the value we create, and there are no upfront fees. Or, you can do it yourself. You get immediate access to our AI builder and the DataInsta marketplace. Just sign up, post your project, and hire from thousands of vetted AI experts by the hour or per project to lead the build on your terms. More Details :

Abu

28,610 просмотров • 6 месяцев назад