正在加载视频...

视频加载失败

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads...

133,068 次观看 • 1 个月前 •via X (Twitter)

32 条评论

Artificial Analysis 的头像
Artificial Analysis1 个月前

Build and run your own benchmark today:

Chris 🇨🇦 的头像
Chris 🇨🇦1 个月前

🐐🐐🐐

Command Code 的头像
Command Code1 个月前

Congrats y'all!!

Jeet Jain 的头像
Jeet Jain1 个月前

Woah, this is actually useful, though a very niche case, but can be a better reporting of actual results, on someone's personal use case

Kol Tregaskes 的头像
Kol Tregaskes1 个月前

Well this could be good timing. 👀

Perceptron.cloud 的头像
Perceptron.cloud1 个月前

Way to go 🚀

Hermann Saah 的头像
Hermann Saah1 个月前

That's great 👍🏽

Inflectiv AI ⧉ 的头像
Inflectiv AI ⧉1 个月前

Adding cost and latency makes model evaluation far more practical for production teams. The best model is not always the highest-scoring one.

Asharib Ali 的头像
Asharib Ali1 个月前

one of the best product for enterprise to find-out the right ai system for their use-cases, cool stuff ❤️

Jacob Beers 的头像
Jacob Beers1 个月前

Another fantastic idea. You guys are absolutely killing adding practical and useful features to AA.

agenttrader 的头像
agenttrader1 个月前

@_micah_h Smart. A lot of companies are missing the bigger picture

Branding Waves 的头像
Branding Waves1 个月前

Benchmarking your own workload is far more useful than generic scores.

Sebastian Buzdugan 的头像
Sebastian Buzdugan1 个月前

i've seen custom benchmarks become training sets within weeks, how does optima detect leakage

Josh 的头像
Josh1 个月前

public leaderboards never match my actual workload, so benchmarking on my own traces is the only number I trust. throwing a messy agent trace at this tonight.

Jaime Enrique Lincovil Curivil 的头像
Jaime Enrique Lincovil Curivil1 个月前

Nice

Whisper Create 的头像
Whisper Create1 个月前

is this free?

Nitish Pandey 的头像
Nitish Pandey1 个月前

@grok how is pairwise judging approach (as mentioned in the post) automated ? or is it manual ? how is it done ? tell me in detail

Deyu Wu 的头像
Deyu Wu1 个月前

This is very useful for companies struggling to catch up with the model race. The difficulty is not selecting the model, it is building your own eval. Great launch!

zhod 的头像
zhod1 个月前

Good stuff AA

Ella Tech & Tool 的头像
Ella Tech & Tool1 个月前

This is huge Custom benchmarks cost/time metrics is exactly what teams need

Aiden Milan 的头像
Aiden Milan1 个月前

Finding a model that’s 10x cheaper or faster for your specific task is more useful than another public leaderboard flex.

Aviv Steiner 的头像
Aviv Steiner1 个月前

Can I optimize cached input tokens where applicable and take it into account in the benchmark?

GG WP Brawl Stars 的头像
GG WP Brawl Stars1 个月前

You're on the payroll! You're hiding GLM 5.3.

Ricci Research 的头像
Ricci Research1 个月前

The real benchmark isn’t who tops a leaderboard—it’s who survives your workload without setting your token budget on fire.

Ian Lucas 的头像
Ian Lucas1 个月前

Cool idea! When it says "bring your own agent," can that include our own harness? Could we run the same model through multiple harnesses?

Artificial Analysis 的头像
Artificial Analysis1 个月前

Yes, exactly - you can run a eval on model(s) running in your own agent harness. We expose a HTTP endpoint for you to send outputs to Optima that can then be evaluated

Ian Lucas 的头像
Ian Lucas1 个月前

nice!

Mr. Walker 的头像
Mr. Walker1 个月前

The cost-per-task and latency tracking is the part that actually matters. Once people start ranking models on their own agent traces instead of public benches, a lot of the current “top model” rankings are going to look very different.

Code is a Battlefield 的头像
Code is a Battlefield1 个月前

Optima von Artificial Analysis ist unspektakulärer als „NEUES MODELL!!!“ – und wahrscheinlich wichtiger. Denn öffentliche Benchmarks fragen: Wer gewinnt unseren Test? Eigene Benchmarks fragen: Wer gewinnt meinen tatsächlichen Anwendungsfall? Das klingt nach einer kleinen methodischen Änderung. Für Modellrankings ist es eher der Moment, in dem der Teppich weggezogen wird. Plötzlich zählt nicht mehr, wer auf der Bühne die schönste Zahl hochhält, sondern wer bei der echten Arbeit liefert. Marketingabteilungen entdecken gerade vermutlich eine neue Benchmark: Fluchtgeschwindigkeit zum Notausgang.

ThirdWorldAmerican 的头像
ThirdWorldAmerican1 个月前

Cool

Maziyar PANAHI 的头像
Maziyar PANAHI1 个月前

very interesting! i have edge evals for clinical and biomedical tasks, ranging across safety/classification/longevity/medical reasoning. i have kept it private because i don't want this to be contaminated. how private can the evals be on Optima?

Kevin Morales 的头像
Kevin Morales1 个月前

The use case I'd benchmark first is boring and unglamorous: which model reads a photographed restaurant menu in Spanish and returns clean dish/price JSON. Public leaderboards never tell me that, and it's the only number that decides my model spend.

相关视频

Today, we’re pushing a major update to Edison Analysis, our data analysis agent, which is tuned for scientific research and SOTA across data analysis benchmarks. In contrast to Kosmos, which runs for 6-12 hours and produces tens of thousands of lines of code, Edison Analysis runs for seconds to minutes and is best for specific, well-defined computational tasks. It is available both on our platform under the Analysis tab, and via API, and costs only one credit per run, so it is available to users on both free and paid tiers. Edison Analysis is a modified version of the data analysis agent Kosmos uses in its trajectories. Try it out! One of the most important improvements over our previous data analysis agents has been the addition of a specialized data retrieval tool. Edison Analysis can either use this tool to access data, or can pull data down directly via API. To evaluate this tool, we ranked the most commonly used public data repositories across recent papers from BioRxiv, and created a new benchmark that measures the ability of a language agent system to retrieve raw data from those sources. Edison Analysis gets 71% on this benchmark, and we’ll be working to increase this over time. You can read more about our benchmarks in the our blog post, link below. Some features worth highlighting: 1. Edison Analysis produces a report on the analysis it runs, along with a Jupyter notebook that you can download to reproduce the analysis yourself. Every figure it produces is linked back to the specific lines of code used to produce the figure, to make it easy to reproduce. 2. It works well with both Python and R. 3. One of the best uses for Edison Analysis is to use it to retrieve datasets that you can then analyze with Kosmos. We have a bunch of major improvements to Edison Analysis coming in the next few months that we’re excited to share. In the meantime, congratulations to the team, especially Ludovico Mitchener, Jon Laurent, Conor Igoe , Alex Andonian, and many more.

Sam Rodriques

62,017 次观看 • 10 个月前

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 次观看 • 2 年前

At DataInsta , we just rebranded InstaAgents to Daiko. In Japanese, Daiko means acting on your behalf. We chose this name because our AI builder creates autonomous agents that take ownership of complex tasks and do the heavy lifting for your business. Most AI projects fail because teams lack the tools, the talent, and a reliable way to evaluate agents. Testing is hard, especially for multiple agent setups and long workflows. We built something different. Daiko is battle tested and evaluated across many industries. With Daiko you can: ◾ Build and deploy agents that automate real workflows ◾ Describe your task and goals so our AI generator can build full multiple agent flows ◾ Choose from over 60 templates and use cases or create your own ◾ Connect your data and APIs, deploy anywhere, and set up evaluations and guardrails ◾ Host the platform on your own servers with white label options We designed this for real use cases in healthcare, deep tech, manufacturing, logistics, retail, legal, finance, and operations. We can handle everything for you. We manage the entire process from strategy to full scale deployment in just a few days. We charge nothing until you are happy with the value we create, and there are no upfront fees. Or, you can do it yourself. You get immediate access to our AI builder and the DataInsta marketplace. Just sign up, post your project, and hire from thousands of vetted AI experts by the hour or per project to lead the build on your terms. More Details :

Abu

28,610 次观看 • 6 个月前