Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads...

133,068 görüntüleme • 1 ay önce •via X (Twitter)

32 Yorum

Artificial Analysis profil fotoğrafı
Artificial Analysis1 ay önce

Build and run your own benchmark today:

Chris 🇨🇦 profil fotoğrafı
Chris 🇨🇦1 ay önce

🐐🐐🐐

Command Code profil fotoğrafı
Command Code1 ay önce

Congrats y'all!!

Jeet Jain profil fotoğrafı
Jeet Jain1 ay önce

Woah, this is actually useful, though a very niche case, but can be a better reporting of actual results, on someone's personal use case

Kol Tregaskes profil fotoğrafı
Kol Tregaskes1 ay önce

Well this could be good timing. 👀

Perceptron.cloud profil fotoğrafı
Perceptron.cloud1 ay önce

Way to go 🚀

Hermann Saah profil fotoğrafı
Hermann Saah1 ay önce

That's great 👍🏽

Inflectiv AI ⧉ profil fotoğrafı
Inflectiv AI ⧉1 ay önce

Adding cost and latency makes model evaluation far more practical for production teams. The best model is not always the highest-scoring one.

Asharib Ali profil fotoğrafı
Asharib Ali1 ay önce

one of the best product for enterprise to find-out the right ai system for their use-cases, cool stuff ❤️

Jacob Beers profil fotoğrafı
Jacob Beers1 ay önce

Another fantastic idea. You guys are absolutely killing adding practical and useful features to AA.

agenttrader profil fotoğrafı
agenttrader1 ay önce

@_micah_h Smart. A lot of companies are missing the bigger picture

Branding Waves profil fotoğrafı
Branding Waves1 ay önce

Benchmarking your own workload is far more useful than generic scores.

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan1 ay önce

i've seen custom benchmarks become training sets within weeks, how does optima detect leakage

Josh profil fotoğrafı
Josh1 ay önce

public leaderboards never match my actual workload, so benchmarking on my own traces is the only number I trust. throwing a messy agent trace at this tonight.

Jaime Enrique Lincovil Curivil profil fotoğrafı
Jaime Enrique Lincovil Curivil1 ay önce

Nice

Whisper Create profil fotoğrafı
Whisper Create1 ay önce

is this free?

Nitish Pandey profil fotoğrafı
Nitish Pandey1 ay önce

@grok how is pairwise judging approach (as mentioned in the post) automated ? or is it manual ? how is it done ? tell me in detail

Deyu Wu profil fotoğrafı
Deyu Wu1 ay önce

This is very useful for companies struggling to catch up with the model race. The difficulty is not selecting the model, it is building your own eval. Great launch!

zhod profil fotoğrafı
zhod1 ay önce

Good stuff AA

Ella Tech & Tool profil fotoğrafı
Ella Tech & Tool1 ay önce

This is huge Custom benchmarks cost/time metrics is exactly what teams need

Aiden Milan profil fotoğrafı
Aiden Milan1 ay önce

Finding a model that’s 10x cheaper or faster for your specific task is more useful than another public leaderboard flex.

Aviv Steiner profil fotoğrafı
Aviv Steiner1 ay önce

Can I optimize cached input tokens where applicable and take it into account in the benchmark?

GG WP Brawl Stars profil fotoğrafı
GG WP Brawl Stars1 ay önce

You're on the payroll! You're hiding GLM 5.3.

Ricci Research profil fotoğrafı
Ricci Research1 ay önce

The real benchmark isn’t who tops a leaderboard—it’s who survives your workload without setting your token budget on fire.

Ian Lucas profil fotoğrafı
Ian Lucas1 ay önce

Cool idea! When it says "bring your own agent," can that include our own harness? Could we run the same model through multiple harnesses?

Artificial Analysis profil fotoğrafı
Artificial Analysis1 ay önce

Yes, exactly - you can run a eval on model(s) running in your own agent harness. We expose a HTTP endpoint for you to send outputs to Optima that can then be evaluated

Ian Lucas profil fotoğrafı
Ian Lucas1 ay önce

nice!

Mr. Walker profil fotoğrafı
Mr. Walker1 ay önce

The cost-per-task and latency tracking is the part that actually matters. Once people start ranking models on their own agent traces instead of public benches, a lot of the current “top model” rankings are going to look very different.

Code is a Battlefield profil fotoğrafı
Code is a Battlefield1 ay önce

Optima von Artificial Analysis ist unspektakulärer als „NEUES MODELL!!!“ – und wahrscheinlich wichtiger. Denn öffentliche Benchmarks fragen: Wer gewinnt unseren Test? Eigene Benchmarks fragen: Wer gewinnt meinen tatsächlichen Anwendungsfall? Das klingt nach einer kleinen methodischen Änderung. Für Modellrankings ist es eher der Moment, in dem der Teppich weggezogen wird. Plötzlich zählt nicht mehr, wer auf der Bühne die schönste Zahl hochhält, sondern wer bei der echten Arbeit liefert. Marketingabteilungen entdecken gerade vermutlich eine neue Benchmark: Fluchtgeschwindigkeit zum Notausgang.

ThirdWorldAmerican profil fotoğrafı
ThirdWorldAmerican1 ay önce

Cool

Maziyar PANAHI profil fotoğrafı
Maziyar PANAHI1 ay önce

very interesting! i have edge evals for clinical and biomedical tasks, ranging across safety/classification/longevity/medical reasoning. i have kept it private because i don't want this to be contaminated. how private can the evals be on Optima?

Kevin Morales profil fotoğrafı
Kevin Morales1 ay önce

The use case I'd benchmark first is boring and unglamorous: which model reads a photographed restaurant menu in Spanish and returns clean dish/price JSON. Public leaderboards never tell me that, and it's the only number that decides my model spend.

Benzer Videolar

Today, we’re pushing a major update to Edison Analysis, our data analysis agent, which is tuned for scientific research and SOTA across data analysis benchmarks. In contrast to Kosmos, which runs for 6-12 hours and produces tens of thousands of lines of code, Edison Analysis runs for seconds to minutes and is best for specific, well-defined computational tasks. It is available both on our platform under the Analysis tab, and via API, and costs only one credit per run, so it is available to users on both free and paid tiers. Edison Analysis is a modified version of the data analysis agent Kosmos uses in its trajectories. Try it out! One of the most important improvements over our previous data analysis agents has been the addition of a specialized data retrieval tool. Edison Analysis can either use this tool to access data, or can pull data down directly via API. To evaluate this tool, we ranked the most commonly used public data repositories across recent papers from BioRxiv, and created a new benchmark that measures the ability of a language agent system to retrieve raw data from those sources. Edison Analysis gets 71% on this benchmark, and we’ll be working to increase this over time. You can read more about our benchmarks in the our blog post, link below. Some features worth highlighting: 1. Edison Analysis produces a report on the analysis it runs, along with a Jupyter notebook that you can download to reproduce the analysis yourself. Every figure it produces is linked back to the specific lines of code used to produce the figure, to make it easy to reproduce. 2. It works well with both Python and R. 3. One of the best uses for Edison Analysis is to use it to retrieve datasets that you can then analyze with Kosmos. We have a bunch of major improvements to Edison Analysis coming in the next few months that we’re excited to share. In the meantime, congratulations to the team, especially Ludovico Mitchener, Jon Laurent, Conor Igoe , Alex Andonian, and many more.

Sam Rodriques

62,017 görüntüleme • 10 ay önce

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 görüntüleme • 2 yıl önce

At DataInsta , we just rebranded InstaAgents to Daiko. In Japanese, Daiko means acting on your behalf. We chose this name because our AI builder creates autonomous agents that take ownership of complex tasks and do the heavy lifting for your business. Most AI projects fail because teams lack the tools, the talent, and a reliable way to evaluate agents. Testing is hard, especially for multiple agent setups and long workflows. We built something different. Daiko is battle tested and evaluated across many industries. With Daiko you can: ◾ Build and deploy agents that automate real workflows ◾ Describe your task and goals so our AI generator can build full multiple agent flows ◾ Choose from over 60 templates and use cases or create your own ◾ Connect your data and APIs, deploy anywhere, and set up evaluations and guardrails ◾ Host the platform on your own servers with white label options We designed this for real use cases in healthcare, deep tech, manufacturing, logistics, retail, legal, finance, and operations. We can handle everything for you. We manage the entire process from strategy to full scale deployment in just a few days. We charge nothing until you are happy with the value we create, and there are no upfront fees. Or, you can do it yourself. You get immediate access to our AI builder and the DataInsta marketplace. Just sign up, post your project, and hire from thousands of vetted AI experts by the hour or per project to lead the build on your terms. More Details :

Abu

28,610 görüntüleme • 6 ay önce