Artificial Analysis's banner
Artificial Analysis's profile picture

Artificial Analysis

@ArtificialAnlys142,554 subscribers

Independent analysis of AI

Shorts

You can test GPT-6 Astra on your own custom benchmarks in Optima Head to to start creating and running your own custom benchmarks today

You can test GPT-6 Astra on your own custom benchmarks in Optima Head to to start creating and running your own custom benchmarks today

17,053 просмотров

You can now benchmark Qwen3.8 27B with Optima. See how a model that can run on your laptop compares on performance, cost and speed for your custom use cases We're also offering credits and free advisory from an Artificial Analysis team member to enterprises looking to build benchmarks for their use cases. Let us know if you’re interested: Build and run your custom benchmarks today at

You can now benchmark Qwen3.8 27B with Optima. See how a model that can run on your laptop compares on performance, cost and speed for your custom use cases We're also offering credits and free advisory from an Artificial Analysis team member to enterprises looking to build benchmarks for their use cases. Let us know if you’re interested: Build and run your custom benchmarks today at

15,764 просмотров

Want to generate images from the world's best models like Nano Banana or GPT Image side by side? Now you can with Image Lab 🚀 Run a single prompt across up to 25 models, with up to 20 images from each, and see results in seconds. You've seen our leaderboards. Now generate and evaluate the models yourself. Link in thread 🧵👇

Want to generate images from the world's best models like Nano Banana or GPT Image side by side? Now you can with Image Lab 🚀 Run a single prompt across up to 25 models, with up to 20 images from each, and see results in seconds. You've seen our leaderboards. Now generate and evaluate the models yourself. Link in thread 🧵👇

12,250 просмотров

Videos

ArtificialAnlys's profile picture

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

132,623 просмотров • 29 дней назад

Больше нет контента для загрузки