
Artificial Analysis
@ArtificialAnlys • 142,554 subscribers
Independent analysis of AI
Shorts
Videos

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at
Artificial Analysis132,623 次观看 • 1 个月前

Last week we launched Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency Watch our new video walkthrough to see how to build and run your own benchmark with Optima, and find the best model for your use case Start building your benchmark today at
Artificial Analysis19,837 次观看 • 24 天前

Example generations from HappyHorse-1.0 compared to Dreamina Seedance 2.0, Kling 3.0 Pro, grok-video-imagine and PixVerse V6 (Text to Video with Audio): Prompt [1/4]: A hula hoop spinning on a kid's waist, gradually climbing to their chest, then dropping to knees, then clattering to the floor. They pick it up to try again.
Artificial Analysis36,791 次观看 • 5 个月前

Gemini 3.5 Flash is a step forward for Google on speed and agentic capabilities but comes at a trade-off of being higher cost than prior models We have measured up to ~280 output tokens/sec, placing it on the speed/intelligence Pareto frontier and well ahead of Gemini 3 Flash. It also shows a major uplift on agentic tasks, reaching ~1650 ELO on GDPVal-AA. The trade-off: cost is up ~5x versus Gemini 3 Flash, driven by higher token prices (3x higher than Gemini 3 Flash) and higher token usage. In this video, Declan Jackson, Member of Technical Staff at Artificial Analysis, breaks it down.
Artificial Analysis16,055 次观看 • 3 个月前

Following up on our Intelligence Index v4.1 release yesterday, in the video below, Daniel from our team shares a short overview of what's changed: 1. Three upgraded evaluations: Terminal-Bench 2.1, τ³-Bench Banking and GDPval-AA v2 2. Cost, time, and tokens per task: Understand the cost, time, and tokens of tasks across our Index and for individual evals, and how these trade off against Intelligence 3. Cached input token reporting: We now report the amount of cached tokens a particular model uses and how this influences cost
Artificial Analysis13,214 次观看 • 2 个月前

We’re unveiling a new look for Artificial Analysis! We’ve come a long way since launching Artificial Analysis over 2 years ago. Today, we benchmark 400+ models, 50+ inference providers, and benchmark not only language models but also image, video, speech, music, hardware, and agents. Our mission to support the AI ecosystem with independent benchmarking remains the same, but our brand and website refresh is designed to better reflect how much we’ve grown and how much further we plan to go. A huge thank you to everyone who has been part of the Artificial Analysis community along the way: from developers choosing models and building agents, to labs, inference and hardware providers, and fellow independent researchers.
Artificial Analysis18,208 次观看 • 5 个月前

Overview of our recent launch of Coding Agent benchmarks on Artificial Analysis and our first Youtube Video! We walk through the performance, cost, token usage and speed differences across different coding agents. This includes looking at Opus 4.7 in Claude Code's leading performance and Composer 2.5's strong positioning on the Coding Agent Index / Cost Pareto frontier. We have also launched our YouTube channel! Come say hi and subscribe:
Artificial Analysis10,702 次观看 • 3 个月前

Announcing Artificial Analysis Video Arena - the first crowdsourced comparison for Text to Video models Text to Video models are accelerating rapidly and crossing quality thresholds every month. We created Video Arena to compare them using the only source of truth for visual media - human preference! Video Arena includes hundreds of videos from the leading video models, including: - Runway's Runway Gen 3 Alpha - Pika's Pika 1.5 - Luma's Dream Machine - MiniMax / Hailuo AI (MiniMax) - Kling AI's Kling 1.0 - Zhipu AI's CogVideoX-5B Voting is open now and we’ll be announcing the first leaderboard results within 24 hours. Any predictions? In the meantime, you can see your own ‘personal leaderboard’ of how you’ve ranked the video models after 30 votes. Link to the Artificial Analysis Video Arena in the below tweet! 👇
Artificial Analysis27,108 次观看 • 1 年前
没有更多内容可加载
