ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

You have a new Dev Playground for testing existing and upcoming FLUX models: ๐—–๐—ผ๐—บ๐—ฝ๐—ฎ๐—ฟ๐—ฒ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€ ๐˜€๐—ถ๐—ฑ๐—ฒ-๐—ฏ๐˜†-๐˜€๐—ถ๐—ฑ๐—ฒ. Run the same prompt across multiple FLUX models at once and see the differences. ๐—–๐—ผ๐—ป๐˜๐—ฟ๐—ผ๐—น ๐—ฝ๐—ฎ๐—ฟ๐—ฎ๐—บ๐—ฒ๐˜๐—ฒ๐—ฟ๐˜€. Adjust API parameters from the UI, so what you test is exactly what you build. ๐—˜๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฎ๐˜๐—ฒ. Test...

15,962 ๆฌก่ง‚็œ‹ โ€ข 3 ไธชๆœˆๅ‰ โ€ขvia X (Twitter)

14 ๆก่ฏ„่ฎบ

Black Forest Labs ็š„ๅคดๅƒ
Black Forest Labs3 ไธชๆœˆๅ‰

Use it at โ†’

Black Forest Labs ็š„ๅคดๅƒ
Black Forest Labs3 ไธชๆœˆๅ‰

Compare multiple results side by side.

Emily ็š„ๅคดๅƒ
Emily3 ไธชๆœˆๅ‰

FLUX 3 when? I thought you guys going to be faster this year ๐Ÿฅฒ๐Ÿฅฒ๐Ÿฅฒ

X Girls ็š„ๅคดๅƒ
X Girls3 ไธชๆœˆๅ‰

dang you have a lot of credits ๐Ÿ˜‚

Vision33X โ™˜ ็š„ๅคดๅƒ
Vision33X โ™˜3 ไธชๆœˆๅ‰

running the same prompt side by side is gonna end so many discord arguments fr

Kol Tregaskes ็š„ๅคดๅƒ
Kol Tregaskes3 ไธชๆœˆๅ‰

This is amazing, thank you.

AI Mastery Guide ็š„ๅคดๅƒ
AI Mastery Guide3 ไธชๆœˆๅ‰

Being able to compare quality, latency, and cost side by side is genuinely useful for picking the right model.

IronRed | SandHive ็š„ๅคดๅƒ
IronRed | SandHive3 ไธชๆœˆๅ‰

GLHF and may the best model win. whatever that means now!

Hamid AI ็š„ๅคดๅƒ
Hamid AI3 ไธชๆœˆๅ‰

Excited to test the new Playgroundโ€”side-by-side model comparison is a great addition. ๐Ÿš€

blankbrain ็š„ๅคดๅƒ
blankbrain3 ไธชๆœˆๅ‰

nothing is ever upcoming with u guys

X-Dimension | GANTZ AI Art ็š„ๅคดๅƒ
X-Dimension | GANTZ AI Art1 ไธชๆœˆๅ‰

when another free week? ๐Ÿ™‚

Puzzle Paws ็š„ๅคดๅƒ
Puzzle Paws3 ไธชๆœˆๅ‰

hear me out: add an open weights toggle so i can benchmark what i actually control. otherwise i'm just optimizing your cloud bill.

Marco ็š„ๅคดๅƒ
Marco3 ไธชๆœˆๅ‰

this is epic, guys! Thanks - will defo test this out

Gulzar Junaid ็š„ๅคดๅƒ
Gulzar Junaid2 ไธชๆœˆๅ‰

Was ist das bitte fรผr eine Magie? ๐Ÿช„โœจ

็›ธๅ…ณ่ง†้ข‘

Weโ€™re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysisโ€™ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysisโ€™ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. Weโ€™ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: โžค Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you โžค Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released โžค Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set โžค Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: โžค Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? โžค Which model best matches the writing style of lawyers for my legal agent? โžค Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,508 ๆฌก่ง‚็œ‹ โ€ข 1 ไธชๆœˆๅ‰