Загрузка видео...

Не удалось загрузить видео

На главную

Are robots ready to accelerate science labs? We created the WetLabs Benchmark to test 9 critical tasks, from easy to hard, and let 3 leading models take 20 attempts per task Astra leads but it comes close...

69,086 просмотров • 14 дней назад •via X (Twitter)

Комментарии: 25

Фото профиля Josh
Josh14 дней назад

Lab work demands delicate and precise movements but Fable 5.1 crushed one of the beakers despite instructions to be gentle Other failure modes showed missed grasps due to shadows, grasps from inside the beaker which lead to cross contamination, and bad planning for the tasks

Фото профиля Josh
Josh14 дней назад

Accidents like a knocked over beaker didn’t end the run Opus 5.5 stood it back up and kept going. In another run, it correctly pulled a stirring rod from a fallen rack. Neither attempt fully succeeded, but both showed signs of mistake recovery and emergent behaviors

Фото профиля Josh
Josh14 дней назад

From successes, models were able to do a mix sequence, place a test tube, and pick a top plate from a stack of 24-well plates Human evaluators recorded 110 successes, 246 partials and 184 failures across 540 attempts Astra led with 25.6% and Opus 5.5 closely followed at 23.3%

Фото профиля Josh
Josh14 дней назад

Hallucination on task completion is still a challenge Models claimed 192 attempts as complete despite human evaluators finding 79 partial completions and 10 failures Astra hallucinated the least, whereas Opus 5.5 had to be corrected by humans on over half the completions

Фото профиля Josh
Josh14 дней назад

Astra exceeds Opus 5.5 on science lab tasks, but it burned ~4x the tokens It had the highest success rate and shortest median attempt time, but the highest mean estimated API cost: $8.64 per attempt Opus 5.5 nearly matched its success rate at $2.39 per attempt

Фото профиля Josh
Josh14 дней назад

Evals are the backbone of understanding and improving model capabilities The WetLab Benchmark will be updated as capabilities progress and we unlock a scale up of scientific work in the real world Explore the deep dive and all runs here:

Фото профиля Josh
Josh14 дней назад

Great work to the team across this effort @itsghvr @PranavKuppili @camrzadki @Kaiwen66lc Zac Gong @jasontheutopian

Фото профиля Josh
Josh14 дней назад

@PranavKuppili @camrzadki @Kaiwen66lc @jasontheutopian + shoutout @ryan_punamiya and friends for the reviews

Фото профиля TK Kong
TK Kong14 дней назад

Incredible work! Congrats Josh, @jasontheutopian, and team

Фото профиля Archie Sengupta
Archie Sengupta14 дней назад

@sohamsankaran

Фото профиля Sanskar Pandey
Sanskar Pandey14 дней назад

@TyneeWorld Great!

Фото профиля Don
Don14 дней назад

so sick!

Фото профиля Peter Haymond
Peter Haymond14 дней назад

@MeckaAI This is why I’m curious what SuperRadiant is doing

Фото профиля Yuri Namikawa
Yuri Namikawa14 дней назад

very needed for progress in robotics x labs. cool stuff!

Фото профиля vinny
vinny14 дней назад

So cool!

Фото профиля Tanvi
Tanvi14 дней назад

really cool stuff!

Фото профиля Rakesh Sahni
Rakesh Sahni14 дней назад

Not yet On the hard lab tasks, Opus 5.5 finished 2 of 60 tries and Astra and Fable 5.1 finished none When Opus 5.5 said a try was finished a person counted that try as finished 38 of 79 times Limit: Graders could see which model ran

Фото профиля steve jang
steve jang14 дней назад

🔥🔥🔥

Фото профиля Vai Viswanathan
Vai Viswanathan14 дней назад

Sick!

Фото профиля Jenny do Forno
Jenny do Forno14 дней назад

🤖 🔥 🚀

Фото профиля Healy Jones
Healy Jones14 дней назад

Very cool! 🤖🤖

Фото профиля Kane
Kane14 дней назад

wow that's dope fen kk

Фото профиля Andrea Wang
Andrea Wang14 дней назад

🚀🚀

Фото профиля Shryas Bhurat
Shryas Bhurat14 дней назад

🚀

Фото профиля Manoela Schmitz
Manoela Schmitz14 дней назад

Which specific tasks among the nine tested proved to be the most challenging for the models? Reviewing the detailed benchmark results shows how they performed across the different difficulty levels.

Похожие видео

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,508 просмотров • 1 месяц назад