Загрузка видео...
Не удалось загрузить видео
Are robots ready to accelerate science labs? We created the WetLabs Benchmark to test 9 critical tasks, from easy to hard, and let 3 leading models take 20 attempts per task Astra leads but it comes close...
69,086 просмотров • 14 дней назад •via X (Twitter)
Комментарии: 25

Lab work demands delicate and precise movements but Fable 5.1 crushed one of the beakers despite instructions to be gentle Other failure modes showed missed grasps due to shadows, grasps from inside the beaker which lead to cross contamination, and bad planning for the tasks

Accidents like a knocked over beaker didn’t end the run Opus 5.5 stood it back up and kept going. In another run, it correctly pulled a stirring rod from a fallen rack. Neither attempt fully succeeded, but both showed signs of mistake recovery and emergent behaviors

From successes, models were able to do a mix sequence, place a test tube, and pick a top plate from a stack of 24-well plates Human evaluators recorded 110 successes, 246 partials and 184 failures across 540 attempts Astra led with 25.6% and Opus 5.5 closely followed at 23.3%

Hallucination on task completion is still a challenge Models claimed 192 attempts as complete despite human evaluators finding 79 partial completions and 10 failures Astra hallucinated the least, whereas Opus 5.5 had to be corrected by humans on over half the completions

Astra exceeds Opus 5.5 on science lab tasks, but it burned ~4x the tokens It had the highest success rate and shortest median attempt time, but the highest mean estimated API cost: $8.64 per attempt Opus 5.5 nearly matched its success rate at $2.39 per attempt

Evals are the backbone of understanding and improving model capabilities The WetLab Benchmark will be updated as capabilities progress and we unlock a scale up of scientific work in the real world Explore the deep dive and all runs here:

Great work to the team across this effort @itsghvr @PranavKuppili @camrzadki @Kaiwen66lc Zac Gong @jasontheutopian

@PranavKuppili @camrzadki @Kaiwen66lc @jasontheutopian + shoutout @ryan_punamiya and friends for the reviews

Incredible work! Congrats Josh, @jasontheutopian, and team

@sohamsankaran

@TyneeWorld Great!

so sick!

@MeckaAI This is why I’m curious what SuperRadiant is doing

very needed for progress in robotics x labs. cool stuff!

So cool!

really cool stuff!

Not yet On the hard lab tasks, Opus 5.5 finished 2 of 60 tries and Astra and Fable 5.1 finished none When Opus 5.5 said a try was finished a person counted that try as finished 38 of 79 times Limit: Graders could see which model ran

🔥🔥🔥

Sick!

🤖 🔥 🚀

Very cool! 🤖🤖

wow that's dope fen kk

🚀🚀

🚀

Which specific tasks among the nine tested proved to be the most challenging for the models? Reviewing the detailed benchmark results shows how they performed across the different difficulty levels.


