Video yükleniyor...
Video Yüklenemedi
Just found an open-source tool that lets you see how well an AI can actually use a computer. it’s called OSWorld 2.0. you give the AI a real task on a computer, then watch to see if it can figure everything out and finish it. it gives you: →... show more
13,429 görüntüleme • 1 gün önce •via X (Twitter)
17 Yorum

repository link:

This looks great fr 👌

thanks for sharing fam

OSWorld-style benchmarks are a useful reality check because they expose the gap between a fluent demo and reliable execution. The next metric I’d watch is recovery: how often does the agent detect a bad action, undo it, and finish without a human reset?

How realistic are the tasks compared to actual daily computer use — is it things like 'book a flight' and 'edit a spreadsheet', or more niche/synthetic benchmarks designed just to be hard?

This open-source repository provides multi-step computer tasks to test AI agents. Would you consider running these tests yourself instead of relying only on published benchmark scores?

Red days test strong hands.

Real world tasks make it way harder

"run the same tests yourself instead of relying on benchmark scores" is good advice, but the current numbers are the more interesting story, a 20.6% completion rate on the current best model means computer-use agents are nowhere near the "figure everything out and finish it" framing implies, this benchmark exists specifically because agents still fail constantly on long real tasks

Will check this out

running the same tests they use is the only way to know if the benchmark holds for your task

Trying this

Another one saved to try 🥲

should be interesting tool for me thanks Mike

just check it out

osworld was already the best benchmark for this, curious what 2.0 adds

wanna see one fail video not just repo

