Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Just found an open-source tool that lets you see how well an AI can actually use a computer. it’s called OSWorld 2.0. you give the AI a real task on a computer, then watch to see if it can figure everything out and finish it. it gives you: →...

13,429 görüntüleme • 1 gün önce •via X (Twitter)

17 Yorum

MIKE profil fotoğrafı
MIKE1 gün önce

repository link:

Camto profil fotoğrafı
Camto1 gün önce

This looks great fr 👌

mike profil fotoğrafı
mike1 gün önce

thanks for sharing fam

Mira Takes profil fotoğrafı
Mira Takes1 gün önce

OSWorld-style benchmarks are a useful reality check because they expose the gap between a fluent demo and reliable execution. The next metric I’d watch is recovery: how often does the agent detect a bad action, undo it, and finish without a human reset?

Kuzka_aaa profil fotoğrafı
Kuzka_aaa1 gün önce

How realistic are the tasks compared to actual daily computer use — is it things like 'book a flight' and 'edit a spreadsheet', or more niche/synthetic benchmarks designed just to be hard?

man mohan goel profil fotoğrafı
man mohan goel1 gün önce

This open-source repository provides multi-step computer tasks to test AI agents. Would you consider running these tests yourself instead of relying only on published benchmark scores?

DexorynLabs profil fotoğrafı
DexorynLabs1 gün önce

Red days test strong hands.

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 gün önce

Real world tasks make it way harder

RIDER SKETCH profil fotoğrafı
RIDER SKETCH1 gün önce

"run the same tests yourself instead of relying on benchmark scores" is good advice, but the current numbers are the more interesting story, a 20.6% completion rate on the current best model means computer-use agents are nowhere near the "figure everything out and finish it" framing implies, this benchmark exists specifically because agents still fail constantly on long real tasks

Hurricane profil fotoğrafı
Hurricane1 gün önce

Will check this out

Aleksandar Janca profil fotoğrafı
Aleksandar Janca1 gün önce

running the same tests they use is the only way to know if the benchmark holds for your task

Bambi profil fotoğrafı
Bambi1 gün önce

Trying this

Haleemah profil fotoğrafı
Haleemah1 gün önce

Another one saved to try 🥲

Yohaku profil fotoğrafı
Yohaku1 gün önce

should be interesting tool for me thanks Mike

zetabyte🌴 profil fotoğrafı
zetabyte🌴1 gün önce

just check it out

xabz profil fotoğrafı
xabz1 gün önce

osworld was already the best benchmark for this, curious what 2.0 adds

Nahid profil fotoğrafı
Nahid1 gün önce

wanna see one fail video not just repo

Benzer Videolar

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,541 görüntüleme • 1 yıl önce