Загрузка видео...
Не удалось загрузить видео
🎮 Computer Use Agent Arena is LIVE! 🚀 🔥 Easiest way to test computer-use agents in the wild without any setup 🌟 Compare top VLMs: OpenAI Operator, Claude 3.7, Gemini 2.5 Pro, Qwen 2.5 vl and more 🕹️ Test agents on 100+ real apps & webs with one-click config... show more
93,273 просмотров • 1 год назад •via X (Twitter)
Комментарии: 28

🖥️ How does it work? 1️⃣ Choose your OS (currently Windows, Ubuntu supported, MacOS coming soon) 2️⃣ Setup the initial desktop environment with one-click configuration 3️⃣ Write your task (e.g. "Upload a CV to Slack") 4️⃣ Observe agents execute step-by-step 5️⃣ Evaluate: Which agent did better? 6️⃣ See revealed agent identities after evaluation Test, compare, and give feedback—all anonymously! [2/🧵]

🛠️ Set up your environment with ease! 🔐 Free & safe access to agents on cloud-hosted machines, fully isolated 🌟 Pre-installed apps & software: LibreOffice, Slack, VSCode, GIMP, PDF editors… 🌐 Web apps/domains: YouTube, Reddit… 🔧 Customize further: Upload files, Open URLs, clone GitHub repos, or upload files Fully operable for any task! [3/🧵]

🏆 Leaderboard Highlights (tentative) 🥇 OpenAI Operator 🥈 Claude 3.7 Sonnet 🥉 Claude 3.5 Sonnet Stay tuned for the official leaderboard in the following weeks. 📊 Check the latest, real-time leaderboard (tentative): [4/🧵]

Use Case #1 Web browser task: please help me find the cheapest man's long-sleeve t-shirt on Amazon, I need a new fit for summer Battle: Gemini 2.5 Pro (Experimental) vs. OpenAI Computer-Use Preview Details at: [5/🧵]

Use Case #2 Personal Use Task: Can you help me export my homepage in Notion to a html file onto my desktop and open it in the browser to preview it? Battle: Gemini 2.0 Flash vs. Claude 3.5 Sonnet (New) - Computer-Use Details at: [6/🧵]

👋Acknowledgement Thanks to the Computer Agent Arena team: @xywang626, @jiaqideng07, @TianbaoX, @RyanLi0802, Gavin Li, @StevenyzZhang, @nikushii_, @istoica05, @infwinston, @Diyi_Yang, @ysu_nlp, Yi Zhang, Zhiguo Wang, @hllo_wrld, @taoyds Also thanks to @gneubig, @dan_fried, @shuyanzhxyc, @pengchengyin, @haozhangml for their helpful discussions. Greatest thanks to @awscloud and @lmarena_ai for their kind support. [7/🧵]

📊 Curious how your favorite computer use agent stacks up? Dive into the leaderboard, explore model performance, and share your feedback to help shape the future of computer-use agents! Data & Code would be open-sourced in a few weeks! 👉 Platform: 🏆 Leaderboard (tentative): 📖 Learn more: 🧑💻Data & Code (coming soon): [8/8]

It's very fun to play with!

Thank you Wenhu, we've tried our best to improve the user experiences when interacting with CUAs

It sometimes gets stuck in infinite loops. Is that expected? Also, what's the user interface on the top? Is that selecting the entry point?

@BowenWangNLP you can check the detailed thoughts from the agents. my experience is infinite loops are usually caused by the agents themselves trying to do a necessary step but failing

Agreed, sometimes models would stuck into loops due to its capabilities, you could click on "See Details" to take a look at its inner thoughts. By "user interface" you may refer to the two windows on top of the page, which is actually a VNC connection for users to interact with the computers.

Congrats @BowenWangNLP this is super neat

Thank you Marco, we're trying to make it better, for more transparent & trustful computer-use agents leaderboard.

very impressive! Are you streaming live environment to the browser or is it several screenshots?

Thank you Siva, the two windows on top are the streaming live of the computer's which users can operate on. the trajectories below are screenshots of each steps,

I realized it later. Curious on how you enabled live interaction through browser. What libraries/software allows you to do that?

We used NoVNC protocol as the library.📚

Super cool idea. Nice work!

Thank you Chris, looking forward to more agents on board.

really cool work we really needed something like this for computer use agents to actually evaluate them in real world scenarios this should help a lot with progress also really fun to play with and compare all the different models

Thank you for your kind comment, and that's the main reason why we want this arena to push the computer-use agents benchmarking forward.

This is exactly what we need - a standardized arena to compare agent capabilities in the wild. Curious how the results might influence future VLM development paths.

Nice work! We need benchmarks like this. @kimmonismus

One prompt, one post, one paycheck. Promptchan x Fanvue.

this is awesome. what's the ETA on the code release for testing/tweaking locally?

Great work. Thanks

Where is Browser Use? ( @gregpr07 @mamagnus00
