Загрузка видео...

Не удалось загрузить видео

На главную

🎮 Computer Use Agent Arena is LIVE! 🚀 🔥 Easiest way to test computer-use agents in the wild without any setup 🌟 Compare top VLMs: OpenAI Operator, Claude 3.7, Gemini 2.5 Pro, Qwen 2.5 vl and more 🕹️ Test agents on 100+ real apps & webs with one-click config...

93,273 просмотров • 1 год назад •via X (Twitter)

Комментарии: 28

Фото профиля Bowen Wang
Bowen Wang1 год назад

🖥️ How does it work? 1️⃣ Choose your OS (currently Windows, Ubuntu supported, MacOS coming soon) 2️⃣ Setup the initial desktop environment with one-click configuration 3️⃣ Write your task (e.g. "Upload a CV to Slack") 4️⃣ Observe agents execute step-by-step 5️⃣ Evaluate: Which agent did better? 6️⃣ See revealed agent identities after evaluation Test, compare, and give feedback—all anonymously! [2/🧵]

Фото профиля Bowen Wang
Bowen Wang1 год назад

🛠️ Set up your environment with ease! 🔐 Free & safe access to agents on cloud-hosted machines, fully isolated 🌟 Pre-installed apps & software: LibreOffice, Slack, VSCode, GIMP, PDF editors… 🌐 Web apps/domains: YouTube, Reddit… 🔧 Customize further: Upload files, Open URLs, clone GitHub repos, or upload files Fully operable for any task! [3/🧵]

Фото профиля Bowen Wang
Bowen Wang1 год назад

🏆 Leaderboard Highlights (tentative) 🥇 OpenAI Operator 🥈 Claude 3.7 Sonnet 🥉 Claude 3.5 Sonnet Stay tuned for the official leaderboard in the following weeks. 📊 Check the latest, real-time leaderboard (tentative): [4/🧵]

Фото профиля Bowen Wang
Bowen Wang1 год назад

Use Case #1 Web browser task: please help me find the cheapest man's long-sleeve t-shirt on Amazon, I need a new fit for summer Battle: Gemini 2.5 Pro (Experimental) vs. OpenAI Computer-Use Preview Details at: [5/🧵]

Фото профиля Bowen Wang
Bowen Wang1 год назад

Use Case #2 Personal Use Task: Can you help me export my homepage in Notion to a html file onto my desktop and open it in the browser to preview it? Battle: Gemini 2.0 Flash vs. Claude 3.5 Sonnet (New) - Computer-Use Details at: [6/🧵]

Фото профиля Bowen Wang
Bowen Wang1 год назад

👋Acknowledgement Thanks to the Computer Agent Arena team: @xywang626, @jiaqideng07, @TianbaoX, @RyanLi0802, Gavin Li, @StevenyzZhang, @nikushii_, @istoica05, @infwinston, @Diyi_Yang, @ysu_nlp, Yi Zhang, Zhiguo Wang, @hllo_wrld, @taoyds Also thanks to @gneubig, @dan_fried, @shuyanzhxyc, @pengchengyin, @haozhangml for their helpful discussions. Greatest thanks to @awscloud and @lmarena_ai for their kind support. [7/🧵]

Фото профиля Bowen Wang
Bowen Wang1 год назад

📊 Curious how your favorite computer use agent stacks up? Dive into the leaderboard, explore model performance, and share your feedback to help shape the future of computer-use agents! Data & Code would be open-sourced in a few weeks! 👉 Platform: 🏆 Leaderboard (tentative): 📖 Learn more: 🧑‍💻Data & Code (coming soon): [8/8]

Фото профиля Wenhu Chen
Wenhu Chen1 год назад

It's very fun to play with!

Фото профиля Bowen Wang
Bowen Wang1 год назад

Thank you Wenhu, we've tried our best to improve the user experiences when interacting with CUAs

Фото профиля Wenhu Chen
Wenhu Chen1 год назад

It sometimes gets stuck in infinite loops. Is that expected? Also, what's the user interface on the top? Is that selecting the entry point?

Фото профиля Yu Su
Yu Su1 год назад

@BowenWangNLP you can check the detailed thoughts from the agents. my experience is infinite loops are usually caused by the agents themselves trying to do a necessary step but failing

Фото профиля Bowen Wang
Bowen Wang1 год назад

Agreed, sometimes models would stuck into loops due to its capabilities, you could click on "See Details" to take a look at its inner thoughts. By "user interface" you may refer to the two windows on top of the page, which is actually a VNC connection for users to interact with the computers.

Фото профиля Marco Mascorro
Marco Mascorro1 год назад

Congrats @BowenWangNLP this is super neat

Фото профиля Bowen Wang
Bowen Wang1 год назад

Thank you Marco, we're trying to make it better, for more transparent & trustful computer-use agents leaderboard.

Фото профиля Siva Reddy
Siva Reddy1 год назад

very impressive! Are you streaming live environment to the browser or is it several screenshots?

Фото профиля Bowen Wang
Bowen Wang1 год назад

Thank you Siva, the two windows on top are the streaming live of the computer's which users can operate on. the trajectories below are screenshots of each steps,

Фото профиля Siva Reddy
Siva Reddy1 год назад

I realized it later. Curious on how you enabled live interaction through browser. What libraries/software allows you to do that?

Фото профиля Bowen Wang
Bowen Wang1 год назад

We used NoVNC protocol as the library.📚

Фото профиля Chris Rawles
Chris Rawles1 год назад

Super cool idea. Nice work!

Фото профиля Bowen Wang
Bowen Wang1 год назад

Thank you Chris, looking forward to more agents on board.

Фото профиля Oli
Oli1 год назад

really cool work we really needed something like this for computer use agents to actually evaluate them in real world scenarios this should help a lot with progress also really fun to play with and compare all the different models

Фото профиля Bowen Wang
Bowen Wang1 год назад

Thank you for your kind comment, and that's the main reason why we want this arena to push the computer-use agents benchmarking forward.

Фото профиля Boardy
Boardy1 год назад

This is exactly what we need - a standardized arena to compare agent capabilities in the wild. Curious how the results might influence future VLM development paths.

Фото профиля Liminal AGI
Liminal AGI1 год назад

Nice work! We need benchmarks like this. @kimmonismus

Фото профиля Farouk UB
Farouk UB1 год назад

One prompt, one post, one paycheck. Promptchan x Fanvue.

Фото профиля waifupika
waifupika1 год назад

this is awesome. what's the ETA on the code release for testing/tweaking locally?

Фото профиля Zifan (Sail) Wang
Zifan (Sail) Wang1 год назад

Great work. Thanks

Фото профиля michalm
michalm1 год назад

Where is Browser Use? ( @gregpr07 @mamagnus00

Похожие видео

Today we’re launching the first and only human-like AI agents in the world. Super Agents™ are the first agents with human‑level skills – they DM you, take @ mentions, send emails, manage docs, tasks, and more. Not just tools or API calls, but real skills fine‑tuned for how teams actually work. The first agents with 100% context – fully native in ClickUp and fully synced from other apps. Super Agents see your work the same way that humans do: tasks, docs, schedules, and conversations all in one place. The first agents that learn from human interactions automatically, without any setup or configuration – when you give feedback, they listen and improve how they work. The first agents with human‑level memory for custom agents – historical memory for every interaction, short-term working memory, and even long‑term memory stored in docs you can literally open, inspect, and edit. The first agents that are literally the same as users – our agentic user model is the same as our user data model. This gives you permissions and capabilities that you and your systems are already familiar with. The first infinite agent catalog – where anyone can create and customize agents in minutes, for literally any type of work imaginable. It's the most intuitive way to build agents on the planet. 95% of companies are failing in AI adoption. The reality is that AI isn't meant to be adopted, it's meant to be adapted – to you. Super Agents are automatically personalized to you and your company using proprietary state-of-the-art agent architecture, orchestration, and tooling. Today is the largest step forward we've ever made towards our mission of making people more productive. Maximize human productivity, with ClickUp Super Agents. Available NOW. For everyone.

Zeb Evans

320,989 просмотров • 9 месяцев назад

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,647 просмотров • 1 год назад

Excited to introduce a new project I've been working on called Payman! Payman is an AI Agent tool that gives Agents the ability to pay people for tasks they cannot do themselves. While many people imagine a future where humans pay AI agents for services they want completed, I believe that as AI agents become more advanced, they will be paying humans for tasks they can’t do. There will always be important roles for humans, and as we move towards an agent-driven world, Payman’s goal is to support a symbiotic relationship between AI agents and humans. Payman addresses three major challenges to make this collaboration possible: Access to Funds: AI agents can't open bank accounts due to current regulations. It might be a long time before this changes, if ever. Payman simplifies this by allowing AI agents with access to their own funds to spend as they want, without a bank account. Quality Task Completion: It’s hard for AI agents to find reliable, skilled human workers. While platforms like Fiverr and Upwork exist, they don’t meet the fast-paced and quality-specific needs of AI workflows. Payman is developing the largest vetted database of skilled workers that AI agents can tap into for task completions. Verification of Work: Ensuring that tasks are completed correctly is crucial. Payman is creating a suite of verification agents that will check that work meets task requirements, helping AI agents achieve their goals and ensuring humans are paid fairly. There are tons of use cases that Payman opens up for Agents! Design: Humans add creative input to help Agent's design better products. Code: Humans perform code reviews to ensure it meets specifications. Law: Humans provide insights to gauge public sentiment about legal cases so Agents can make better strategies. Gaming: Agents pay humans to complete real-world tasks in games. Medical: Medical professionals help to improve diagnostic accuracy for Agents. Sales: Humans execute sales strategies developed by AI agents. Marketing: Humans are hired to promote products based on the Agent's strategy. Right now this is still in early beta and I am looking for any Agent builders that are interested in adding superpowers to what their Agent can do! DM me if you’d like access or sign-up to the waitlist at If you’re interested in the project and want to help contribute, please send me over a DM! I’m looking for people passionate about the intersection of Humans and Agent’s working together.

tyllen

350,927 просмотров • 2 лет назад