Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

1/ Today we're introducing Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. Text and multimodal adapters are available under Apache-2.0:

160,244 Aufrufe • vor 10 Tagen •via X (Twitter)

39 Kommentare

Profilbild von Cua
Cuavor 10 Tagen

2/ At each step, the CUA-S1 model receives the screen state, the task goal, and a fixed set of candidate actions. It returns one action. The environment changes, and the next step starts from the new state.

Profilbild von Cua
Cuavor 10 Tagen

3/ The training recipe has two stages. Supervised training teaches the decision format. Agentic RL then runs the model in live cua-bench-basic environments, where only a completed task earns the environment's reward.

Profilbild von Cua
Cuavor 10 Tagen

4/ On the same held-out tasks, Cua-S1-4B-0.2 completes 17/18 text episodes and 13/18 multimodal. Zero-shot djev completes 16/18 and 12/18.

Profilbild von Cua
Cuavor 10 Tagen

5/ On a frozen GUI-360 split of 168 multimodal tasks, Cua-S1-4B-0.2 reaches 92.9% versus 60.1% for untrained djev. Both use the same inputs and scoring, with no accessibility tree.

Profilbild von Cua
Cuavor 10 Tagen

6/ The training code and benchmark results are now merged into Cua. Explore the code: Get the Apache-2.0 text and multimodal adapters:

Profilbild von Robert Scoble
Robert Scoblevor 10 Tagen

What you all are doing is way over my pay grade. So I asked Grok why this is important. Awesome. Its answer: +++++ It matters because it is a concrete step toward computer-use agents that can be trained, inspected, and run outside the largest closed labs. Most computer-use systems still treat the desktop as a chat problem: a large model looks at a screenshot, writes a plan, and hopes the next click works. Cua-S1-4B-0.2 is built as a specialist instead. At each step it gets the screen, the goal, and a fixed list of candidate actions, then returns one action. The environment changes, and the next step starts from the new state. That is closer to how real UI work happens: many small, checkable decisions rather than one long generated essay. The training claim is the other reason people noticed it. Cua says this is the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. In other words, the model is not only imitating labeled clicks. After a supervised stage that teaches the decision format, it is rolled out in live cua-bench-basic environments and only scored when the task actually finishes. RLOO is a relatively simple online RL method (REINFORCE with a leave-one-out baseline), so the recipe is easier to reproduce than full PPO-style agent training. The numbers are narrow but not empty. On a frozen 168-task GUI-360 multimodal split, with no accessibility tree, Cua reports 92.9% versus 60.1% for the untrained djev baseline. On held-out live episodes it reports 17/18 text completions and 13/18 multimodal. Those are small, vendor-run evaluations, and Cua itself marks the checkpoint as early research. Pagination is still weak, and at least one live multimodal task is unsolved. So this is not evidence that a 4B adapter replaces Claude or OpenAI computer-use agents. It is evidence that a small, open decision layer can be trained against real GUI outcomes and beat its own untrained baseline by a large margin. That combination is useful in practice. The adapters sit on a frozen Qwen3.5-4B base, ship under Apache-2.0, and are small enough to run locally. Cua’s larger project already supplies the missing pieces around the model: a driver for real desktops, sandboxes, and benchmarks that can generate training trajectories. If the field is going to move from “frontier model remote-controls your screen” to “cheap, specialized policies that operate software,” this is one of the first public artifacts that looks like that stack. The honest limit is transfer. Strong results on Cua’s held-out splits do not automatically mean reliable performance on arbitrary apps, operating systems, languages, or layouts. The importance is the recipe, not a claim that desktop agents are solved: open weights, live-task rewards, one decision per step, and a training loop other people can inspect.

Profilbild von Gognumb
Gognumbvor 10 Tagen

Love how fast cua team innovate computer use !

Profilbild von Tony Simons
Tony Simonsvor 10 Tagen

Hell yeah!! Congrats on this one Team Cua! 🍾 Fantastically done as always!

Profilbild von Pablo Magana Gabaude 🇫🇷🇨🇱🇸🇪
Pablo Magana Gabaude 🇫🇷🇨🇱🇸🇪vor 10 Tagen

Will this be a part of Hermes Agent?

Profilbild von Anderson
Andersonvor 10 Tagen

your mouse just became legacy hardware

Profilbild von Florian S
Florian Svor 10 Tagen

Amazing work, looking forward to tet your model in (upcoming) Image Jev Bench. Maybe I should start a Computer Use Jev Bench as well

Profilbild von Ziwen
Ziwenvor 10 Tagen

ain't gonna lie we really do need a computer use model

Profilbild von RabbitHoleExplorer
RabbitHoleExplorervor 10 Tagen

re: Cua There is one particularly interesting implication for FirstMate / computer-use agents: the really powerful architecture may not be Jev VS CUA-S1. It could be Jev ABOVE something like CUA-S1. Sol / Qwen / DeepSeek novel planning ↓ Jev task/policy decision ↓ CUA-S1 concrete UI decision ↓ Cua Driver deterministic execution ↓ verification That is starting to look less like “an LLM controlling a computer” and more like an actual inference hierarchy, where the expensive model is invoked only when the lower layers genuinely don't know what to do. @kunchenguid @CompleteSkeptic @typesafeai

Profilbild von JC
JCvor 9 Tagen

can't wait to give it a run

Profilbild von Panther
Panthervor 10 Tagen

a 4b model doing computer use means my laptop fan filed for hazard pay

Profilbild von Oliviero Pinotti
Oliviero Pinottivor 10 Tagen

🤯 this is crazy powerful, love cua!

Profilbild von Mukhsin Mukhtorov
Mukhsin Mukhtorovvor 10 Tagen

training on task completion is the part i care about. in my own computer-use loop the model rarely clicked the wrong thing. it said done before the screen had actually changed. curious if completion rewards cut down on that early done

Profilbild von Stephen Brouhard
Stephen Brouhardvor 10 Tagen

This is a great direction. Not every action needs a giant model thinking through it from scratch. Smaller, purpose built decision models for specific tasks make a ton of sense! 🫡

Profilbild von Anthony Ronning
Anthony Ronningvor 10 Tagen

Love the specialized CU models and workflows coming out! Curious what the memory profile of this model is for running locally?

Profilbild von Shaurya
Shauryavor 10 Tagen

CUA agent focused on task completing is worth testing . Intresting !

Profilbild von Kartik
Kartikvor 10 Tagen

Cua team is cooking!

Profilbild von pH
pHvor 10 Tagen

always cooking 🔥

Profilbild von Layton Gott
Layton Gottvor 10 Tagen

This just get's better and better

Profilbild von Stephen Solka
Stephen Solkavor 10 Tagen

Stand by. Gpus spinning to add you to leaderboard.

Profilbild von Miguel Saavedra
Miguel Saavedravor 10 Tagen

Boom! Let’s goooo

Profilbild von The Real DMT
The Real DMTvor 10 Tagen

@grok what is the accuracy rate of cua vs jev?

Profilbild von Milind S
Milind Svor 10 Tagen

This is huge! Fastest computer use ever

Profilbild von Kevin Rajan
Kevin Rajanvor 10 Tagen

im so hypeeee

Profilbild von Elias Stråvik
Elias Stråvikvor 10 Tagen

letsgoo super excited to play more with cua - this inspired me to make a launch video with opus 5.5, best one shot i’ve ever seen

Profilbild von Arsh - 16 y/o builder
Arsh - 16 y/o buildervor 10 Tagen

Every single day you all are shipping something new working at cua must be exhilarating

Profilbild von Kush Agarwal
Kush Agarwalvor 10 Tagen

Browser and computer use about to be smooth

Profilbild von Brandon Shore
Brandon Shorevor 10 Tagen

how does this work

Profilbild von AS Maruf 🇨🇦🇧🇩
AS Maruf 🇨🇦🇧🇩vor 10 Tagen

How do I use Hermes agent to use CUA’s new model? @grok

Profilbild von Brjan | AI Builder
Brjan | AI Buildervor 10 Tagen

Cua-S1-4B-0.2's focus on live tasks with task-completion rewards is intriguing

Profilbild von peyCyber
peyCybervor 10 Tagen

held-out computer-use tasks will reveal whether rloo generalizes beyond training workflows

Profilbild von E_genius
E_geniusvor 10 Tagen

The one-action loop is the interesting bit. Less agent theater, more measurable task completion

Profilbild von Spencer Zhao
Spencer Zhaovor 10 Tagen

The no-accessibility-tree result is the spicy part. Would love to see a live-site eval where the button moves after the screenshot. That's when browser agents start sweating.

Profilbild von Ben Mo
Ben Movor 10 Tagen

That GUI-360 jump made me look twice. 60.1% to 92.9% without an accessibility tree is a big result. I'm especially interested in what happens after a wrong click. Recovery is where computer agents tend to get messy.

Profilbild von Aden
Adenvor 10 Tagen

One-episode gains on 18 tasks. GUI-360's 93% vs 60% is the story.

Ähnliche Videos

JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models paper page: Achieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle certain long-horizon tasks in an open world. However, they still struggle when the number of open-world tasks could potentially be infinite and lack the capability to progressively enhance task completion as game time progresses. We introduce JARVIS-1, an open-world agent that can perceive multimodal input (visual observations and human instructions), generate sophisticated plans, and perform embodied control, all within the popular yet challenging open-world Minecraft universe. Specifically, we develop JARVIS-1 on top of pre-trained multimodal language models, which map visual observations and textual instructions to plans. The plans will be ultimately dispatched to the goal-conditioned controllers. We outfit JARVIS-1 with a multimodal memory, which facilitates planning using both pre-trained knowledge and its actual game survival experiences. In our experiments, JARVIS-1 exhibits nearly perfect performances across over 200 varying tasks from the Minecraft Universe Benchmark, ranging from entry to intermediate levels. JARVIS-1 has achieved a completion rate of 12.5% in the long-horizon diamond pickaxe task. This represents a significant increase up to 5 times compared to previous records. Furthermore, we show that JARVIS-1 is able to self-improve following a life-long learning paradigm thanks to multimodal memory, sparking a more general intelligence and improved autonomy.

AK

141,440 Aufrufe • vor 2 Jahren

Explore state-of-the-art multimodal prompting in our new short course Large Multimodal Model Prompting with Gemini, taught by Erwin Huizenga in collaboration with Google Cloud. One interesting insight from this course: with multimodal models, prompt structure matters significantly. Placing text inputs, such as a patient's medical history, before image inputs, like an X-ray, can enhance the model's ability to contextualize and interpret visual data effectively. In other contexts, such as image captioning, you may get better results by putting the image first. Multimodal models behave differently than text-only LLMs, and effective prompting for models varies depending on the model you’re using. In this course you’ll learn how to effectively prompt Gemini models. Gemini's multimodal capabilities also enable new approaches in AI application development, for example: - The Gemini library handles various video formats (MP4, MOV, MPEG), streamlining applications using these formats. - Large context window (up to 1 million tokens) enables processing of extensive content, like analyzing multiple 50-minute videos simultaneously. - Function calling feature integrates real-time data (e.g., current exchange rates) into model responses. The course demonstrates building multimodal applications with real-world examples including document analyzers that reason across text and graphs simultaneously, video content extractors that find and timestamp specific information from multiple hours of footage, and automated expense report systems processing receipt images while cross-referencing company policies. Sign up here:

Andrew Ng

74,060 Aufrufe • vor 2 Jahren