Loading video...
Video Failed to Load
Hermes has twelve browser tools. Browser Use mode replaces them with a single one, driven by Browser Use's CLI 3.0. Instead of a dozen schemas in every request and a tool call per click, the agent writes a script. In our tests that cut token use 48-66% with no... show more
839,175 views • 1 month ago •via X (Twitter)
35 Comments

Simply run 'browser.backend: browser-use' to get started. More info on the browser automation toolset in Hermes:

We achieve this savings in two ways - the tool schema went from 8 tools, using a lot of context, to one, and the new tool has the agent drive the CLI with code, instead of a variety of individual actions. In our in house tests the total trajectory uses on average ~60% less tokens per task, with no accuracy drop!

@browser_use Also FYI - this works FOR EVERYONE. Local browser, browserbase, browser use, whatever (except camofox local- doesn't work with that)

@browser_use Testing this today :D

@browser_use Hermes 🤝 Browser Use

@browser_use Nousss not playing any games I see!! smokinnn fire 🔥

@browser_use Hey Hermes. Get your Doom Scroll on. Alert me when you find something interesting

@browser_use 🔥🔥🔥

@browser_use keep it up! great update

Collapsing twelve tools into one script changes more than token cost, it gives the agent a continuous execution model instead of forcing it through lossy tool boundaries. The real benchmark is recovery after partial failure: can it inspect state and resume without replaying everything?

@browser_use YES!!!! excellent work

@browser_use No accuracy drop for 48-66% less token is huge!

@browser_use Does this only work if Browser Use (via a Nous Portal subscription) is used as the backend? Or does all of this work, for example, with the local Camofox?

@browser_use Will test when I get back home! Any results regarding the speed?

@browser_use smart abstraction fewer schemas in context and fewer interaction turns is exactly where agent efficiency should improve

this is gonna be soooo helpful for running tests. has anyone figured out the process of automating test generation based on mapping out the user journey? i've tried using tools like but that's more for capturing user sessions and regression testing rather than e2e. it's annoying to have to find bugs manually after my llm makes a website for me or something. the tests it writes while long as fuck are not robust enough to test properly for functionality.

@browser_use i use hermes too, and i'd want the generated script saved alongside the run. when a browser task fails, replaying the exact code is much more useful than reconstructing the click sequence

@browser_use The video says for cdp backends. Is there a plan to make it work for camofox as well?

@browser_use Best agent to use. I know many people trying to make their own harness and such but this is by far the best agent harness if you want to actual get some work done.

@browser_use question, how did hermes cam up with its logo and brand.

@browser_use I might have feelings

@browser_use how can i get hermes to ask permission?

@browser_use Good, now work on overall token consumption based on @CommandCodeAI

@browser_use isn't browser use justanother sub?

@browser_use 🔥❤️

@browser_use If you're looking for alternatives, check @skyvernai out!

@browser_use there is no `browser.backend` key in config.yaml, fam. do you mean `web.backend`, or ` ? or something else? how do we set this?

@browser_use gm 😇

@browser_use one browser tool. good. I was running out of different ways to click the wrong thing.

@browser_use This entire comment section is just ai replies probably done by hermes.

@browser_use Nice, please do the same with Computer Use.

@browser_use The lower tool overhead is interesting. I would still want the script to leave a readable trace when it stops: what it tried, what state it reached, and what a human should check next. That recovery record matters just as much as the token savings.

@browser_use Browser use was very expensive, how is it now viable ?

@browser_use the real bottleneck was always token count, not accuracy

This is the same tradeoff every "agent writes code instead of discrete calls" post makes: 12 schema-validated tools each do exactly one narrow thing and nothing else. One tool that lets the agent write a script can do anything the CLI can do. The token savings are real. Worth knowing that the blast radius of a single bad generation just got wider than it was with 12 separate guardrails.
