Video wird geladen...
Video konnte nicht geladen werden
alright agent nerds, if you care about your tokens and usage limits, pay attention to the tools you give to your agents. i built a benchmark that compared various browser tools for agents, and here's an example of their massive difference in cost and latency doing the same task
571,689 Aufrufe • vor 6 Monaten •via X (Twitter)
62 Kommentare

full benchmark published at and you can try the most efficient option "chrome-devtools-axi" by simply telling your agent: "Run `npx -y chrome-devtools-axi` for browser automation"

Dude you’ve just come out of nowhere and dropped some of best agent tech I’ve seen

haha thanks. was previously working in big tech. only recently left to be a solo builder :)

Legendary

try and , much faster than all.

dev-browser was included in my benchmark report

codex with aegis just did this in 5 secs…

any idea why the web navigation took so little time? cached responses?

388 tokens in total also feels quite odd

this is bc aegis is fully scriptable — codex completed the task in the exact same fashion it writes/execs a bash script

interesting - that sounds similar to dev-browser which i evaluated in the benchmark and it could not get close to the numbers you are showing. i wonder what’s the delta there.. when i get a chance i’ll give aegis a go as well!

i don’t think that sounds like dev browser at all — granted i’ll have to go jump into their code but i would be extremely surprised if there was anything similar going on truth be told yes check it out and repro — full linux and macos support (macos is much more prod ready than linux as it stands currently)

wish you'd include agent-browser, playwriter-cli and other in the benchmark too so I finally know what to pick.

i did include agent-browser. it's right there in the video and benchmark report. i also included another popular option dev-browser which is doing some really interesting things

oh my bad, missed that. i'll give axi a spin for sure cause this workflow is actually something I use a lot.

BTW, have you tried getting agents to execute JS directly in the browser console? Most browser harnesses (like Claude Chrome or just writing playwright) are pretty bad for repetitive tasks - like scanning 50 Facebook posts related to some topic and summarizing them.

yes - all the conditions evaluated in the benchmark had a tool for executing JS in the browser environment, and that’s heavily used in many tasks during the benchmark

You have my attention, great work. I'm looking forward to adding it to my Hermes Agent CLI @NousResearch

I get the idea, but why are we trying to make TOON a thing? Why not just the standard TSV CSV or whatever that a CLI already outputs

csv/tsv is good with tabular data, but sometimes we need nested structures json is good at nested structures, but is designed for machines, not agents. its token efficiency is very poor

@cipherstein i think it becomes complex for an llm to learn a new thing and emit it at the same time. i think it will make the output less in quality

@cipherstein in AXI, TOON is only used for inputs to LLM. it never asks LLM to write outputs in TOON, exactly for that reason

Actionbook CLI went through the benchmark. 99% success, $0.0656/task, 24.1s avg. cheap and fast 🤔

saw your PR! will review soon

Hermes said it is powerful. We implemented. Thank you!

great job!! Would love to test libretto against this. Give me a moment.

why is this not front page news?! integrating as we speak, but for our usecase. Unreal guy, what a well done job!!

love it! thank you for sharing

bogus claims

yeah clearly something’s off in your setup there adding extra overhead, if every operation took 1+ seconds

tooling layer for agents is going through the exact same optimization cycle the frontend web did 10 years ago. we're focusing on hyperoptimized, minimal payloads.

pretty neat! The interface between agent and web is the next bottleneck.

@kunchenguid great work on the benchmark. what about also a score/criteria for sessions behave most indistinguishably from genuine human browsing? (timing patterns, input simulation, header/fingerprint management, etc.)?

hmm that feels like a goal i don’t want to see people optimize for

This is also a great ux to audit tool calling

can you also try @dokobot?

wish this existed six months ago. tested three browser tools for my agents — token cost gap was 10-20x between text snapshots and screenshot vision. and the cheap one wasn't even cheaper once you factor in retry loops from misread elements

Exceptional @kunchenguid driving it right now and it feels like cookie

Damn more of this content thanks

Great timing on this, I've been working on some MCPs and going to give this a shot in their place. 🔥

It's an interesting experiment and awesome to try, but after I installed it, it loads every-f-where in every agent as a hook with non-relevant URL details.

it’s just like a mcp server and skills that loads minimal instructions but good point maybe we should add a parameter to run one-off only

were you seeing it load a lot of data? it’s supposed to be just instructions for how to run the axi

Not lots, but in codex session start it gave few paragraph start hook with non relevant page info which was opened by another agent in another session

ok the "non relevant page info" part might be an oversight - let me take a look! thanks for flagging this Alex

fixed and released - page snapshot should not have been rendered in the default view

wild

hey a question how do you know what to build exactly so that problems are solved like the axi.md?

you mean how do I know what problems to solve? by doing a lot of things and talking to people, where i'd run into all those problems

so like currently i am deploying an agentic ai in kagent. how to know if i am right also i am building karpathy's llm_wiki for blog posts

ah that's a great question.. i actually don't think there's a good universal way right now to know "is my agent maximally efficient". i think you'd need to start by instrumenting where your tokens got spent and identify hypotheses for improvements

cool thanks for your answer. also would love to hear some of your insights on my builds

Holy! We need more people working with this principle-based mindset The next step is to make this work under auth websites

im sorry whats an AXI?

Nice benchmark, great heads-up. Y'all using the wrong browser tools are literally torching tokens and adding latency, lol

wow. axi is the way to go… more efficient and live round-trip editing and git optimization. i will primarily be using codex. hope it works

I use browser use

There's been many examples of superior tech that were never adopted for various reasons, the LaserDisc comes to mind. The key is adoption not just better metrics.

Can you add codemode to comparison?

code mode was evaluated and covered in the full report

why people care about token usage if in the future most likely the charge is per api call? This most likely will happen in the next few months..

@grok summarize this tool and double check the metrics here please

