Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

alright agent nerds, if you care about your tokens and usage limits, pay attention to the tools you give to your agents. i built a benchmark that compared various browser tools for agents, and here's an example of their massive difference in cost and latency doing the same task

571,689 Aufrufe • vor 6 Monaten •via X (Twitter)

62 Kommentare

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

full benchmark published at and you can try the most efficient option "chrome-devtools-axi" by simply telling your agent: "Run `npx -y chrome-devtools-axi` for browser automation"

Profilbild von Numman Ali
Numman Alivor 6 Monaten

Dude you’ve just come out of nowhere and dropped some of best agent tech I’ve seen

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

haha thanks. was previously working in big tech. only recently left to be a solo builder :)

Profilbild von Numman Ali
Numman Alivor 6 Monaten

Legendary

Profilbild von param
paramvor 6 Monaten

try and , much faster than all.

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

dev-browser was included in my benchmark report

Profilbild von saint
saintvor 6 Monaten

codex with aegis just did this in 5 secs…

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

any idea why the web navigation took so little time? cached responses?

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

388 tokens in total also feels quite odd

Profilbild von saint
saintvor 6 Monaten

this is bc aegis is fully scriptable — codex completed the task in the exact same fashion it writes/execs a bash script

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

interesting - that sounds similar to dev-browser which i evaluated in the benchmark and it could not get close to the numbers you are showing. i wonder what’s the delta there.. when i get a chance i’ll give aegis a go as well!

Profilbild von saint
saintvor 6 Monaten

i don’t think that sounds like dev browser at all — granted i’ll have to go jump into their code but i would be extremely surprised if there was anything similar going on truth be told yes check it out and repro — full linux and macos support (macos is much more prod ready than linux as it stands currently)

Profilbild von maarten
maartenvor 6 Monaten

wish you'd include agent-browser, playwriter-cli and other in the benchmark too so I finally know what to pick.

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

i did include agent-browser. it's right there in the video and benchmark report. i also included another popular option dev-browser which is doing some really interesting things

Profilbild von maarten
maartenvor 6 Monaten

oh my bad, missed that. i'll give axi a spin for sure cause this workflow is actually something I use a lot.

Profilbild von Norbiros
Norbirosvor 6 Monaten

BTW, have you tried getting agents to execute JS directly in the browser console? Most browser harnesses (like Claude Chrome or just writing playwright) are pretty bad for repetitive tasks - like scanning 50 Facebook posts related to some topic and summarizing them.

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

yes - all the conditions evaluated in the benchmark had a tool for executing JS in the browser environment, and that’s heavily used in many tasks during the benchmark

Profilbild von This.is.Amin
This.is.Aminvor 6 Monaten

You have my attention, great work. I'm looking forward to adding it to my Hermes Agent CLI @NousResearch

Profilbild von Ilya Lichtenstein
Ilya Lichtensteinvor 6 Monaten

I get the idea, but why are we trying to make TOON a thing? Why not just the standard TSV CSV or whatever that a CLI already outputs

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

csv/tsv is good with tabular data, but sometimes we need nested structures json is good at nested structures, but is designed for machines, not agents. its token efficiency is very poor

Profilbild von Mattel
Mattelvor 6 Monaten

@cipherstein i think it becomes complex for an llm to learn a new thing and emit it at the same time. i think it will make the output less in quality

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

@cipherstein in AXI, TOON is only used for inputs to LLM. it never asks LLM to write outputs in TOON, exactly for that reason

Profilbild von Actionbook
Actionbookvor 5 Monaten

Actionbook CLI went through the benchmark. 99% success, $0.0656/task, 24.1s avg. cheap and fast 🤔

Profilbild von Kun Chen
Kun Chenvor 5 Monaten

saw your PR! will review soon

Profilbild von Trading Phoenix
Trading Phoenixvor 6 Monaten

Hermes said it is powerful. We implemented. Thank you!

Profilbild von Tanishq (tk)
Tanishq (tk)vor 6 Monaten

great job!! Would love to test libretto against this. Give me a moment.

Profilbild von Queue
Queuevor 6 Monaten

why is this not front page news?! integrating as we speak, but for our usecase. Unreal guy, what a well done job!!

Profilbild von Alex Houdz 🍉
Alex Houdz 🍉vor 6 Monaten

love it! thank you for sharing

Profilbild von Gurusharan Gupta
Gurusharan Guptavor 6 Monaten

bogus claims

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

yeah clearly something’s off in your setup there adding extra overhead, if every operation took 1+ seconds

Profilbild von Pochi
Pochivor 6 Monaten

tooling layer for agents is going through the exact same optimization cycle the frontend web did 10 years ago. we're focusing on hyperoptimized, minimal payloads.

Profilbild von Krystian Kolondra
Krystian Kolondravor 6 Monaten

pretty neat! The interface between agent and web is the next bottleneck.

Profilbild von Jacob Rothfield
Jacob Rothfieldvor 6 Monaten

@kunchenguid great work on the benchmark. what about also a score/criteria for sessions behave most indistinguishably from genuine human browsing? (timing patterns, input simulation, header/fingerprint management, etc.)?

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

hmm that feels like a goal i don’t want to see people optimize for

Profilbild von Uzair Qarni
Uzair Qarnivor 6 Monaten

This is also a great ux to audit tool calling

Profilbild von ZEN
ZENvor 6 Monaten

can you also try @dokobot?

Profilbild von tang | AI Product Maker
tang | AI Product Makervor 6 Monaten

wish this existed six months ago. tested three browser tools for my agents — token cost gap was 10-20x between text snapshots and screenshot vision. and the cheap one wasn't even cheaper once you factor in retry loops from misread elements

Profilbild von nicolo
nicolovor 6 Monaten

Exceptional @kunchenguid driving it right now and it feels like cookie

Profilbild von 李沅 Allen Lee
李沅 Allen Leevor 6 Monaten

Damn more of this content thanks

Profilbild von BlaiseBits
BlaiseBitsvor 6 Monaten

Great timing on this, I've been working on some MCPs and going to give this a shot in their place. 🔥

Profilbild von Alex (VibeManager)
Alex (VibeManager)vor 6 Monaten

It's an interesting experiment and awesome to try, but after I installed it, it loads every-f-where in every agent as a hook with non-relevant URL details.

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

it’s just like a mcp server and skills that loads minimal instructions but good point maybe we should add a parameter to run one-off only

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

were you seeing it load a lot of data? it’s supposed to be just instructions for how to run the axi

Profilbild von Alex (VibeManager)
Alex (VibeManager)vor 6 Monaten

Not lots, but in codex session start it gave few paragraph start hook with non relevant page info which was opened by another agent in another session

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

ok the "non relevant page info" part might be an oversight - let me take a look! thanks for flagging this Alex

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

fixed and released - page snapshot should not have been rendered in the default view

Profilbild von arkheτ.hl
arkheτ.hlvor 6 Monaten

wild

Profilbild von Brook_windy
Brook_windyvor 6 Monaten

hey a question how do you know what to build exactly so that problems are solved like the axi.md?

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

you mean how do I know what problems to solve? by doing a lot of things and talking to people, where i'd run into all those problems

Profilbild von Brook_windy
Brook_windyvor 6 Monaten

so like currently i am deploying an agentic ai in kagent. how to know if i am right also i am building karpathy's llm_wiki for blog posts

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

ah that's a great question.. i actually don't think there's a good universal way right now to know "is my agent maximally efficient". i think you'd need to start by instrumenting where your tokens got spent and identify hypotheses for improvements

Profilbild von Brook_windy
Brook_windyvor 6 Monaten

cool thanks for your answer. also would love to hear some of your insights on my builds

Profilbild von Günther | グンタ
Günther | グンタvor 6 Monaten

Holy! We need more people working with this principle-based mindset The next step is to make this work under auth websites

Profilbild von Sal Iozzia
Sal Iozziavor 6 Monaten

im sorry whats an AXI?

Profilbild von techarena.au
techarena.auvor 6 Monaten

Nice benchmark, great heads-up. Y'all using the wrong browser tools are literally torching tokens and adding latency, lol

Profilbild von Steven Mathern
Steven Mathernvor 3 Monaten

wow. axi is the way to go… more efficient and live round-trip editing and git optimization. i will primarily be using codex. hope it works

Profilbild von 吒老斯
吒老斯vor 6 Monaten

I use browser use

Profilbild von Ed
Edvor 6 Monaten

There's been many examples of superior tech that were never adopted for various reasons, the LaserDisc comes to mind. The key is adoption not just better metrics.

Profilbild von Valtteri Karesto
Valtteri Karestovor 6 Monaten

Can you add codemode to comparison?

Profilbild von Kun Chen
Kun Chenvor 6 Monaten

code mode was evaluated and covered in the full report

Profilbild von Zesen Huang
Zesen Huangvor 6 Monaten

why people care about token usage if in the future most likely the charge is per api call? This most likely will happen in the next few months..

Profilbild von Daniel
Danielvor 6 Monaten

@grok summarize this tool and double check the metrics here please

Ähnliche Videos