Video wird geladen...
Video konnte nicht geladen werden
Combined Jev with Playwright-controlled Chrome to build jev-browser, a general browser automation skill. Works across agents. Two demos in Antigravity CLI and Codex: finding related articles, then job openings on a site. Both ran fast. This could handle many more browser tasks.
45,053 Aufrufe • vor 16 Tagen •via X (Twitter)
41 Kommentare

I built this as a non-coder. Imagine what the nerds do.

This looks pretty cool.

Thanks.

Jev supplies the agent's reasoning while Playwright supplies reliable browser control, so the same skill can transfer across agents instead of being locked to one tool.

Exactly. It can handle pretty general browser tasks this way.

when you say “general browser tasks,” do you mean multi-step flows like research + form filling, or mostly single-site actions?

Running systematic latency and resource‑usage benchmarks across varied page loads will reveal how jev-browser scales, while security checks on script injection and cookie handling add confidence for broader deployment.

It’s still more of a personal tool for now. There are definitely a lot more details to think through if you want to use it at scale.

Combining browser control with agents allows them to find articles and job openings. This demonstrates how agents can be equipped with general skills to execute diverse online tasks.

2fa is exactly where this needs a human handoff state, not a stronger model. pause at the code field, preserve the page and resume with a fresh screenshot after the user enters it.

Why not run the model locally as well Meet Kevin 🤝

Job hunting with this sounds useful.

How does it handle interactive elements that are not obvious by inspecting the DOM itself? Like a div that triggers some JS on click.

nice, the cross-agent angle is cool. i went the existing Codex browser-tools route with Jev choosing the next click. how are you verifying each hop in yours?

@hqmank related articles is the warmup. job openings is why Playwright's holding Chrome.

What does the skill hand Jev on each step? TypeSafe's own writeup says it takes structured program state as text rather than images, so I am curious whether you feed it the accessibility tree, a DOM digest, or something you rolled yourself.

First, it parses the DOM to get all the actionable elements on the page. Jev then picks the most likely element + action, like clicking a link or typing into an input. It executes that action and repeats the loop.

A real page's actionable set blows past Jev's 32k state-plus-longest-question budget fast. Do you rank the candidates down before they go in, or does the whole list ride in the Choice options?

Love the Jev + real Chrome angle. Once you have a decision layer, the boring part is still the browser control plane: keep a real session, see what it sees, and give the agent deterministic actions. I went the "desktop Chromium with a localhost REST API" route for that. page_map / click / type / wait / screenshot on 127.0.0.1:9090, so any agent runtime can drive it without bolting on a full Playwright stack every time. Shipped as NTXM Browser (Windows, CEF). Early, but the loop feels solid for demos like related articles / job pages.

does this require the actual jev model?

I’m using TypeSafe’s official Jev API.

so it would require an API key?

give me github link bro

I’ll open source it once it’s ready and post an update. Feel free to follow me so you don’t miss it.

i did not expect you this man you did not offer me dm me

еще один фэйк проект

how does it behave on a rerun? we point jev at approve or reject on drafts and the thing i watch is how often the same state flips the answer.

I’ve rerun the same task 2–3 times and haven’t seen it flip or behave unexpectedly so far.

more stable than i expected. i'd still want a few hundred runs before it gets approve or reject rights. did you keep the page identical between runs, or let it change?

can we get access

Still polishing it. I’ll share it soon, so keep an eye on my posts.

'Works across agents' is the line I'd pull on. Two harnesses driving one Chrome look like one actor to the site and to your logs; nothing on the wire says whether Codex or Antigravity clicked. We stamp the calling agent on each CDP attach for that; per-agent consent came free.

Browser automation as a portable skill is the right shape. The bit that decides whether it survives a week is what it returns on a bad run: finding zero related articles and failing to load the page should not be the same output. How does yours signal that?

Handling session persistence across multiple agent runs is usually where these break down.

Can this solve captchas?

I’ve seen GPT-6 handle this, so I’d probably leave it to a more capable model. Should be doable.

Finding related articles then job openings is a great demo. Very practical use of browser automation.

Jev gave my job crawler a pretty big speed boost. Today I took it a step further and turned it into a more general skill that can handle all kinds of browser tasks.

great u think it can handle auth flows with 2fa?

I think it could. For more complex steps, you could hand them off to a stronger model like GPT-6.

yeah smart, no point wasting astra on clicking buttons lol
