Video wird geladen...
Video konnte nicht geladen werden
Delete all your E2E tests, seriously Okay, not the coverage, the code I’ve written Playwright tests, they’re great Maintaining hundreds of browser scripts while the UI keeps changing is not so much Honestly, AI feels more naturally suited to this than half the coding use cases we keep forcing... show more
90,340 Aufrufe • vor 1 Monat •via X (Twitter)
23 Kommentare

Tests are written as scripts because they need to be deterministic. Using AI introduces a variable you don’t fully control. I see them as complementary approaches, not as one replacing the other.

The control part is the part that you need to weight, since full control comes with a cost as well. If you don't want to pay by hour for developers to maintain tests but want to have a regression testing path. It's an easy way to get it done without even thinking much about it :)

Have the AI write playwright tests Run them on a schedule If failed, have an agent fix them Run tests for free instead of a paid sub

You could do that but it's still not free, takes your time, validation and you have to physically be responsible for them It's a decision between your own time or just having it handled by a third party where all you maintain are written user stories

The maintenance cost is real and it comes from the same place every time, selectors encode layout instead of intent. A test that says click the button labeled Save survives a redesign, one pinned to a generated class name does not. Agents help most when the step is described that way, and the flaky part left over is usually waiting, asserting on state that has actually settled rather than a fixed timeout.

Ok, but what about costs? I assume that TesterArmy agent can interact with website by sending screenshot/html/a11y tree to LLM to identify crucial elements to interact with. If so, then it is expensive in terms of costs. In that case I would keep e2e tests.

From what I understand from docs it's mostly based on visuals and it charges you per run with a monthly subscription With regular E2E it all depends how much maintenance you have around them. If you have an app that doesn't change that much, you don't feel the pain. In fast paced env tools like TesterArmy remove the burden of updating tests and adding new is super simple Gonna dig around how actual internals work other than screenshots :)

I actually found out they’re not relying on vision by default. They use accessibility APIs first (DOM on the web, native APIs on mobile), and only fall back to vision when needed since it’s much slower and more expensive However you are paying per run so they need to care about the tokens not you :)

The maintenance burden is the real killer — AI handling intent while the UI changes underneath feels like the right split

So far I had a really good and consistent results :)

Plain English tests instead of babysitting selectors sounds so much easier.

Indeed! it's much more convenient not to think about all the selectors and code around them when you don't have to :)

Less brittle scripts, more time actually shipping features.

Looks interesting

Give it a try :)

Agree. Maintaining selectors is probably one of the worst parts of E2E tests. These days, agents are smart enough, and computer use has become so much better. Check out also what Cloudflare is doing. They released a lot of infra in the past two weeks, to build this at scale.

at leats playwright is free

yes it is, but your time is not :) Gotta decide what is worth more

The issue is the latency and cost plus unpredictability. Will the model always find the right button to click and test 😅

Give it a try :) it comes with a free trial, cost is per run - pretty convenient as well. In my experience it does a pretty damn good job at clicking the right things. Not sure about the latency? What's the issue there?

Take screenshot, call llm to find the right element and then simulate click - this would have higher latency than a data test id click

Quoting a VC as an authority on building is like quoting a food critic as an authority on cooking.

Deleting brittle scripts while preserving meaningful coverage is the right distinction. AI can reduce maintenance, but teams still need deterministic checks and ownership of release risk. Related QA view:
