Загрузка видео...

Не удалось загрузить видео

На главную

Delete all your E2E tests, seriously Okay, not the coverage, the code I’ve written Playwright tests, they’re great Maintaining hundreds of browser scripts while the UI keeps changing is not so much Honestly, AI feels more naturally suited to this than half the coding use cases we keep forcing...

90,340 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 23

Фото профиля Éverton Toffanetto
Éverton Toffanetto1 месяц назад

Tests are written as scripts because they need to be deterministic. Using AI introduces a variable you don’t fully control. I see them as complementary approaches, not as one replacing the other.

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

The control part is the part that you need to weight, since full control comes with a cost as well. If you don't want to pay by hour for developers to maintain tests but want to have a regression testing path. It's an easy way to get it done without even thinking much about it :)

Фото профиля Kevin Möchel
Kevin Möchel1 месяц назад

Have the AI write playwright tests Run them on a schedule If failed, have an agent fix them Run tests for free instead of a paid sub

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

You could do that but it's still not free, takes your time, validation and you have to physically be responsible for them It's a decision between your own time or just having it handled by a third party where all you maintain are written user stories

Фото профиля AI Apps API
AI Apps API1 месяц назад

The maintenance cost is real and it comes from the same place every time, selectors encode layout instead of intent. A test that says click the button labeled Save survives a redesign, one pinned to a generated class name does not. Agents help most when the step is described that way, and the flaky part left over is usually waiting, asserting on state that has actually settled rather than a fixed timeout.

Фото профиля Paweł Kubiak 🅰️
Paweł Kubiak 🅰️1 месяц назад

Ok, but what about costs? I assume that TesterArmy agent can interact with website by sending screenshot/html/a11y tree to LLM to identify crucial elements to interact with. If so, then it is expensive in terms of costs. In that case I would keep e2e tests.

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

From what I understand from docs it's mostly based on visuals and it charges you per run with a monthly subscription With regular E2E it all depends how much maintenance you have around them. If you have an app that doesn't change that much, you don't feel the pain. In fast paced env tools like TesterArmy remove the burden of updating tests and adding new is super simple Gonna dig around how actual internals work other than screenshots :)

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

I actually found out they’re not relying on vision by default. They use accessibility APIs first (DOM on the web, native APIs on mobile), and only fall back to vision when needed since it’s much slower and more expensive However you are paying per run so they need to care about the tokens not you :)

Фото профиля AI Alpha By Mario
AI Alpha By Mario1 месяц назад

The maintenance burden is the real killer — AI handling intent while the UI changes underneath feels like the right split

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

So far I had a really good and consistent results :)

Фото профиля AI Mastery Guide
AI Mastery Guide1 месяц назад

Plain English tests instead of babysitting selectors sounds so much easier.

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

Indeed! it's much more convenient not to think about all the selectors and code around them when you don't have to :)

Фото профиля AI Mastery Guide
AI Mastery Guide1 месяц назад

Less brittle scripts, more time actually shipping features.

Фото профиля Pavan K Jadda
Pavan K Jadda1 месяц назад

Looks interesting

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

Give it a try :)

Фото профиля Shahar Nechmad
Shahar Nechmad1 месяц назад

Agree. Maintaining selectors is probably one of the worst parts of E2E tests. These days, agents are smart enough, and computer use has become so much better. Check out also what Cloudflare is doing. They released a lot of infra in the past two weeks, to build this at scale.

Фото профиля incpo
incpo1 месяц назад

at leats playwright is free

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

yes it is, but your time is not :) Gotta decide what is worth more

Фото профиля Vivek Nayyar
Vivek Nayyar1 месяц назад

The issue is the latency and cost plus unpredictability. Will the model always find the right button to click and test 😅

Фото профиля Daniel Glejzner
Daniel Glejzner1 месяц назад

Give it a try :) it comes with a free trial, cost is per run - pretty convenient as well. In my experience it does a pretty damn good job at clicking the right things. Not sure about the latency? What's the issue there?

Фото профиля Vivek Nayyar
Vivek Nayyar1 месяц назад

Take screenshot, call llm to find the right element and then simulate click - this would have higher latency than a data test id click

Фото профиля fudes
fudes1 месяц назад

Quoting a VC as an authority on building is like quoting a food critic as an authority on cooking.

Фото профиля Mitchell Agoma
Mitchell Agoma1 месяц назад

Deleting brittle scripts while preserving meaningful coverage is the right distinction. AI can reduce maintenance, but teams still need deterministic checks and ownership of release risk. Related QA view:

Похожие видео

everyone in iOS development should watch this. seriously, it might change the whole industry. i pointed claude code at a live ios device running on revyl, typed "test everything," and walked away. here's what's actually happening: ① you don't write the tests. no scripts, no selectors, no test plan. i never told it which screens to open or what to check. it read the app, decided what mattered, and tested it. the entire instruction was "test everything." ② it built its own test team. it looked at the app, clocked that it's basically four mini apps (rides, delivery, services, account), and split itself into 4 agents, one per surface. scoping coverage like that is usually a person's whole afternoon. it did it in seconds, unprompted. ③ all four ran at the same time, each on its own live device. this is where revyl comes in. every agent gets its own live ios session in the cloud, so four running apps get tested in parallel instead of taking turns on one simulator. serial testing turns coverage into a time tax. running all of it at once removes the tax. ④ it tests like a person, not like a script. each agent drives the app the way a user would, taps through the flows, and visually checks each screen against what it expected to see. nothing is pinned to a brittle element id, so renaming a button doesn't take down half your suite. that one detail is the most annoying thing about how we test today, and it just quietly goes away. ⑤ no xcuitest, no sims melting your laptop. i didn't write a single xcuitest script, and there were no simulators booting on my machine. the agents run on cloud devices, so coverage stops being capped by what your laptop can handle. the part that got me isn't that an agent tested an app. it's that i never told it how. i handed it a device and an intent, and it figured out the scoping, the parallelizing, and the driving on its own. if you still write and maintain mobile ui tests by hand, i'm not sure that lasts the year.

Landseer Enga

23,963 просмотров • 4 месяцев назад

Bash is all you need! Which is why I'm introducing my holiday project: just-bash just-bash is a pretty complete implementation of bash in TypeScript designed to be used as a bash tool by AI agents. Because it turns out agents love exploring data via shell scripts, even beyond coding. It comes with grep, sed, awk and the 99th percentile features that an agent like Claude Code or Cursor would use. In fact, Claude Code can use it for secure bash execution. In the package - A bash-tool for AI SDK - A binary for use by yourself or your coding agents - An overlay filesystem to feed files to your agent securely - A Vercel Sandbox compatible API, so you can quickly upgrade to a real VM if you need to run binaries - An example AI agent that explores the just-bash code base using just-bash - I imported the Oils shell bash compatibility suite and just-bash passes a very good chunk What is interesting about this codebase: It was essentially entirely written by Opus 4.5. Coding agents love bash and they are good at reproducing it. They are also great at text-book recursive descent parsers and AST tweet-walk interpreters. That said, it is, like, a lot of code and I didn't read it all 😅. This is very much a hack, but it also seems to be _really_ useful. I haven't really found anything agents want to use that it doesn't support and it's fast and secure (caveats apply). It doesn't have write access to your computer and the filesystem is given a root that the agent cannot escape from. Find it at Related: Our recent blog post how we migrated our data analysis agent to bash tools and achieved incredible quality improvements The video shows the example agent investigating the just-bash code base

Malte Ubl

125,326 просмотров • 9 месяцев назад

Ever since I wired Claude Code to WhatsApp 3 weeks ago, I built a stupidly large infra around it. I mean, opus built it. No clue how the code even looks. The entire thing was vibe coded using my phone. I wanted to see how far I could push it without touching the computer. Everything via WhatsApp. Build what I need on the fly. So the resulting infrastructure will already be battle tested for software development. The entire thing was streamlined with nearly no manual interventions, everything was communicated via WhatsApp using a single script establishing this connection. If the script is down, I need to get home to start it again to resume the development. Claude was upgrading it, debugging it, restarting it while maintaining constant uptime so it could keep communicating with me. I stressed Claude about it, telling it that it will be “in the dark” and other words that deliberately sound scary about losing communications if the script dies. I also refused git and refused cloning the code, I wanted to see Claude adapting to work on a *LIVING* system. The way this whole thing works: Claude has its own dedicated phone number that I am paying for. A real WhatsApp account for it is installed on a real iPhone that is sitting on my desk. All is registered under my name, this is legit setup with no hacks and tricks. I’ve set up a WhatsApp “Community” and multiple different groups under it. Both me and Claude are the admins, so Claude could edit it on my behalf. Each group is a project I am working on and has its own isolated context. The Group description is a system prompt that gets auto-appended to the larger system prompt explaining this setup in general. When I send a message it’s an instant interrupt to Claude Code’s process, just like in the terminal. Voice notes are seamlessly transcribed with a local Whisper model. Images are used with multimodal reading in an isolated parallel session. Multiple groups running in parallel so I can work on all projects at the same time. No cross-talking, everything has an isolated context and history. And because it’s local on my own machine: Everything is REAL. The browser is REAL. I am connected as myself on it to all services because I actually use it in real life. Claude has unlimited internet access, just like humans who use actual browsers. It utilizes custom-made browser tools that I made to control any browser session it wants. Depending on the situation, it can either connect to my existing session or create one for its own. (You can tell it ‘look at my browser for a sec’ then talk about the current page you are on and it just works, pretty cool) My custom browser tools are not perfect (not by a long shot) but I managed to make them work well to the point they are somewhat reliable. This gives Claude full access to my real creds and all the services I actually use. I’m productive AS HELL with this. It really feels like a personal assistant. I ask it to read my emails and msgs, check x .com for news, research arxiv papers, write code, run experiments for me, investigate and reverse engineer github repos, even use my credit card and order things. [I try not to do this one a lot lol so far no disasters]. All from my phone. Super convenient. This is not a product or an open source project (maybe soon of it will make sense). This is just an ugly script I hacked the entire thing is ~600 lines. (ok maybe i did look at the code, but i swear i didn’t edit!) You can also vibe code this from scratch pretty fast and it will probably even end up better. This is just a cool thing so I’m sharing. It is a real speed booster for many things I do on daily basis, mostly boring things. Forcing my routine into some new “agent platform” just didn’t feel right for me. WhatsApp is where I already communicate and look for messages, so I decided that my agents will live there too. AGI in my pocket 24/7.

Yam Peleg

420,390 просмотров • 9 месяцев назад

i watched gemma 4 12b build something genuinely impressive today, and then loop itself to death right in front of me. the full run is in the video, sped up but completely uncut, watch it to the end and you will catch the exact moment it stops building and starts looping right in the middle of the work. the task was clean, build a single file gravity simulator, n-body physics, orbits, collisions, running locally on one 3090 through an agent. and for ten minutes it was a joy to watch. it reached for a symplectic integrator on its own, the correct one, the kind that keeps orbits stable instead of spiralling out. real gravity with softening, proper orbital velocities, momentum conserved on collision. the physics was right. the thing actually worked. then on the very last step, writing a few tests to prove its own code, it fell into a loop. not a crash, a loop. it started repeating itself and would not stop. ten more minutes, thirty four thousand tokens into a single answer, the same fragments over and over, until i killed it myself. so it's not that gemma can't code. it did the hard part beautifully. it cannot finish. it cannot hold a long task together without unravelling, and finishing is the entire job in agentic work. here's the part that stings. i run this exact task, same harness, same card, on the chinese open models, qwen especially, and i never see this. they build it, they test it, they stop. every single time. google has the raw capability, you can see it sitting right there in the code, and then the model loops itself to death on a task a 27b from alibaba finishes clean. open weights, apache 2.0, so much to love on paper. i just need it to know when to stop talking.

Sudo su

39,764 просмотров • 3 месяцев назад