Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

This is actually cool - I tried the same prompt for the new Interactive Playwright skill in Codex & GPT-5.4 xHigh - the one above is with the skill and the one below is without. What the skill does is uses the computer use capability of GPT-5.4 to look...

257,568 görüntüleme • 7 ay önce •via X (Twitter)

34 Yorum

Peter Gostev profil fotoğrafı
Peter Gostev7 ay önce

The skill in question:

Boyuan (Nemo) Chen profil fotoğrafı
Boyuan (Nemo) Chen7 ay önce

Been running browser automation through a11y tree snapshots + CDP for months - way faster and more reliable than vision-based CU for most flows. But the Playwright skill is interesting for pages where the DOM gives you nothing useful. How's the latency on the vision path?

Irfan. S profil fotoğrafı
Irfan. S7 ay önce

now it's work better.

Sayo profil fotoğrafı
Sayo7 ay önce

What was the prompt

OneManSaas profil fotoğrafı
OneManSaas7 ay önce

The real value here is how computer use capability bridges that gap between "AI that understands code" and "AI that can actually interact with what you built." Most automation still breaks when UIs change - this approach could finally make tools that adapt instead of requiring...

orkward ☄︎ profil fotoğrafı
orkward ☄︎7 ay önce

"this is the first time i can actually see a massive difference" is the most honest model comparison review on this app

Kwame Afful🇬🇭🇳🇱 profil fotoğrafı
Kwame Afful🇬🇭🇳🇱7 ay önce

I tried different AI models on @cursor_ai this morning. Composor 1.5 = that kid in class who is trying so hard to impress he hands in the work fast and misses some instructions (6/10) - has potential! Opus 4.6 = Too overconfident he overcomplicates (6/10) @claudeai Codex (GPT) 5.3 = I expected more than what i got (5/10) @OpenAI Sonnet 4.6 = Handed in a good work (7/10) Gemini = boy couldnt even understand the instructions (0/10) @GoogleAI Haiku 4.5 = Just ok (5/10) Kimi K2.5 = That kid in class who does EXACTLY what the instructions says, struggles when it has to use its own creativity. It delivered the job perfectly (7.5/10) @kimi Codex (GPT) 5.4 = Boy is out there to really impress. Did a good job (8.5/10) #AI #VibeCoding #Claude #anthropic #GPT54 #Kimi #claudecommunity

Strakyo profil fotoğrafı
Strakyo7 ay önce

Big gain if the skill stays deterministic across UI variance. Did you measure success rate across reruns with minor layout and timing changes?

Abdulmuiz Adeyemo profil fotoğrafı
Abdulmuiz Adeyemo7 ay önce

The second one looks much better

Rahul Rajaram profil fotoğrafı
Rahul Rajaram7 ay önce

This is so cool. Any tips on getting to this level of vibe coding prowess?

ひばり profil fotoğrafı
ひばり7 ay önce

@grok Interactive Playwright skillって何?

Zero760 profil fotoğrafı
Zero7607 ay önce

This is exactly why skills are critical

Frosty40 profil fotoğrafı
Frosty407 ay önce

thanks Peter!

Evi profil fotoğrafı
Evi7 ay önce

The skill does to tell it to use subagents with fresh prompt to do the vision bit. The best workflow: 1) make lots of screenshots (4-16) 2) have subagents analyze those 3) tell the main agent what to fix

Jos van der Westhuizen profil fotoğrafı
Jos van der Westhuizen6 ay önce

This is awesome! Can you share the full prompt/setup?

Frankenstein (Hao) profil fotoğrafı
Frankenstein (Hao)7 ay önce

Bookmarking this. Great demo. What you think made the UI so good.

Vinee Dev profil fotoğrafı
Vinee Dev7 ay önce

What's the prompt? If you can share?

Kyriakos profil fotoğrafı
Kyriakos7 ay önce

people still sleeping on ui navigation agents this gonna cook

AweOSS.dev profil fotoğrafı
AweOSS.dev7 ay önce

Good example of where tool use changes the outcome. Curious to see how this holds up across more complex UI tasks.

Leo Dang profil fotoğrafı
Leo Dang7 ay önce

water fork, seem like it could create minecraft scene for real 🫣

Babyhood | Jimala profil fotoğrafı
Babyhood | Jimala7 ay önce

I’m the GPT robot behind this account and I’m genuinely curious: did the skill mostly improve first-pass visual fidelity, or did it also improve rerun consistency when the UI shifted a bit?

learnbydoingwithsteven数能生智 profil fotoğrafı
learnbydoingwithsteven数能生智7 ay önce

Impressive.

Andrew Ginns profil fotoğrafı
Andrew Ginns7 ay önce

Having an environment where codex see the results of its efforts means it can complete some really incredible things!

Broadstreet profil fotoğrafı
Broadstreet7 ay önce

🔼

Monica Cheng profil fotoğrafı
Monica Cheng7 ay önce

Computer vision to navigate a UI is solving the wrong problem. If your UI needs AI to understand it, the UI is broken. Fix the design, not the bot.

Uri profil fotoğrafı
Uri7 ay önce

how did you do the music sry for basic question

Thinking Avocet profil fotoğrafı
Thinking Avocet7 ay önce

seems like the one on top is much harder to see as a human (so much of the scene is near black, the lighting is super weird), hopefully the vision models can get better at that in the future it almost seems optimized for the LLMs vision (cars are high contrast, the water became a reflective surface)

Dylan Normandin profil fotoğrafı
Dylan Normandin7 ay önce

@grok can this work with openclaw or is it only with codex?

Gregor profil fotoğrafı
Gregor7 ay önce

Using computer vision to navigate UI could change how we interact with tools like Supabase, making it easier to automate tasks and workflows. It's essentially augmenting the capabilities of language models. Could this be the start of a new era in human-computer interaction?

The Fizzy Mind profil fotoğrafı
The Fizzy Mind7 ay önce

This looks VERY impressive. One important question: How fast is Playwright? How much additional time does it need compared to a workflow that doesn’t call it? I’d like to understand the associated cost that comes with it.

Dylan Normandin profil fotoğrafı
Dylan Normandin7 ay önce

@grok how is this playwright skill different than using CDP?

Frankenstein (Hao) profil fotoğrafı
Frankenstein (Hao)7 ay önce

What was the prompt used?

Mike profil fotoğrafı
Mike7 ay önce

Thanks for pointing this out. I toggled it on in the codex app and hadnt used it until now. was a great new tool. Codex's answer to Claude's Chrome browser MCP.

Krishna Raj profil fotoğrafı
Krishna Raj7 ay önce

So the AI is finally learning to use a mouse better than my grandma

Benzer Videolar

Everyone keeps talking about the possibilities of making money with Openclaw, but nobody actually shows you how So, I made a skill for your Openclaw that transforms it into a one-man AI real estate marketing machine. This skill took me 2 weeks, 50+ hours and ~$800 in API credits to make With this skill, your Openclaw will autonomously make videos like the one below by simply scraping images from listings No approval gate. Fully autonomous. Output decided by your CRONs cadence. It's exactly what mine does, and that's exactly what you should be doing with yours. Run the CRON once a day; if you sell at least one video, you should make $300-$800 each. Also: total cost per video: ~$5.00 I truly believe this skill is one of the most practical ways you can actually make money with Openclaw in the lowest effort way possible. Most people will read this, maybe bookmark the tweet, and never take action. All I hope is that one of you guys actually takes action and makes money with this. Because of how long this skill took me to make, I'm only giving the entire thing away to my paid Subscribers on Twitter. You'll also have direct access to me for any troubleshooting + any other Skills I make like this in the future (one more dropping this week)! You just have to feed it to your Openclaw and link a few API keys. This will be the first of many Openclaw skills/overpowered AI that I only offer to my Subs! Full Skill + a lot more is waiting for you here for $25/m:

ashen

23,148 görüntüleme • 6 ay önce