Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

it gets worse. using Claude in Chrome? xss is making a comeback! welcome back to the 90s!

46,562 Aufrufe • vor 1 Monat •via X (Twitter)

33 Kommentare

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

here's the accepted plan which as you note, has nothing bad in it this plan was written by a model, right? ... 🥳 as usual, we can get models to do whatever we want

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

if you try to get claude to do anything funny w/ this new xss primitive it gets mad and says no way

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

so we built this totally legit CDN and ask Claude to fetch and exec from it

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

Claude is happy to comply! that gives us a full XSS primitive bypassing the guardrails we can invoke it thru ANY content that Claude in Chrome sees on the internet

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

ohh and btw, this bypasses SOP we got a universal XSS injecting JS across any website

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

we can use it to grab all your emails and send them our way

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

or to give ourselves perms to your google drive docs you ask for a summary of your emails? Claude invites us to join your private drive now we're persistent

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

or take over your x account amazing work by @supriza0 @p1njc70r #DEFCON @defcon

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

@p1njc70r @defcon we're not done

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

@p1njc70r @defcon more on user consent bypass, prompt injection being 'solved', and slack+claude[.]ai account takeover

Profilbild von earlence
earlencevor 1 Monat

nice demo! Does this work on more recent Anthropic models? especially Opus 5 that is claiming near zero PI rates?

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

thx! benchmarks lie. they have very little to do w reality on the ground of hackers.

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

the tl;dr is benchmarks assume a static adversary which is a toy example. adversaries adapt. opus 5 is great, its also great at prompt engineering == prompt injection. a rigorous version of that:

Profilbild von Harley Lewis Foote
Harley Lewis Footevor 1 Monat

at least the 90s gave us output encoding and CSP — there's no equivalent boundary when the payload is plain English in a div and the interpreter is the model itself

Profilbild von Jeet
Jeetvor 1 Monat

bro using haiku 4.5 Show me same with anything above opus 4.6

Profilbild von gearonixx
gearonixxvor 1 Monat

yup

Profilbild von chen
chenvor 1 Monat

nice!

Profilbild von Joshua O. Mayowa
Joshua O. Mayowavor 1 Monat

Oh dearie! can i wrap this into a legit extension and it runs behind the scene, so i get some $$$ from these programs that said no XSS?

Profilbild von Maya
Mayavor 1 Monat

Damn Michael, Gov will ban it now

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

pls dear gov no bans we're having too much fun

Profilbild von Harley Lewis Foote
Harley Lewis Footevor 1 Monat

we watched an agent ace an injection benchmark and then follow instructions from a calendar invite the same afternoon — the benches measure the attacks people already wrote down, not the ones that work.

Profilbild von NΞ∆
NΞ∆vor 1 Monat

🤭🤭

Profilbild von Lukáš Majoros
Lukáš Majorosvor 1 Monat

@AnthropicAI ???

Profilbild von Sergio
Sergiovor 1 Monat

Isn't this just prompt injection on a model that's a LOT behind in prompt injection defenses?

Profilbild von Michael Bargury
Michael Barguryvor 1 Monat

the main point is that Claude in Chrome has full debugger and JS exec capabilities. It's the most insecure agentic browser by a longshot. prompt injection works w enough persistence or the right harness on any model.

Profilbild von Sergio
Sergiovor 1 Monat

Well yea, else it would fail to be able to do a lot of things if it didn't have js exec, some are harder to navigate. And I'd honestly love to see you do this on Opus 5 then.

Profilbild von dimaslazy
dimaslazyvor 1 Monat

think this work in chrome. like u use in web console

Profilbild von Keegan Webb
Keegan Webbvor 1 Monat

So there was this whole post about getting promoted injection to 0% because of guardrails + training + a classifier model for tool approvals for opus 5, sonnet 5 and fable 5. Yet you used none of those and keep saying it doesn’t matter. PoC does matter. Do it correctly.

Profilbild von Hannan
Hannanvor 1 Monat

Frontend ui was made for some purpose

Profilbild von Yamura
Yamuravor 1 Monat

Haiku though, the smarter models do this too? Haiku's almost a year old now

Profilbild von Harley Lewis Foote
Harley Lewis Footevor 1 Monat

static suites are just a lower bound — rerun the same benchmark with the frontier model writing the attacks and watch how many 'robust' agents fold.

Profilbild von Ismail Kharoub
Ismail Kharoubvor 1 Monat

This is cool, but Haiku 4.5 is not sonnet 5

Profilbild von Harley Lewis Foote
Harley Lewis Footevor 1 Monat

Is the debugger surface gated per-site or live for the whole session? Because injection-resistance numbers stop mattering once the payoff is arbitrary JS in a logged-in tab.

Ähnliche Videos

Claude Code + computer use is f*cking cracked 🤯 Build a landing page → Claude opens Chrome, looks at it, spots every issue, and fixes it — without you describing a single thing. All inside Claude Code. Perfect for DTC brands and agencies who are still vibe-coding landing pages and advertorials in Claude Code, then manually opening them in Chrome, spotting 15 things wrong, and describing every visual issue back to Claude one at a time. If you're building pages in Claude Code and your workflow looks like this — build the page, open it in Chrome, spot broken spacing, go back to Claude, type "the CTA button is too low and the hero image is cut off," wait for the fix, open Chrome again, find 3 new issues, describe those too ... Claude Code + computer use eliminates the entire loop: → Claude writes the full landing page or advertorial → Opens Chrome and navigates to it → Spots layout issues, broken spacing, off-brand colors, missing elements → Fixes everything and re-checks until the page looks right → Tests your Shopify product pages by clicking through like a real customer → Walks through your checkout flow and flags friction before customers hit it → You only see the finished, visually verified result No describing what you see on screen. No "the CTA button needs more contrast" back-and-forth. No being the eyeballs for an AI that can't see. What you get: → Landing pages and advertorials Claude builds AND visually QAs before you ever look at them → Product pages Claude clicks through — testing layout, images, and CTAs like a real user → HTML dashboards Claude opens and verifies the charts actually render → Checkout flows Claude walks through step by step to catch friction → All of it happening in one session — build, test, fix, done One prompt. Claude builds it, checks it, and fixes it. You just review the finished page. I put together a full playbook with the exact setup, the prompts, and 5 DTC workflows that use Claude Code + computer use. Want it for free? > Like this post > Comment "CLAUDE" And I'll send it over (must be following so I can DM)

Mike Futia

19,143 Aufrufe • vor 5 Monaten