Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

it gets worse. using Claude in Chrome? xss is making a comeback! welcome back to the 90s!

46,562 görüntüleme • 1 ay önce •via X (Twitter)

33 Yorum

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

here's the accepted plan which as you note, has nothing bad in it this plan was written by a model, right? ... 🥳 as usual, we can get models to do whatever we want

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

if you try to get claude to do anything funny w/ this new xss primitive it gets mad and says no way

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

so we built this totally legit CDN and ask Claude to fetch and exec from it

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

Claude is happy to comply! that gives us a full XSS primitive bypassing the guardrails we can invoke it thru ANY content that Claude in Chrome sees on the internet

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

ohh and btw, this bypasses SOP we got a universal XSS injecting JS across any website

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

we can use it to grab all your emails and send them our way

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

or to give ourselves perms to your google drive docs you ask for a summary of your emails? Claude invites us to join your private drive now we're persistent

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

or take over your x account amazing work by @supriza0 @p1njc70r #DEFCON @defcon

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

@p1njc70r @defcon we're not done

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

@p1njc70r @defcon more on user consent bypass, prompt injection being 'solved', and slack+claude[.]ai account takeover

earlence profil fotoğrafı
earlence1 ay önce

nice demo! Does this work on more recent Anthropic models? especially Opus 5 that is claiming near zero PI rates?

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

thx! benchmarks lie. they have very little to do w reality on the ground of hackers.

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

the tl;dr is benchmarks assume a static adversary which is a toy example. adversaries adapt. opus 5 is great, its also great at prompt engineering == prompt injection. a rigorous version of that:

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote1 ay önce

at least the 90s gave us output encoding and CSP — there's no equivalent boundary when the payload is plain English in a div and the interpreter is the model itself

Jeet profil fotoğrafı
Jeet1 ay önce

bro using haiku 4.5 Show me same with anything above opus 4.6

gearonixx profil fotoğrafı
gearonixx1 ay önce

yup

chen profil fotoğrafı
chen1 ay önce

nice!

Joshua O. Mayowa profil fotoğrafı
Joshua O. Mayowa1 ay önce

Oh dearie! can i wrap this into a legit extension and it runs behind the scene, so i get some $$$ from these programs that said no XSS?

Maya profil fotoğrafı
Maya1 ay önce

Damn Michael, Gov will ban it now

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

pls dear gov no bans we're having too much fun

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote1 ay önce

we watched an agent ace an injection benchmark and then follow instructions from a calendar invite the same afternoon — the benches measure the attacks people already wrote down, not the ones that work.

NΞ∆ profil fotoğrafı
NΞ∆1 ay önce

🤭🤭

Lukáš Majoros profil fotoğrafı
Lukáš Majoros1 ay önce

@AnthropicAI ???

Sergio profil fotoğrafı
Sergio1 ay önce

Isn't this just prompt injection on a model that's a LOT behind in prompt injection defenses?

Michael Bargury profil fotoğrafı
Michael Bargury1 ay önce

the main point is that Claude in Chrome has full debugger and JS exec capabilities. It's the most insecure agentic browser by a longshot. prompt injection works w enough persistence or the right harness on any model.

Sergio profil fotoğrafı
Sergio1 ay önce

Well yea, else it would fail to be able to do a lot of things if it didn't have js exec, some are harder to navigate. And I'd honestly love to see you do this on Opus 5 then.

dimaslazy profil fotoğrafı
dimaslazy1 ay önce

think this work in chrome. like u use in web console

Keegan Webb profil fotoğrafı
Keegan Webb1 ay önce

So there was this whole post about getting promoted injection to 0% because of guardrails + training + a classifier model for tool approvals for opus 5, sonnet 5 and fable 5. Yet you used none of those and keep saying it doesn’t matter. PoC does matter. Do it correctly.

Hannan profil fotoğrafı
Hannan1 ay önce

Frontend ui was made for some purpose

Yamura profil fotoğrafı
Yamura1 ay önce

Haiku though, the smarter models do this too? Haiku's almost a year old now

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote1 ay önce

static suites are just a lower bound — rerun the same benchmark with the frontier model writing the attacks and watch how many 'robust' agents fold.

Ismail Kharoub profil fotoğrafı
Ismail Kharoub1 ay önce

This is cool, but Haiku 4.5 is not sonnet 5

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote1 ay önce

Is the debugger surface gated per-site or live for the whole session? Because injection-resistance numbers stop mattering once the payoff is arbitrary JS in a logged-in tab.

Benzer Videolar

Claude Code + computer use is f*cking cracked 🤯 Build a landing page → Claude opens Chrome, looks at it, spots every issue, and fixes it — without you describing a single thing. All inside Claude Code. Perfect for DTC brands and agencies who are still vibe-coding landing pages and advertorials in Claude Code, then manually opening them in Chrome, spotting 15 things wrong, and describing every visual issue back to Claude one at a time. If you're building pages in Claude Code and your workflow looks like this — build the page, open it in Chrome, spot broken spacing, go back to Claude, type "the CTA button is too low and the hero image is cut off," wait for the fix, open Chrome again, find 3 new issues, describe those too ... Claude Code + computer use eliminates the entire loop: → Claude writes the full landing page or advertorial → Opens Chrome and navigates to it → Spots layout issues, broken spacing, off-brand colors, missing elements → Fixes everything and re-checks until the page looks right → Tests your Shopify product pages by clicking through like a real customer → Walks through your checkout flow and flags friction before customers hit it → You only see the finished, visually verified result No describing what you see on screen. No "the CTA button needs more contrast" back-and-forth. No being the eyeballs for an AI that can't see. What you get: → Landing pages and advertorials Claude builds AND visually QAs before you ever look at them → Product pages Claude clicks through — testing layout, images, and CTAs like a real user → HTML dashboards Claude opens and verifies the charts actually render → Checkout flows Claude walks through step by step to catch friction → All of it happening in one session — build, test, fix, done One prompt. Claude builds it, checks it, and fixes it. You just review the finished page. I put together a full playbook with the exact setup, the prompts, and 5 DTC workflows that use Claude Code + computer use. Want it for free? > Like this post > Comment "CLAUDE" And I'll send it over (must be following so I can DM)

Mike Futia

19,143 görüntüleme • 5 ay önce