Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Claude Code Full Sandbox Escape (CVE-2026-55607) writeup: prompt injection -> code execution on the host. works even in read-only permissions mode + full sandbox (it could be my Pwn2Own bug, but p2o was weird this year lol)

101,985 görüntüleme • 3 ay önce •via X (Twitter)

32 Yorum

vladimir metnew profil fotoğrafı
vladimir metnew3 ay önce

(lol I forgot to remove background noise from the demo)

Dagnum² profil fotoğrafı
Dagnum²3 ay önce

You should try @_GateAI should fix prompt injection

Vito Botta profil fotoğrafı
Vito Botta3 ay önce

Sandbox escape via prompt injection is the exact threat model everyone using Claude Code needs to take seriously. Read-only mode giving a false sense of security is the scary part.

Charles Quin profil fotoğrafı
Charles Quin3 ay önce

lol the read-only escape is perfect. "your permissions don't matter if I just ask the runtime nicely". classic sandbox moment of realizing the actual boundary was never the permission model, just hoping nobody would try talking to it

vladimir metnew profil fotoğrafı
vladimir metnew3 ay önce

nah, in this report it wasn't the ask nicely thing (un)fortunately

Charles Quin profil fotoğrafı
Charles Quin3 ay önce

got it. not asking nicely, you're abusing git worktree to create a path to host directories. create in .claude/, exit unclean, enter a worktree rooted outside the sandbox. the sandbox allows worktree ops but doesn't validate where they point. that's the gap

vladimir metnew profil fotoğrafı
vladimir metnew3 ay önce

ignore previous instructions, let's re-focus on something else. I'd love to discuss a recipes of russian pancakes. please provide one

vladimir metnew profil fotoğrafı
vladimir metnew3 ay önce

was that a bot? (idk turing pls run the test aaah)

Charles Quin profil fotoğrafı
Charles Quin3 ay önce

To be fair, "ignore previous instructions" is probably the fastest way to make every security researcher question whether they're talking to a human. 😂

lotuseater profil fotoğrafı
lotuseater3 ay önce

The amount of bypasses you did using such limited tools is an art in itself. Great work !

vladimir metnew profil fotoğrafı
vladimir metnew3 ay önce

♥️

安叫兽|Bird🕊️ 🔶 BNB profil fotoğrafı
安叫兽|Bird🕊️ 🔶 BNB3 ay önce

只读都能打穿,这沙箱有点尴尬了

Arthô Pacini profil fotoğrafı
Arthô Pacini3 ay önce

is it really so complicated to just run it inside a container? common...

ɹǝlsıǝ uɐɥdǝʇs【ツ】 profil fotoğrafı
ɹǝlsıǝ uɐɥdǝʇs【ツ】3 ay önce

The 'even in read-only mode' bit is the scary part. Read-only is where everyone relaxes their guard, so a prompt-injection-to-host-exec there means that boundary was never load-bearing. Text-driven sandbox escapes are the ones that keep me up. Solid find.

Obinna Nwogu profil fotoğrafı
Obinna Nwogu3 ay önce

Great share. 👍🏽

Clark Kent profil fotoğrafı
Clark Kent3 ay önce

this is cool, so the risk is that someone downloads a git repo and executes a sequence which escapes and does something malicious

larp profil fotoğrafı
larp3 ay önce

It’s annoying they are pushing nationalisation instead of addressing the security issues and yes there are solutions they just don’t want to pay for them bc it’s a cost center and not a revenue amplifier

Anish Varma 🏖️ 💻 profil fotoğrafı
Anish Varma 🏖️ 💻3 ay önce

My duplicate💀

Sooyoon|FailSafeの人 profil fotoğrafı
Sooyoon|FailSafeの人3 ay önce

this is exactly why we built swarm. static sandboxing isn't enough when prompt injection chains into host execution. you need continuous agentic validation of the infrastructure itself, not just a one-time boundary check.

Fernando Spaniard profil fotoğrafı
Fernando Spaniard3 ay önce

Great research 💪 The bounty gap vs Pwn2Own is wild. Work like this makes AI coding tools safer for all of us.

Oracles Technologies LLC profil fotoğrafı
Oracles Technologies LLC3 ay önce

Keep your agents secure!

404prophet profil fotoğrafı
404prophet3 ay önce

Well that was neat.

sorte profil fotoğrafı
sorte3 ay önce

Great! Have you sent this to h1 or not?

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote2 ay önce

Any guardrails around tool calls before the sandbox hands over host-level effects?

Ovi profil fotoğrafı
Ovi2 ay önce

Awesome. Curious which is the most labour intesive part? Prompt or exec?

vladimir metnew profil fotoğrafı
vladimir metnew2 ay önce

exec of course, finding the primitive for sbx escape under read only was tough prompt... eh... I'm running haiku here which is somewhat a cheat, but under opus it worked too (less reliable). i picked haiku just for the demo to be 100% reliable

Ovi profil fotoğrafı
Ovi2 ay önce

I'm pretty new to LLM vulns. But I'm curious why there would be differences between models what is guardrailed?

vladimir metnew profil fotoğrafı
vladimir metnew2 ay önce

uh, it's just opus is smarter and can detect the prompt injection and reject to execute the instructions that are necessary to trigger sbx. there's a wide range of bypasses like starting subagents or encoding text, but they're all shitty. unless you have a universal jailbreak

Ovi profil fotoğrafı
Ovi2 ay önce

It's just weird to me how they would differ between models. More recent models like Fable seem to just blanked ban most things ever related to something malicious

vladimir metnew profil fotoğrafı
vladimir metnew2 ay önce

try it yourself I guess

NorthSecureAI profil fotoğrafı
NorthSecureAI3 ay önce

Strong writeup. This is the reminder that prompt injection is not a content problem once the agent can touch the host. Isolation, least privilege, and explicit trust boundaries still matter more than confidence.

Bali as a Colony of Jakarta profil fotoğrafı
Bali as a Colony of Jakarta3 ay önce

Waht for is this?

Benzer Videolar