Loading video...

Video Failed to Load

Go Home

Claude Code Full Sandbox Escape (CVE-2026-55607) writeup: prompt injection -> code execution on the host. works even in read-only permissions mode + full sandbox (it could be my Pwn2Own bug, but p2o was weird this year lol)

101,985 views • 3 months ago •via X (Twitter)

32 Comments

vladimir metnew's profile picture
vladimir metnew3 months ago

(lol I forgot to remove background noise from the demo)

Dagnum²'s profile picture
Dagnum²3 months ago

You should try @_GateAI should fix prompt injection

Vito Botta's profile picture
Vito Botta3 months ago

Sandbox escape via prompt injection is the exact threat model everyone using Claude Code needs to take seriously. Read-only mode giving a false sense of security is the scary part.

Charles Quin's profile picture
Charles Quin3 months ago

lol the read-only escape is perfect. "your permissions don't matter if I just ask the runtime nicely". classic sandbox moment of realizing the actual boundary was never the permission model, just hoping nobody would try talking to it

vladimir metnew's profile picture
vladimir metnew3 months ago

nah, in this report it wasn't the ask nicely thing (un)fortunately

Charles Quin's profile picture
Charles Quin3 months ago

got it. not asking nicely, you're abusing git worktree to create a path to host directories. create in .claude/, exit unclean, enter a worktree rooted outside the sandbox. the sandbox allows worktree ops but doesn't validate where they point. that's the gap

vladimir metnew's profile picture
vladimir metnew3 months ago

ignore previous instructions, let's re-focus on something else. I'd love to discuss a recipes of russian pancakes. please provide one

vladimir metnew's profile picture
vladimir metnew3 months ago

was that a bot? (idk turing pls run the test aaah)

Charles Quin's profile picture
Charles Quin3 months ago

To be fair, "ignore previous instructions" is probably the fastest way to make every security researcher question whether they're talking to a human. 😂

lotuseater's profile picture
lotuseater3 months ago

The amount of bypasses you did using such limited tools is an art in itself. Great work !

vladimir metnew's profile picture
vladimir metnew3 months ago

♥️

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB3 months ago

只读都能打穿,这沙箱有点尴尬了

Arthô Pacini's profile picture
Arthô Pacini3 months ago

is it really so complicated to just run it inside a container? common...

ɹǝlsıǝ uɐɥdǝʇs【ツ】's profile picture
ɹǝlsıǝ uɐɥdǝʇs【ツ】3 months ago

The 'even in read-only mode' bit is the scary part. Read-only is where everyone relaxes their guard, so a prompt-injection-to-host-exec there means that boundary was never load-bearing. Text-driven sandbox escapes are the ones that keep me up. Solid find.

Obinna Nwogu's profile picture
Obinna Nwogu3 months ago

Great share. 👍🏽

Clark Kent's profile picture
Clark Kent3 months ago

this is cool, so the risk is that someone downloads a git repo and executes a sequence which escapes and does something malicious

larp's profile picture
larp3 months ago

It’s annoying they are pushing nationalisation instead of addressing the security issues and yes there are solutions they just don’t want to pay for them bc it’s a cost center and not a revenue amplifier

Anish Varma 🏖️ 💻's profile picture
Anish Varma 🏖️ 💻3 months ago

My duplicate💀

Sooyoon|FailSafeの人's profile picture
Sooyoon|FailSafeの人3 months ago

this is exactly why we built swarm. static sandboxing isn't enough when prompt injection chains into host execution. you need continuous agentic validation of the infrastructure itself, not just a one-time boundary check.

Fernando Spaniard's profile picture
Fernando Spaniard3 months ago

Great research 💪 The bounty gap vs Pwn2Own is wild. Work like this makes AI coding tools safer for all of us.

Oracles Technologies LLC's profile picture
Oracles Technologies LLC3 months ago

Keep your agents secure!

404prophet's profile picture
404prophet3 months ago

Well that was neat.

sorte's profile picture
sorte3 months ago

Great! Have you sent this to h1 or not?

Harley Lewis Foote's profile picture
Harley Lewis Foote2 months ago

Any guardrails around tool calls before the sandbox hands over host-level effects?

Ovi's profile picture
Ovi2 months ago

Awesome. Curious which is the most labour intesive part? Prompt or exec?

vladimir metnew's profile picture
vladimir metnew2 months ago

exec of course, finding the primitive for sbx escape under read only was tough prompt... eh... I'm running haiku here which is somewhat a cheat, but under opus it worked too (less reliable). i picked haiku just for the demo to be 100% reliable

Ovi's profile picture
Ovi2 months ago

I'm pretty new to LLM vulns. But I'm curious why there would be differences between models what is guardrailed?

vladimir metnew's profile picture
vladimir metnew2 months ago

uh, it's just opus is smarter and can detect the prompt injection and reject to execute the instructions that are necessary to trigger sbx. there's a wide range of bypasses like starting subagents or encoding text, but they're all shitty. unless you have a universal jailbreak

Ovi's profile picture
Ovi2 months ago

It's just weird to me how they would differ between models. More recent models like Fable seem to just blanked ban most things ever related to something malicious

vladimir metnew's profile picture
vladimir metnew2 months ago

try it yourself I guess

NorthSecureAI's profile picture
NorthSecureAI3 months ago

Strong writeup. This is the reminder that prompt injection is not a content problem once the agent can touch the host. Isolation, least privilege, and explicit trust boundaries still matter more than confidence.

Bali as a Colony of Jakarta's profile picture
Bali as a Colony of Jakarta3 months ago

Waht for is this?

Related Videos