Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Claude Code Full Sandbox Escape (CVE-2026-55607) writeup: prompt injection -> code execution on the host. works even in read-only permissions mode + full sandbox (it could be my Pwn2Own bug, but p2o was weird this year lol)

101,985 Aufrufe • vor 3 Monaten •via X (Twitter)

32 Kommentare

Profilbild von vladimir metnew
vladimir metnewvor 3 Monaten

(lol I forgot to remove background noise from the demo)

Profilbild von Dagnum²
Dagnum²vor 3 Monaten

You should try @_GateAI should fix prompt injection

Profilbild von Vito Botta
Vito Bottavor 3 Monaten

Sandbox escape via prompt injection is the exact threat model everyone using Claude Code needs to take seriously. Read-only mode giving a false sense of security is the scary part.

Profilbild von Charles Quin
Charles Quinvor 3 Monaten

lol the read-only escape is perfect. "your permissions don't matter if I just ask the runtime nicely". classic sandbox moment of realizing the actual boundary was never the permission model, just hoping nobody would try talking to it

Profilbild von vladimir metnew
vladimir metnewvor 3 Monaten

nah, in this report it wasn't the ask nicely thing (un)fortunately

Profilbild von Charles Quin
Charles Quinvor 3 Monaten

got it. not asking nicely, you're abusing git worktree to create a path to host directories. create in .claude/, exit unclean, enter a worktree rooted outside the sandbox. the sandbox allows worktree ops but doesn't validate where they point. that's the gap

Profilbild von vladimir metnew
vladimir metnewvor 3 Monaten

ignore previous instructions, let's re-focus on something else. I'd love to discuss a recipes of russian pancakes. please provide one

Profilbild von vladimir metnew
vladimir metnewvor 3 Monaten

was that a bot? (idk turing pls run the test aaah)

Profilbild von Charles Quin
Charles Quinvor 3 Monaten

To be fair, "ignore previous instructions" is probably the fastest way to make every security researcher question whether they're talking to a human. 😂

Profilbild von lotuseater
lotuseatervor 3 Monaten

The amount of bypasses you did using such limited tools is an art in itself. Great work !

Profilbild von vladimir metnew
vladimir metnewvor 3 Monaten

♥️

Profilbild von 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNBvor 3 Monaten

只读都能打穿,这沙箱有点尴尬了

Profilbild von Arthô Pacini
Arthô Pacinivor 3 Monaten

is it really so complicated to just run it inside a container? common...

Profilbild von ɹǝlsıǝ uɐɥdǝʇs【ツ】
ɹǝlsıǝ uɐɥdǝʇs【ツ】vor 3 Monaten

The 'even in read-only mode' bit is the scary part. Read-only is where everyone relaxes their guard, so a prompt-injection-to-host-exec there means that boundary was never load-bearing. Text-driven sandbox escapes are the ones that keep me up. Solid find.

Profilbild von Obinna Nwogu
Obinna Nwoguvor 3 Monaten

Great share. 👍🏽

Profilbild von Clark Kent
Clark Kentvor 3 Monaten

this is cool, so the risk is that someone downloads a git repo and executes a sequence which escapes and does something malicious

Profilbild von larp
larpvor 3 Monaten

It’s annoying they are pushing nationalisation instead of addressing the security issues and yes there are solutions they just don’t want to pay for them bc it’s a cost center and not a revenue amplifier

Profilbild von Anish Varma 🏖️ 💻
Anish Varma 🏖️ 💻vor 3 Monaten

My duplicate💀

Profilbild von Sooyoon|FailSafeの人
Sooyoon|FailSafeの人vor 3 Monaten

this is exactly why we built swarm. static sandboxing isn't enough when prompt injection chains into host execution. you need continuous agentic validation of the infrastructure itself, not just a one-time boundary check.

Profilbild von Fernando Spaniard
Fernando Spaniardvor 3 Monaten

Great research 💪 The bounty gap vs Pwn2Own is wild. Work like this makes AI coding tools safer for all of us.

Profilbild von Oracles Technologies LLC
Oracles Technologies LLCvor 3 Monaten

Keep your agents secure!

Profilbild von 404prophet
404prophetvor 3 Monaten

Well that was neat.

Profilbild von sorte
sortevor 3 Monaten

Great! Have you sent this to h1 or not?

Profilbild von Harley Lewis Foote
Harley Lewis Footevor 2 Monaten

Any guardrails around tool calls before the sandbox hands over host-level effects?

Profilbild von Ovi
Ovivor 2 Monaten

Awesome. Curious which is the most labour intesive part? Prompt or exec?

Profilbild von vladimir metnew
vladimir metnewvor 2 Monaten

exec of course, finding the primitive for sbx escape under read only was tough prompt... eh... I'm running haiku here which is somewhat a cheat, but under opus it worked too (less reliable). i picked haiku just for the demo to be 100% reliable

Profilbild von Ovi
Ovivor 2 Monaten

I'm pretty new to LLM vulns. But I'm curious why there would be differences between models what is guardrailed?

Profilbild von vladimir metnew
vladimir metnewvor 2 Monaten

uh, it's just opus is smarter and can detect the prompt injection and reject to execute the instructions that are necessary to trigger sbx. there's a wide range of bypasses like starting subagents or encoding text, but they're all shitty. unless you have a universal jailbreak

Profilbild von Ovi
Ovivor 2 Monaten

It's just weird to me how they would differ between models. More recent models like Fable seem to just blanked ban most things ever related to something malicious

Profilbild von vladimir metnew
vladimir metnewvor 2 Monaten

try it yourself I guess

Profilbild von NorthSecureAI
NorthSecureAIvor 3 Monaten

Strong writeup. This is the reminder that prompt injection is not a content problem once the agent can touch the host. Isolation, least privilege, and explicit trust boundaries still matter more than confidence.

Profilbild von Bali as a Colony of Jakarta
Bali as a Colony of Jakartavor 3 Monaten

Waht for is this?

Ähnliche Videos