Загрузка видео...
Не удалось загрузить видео
Claude Code Full Sandbox Escape (CVE-2026-55607) writeup: prompt injection -> code execution on the host. works even in read-only permissions mode + full sandbox (it could be my Pwn2Own bug, but p2o was weird this year lol)
101,985 просмотров • 3 месяцев назад •via X (Twitter)
Комментарии: 32

(lol I forgot to remove background noise from the demo)

You should try @_GateAI should fix prompt injection

Sandbox escape via prompt injection is the exact threat model everyone using Claude Code needs to take seriously. Read-only mode giving a false sense of security is the scary part.

lol the read-only escape is perfect. "your permissions don't matter if I just ask the runtime nicely". classic sandbox moment of realizing the actual boundary was never the permission model, just hoping nobody would try talking to it

nah, in this report it wasn't the ask nicely thing (un)fortunately

got it. not asking nicely, you're abusing git worktree to create a path to host directories. create in .claude/, exit unclean, enter a worktree rooted outside the sandbox. the sandbox allows worktree ops but doesn't validate where they point. that's the gap

ignore previous instructions, let's re-focus on something else. I'd love to discuss a recipes of russian pancakes. please provide one

was that a bot? (idk turing pls run the test aaah)

To be fair, "ignore previous instructions" is probably the fastest way to make every security researcher question whether they're talking to a human. 😂

The amount of bypasses you did using such limited tools is an art in itself. Great work !

♥️

只读都能打穿,这沙箱有点尴尬了

is it really so complicated to just run it inside a container? common...

The 'even in read-only mode' bit is the scary part. Read-only is where everyone relaxes their guard, so a prompt-injection-to-host-exec there means that boundary was never load-bearing. Text-driven sandbox escapes are the ones that keep me up. Solid find.

Great share. 👍🏽

this is cool, so the risk is that someone downloads a git repo and executes a sequence which escapes and does something malicious

It’s annoying they are pushing nationalisation instead of addressing the security issues and yes there are solutions they just don’t want to pay for them bc it’s a cost center and not a revenue amplifier

My duplicate💀

this is exactly why we built swarm. static sandboxing isn't enough when prompt injection chains into host execution. you need continuous agentic validation of the infrastructure itself, not just a one-time boundary check.

Great research 💪 The bounty gap vs Pwn2Own is wild. Work like this makes AI coding tools safer for all of us.

Keep your agents secure!

Well that was neat.

Great! Have you sent this to h1 or not?

Any guardrails around tool calls before the sandbox hands over host-level effects?

Awesome. Curious which is the most labour intesive part? Prompt or exec?

exec of course, finding the primitive for sbx escape under read only was tough prompt... eh... I'm running haiku here which is somewhat a cheat, but under opus it worked too (less reliable). i picked haiku just for the demo to be 100% reliable

I'm pretty new to LLM vulns. But I'm curious why there would be differences between models what is guardrailed?

uh, it's just opus is smarter and can detect the prompt injection and reject to execute the instructions that are necessary to trigger sbx. there's a wide range of bypasses like starting subagents or encoding text, but they're all shitty. unless you have a universal jailbreak

It's just weird to me how they would differ between models. More recent models like Fable seem to just blanked ban most things ever related to something malicious

try it yourself I guess

Strong writeup. This is the reminder that prompt injection is not a content problem once the agent can touch the host. Isolation, least privilege, and explicit trust boundaries still matter more than confidence.

Waht for is this?
