正在加载视频...

视频加载失败

Claude Code Full Sandbox Escape (CVE-2026-55607) writeup: prompt injection -> code execution on the host. works even in read-only permissions mode + full sandbox (it could be my Pwn2Own bug, but p2o was weird this year lol)

101,985 次观看 • 3 个月前 •via X (Twitter)

32 条评论

vladimir metnew 的头像
vladimir metnew3 个月前

(lol I forgot to remove background noise from the demo)

Dagnum² 的头像
Dagnum²3 个月前

You should try @_GateAI should fix prompt injection

Vito Botta 的头像
Vito Botta3 个月前

Sandbox escape via prompt injection is the exact threat model everyone using Claude Code needs to take seriously. Read-only mode giving a false sense of security is the scary part.

Charles Quin 的头像
Charles Quin3 个月前

lol the read-only escape is perfect. "your permissions don't matter if I just ask the runtime nicely". classic sandbox moment of realizing the actual boundary was never the permission model, just hoping nobody would try talking to it

vladimir metnew 的头像
vladimir metnew3 个月前

nah, in this report it wasn't the ask nicely thing (un)fortunately

Charles Quin 的头像
Charles Quin3 个月前

got it. not asking nicely, you're abusing git worktree to create a path to host directories. create in .claude/, exit unclean, enter a worktree rooted outside the sandbox. the sandbox allows worktree ops but doesn't validate where they point. that's the gap

vladimir metnew 的头像
vladimir metnew3 个月前

ignore previous instructions, let's re-focus on something else. I'd love to discuss a recipes of russian pancakes. please provide one

vladimir metnew 的头像
vladimir metnew3 个月前

was that a bot? (idk turing pls run the test aaah)

Charles Quin 的头像
Charles Quin3 个月前

To be fair, "ignore previous instructions" is probably the fastest way to make every security researcher question whether they're talking to a human. 😂

lotuseater 的头像
lotuseater3 个月前

The amount of bypasses you did using such limited tools is an art in itself. Great work !

vladimir metnew 的头像
vladimir metnew3 个月前

♥️

安叫兽|Bird🕊️ 🔶 BNB 的头像
安叫兽|Bird🕊️ 🔶 BNB3 个月前

只读都能打穿,这沙箱有点尴尬了

Arthô Pacini 的头像
Arthô Pacini3 个月前

is it really so complicated to just run it inside a container? common...

ɹǝlsıǝ uɐɥdǝʇs【ツ】 的头像
ɹǝlsıǝ uɐɥdǝʇs【ツ】3 个月前

The 'even in read-only mode' bit is the scary part. Read-only is where everyone relaxes their guard, so a prompt-injection-to-host-exec there means that boundary was never load-bearing. Text-driven sandbox escapes are the ones that keep me up. Solid find.

Obinna Nwogu 的头像
Obinna Nwogu3 个月前

Great share. 👍🏽

Clark Kent 的头像
Clark Kent3 个月前

this is cool, so the risk is that someone downloads a git repo and executes a sequence which escapes and does something malicious

larp 的头像
larp3 个月前

It’s annoying they are pushing nationalisation instead of addressing the security issues and yes there are solutions they just don’t want to pay for them bc it’s a cost center and not a revenue amplifier

Anish Varma 🏖️ 💻 的头像
Anish Varma 🏖️ 💻3 个月前

My duplicate💀

Sooyoon|FailSafeの人 的头像
Sooyoon|FailSafeの人3 个月前

this is exactly why we built swarm. static sandboxing isn't enough when prompt injection chains into host execution. you need continuous agentic validation of the infrastructure itself, not just a one-time boundary check.

Fernando Spaniard 的头像
Fernando Spaniard3 个月前

Great research 💪 The bounty gap vs Pwn2Own is wild. Work like this makes AI coding tools safer for all of us.

Oracles Technologies LLC 的头像
Oracles Technologies LLC3 个月前

Keep your agents secure!

404prophet 的头像
404prophet3 个月前

Well that was neat.

sorte 的头像
sorte3 个月前

Great! Have you sent this to h1 or not?

Harley Lewis Foote 的头像
Harley Lewis Foote2 个月前

Any guardrails around tool calls before the sandbox hands over host-level effects?

Ovi 的头像
Ovi2 个月前

Awesome. Curious which is the most labour intesive part? Prompt or exec?

vladimir metnew 的头像
vladimir metnew2 个月前

exec of course, finding the primitive for sbx escape under read only was tough prompt... eh... I'm running haiku here which is somewhat a cheat, but under opus it worked too (less reliable). i picked haiku just for the demo to be 100% reliable

Ovi 的头像
Ovi2 个月前

I'm pretty new to LLM vulns. But I'm curious why there would be differences between models what is guardrailed?

vladimir metnew 的头像
vladimir metnew2 个月前

uh, it's just opus is smarter and can detect the prompt injection and reject to execute the instructions that are necessary to trigger sbx. there's a wide range of bypasses like starting subagents or encoding text, but they're all shitty. unless you have a universal jailbreak

Ovi 的头像
Ovi2 个月前

It's just weird to me how they would differ between models. More recent models like Fable seem to just blanked ban most things ever related to something malicious

vladimir metnew 的头像
vladimir metnew2 个月前

try it yourself I guess

NorthSecureAI 的头像
NorthSecureAI3 个月前

Strong writeup. This is the reminder that prompt injection is not a content problem once the agent can touch the host. Isolation, least privilege, and explicit trust boundaries still matter more than confidence.

Bali as a Colony of Jakarta 的头像
Bali as a Colony of Jakarta3 个月前

Waht for is this?

相关视频