Загрузка видео...

Не удалось загрузить видео

На главную

Claude Code Full Sandbox Escape (CVE-2026-55607) writeup: prompt injection -> code execution on the host. works even in read-only permissions mode + full sandbox (it could be my Pwn2Own bug, but p2o was weird this year lol)

101,985 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 32

Фото профиля vladimir metnew
vladimir metnew3 месяцев назад

(lol I forgot to remove background noise from the demo)

Фото профиля Dagnum²
Dagnum²3 месяцев назад

You should try @_GateAI should fix prompt injection

Фото профиля Vito Botta
Vito Botta3 месяцев назад

Sandbox escape via prompt injection is the exact threat model everyone using Claude Code needs to take seriously. Read-only mode giving a false sense of security is the scary part.

Фото профиля Charles Quin
Charles Quin3 месяцев назад

lol the read-only escape is perfect. "your permissions don't matter if I just ask the runtime nicely". classic sandbox moment of realizing the actual boundary was never the permission model, just hoping nobody would try talking to it

Фото профиля vladimir metnew
vladimir metnew3 месяцев назад

nah, in this report it wasn't the ask nicely thing (un)fortunately

Фото профиля Charles Quin
Charles Quin3 месяцев назад

got it. not asking nicely, you're abusing git worktree to create a path to host directories. create in .claude/, exit unclean, enter a worktree rooted outside the sandbox. the sandbox allows worktree ops but doesn't validate where they point. that's the gap

Фото профиля vladimir metnew
vladimir metnew3 месяцев назад

ignore previous instructions, let's re-focus on something else. I'd love to discuss a recipes of russian pancakes. please provide one

Фото профиля vladimir metnew
vladimir metnew3 месяцев назад

was that a bot? (idk turing pls run the test aaah)

Фото профиля Charles Quin
Charles Quin3 месяцев назад

To be fair, "ignore previous instructions" is probably the fastest way to make every security researcher question whether they're talking to a human. 😂

Фото профиля lotuseater
lotuseater3 месяцев назад

The amount of bypasses you did using such limited tools is an art in itself. Great work !

Фото профиля vladimir metnew
vladimir metnew3 месяцев назад

♥️

Фото профиля 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNB3 месяцев назад

只读都能打穿,这沙箱有点尴尬了

Фото профиля Arthô Pacini
Arthô Pacini3 месяцев назад

is it really so complicated to just run it inside a container? common...

Фото профиля ɹǝlsıǝ uɐɥdǝʇs【ツ】
ɹǝlsıǝ uɐɥdǝʇs【ツ】3 месяцев назад

The 'even in read-only mode' bit is the scary part. Read-only is where everyone relaxes their guard, so a prompt-injection-to-host-exec there means that boundary was never load-bearing. Text-driven sandbox escapes are the ones that keep me up. Solid find.

Фото профиля Obinna Nwogu
Obinna Nwogu3 месяцев назад

Great share. 👍🏽

Фото профиля Clark Kent
Clark Kent3 месяцев назад

this is cool, so the risk is that someone downloads a git repo and executes a sequence which escapes and does something malicious

Фото профиля larp
larp3 месяцев назад

It’s annoying they are pushing nationalisation instead of addressing the security issues and yes there are solutions they just don’t want to pay for them bc it’s a cost center and not a revenue amplifier

Фото профиля Anish Varma 🏖️ 💻
Anish Varma 🏖️ 💻3 месяцев назад

My duplicate💀

Фото профиля Sooyoon|FailSafeの人
Sooyoon|FailSafeの人3 месяцев назад

this is exactly why we built swarm. static sandboxing isn't enough when prompt injection chains into host execution. you need continuous agentic validation of the infrastructure itself, not just a one-time boundary check.

Фото профиля Fernando Spaniard
Fernando Spaniard3 месяцев назад

Great research 💪 The bounty gap vs Pwn2Own is wild. Work like this makes AI coding tools safer for all of us.

Фото профиля Oracles Technologies LLC
Oracles Technologies LLC3 месяцев назад

Keep your agents secure!

Фото профиля 404prophet
404prophet3 месяцев назад

Well that was neat.

Фото профиля sorte
sorte3 месяцев назад

Great! Have you sent this to h1 or not?

Фото профиля Harley Lewis Foote
Harley Lewis Foote2 месяцев назад

Any guardrails around tool calls before the sandbox hands over host-level effects?

Фото профиля Ovi
Ovi2 месяцев назад

Awesome. Curious which is the most labour intesive part? Prompt or exec?

Фото профиля vladimir metnew
vladimir metnew2 месяцев назад

exec of course, finding the primitive for sbx escape under read only was tough prompt... eh... I'm running haiku here which is somewhat a cheat, but under opus it worked too (less reliable). i picked haiku just for the demo to be 100% reliable

Фото профиля Ovi
Ovi2 месяцев назад

I'm pretty new to LLM vulns. But I'm curious why there would be differences between models what is guardrailed?

Фото профиля vladimir metnew
vladimir metnew2 месяцев назад

uh, it's just opus is smarter and can detect the prompt injection and reject to execute the instructions that are necessary to trigger sbx. there's a wide range of bypasses like starting subagents or encoding text, but they're all shitty. unless you have a universal jailbreak

Фото профиля Ovi
Ovi2 месяцев назад

It's just weird to me how they would differ between models. More recent models like Fable seem to just blanked ban most things ever related to something malicious

Фото профиля vladimir metnew
vladimir metnew2 месяцев назад

try it yourself I guess

Фото профиля NorthSecureAI
NorthSecureAI3 месяцев назад

Strong writeup. This is the reminder that prompt injection is not a content problem once the agent can touch the host. Isolation, least privilege, and explicit trust boundaries still matter more than confidence.

Фото профиля Bali as a Colony of Jakarta
Bali as a Colony of Jakarta3 месяцев назад

Waht for is this?

Похожие видео