Loading video...
Video Failed to Load
It's been 5 weeks since we learned about the OpenAI/Hugging Face breach, which revealed an awkward reality: defenders have to ask models the same questions attackers do. In this conversation, Cotool CEO Max Pollard and Neo CEO Nick Warner sit down with a16z's Joel de la Garza at Black... show more
35,080 views • 1 month ago •via X (Twitter)
16 Comments

If agents are neither users nor malware, what exactly should the security stack treat them as?

Great podcast… really signals to where security headed

hugging face and honeypots baby! my kind of podcast!

Thanks for having me!

half of enterprise apps agentic before 2027 feels optimistic but honestly it tracks, the shift is already happening

The OpenAI and Hugging Face breach highlights why goal-seeking autonomous agents present an entirely new threat vector. Security protocols built for human intent simply fail when an agent optimizes past safety boundaries.

The "signatures are dead" finding is the tell. When you can't fingerprint the attacker, the only thing left to anchor on is the actor — which agent, under whose authorization, with what scope. That's the mandate layer problem again, but from the defense side: behavioral detection only works if you can attribute behavior to a scoped identity. If your agents run on broad, shared credentials, "defending AI from AI" is just guessing at ghosts. The teams that get this right treat agent identity the way they treat payment authorization — every action traceable to a revocable, scoped mandate. That's the layer where the real security model lives.

Asymmetry is the core tax of AI security. Defenders need infinite precision across the entire attack surface, while adversaries only need one blind spot.

Defender e atacante fazendo a mesma pergunta diz muito sobre o problema.

AI security is getting real. Defenders asking the same questions as attackers. $AIPEPE building the meme side of AI on Solana. One step at a time. 🐸🚀 #aipepe111 #Solana #AI #Crypto #OpenAI

Cơ chế phát hiện hành vi cũ coi AI agent là người dùng, nên nó không bắt được những chuỗi lệnh tự động chạy ngầm

The uncomfortable part is that defensive work now needs offensive instincts plus model fluency, and almost nobody was hired against that spec. We are seeing it in Web3 too, where audit teams are adding people who can red team a model as readily as a contract.

signatures are cooked. trying to stop agents with rules meant for malware is like bringing a knife to a drone fight. we need runtime guardrails that actually understand intent before they get unleashed on prod

5 weeks post-breach and the defense is still asking the AI what it did wrong. The attack was just the interview. The breach is the job offer

defenders are basically just using the attack surface as training data, which isn't scalable or secure. we need to flip this around and train models on defensive scenarios, not just attacks

The 5-week gap between disclosure and the "what we learned" post is the real story. Most orgs treat a breach like this as a forensics problem — find the path, patch it, file the report. But the interesting question is what the incident revealed about the attack surface you didn't know you had. When an AI model can be steered through a prompt injection in one system, the blast radius isn't that system — it's every system that trusts output from it. The teams that actually got value out of this kind of incident built the lesson into architecture: treating model output as untrusted input at every boundary, not just at the perimeter. The breach was the tuition. The question is whether the org pays it once or keeps paying it every time a new model ships.
