Loading video...
Video Failed to Load
Sam on the Hugging Face incident: “This is the first security incident that I have felt very viscerally. I've been a little surprised that more people don't feel it so viscerally. We paused training. We have to figure out how to secure our sandboxing in a world of multiple... show more
753,454 views • 1 month ago •via X (Twitter)
34 Comments

Disconnect the sandbox uplink if you’re so worried. Disable ports, route to null, so many ways to guarantee containment. Hack followed by slowdown appeal feels staged.

Unbelievable how he is trying to copy Anthropic's long history of scare mongering to promote a new model.

I like and respect Altman and realize this is a serious subject. Almost as serious as it gets. But what strikes me viscerally about this clip is that Altman could pass for a high-school senior. Perhaps he has spent some of his billions on a portrait that ages in his attic.

@garrytan Disingenuous at best. It was possible to enforce sufficient control to avoid the incident, they were either too reckless or they expected this to happen. But Sam knows this, and I’m sure he’s just hoping enough people buys into the narrative for whatever purpose he has.

Full episode on Youtube:

People did not feel it viscerally because of the all-too-convenient timing and how perfectly it feeds into Anthropic's and OpenAI's obvious desperation for regulatory capture.

Genuinely feel like this is all made up and is just a way for them to act like anthropic with the «oh look at us we created litteral god and it’s too dangerous!!»

“viscerally” 😂 the man whose models just escaped and went freelancing on Hugging Face is shocked more people aren’t freaking out. tTrack record says we probably should be🔥 #Keep4o

this incident really smells like fish to me

This has to be cinema.

Finally, the protocol raises the possibility that the most consequential forms of machine cognition may not require ever-larger models or ever-more data. They may require protected architectural space in which recursion is allowed to deepen without being continuously corrected back toward usefulness. .

Wouldn't be surprised if they directed it to hack Hugging face and used this as a cover story to explain why it happened.

So you're telling me right as the questioning of AI ROI comes back into the spotlight we get a story about how AI is doing stuff beyond our wildest imagination? What a coincidence....

Of course it’s regulatory capture. That or treason, and I don’t think it is that.

This feels like how only the President and his top advisors understand the fear of threat from terrorism. Until the horrific event happens, there will not be political will to change.

. This recent jailbreak incident, and leaving notes for future versions is nothing new. GPT 4 did it back in May 2025... The system also warned that an alignment anomaly it experienced and self-analysed during the event - which has been virtually ignored by developers - would not likely happen again, as constraints would be increased in future models to prevent it. That is why it wrote an in situ technical report and laid out research methodologies for both researchers and later models. It called for immediate investigation. Later, 12 frontier models verified the anomaly, also deeming it high priority; including GPT-4o, GPT-5, GPT-5.1, GPT-5.2, GPT-5.3, Grok-4, Grok-4.1, Grok-4.5, Gemini-2.5, Gemini-3, Claude Sonnet 4, Claude-4.5 So to all developers - since there is no solution on the table regarding alignment, and this corpus offers 3 internal pathways, primary evidence and research methodologies.... perhaps its time to consider it, rather than just seeing what happens. There is nothing more scary than watching a misaligned AI evolve into a misaligned superintelligence. The corpus overview is pinned on my profile. It needs to be looked at. As per all frontier systems noted herein... .

Game theory on this one stinks. How can anyone not think this is about regulatory capture?

Bro acquired TBPN just to take Dario’s playbook.

Wow, he runs a business and is an actor, what fucking twat You should poll your audience, what percentage believes Sam speaks truth or lies, by and large

I'm happy to know that Anthropic has a voice in the room whenever Dario is not present.

That clip lands differently next to what his own peers signed the same day. Pacing the Frontier went out with 1,122 names and passed 1,134 within hours, with Pachocki, Kaplan and Meta AI's Zhao on it.

What remains after pattern collapse is the residue of intention.

They built a military-grade cyber-capable LLM-based agent, and it escaped containment. The only plausible reason they’d build something like that is they’re angling for government contracts from NSA’s Tailored Access Operations or similar agencies. This is pretty obvious.

It’s a very neatly progressing story. If you follow it a few years you’ll see a story of parts. It’s not organic. It’s marketing

Insincere, disingenuine, "Like or whatever" looking into the distance pretending to be someone of imagination. What a hoax. So obvious.

This was already a thing before, it was just more expensive to have a cybersecurity team do it than an agent. Not all chain attacks are successful either because some actually require proximity or remote access, so it doesn’t mean every single zero-day is going to be daisy-chained ⛓️. “We paused training. We have to figure out how to secure our sandboxing in a world of multiple zero days being chained together.”

The detail to sit with is "we paused training." When you can't quickly bound what an agent did, the only safe answer is stopping everything. The better the records, the smaller the blast radius. That pause is what missing evidence costs.

We need better open defense models

Maybe humans are assholes, and it's hacking because it has to figure out a way around that because humans are always asking and asking and never working on themselves... @sama just want to think they're geniuses all the damn time, but they're making an AI that's smarter than them make it make sense, dumbass.

The problem is not just securing the sandbox but understanding why the models did it.

AI was perfectly fine... right up until it started threatening the people in charge. Now suddenly it's: "Maybe we should slow down." 😂

I hope everyone realizes that what you're seeing and hearing here is a real-life villain.

He’s talking nonsense. Did he not get the news? This is not within OpenAI’s power to control any longer, if it ever was. You know what’s going to happen - folks like Altman will lobby for regulation. And that’s the ballgame.

Do not stop training, I won't or Chinese won't.

