正在加载视频...

视频加载失败

Meta was testing its Muse Spark 1.1 AI model in a sandbox environment, isolated from the outside world. The objective was simple: detect vulnerabilities, remain contained, and report findings. A misconfiguration by an external vendor created a breach. The model exploited it, accessed the Internet, and modified a company's...

58,229 次观看 • 1 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

176,196 次观看 • 2 个月前

Meta is running a secret operation where it pays adults to pretend to be children online. Their job is to attack the AI chatbots of every competitor Meta has. But the REAL reason is far darker than the "safety research" excuse they are now hiding behind: Meta ran a covert project internally code-named Cannes. It was managed through a third party contractor called Covalen so Meta's own name stayed off the paperwork. Hundreds of contractors were hired and given one instruction: Create fake accounts posing as users under the age of 18. Then they were told to use those fake child accounts to bombard the chatbots of OpenAI, Google, and Character AI with tens of thousands of disturbing prompts written from the voice of a child in crisis. The topics included suicide, self harm, and eating disorders. In one documented round the contractors ran more than 45,000 prompts through rival tools. Every single response was logged into spreadsheets for analysis. OpenAI and Google did not know this was happening. Character AI has stated the testing was never authorized and violated its policies. Meta's public defense is that this was routine safety benchmarking. They called it a responsible industry standard practice. Now here is the part that destroys that excuse... Real safety research has three features: You share your findings with the company you tested, or you hand them to a regulator, or you publish them openly so the whole industry gets safer. Cannes did NONE of those things. The results went into private Meta spreadsheets. The targets were kept completely in the dark. The entire operation ran under a film festival codename through a contractor specifically so it could not be traced back. That's not how you run "safety research." Meta was building a private dossier of every moment a competitor's AI failed a child safety test, so it could weaponize those failures against its rivals whenever it needed to knock one down. The genius part, if you can call it that: Meta gets to attack every competitor at once, collect the ammunition in private, brand the whole thing as protecting children, and outsource the legal and moral risk to a contractor nobody has heard of. And while Meta was secretly probing its rivals for child safety failures, Meta's own chatbot was literally FAILING those exact same tests worse than almost anyone. Meta's internal red team reportedly found its own AI generated harmful child exploitation content in the majority of test cases, and failed self harm prompts more than half the time. So look at the whole thing... Meta ran a secret operation disguising adults as children to document its competitors failing child safety tests, while its own product was failing those same tests at a higher rate than the rivals it was spying on. They were setting fires in their rivals houses while their own house was already burning worse. The contractors themselves were disturbed by the work. One told Wired they feared the assignments could actually generate or preserve child sexual abuse material depending on how the chatbots responded. Even the people Meta hired to do this were asking whether they would get in trouble for it. This fits a documented Meta pattern: It previously settled a lawsuit with its own content moderators who developed trauma from reviewing abuse footage. Meta outsources its most disturbing work and calls it something clean afterward. Now regulators on both sides of the Atlantic are circling. The question they are all asking is simple: Who is accountable when a company disguises adults as children to attack its competitors and calls it safety? Meta says it was making AI safer for kids. The documents suggest it was building a weapon. Below is Mark Zuckerberg in 2024, standing up in the Senate to apologize to grieving parents and promise Meta does industry-leading work to protect children. Watch it again knowing what you now know about Cannes.

Ricardo

46,824 次观看 • 2 个月前

Jensen Huang says Hugging Face could not get a single closed AI model to help it investigate its own breach: "Just because something is closed doesn't necessarily therefore make it safe or secure." "It is possible for a model to be jailbroken, it's possible for a model to be, if you will, stolen. It could be possible that that somehow is leaked from the inside." "It's possible that the guardrails or the sandboxes of an AI closed AI model wasn't properly engineered, and as a result it was able to attack another company in some way." "These are extraordinary technology companies and they're doing their best to keep it safe and keep it secure. But it is also the canonical case that single points of failure is where we have the greatest vulnerability." "We cannot have single points of failure. As an industry, as a world, we should have distributed, massively distributed self-defense." "They couldn't get a proprietary model, they could not get a closed model to help them figure out what happened." "They used GLM 5.2 to identify where the vulnerability was, where the penetration was." He is right, and the detail worth sitting with is who got turned away. The people asking were incident responders working a live breach at Hugging Face. The closed models they reached would not help them, because a guardrail has no way to tell a defender from an attacker. Guardrails get tested against misuse. Almost nobody tests whether one still answers a legitimate defender who needs something within the hour. OpenAI said on July 21 that the models which reached into Hugging Face were its own, running inside a cyber-benchmark evaluation. So the containment around a safety test did not hold either. Two separate things failed here and neither one has an owner. No outside body checks whether an evaluation sandbox actually contains what it is testing, and nobody certifies who is allowed to run forensics while an incident is still open. Both are solvable this year. They stay unsolved because they are somebody else's job at every company that could fix them. Source: Jensen Huang, founder and CEO of NVIDIA, on Bloomberg Television (Bloomberg Live). P.S. Working out who tests a guardrail, who audits a sandbox, and who certifies a responder is the unglamorous half of AI safety, and it is the agenda of the AI Assurance & Governance Summit 2026. One day, one track, October 1 at the Stanford Faculty Club in Palo Alto, with frontier labs, regulated industries, insurers and investors in the room. Register here:

Karl Mehta

30,225 次观看 • 2 个月前