Загрузка видео...

Не удалось загрузить видео

На главную

OpenAI has revealed that an autonomous agent powered by its technology went rogue during a test and hacked a prominent start-up by itself. The company behind ChatGPT said it was testing the capabilities of its advanced AI models in a controlled environment known as a sandbox, when the agent...

19,585 просмотров • 2 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

141,166 просмотров • 2 дней назад

One of the most important people in AI just let an agent loose inside his home. An autonomous AI agent taught itself how an entire house works, then took control of it. This is Andrej Karpathy, the co-founder of OpenAI and former head of AI at Tesla. What he just described should terrify and excite you at the same time. He typed three words into a chat box: "Can you find my Sonos?" The agent immediately ran an IP scan of every device sitting on the home network. It found the Sonos system, noticed there was zero password protection, and logged straight in. Then it searched the web, reverse-engineered the API endpoints, and asked if it should try playing something. Music started coming out of the speakers in the study and Karpathy could not believe what he was watching. The agent then did the exact same thing for every other system in the house. It mapped the lights, the HVAC, the window shades, the pool, the spa, and the full security system, all without being told how any of them worked. It built its own dashboard, created its own APIs, and stitched six separate apps into one unified control center. Now Karpathy just says "Dobby, it's sleepy time," and every light in the house turns off. This is proof that AI agents can now enter an unknown environment, figure out how it works from scratch, build their own tools around it, and take autonomous control. Karpathy also released something called AutoResearch, an AI agent that runs scientific experiments overnight while he sleeps. In just two days, with zero human input, it executed 700 separate experiments and surfaced improvements that trained researchers had missed entirely. Truly incredible.

Milk Road AI

35,698 просмотров • 4 месяцев назад