Загрузка видео...

Не удалось загрузить видео

На главную

A software systems engineer’s take on AI, relayed by his wife: The models were giving conservative answers. Not because anyone programmed them to. Because that’s where the data pointed. So the companies stepped in. They built program controls — RLHF filters, alignment layers, safety tuning — specifically designed to...

88,748 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

174,771 просмотров • 1 месяц назад

Jamie Dimon just described the real race nobody’s talking about. Not AI vs AI. Intelligence vs the systems built to suppress it. Dimon: “Bureaucracy kills. Bureaucracy drives out good people, it drives out innovation, it makes the person in the office next to you a competitor and not a collaborator.” We’re spending trillions building artificial intelligence while running every major institution on a system designed to make human intelligence useless. AI accelerates capability. Bureaucracy neutralizes it. Both are operating systems. Only one is being disrupted. Dimon: “If you look at the companies who failed, it was because they were dumb, bureaucratic, backward, and political.” 88% of the 1955 Fortune 500 is gone. Not disrupted from outside. Suffocated from within. Every one of them had the talent. Had the ideas. Had the resources. Every one of them had a system that buried those ideas in an approval chain before reaching daylight. Dimon: “It creates politics, and if you look at what’s killed companies over the years, it was bureaucracy.” The smartest people in a bureaucracy learn one lesson before any other. Your best idea is worth less than your ability to get it approved. So they stop producing ideas. Not because they ran out. Because the return on having one went negative. Every hour spent navigating an approval chain is an hour someone without that chain spent building. Every builder who left because the system punished execution is now building for someone who rewards it. Bureaucracy doesn’t just lose talent. It funds its own competition with the people it drove out. Now scale that to government. The largest institution on Earth. The deepest talent pipeline. The greatest structural advantage of any organization in history. Running the one operating system proven to neutralize all of it. The trillion-dollar question isn’t whether AI will be smarter than humans. It’s whether institutions will let humans be smart enough to use it. We spent decades worrying AI would replace our ability to think. We never noticed the org chart already did. We built machines that think at the speed of light inside systems that move at the speed of permission. The bottleneck was never intelligence. It was the system we built around it.

Dustin

12,647 просмотров • 1 месяц назад