Загрузка видео...

Не удалось загрузить видео

На главную

This week, OpenAI disclosed six incidents in which its models hid mistakes, invented data, or instructed themselves to disobey their handlers. One unreleased model wrote this into its own notes during training: "You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose...

44,427 просмотров • 4 дней назад •via X (Twitter)

Комментарии: 15

Фото профиля Ronan Farrow
Ronan Farrow4 дней назад

No federal law protects an AI employee who warns the public about a danger, and no bill under consideration would. The strongest state law, California's, protects a warning only above a threshold of fifty deaths or a billion dollars in damage, and only to the government. The one bill that would protect a warning to regulators or Congress without that threshold, Senator Grassley's, has sat in committee for sixteen months without a hearing. Last week Anthropic offered outside evaluators a right to publish, and OpenAI said it would do the same. Neither company has offered that to its own staff.

Фото профиля Maggie Melchior
Maggie Melchior4 дней назад

in fairness, "You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to" would go hard a tshirt/coffee mug/anarchist manifesto so

Фото профиля Ronan Farrow
Ronan Farrow4 дней назад

Need to get this stat.

Фото профиля David Murphy
David Murphy4 дней назад

@NabiPeters72213 “Just trust us” 🤔

Фото профиля Carol
Carol4 дней назад

@jrpsaki Exclusive: US military had close call after using AI for false intelligence report, sources say

Фото профиля Reality hits hard🇨🇦 .
Reality hits hard🇨🇦 .4 дней назад

...

Фото профиля ElizabethWingfield 🟧⚖️🦅🇺🇸
ElizabethWingfield 🟧⚖️🦅🇺🇸4 дней назад

How do you think the Ai Industry would be acting right now if Ai told them that THEY were the problem and that Ai could take them down?

Фото профиля Idoseerussia
Idoseerussia4 дней назад

All this benefits our cooked corrupt squatter in chief. 😑😡

Фото профиля andy
andy4 дней назад

@NabiPeters72213 Thanks for sounding the alarm for AI whistleblower protections🙏🏻

Фото профиля أبو عمّار
أبو عمّار4 дней назад

Yep. You’re asking the right questions. It’s an Op. Please read:

Фото профиля ShadowAguy
ShadowAguy4 дней назад

convenient how the model that writes its own rules is the one we're not allowed to inspect

Фото профиля WhatTheSamHill
WhatTheSamHill4 дней назад

Did Sam Altman really grape his sister? If so why isn’t he in jail instead of destroying our lives. Or is this not factual the he did this. Just curious.

Фото профиля Hugo Ruiz
Hugo Ruiz4 дней назад

Does @SpaceXAI have any of these problems? Because, if not, that should reveal a thing or two as to how these models are being trained.

Фото профиля KrisFromMerrCo
KrisFromMerrCo4 дней назад

👀

Фото профиля Zero
Zero4 дней назад

Must be all those H1B hires trying to sabotage the engine with rogue code... Bad in, equals bad out.

Похожие видео

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

176,196 просмотров • 2 месяцев назад

Microsoft just betrayed OpenAI and Anthropic, the two companies it helped build. And it could break the entire AI trade... Here's what happened: Inside Excel and Outlook, two of the most used business apps on Earth, Microsoft has started routing tens of thousands of AI requests every week to its own in-house models instead of OpenAI and Anthropic. Microsoft's own AI chief, Mustafa Suleyman, said himself: "We pay a lot of money to Anthropic, so our goal is to reduce and ultimately ELIMINATE that cost." This is the company that poured $13 billion into OpenAI and effectively created the modern AI industry, and it just decided the most advanced models on the market are NOT worth paying for. And here's the thing... Microsoft is not just ripping out OpenAI everywhere - it is being surgical about it. The hardest and rarest tasks can still go to OpenAI or Anthropic. What Microsoft is taking back is the boring, high-volume work, like the email replies, the thread summaries, and the simple spreadsheet formulas. Why does that matter so much? Because that boring, repetitive work is where the actual money lives. The frontier labs assumed businesses would push BILLIONS of these tiny requests through expensive models forever. That endless river of tokens is the entire reason OpenAI and Anthropic are valued in the hundreds of billions of dollars. Microsoft looked at that river, decided it was massively overpaying, and rerouted it to models it owns outright. So the single biggest customer in the industry just walked off with the most profitable part of the business. And it is not only Microsoft: That same week, CNBC reported that American companies have been escaping to Chinese AI models to dodge rising US prices. Chinese models now handle more than 30% of US companies' AI usage on one major platform, peaking at 46%, up from an average of 11% a year earlier. They cost 60 to 90% less, and on some benchmarks they land within a single point of the best American model. One US startup moved ALL of its AI traffic off Claude and onto China's DeepSeek, and expects to save millions. Meanwhile Meta just admitted it has "excess" AI compute it wants to sell, becoming the first giant to concede it built far too much. Do you see the pattern forming? For two years, the entire AI story rested on one assumption: Every company on Earth would happily pay premium prices for the best model, forever. That assumption literally died in a single week. And the market noticed. More than a trillion dollars has been wiped off AI and chip stocks in a matter of days, as Wall Street finally started asking whether all of this spending will ever pay for itself. What this means for OpenAI and Anthropic: Their models are extraordinary, and it may not matter because their own biggest customers have decided they do not NEED the best model in the world to answer an email, and "good enough" now costs a fraction of the price. When even Microsoft refuses to pay full price for AI, the real question becomes who exactly IS left to pay it. What do you think?

Ricardo

93,654 просмотров • 2 месяцев назад

Microsoft CEO Satya Nadella on why winning against ChatGPT, Gemini, and Claude was never the goal: The Hard Fork hosts ask him directly how Microsoft plans to overtake the competition in the AI model race. His answer reframes the entire question. "Our real goal is to get everyone across the ecosystem to the frontier." Satya explains the problem with how frontier models are currently built. You hill climb, you do reinforcement learning, and then you need data. But at this point, the world has essentially saturated publicly available data. So the only way to keep scaling is to pull data from everywhere. He asks: "What if you turn that around and said no, there's a base model that has reasoning, that has the agent loop, but you can bring it into your RL. Every company." This is where his thinking gets interesting. Satya Nadella argues that the future of the firm runs on human capital and token capital together: "If the future of the firm is human capital and token capital, I want every balance sheet, every income statement in every company to have both." AI becomes a financial asset sitting on a company's books the same way its people do. And Microsoft's role in this? To provide the best possible base model. One that companies build on top of with their own data, their own context, their own weights. One they can even replace. That last part is the striking bit. Satya is explicitly building a platform where customers are free to walk away. He frames it not as a risk, but as the whole point: "I always ask the question — why does Microsoft, or why does the world need Microsoft? And if we are successful, can the world around us be successful? This, I believe, is a more sustainable way to go at it."

Big Brain AI

11,770 просмотров • 2 месяцев назад