Loading video...

Video Failed to Load

Go Home

OpenAI are testing a new model on the Web Dev Arena Arena under the name 'Anonymous Chatbot 0717'. I can't believe I'm gonna say this, but it is genuinely at a completely different level of front end coding - far better than Sonnet, o3, Gemini 2.5 Pro, or Grok...

495,815 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

173,775 views • 1 month ago

Q: Esteban, new year, new rules, new car. You've had the first run out this morning at the Shakedown. What are the first impressions? How did the program go for you? Esteban Ocon: Yeah, feeling good. I think, frist of all, an unbelievable effort from the team really to put the car down, you know, at 9:20 this morning. But the was ready at 9:00 you know. We were waiting on a bit of a better track condition and a few things that we wanted to be perfect before we went out. But yeah, from Fiorano testing with Ollie to here there's been moving like people have climbed mountains really to make this car work and it's been really good. So, we are dealing with the plan, learning as it goes. Of course, it's a busy program that we have for the day. So you know, its gonna be difficult to complete it. But for the first real day of driving, I think so far it's going really well and we'll keep pushing to make sure that all the details are covered. But we have more days than normal which is a good thing. Q: Very early days, of course, but just how different are these cars to drive? Have you had a chance to play around a little bit with some of the new modes? Ocon: Yeah, it's very difficult, very complicated. I got lucky be able to do a lot of simulator days before we started the year, so we are pretty well set on that. Everything is clear, but yes, it's very complicated you know, for all of us. But I hope that this will be the same for everyone, because if it is we're in the same boat, so we'll see. Q: But I get the sense you're really relishing that challange and what about the priorities from here, the rest of the weel here at the shakedown? Ocon: Yeah, the aim is really to learn, to get mileage under the car, you know, see the weak points, what we have to improve really. First feel of things, so we are sure that we take the right development path and we are sure that we put the resources where it matters the most. Where it's the most bothering us, so you know, we'll try and put all that together for that end of the test. It's a long week, which is very good and then we have the chance to go back to Bahrain with hopefully, further step made, so that's the aim. #HaasF1 #F1Testing #F1

Lucho Yoma

12,593 views • 7 months ago