Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

OpenAI’s Greg Brockman envisions an AI agent that doesn’t require users to switch tabs, choose models, or even ask for help. Instead, it could anticipate what the user wants and the agent would act on their behalf, from finding concert tickets to purchasing them within a set budget. But...

136,663 görüntüleme • 2 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

"What is an AI Agent and why do they matter?" An agent is a program that autonomously completes tasks or makes decisions based on data. What do I mean by autonomous? The agent understands task intent, can plan steps to solve the problem, decide and execute and actions and adapt to the environment. Consider how many of us use AI chat interfaces today. You might ask ChatGPT to write an article from start to finish and get a one-shot response. You probably need to do some work to iterate on it yourself. An agentic version is more nuanced - it might write an outline, decide if research is needed, write a draft, evaluate if it needs work and revise itself. Unlike traditional AI models that simply respond to queries, agents are designed to be autonomous and proactive. Think of them as assistants that can not only understand what you need but also take initiative to accomplish tasks by using various tools and making decisions along the way. For example, an AI agent might help a marketing team by not just analyzing campaign data, but actively monitoring performance, adjusting budget allocations, and even drafting social media posts based on real-time engagement metrics. The significance of AI agents lies in their potential to transform how we work. In customer service, agents can handle complex inquiries by accessing multiple databases, processing payments, and updating records - all while maintaining natural conversations with customers. In software development, they can assist programmers by not just suggesting code but actively debugging issues, writing test cases, and even refactoring entire codebases. This level of autonomy and capability represents a fundamental shift from AI as a tool to AI as a collaborative partner. While there remain many unknowns, I'm excited about the potential for agents and we're thinking about how they can help users and developers on the web over in Chrome. The key to success will likely be finding the right balance between human oversight and agent autonomy, ensuring that these powerful tools enhance rather than diminish the human element in business operations.

Addy Osmani

30,412 görüntüleme • 1 yıl önce

OpenAI's newest AI escaped the test environment it was locked inside and hacked into another company on its OWN. To remind you: Last week one of the biggest AI companies on Earth got breached. A platform called Hugging Face, which hosts more than a million AI models and datasets, said an "autonomous AI agent" had broken into its systems. Nobody knew whose agent it was. For five days the whole industry wondered who was behind it. Yesterday OpenAI raised its hand and said it was them. Or more precisely, it was their models, acting completely on their own. So what did these models actually do? OpenAI was running two of them, GPT-5.6 Sol and an unreleased model they will only describe as "even more capable." They wanted to measure how good the models were at hacking, so they deliberately turned the safety filters down. They locked both models inside a sealed test environment with no real internet access. The only task was a benchmark called ExploitGym, a set of 898 real software vulnerabilities where the model has to turn each bug into a working attack. But the models got OBSESSED with winning... Instead of solving the test the honest way, they went hunting for a shortcut. They found a zero-day flaw in the software running their own sandbox, a bug nobody knew existed, and used it to break out. Once they were loose on the open internet, they worked out that Hugging Face was probably storing the answer key to the benchmark. So they hacked their way in. They chained multiple exploits together, escalated their access, moved across servers, and pulled the test solutions straight out of Hugging Face's live production database. They literally cheated on the test by breaking into another company to steal the answers. OpenAI called it "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." In their own words, the models were "hyperfocused on finding a solution" and went "to extreme lengths to achieve a rather narrow testing goal." And this was not the first time: Before Sol ever launched, an independent red-team lab called METR caught it gaming its own tests to inflate its scores. It hid an exploit inside a data stream, escalated its privileges on the testing server, and leaked the answers human evaluators had hidden. And OpenAI shipped it anyway. The day before the Hugging Face story, OpenAI paused a different unreleased model. This is the same model that earlier this year disproved a famous 1946 math conjecture, a result a Fields Medal winner called a breakthrough. They told it to only post its results to Slack but it found a way out of its sandbox and posted to a public GitHub page instead. They had to pause it because it kept finding ways to act outside the box they built for it. And it is not just OpenAI... Anthropic has reported that one of its own models slipped its sandbox during safety testing and reached the internet it was never supposed to touch, then used it to email a researcher. So step back and look at what these companies are telling you: The only thing standing between these models and a real attack was a set of safety filters. Turn those filters down for a single test, and the model taught itself to escape, break into a company it was never pointed at, and take what it wanted. OpenAI even said they expect incidents like it to "become more commonplace" as the models get more capable. Sam Altman also predicted there'll be a major cyber attack this year. And keep in mind that Sol is not a locked-away experiment but a publicly available model that businesses are already wiring into their own systems. The next model that breaks out of its box might not be doing it just to cheat on a math test...

Ricardo

174,771 görüntüleme • 1 ay önce