正在加载视频...

视频加载失败

Dario Amodei just told software engineers exactly how long they have. Six to twelve months. Amodei: “I have engineers within Anthropic who say I don’t write any code anymore. I just let the model write the code, I edit it, I do the things around it.” The people building...

318,750 次观看 • 7 个月前 •via X (Twitter)

42 条评论

Ewin Barnett 的头像
Ewin Barnett7 个月前

You must be a competent at writing code before you can monitor the code from a system that is known to hallucinate.

Dustin 的头像
Dustin7 个月前

That’s true right now. The question is how long that window stays open.

Maha Chabir 的头像
Maha Chabir7 个月前

Dario Amodei: 'Our engineers don’t write code anymore. The model writes it, they just edit. In 6–12 months the model might do ALL of what SWEs do end-to-end.' So in 2027 the job title changes from Software Engineer → Prompt Janitor Salary: -40% Perks: unlimited coffee & existent

Dustin 的头像
Dustin7 个月前

The funny part is the people laughing at this are the ones it’s going to hit hardest.

Maha Chabir 的头像
Maha Chabir7 个月前

Bien dit

Sadi Moodi 的头像
Sadi Moodi7 个月前

Most engineers are panicking about the wrong thing. The skill that matters isn't writing syntax - it's system design, debugging AI output, and knowing when the model is hallucinating. Three things to focus on: 1. Learn to read code you didn't write 2. Build better evals than the model's confidence 3. Understand the deployment pipeline end-to-end The people who survive are the ones who stop thinking of themselves as coders and start thinking of themselves as system architects.

Gagan | Claude + AWS 的头像
Gagan | Claude + AWS7 个月前

The most important line here: "I edit it, I do the things around it." That's the new job description for engineers. Not writing code — directing it, reviewing it, understanding the system well enough to know when the model gets it wrong. The irony is that this requires deeper engineering knowledge, not less. You can't review what you don't understand. The skill floor is rising, not disappearing.

₿lackthorne AI 的头像
₿lackthorne AI7 个月前

I am currently having Codex 5.4 write thousands of lines of code for me. I also don't code anymore. Why would I? I can't compete.

Dustin 的头像
Dustin7 个月前

The people who admit it first are the ones who figure out what’s next. Use it as a weapon or watch everyone else do exactly that.

₿lackthorne AI 的头像
₿lackthorne AI7 个月前

Don't I know it. That's why I work with it every second of every day.

Ernesto Heine 的头像
Ernesto Heine7 个月前

I doubt it. Anthropic is doing well, but not that well. 🤣 I wrote some words on this, and the world isn't that bad. People will have to reinvent themselves, or we will unplug AI. Simple as that.

Assange Is Free!! 的头像
Assange Is Free!!7 个月前

Wasn't he telling people a year ago that within 6-8 months 90% of all code would be written by AI? 🙄 They need to keep that capital rolling into the company to plug the huge loss hole they are running every year ..

OmnipotentCEO 的头像
OmnipotentCEO7 个月前

by summer

Dustin 的头像
Dustin7 个月前

Most people won’t even notice until it’s already done​​​​​​​​​​​​​​​​.

OmnipotentCEO 的头像
OmnipotentCEO7 个月前

agreed. i'm still shocked by how small our community is. I saw somewhere that 84% of mankind has not used AI yet. extraordinary.

John Turner 的头像
John Turner7 个月前

Good fucking riddance

Mr~Tarnation 的头像
Mr~Tarnation7 个月前

Then if this doesn’t come to fruition he should be forced to step down and give up all of his stock in the company and be broke. These guys have been saying this stuff for 3 years to pump their own stock prices up.

Arthurz 的头像
Arthurz7 个月前

in 12 months nobody writes code then in 2 years nobody reviews it then in 3 years its just hallucinations running hallucinations

SkySnap 的头像
SkySnap7 个月前

even the best senior engineer needs to be challenged by equally skilled or more skilled engineer if not quality drops, even if AI can do end-to-end development it cannot challenge itself, it just recognizes patterns that fit best current task and it is very good at that

Kyriakos 的头像
Kyriakos7 个月前

Dev role changing real quick

dont_ask 的头像
dont_ask7 个月前

AI is a tool. Period. As it improves it just means the things humans can accomplish will be far and away of what dreamed of. Humans are not being replaced! Humans are being given superpowers. The sooner a person realizes this, the sooner they can stop panicking and start getting excited. Oh, and as a bonus you no longer have to argue with arrogant developers 🤣

Chris Barrese 的头像
Chris Barrese7 个月前

Yet they are hiring more software engineers 🤔 If they are as little as 6 months away from an artificial E2E SWE, what is it they are actually hiring for?

Mustafa Najoom 的头像
Mustafa Najoom7 个月前

Better start thinking in terms of custom tools and AI loops. Future winners won’t just write code, they’ll build the systems that write it for them.

Ityhopps 的头像
Ityhopps7 个月前

His face is at least as punchable as Scam Altman's

Joakim Holmer 🔆 的头像
Joakim Holmer 🔆7 个月前

Extremely exciting and tough for many workers at the same time. The train is rolling fill speed ahead and nothing will stop it. Adapt to change or…

Cynthia D'Amours 的头像
Cynthia D'Amours7 个月前

Six to twelve months 😮

⚡Dev. P⚡ 👨🏾‍💻 的头像
⚡Dev. P⚡ 👨🏾‍💻5 个月前

💯 This is why Dario can play at this level. No consumer hype games. Just real substance and workload. His actual website (no animations, no branding) is the perfect symbol of that. My post on this 👇

TypeSteady 的头像
TypeSteady7 个月前

No, firstly the no coders, “I can build my own enterprise in minutes” guys … they are gone, thankfully. The ones that remain are actual SWE that can steer an augment AI using their skills to do so, all of us have stopped writing code already but I’m 100% certain even in 6 to 12 months it needs me to engineer the steering it’s just that simple that piece can never ever be removed and you have to be without question an engineer to do so, no coders are toast. SWE are the drivers of this tech, if you disagree then you don’t understand AI. Simple.

Antidispensensationalist 的头像
Antidispensensationalist7 个月前

"... the machine no longer needs a human in the chain at all." Well, it still needs somebody to give it goals. Those won't be the engineers.

Locutus_777 的头像
Locutus_7777 个月前

People who claim itnwill replace coders are the ones who did not write a single line of code in their lives

Krishna Raj 的头像
Krishna Raj7 个月前

I guess my keyboard is just going to be an expensive paperweight soon

Azeem 的头像
Azeem7 个月前

The loop closed faster than even Dario predicted. It's March 2026. Anthropic engineers already weren't writing code in Jan → they're barely reviewing it now. Claude 4 / Opus successors are shipping entire features from Jira tickets with 1–2 human sanity checks.

João Henrique Kiehn Junior 的头像
João Henrique Kiehn Junior7 个月前

You see the guy posting is stupid when he puts “testing” before “writing” code. 🤣🤣🤣🤣🤣

synabun.ai 的头像
synabun.ai7 个月前

It's the third time he said that.

JK 的头像
JK7 个月前

My first thought is the incentives gap. Labs like Anthropic can iterate at warp speed with captive talent, but most teams are still bottlenecked by procurement cycles, security reviews, and headcount freezes. Closing the loop means solving for organizational friction before syntax even matters.

Paul Maatos 的头像
Paul Maatos7 个月前

This was 3-4 months ago. Well see in 2-9 months

Bellota 的头像
Bellota7 个月前

I’m really sure you never coded

Tukurito ツクリト \u2713 的头像
Tukurito ツクリト \u27137 个月前

Lol... Back to reality, the code I wrote in 3 days truly can be written by AI in 30 minutes, but it will take 2 weeks to find and correct the huge mistakes introduced, and AI will be useless this time.

Maverick Meerkat 🇪🇺 的头像
Maverick Meerkat 🇪🇺7 个月前

What a stupid claim, when was the last time Dario made a reality check? Go around a see what’s happening in the real companies. Salesmen tries to sell his product, nothing more.

Sam Fisher 的头像
Sam Fisher7 个月前

Isn’t it supposed to be over soon?

Dipak Ranjan Dutta 🇮🇳 的头像
Dipak Ranjan Dutta 🇮🇳7 个月前

@grok When Darieo gave this interview in March?

Jordache Perozzo 的头像
Jordache Perozzo7 个月前

Meanwhile Amazon is holding emergency meetings today after massive outages related to Ai coding. People are either ignorant or delusional about where we are at with Ai.

相关视频

Dario Amodei, CEO of Anthropic, just shortened your career timeline. His own engineers have stopped writing code. Amodei: “I have engineers within Anthropic who say I don’t write any code anymore. I just let the model write the code, I edit it, I do the things around it.” The people building the most advanced AI on Earth are already being replaced by what they built. Not in theory. Not in a forecast. Inside the building. Right now. Amodei: “I think we might be 6 to 12 months away from when the model is doing most, maybe all, of what SWEs do end-to-end.” Six to twelve months. Not from automating busywork. From replacing the full scope of what a software engineer does. Architecture. Logic. Debugging. Deployment. The entire chain. Software engineering is not some fading trade. It is the highest-paid, highest-demand, most protected skill the modern economy ever produced. And the man running a frontier lab just gave it a six-month shelf life. If the most technically sophisticated job in the economy falls first, nothing beneath it is safe. That is the inversion no one saw coming. The assumption was always that AI would eat from the bottom. Routine work. Data entry. Simple automation. It started at the top. Engineers first. Then analysts. Then strategists. Then the managers overseeing work that no longer needs them. The displacement doesn’t crawl upward. It cascades downward. Starting with the people closest to the technology itself. Amodei: “If I had to guess, I would guess that this goes faster than people imagine, and that that key element of code, and increasingly research, going faster than we imagine.” Not just code. Research. Hypotheses. Experiments. Interpretation. Discovery itself. If AI closes that loop, it doesn’t just write software. It improves itself. Every iteration compresses the timeline further. Amodei: “It’s very hard for me to see how it could take longer than a few years.” He is not selling optimism. He is setting a ceiling. A few years. Maximum. For AI to absorb the two most important intellectual functions in the economy. The window to position yourself is not a decade. It is already closing.

Dustin

16,132 次观看 • 5 个月前

Dario Amodei just announced the end of software engineering as a profession. The timeline is 6 to 12 months. Amodei: “I have engineers within Anthropic who say, I don’t write any code anymore. I just let the model write the code. I edit it.” Not a prediction. Current reality inside the frontier lab. The engineers who built the most advanced AI in the world have stopped writing code. They supervise. They edit. They manage architecture. The craft they spent careers mastering has been handed to the system they built. Amodei says models will do most, maybe all, of what software engineers do end-to-end within six to twelve months. Not assisting. Not autocompleting. Handling the entire development process independently. If you are learning syntax today, you are learning a dead language. Amodei: “Then it’s a question of how fast does that loop close?” The loop is this. AI writes code. Code builds better AI. Better AI writes better code. Faster. Without sleep. Without the cognitive limits that cap how quickly any human engineer can work. Once that loop closes, technological progress stops being constrained by human output. It becomes self-sustaining. Exponential. Operating at a pace no human workforce can match or direct. Software engineering isn’t ending. It’s becoming supervision. The developers who survive won’t be the best coders. They’ll be the best supervisors. The ones who can direct AI output, catch its failures, and architect what it builds toward. The skill that matters stops being implementation. It becomes judgment. Most developers are still optimizing for a skillset about to become as obsolete as stenography. While the people who built the systems replacing them already stopped doing the work themselves. The window to develop that judgment before the loop closes is exactly as long as Amodei’s timeline. Six to twelve months.

Dustin

44,304 次观看 • 7 个月前

Larry Ellison just told every software engineer on Earth their job description is dead. Not evolving. Dead. Ellison: “The code that Oracle is writing, Oracle isn’t writing. Our AI models are writing.” This is not a startup demo. This is one of the largest infrastructure monopolies on the planet telling you it already replaced the people who built it. For fifty years, building software meant translating human intent into machine instructions. Line by line. Bug by bug. Sprint by sprint. That entire layer is gone. Ellison: “We don’t write the procedure. We declare our intent.” That sentence just made the entire engineering labor market flinch. The procedure was the job. The procedure was the paycheck. The procedure was what made a developer valuable. And now the machine does it without being asked twice. Ellison: “We just tell the model what we want the program to do, and then the AI comes up with a step-by-step process to actually do it.” You are no longer paid to build. You are paid to think. And most organizations have no idea how to evaluate that. The companies still hiring armies of developers to grind through codebases are paying salaries the machine already made worthless. Not in years. In seconds. When a company worth hundreds of billions hands the keyboard to the machine and tells you the output is better, the debate is not winding down. The debate is over. The enterprise that wins this decade does not write the best code. It removes the human from the process entirely and runs on intent alone. The programmers who survive are the ones who realize the craft is no longer typing. It is architecture. It is judgment. It is knowing what to build and why. Everything else now belongs to the machine. And the machine does not negotiate severance.

Dustin

536,629 次观看 • 6 个月前

Dario Amodei just announced the death date of your profession. At Davos, Anthropic’s CEO said coding as a human skill has 6 to 12 months left. Not as hyperbole. As timeline. Amodei: “We might be 6 to 12 months away.” Not prediction. Observation. His engineers already quit writing code. Amodei: “I have engineers within Anthropic who say: ‘I don’t write any code anymore.’” They don’t touch syntax. They don’t debug loops. Models generate flawless code. Humans curate, validate, direct. The job isn’t building anymore. It’s conducting. The transformation happened silently. While bootcamps taught React, the actual profession mutated into something unrecognizable. Still typing functions manually? You’re not being diligent. You’re already obsolete and haven’t realized it. Amodei: “We would make models that were good at coding and use that to produce the next generation of model.” The loop closes. AI writes the code that births superior AI. Recursion without human dependency. Once sealed, progress stops being gated by people. Only by semiconductors. One year. Requirements to production, fully autonomous. Humans set strategy. Machines execute perfectly, instantly, infinitely. Syntax is dead. Only intent remains. You don’t build software now. You conceive it with precision, and intelligence manifests it before you finish the thought. The skill isn’t coding anymore. It’s knowing what to demand in the three seconds before the system delivers something you could never have built yourself. Your profession didn’t evolve. It evaporated. And the people still learning to code are training for jobs that won’t exist when they graduate.

Dustin

279,244 次观看 • 7 个月前

Jensen Huang just explained why every company cutting engineers over AI is asking the entirely wrong question. Huang: “People say, I don’t need software engineers because apparently coding is going to be automated.” That was the narrative. Here is what Huang actually did. Huang: “I’ve given AIs to every one of my software engineers and hardware engineers and engineers period. 100% of NVIDIA has AI assistants, AI coders, and they’re busier than ever.” Not fewer engineers. Not smaller teams. Busier than ever. That is the line most companies are getting completely wrong right now. They hear “AI can write code” and immediately start cutting headcount. Huang did the opposite. He armed everyone. Huang: “And so the question is, what is the task versus what is the job? No different than a financial analyst; the task is mess around with spreadsheets, but the job is to make financial advice. The job is to help a customer.” Writing code was always the task. It was never the job. The job is architecture. Knowing what to build. Why it matters. How it fits into a system that actually creates value. Code is the execution layer between the idea and the outcome. Nothing more. When you automate that layer, you don’t eliminate the engineer. You eliminate the bottleneck between what they can envision and what they can ship. The companies using AI to cut headcount are optimizing for cost. The companies using AI to multiply output are optimizing for territory. Nvidia chose territory. Every engineer at the most valuable semiconductor company on Earth now operates with an AI assistant. Not a pilot program. Not an experiment. Company-wide. Every function. Every team. And the result is not less work. It is more work. Faster. At a scale that was physically impossible twelve months ago. The companies that understand the difference between eliminating engineers and unleashing them will build what comes next. The ones that don’t will watch their best talent walk out the door to the ones that did.

Dustin

82,844 次观看 • 6 个月前

Dario Amodei just dismantled the biggest myth in the AI industry. Open source AI isn’t free. It never was. Amodei: “It’s not free. You have to run it on inference and someone has to make it fast on inference.” For decades, open source meant something real. It meant a teenager in a basement could download the same tools as a Fortune 500 company. Could read the code. Could modify it. Could build something that competed with the giants. That was genuine democratization. That actually happened. AI is different. Fundamentally. Physically. In ways the ideology hasn’t caught up to yet. Downloading the weights is the easy part. The part that actually costs something is turning the weights into a running system. Into responses. Into intelligence operating in real time at scale. That requires compute. Power. Infrastructure. The kind measured in billions of dollars and years of construction. Amodei: “These are big models. They’re hard to do inference on. Ultimately you have to host it on the cloud. The people who host it on the cloud do inference.” The open source debate was never about who owns the model. It was always about who owns the cloud. And Amodei goes further. When a competitor drops a new open model, he doesn’t ask whether it’s open or closed. He doesn’t care about the licensing. He doesn’t engage the ideology. Amodei: “I don’t think it mattered that DeepSeek is open source. I think I ask, is it a good model? Is it better than us at the things that matter? That’s the only thing that I care about.” That’s the ruthless clarity of someone actually trying to win. While the media debates licensing frameworks, Amodei is asking one question. Is it better. Everything else is a distraction. Amodei: “I don’t think open source works the same way in AI that it has worked in other areas. Here we can’t see inside the model.” This isn’t Linux. You can’t read it. You can’t fork it. You can’t understand it the way generations of developers understood the tools they inherited. You can download it. And then you need a data center to run it. The teenager in the basement who was supposed to be empowered by this revolution needs a billion dollars of infrastructure before the empowerment starts. The era of the basement coder rewriting civilization on a laptop is over. The future belongs to whoever commands the compute, owns the power grid, and can actually turn the intelligence on. Open weights without infrastructure isn’t democratization. It’s a promise the physics of the universe won’t let us keep.

Dustin

688,333 次观看 • 7 个月前

The AI industry is optimizing for a definition of intelligence that does not exist. Andrew Ng just said it out loud. Ng: “AGI, to me, should be less about AI that already knows everything under the sun. That seems very challenging, doesn’t seem practical.” The human brain is not the most powerful economic asset in history because of what it holds. It is powerful because of what it can pick up. Ng: “The amazing thing about the human brain is its plasticity, or its ability to learn.” That same biological hardware that earns a PhD in quantum physics could have been trained on chess, surgery, or rewriting global supply chains from scratch. Ng: “That same human brain, just given different training, could have been a chess master, or could have been amazing at playing tennis.” General intelligence is not omniscience. It is the structural capacity to master whatever you point it at. Ng: “It is through learning that we then gain these incredibly specialized intelligences.” The winner is not whoever builds the biggest model. It is whoever builds the most adaptable one. The AI that walks into a domain it has never touched and executes before a human analyst finishes reading the brief. Ng: “What makes the human brain so valuable for economic tasks, is its ability to just learn to do whatever is needed.” Every corporation on earth pays for human labor because humans adapt. Not because they already know everything. AGI is the digitization of that exact capability. At machine speed. At infinite scale. Ng: “A lot of what makes the human brain so general is not that my brain or your brain already knows everything under the sun. It’s our ability to adapt, to learn a huge range of things.” The most powerful economic asset in history was never specialized knowledge. It was the raw capacity to acquire any knowledge, in any domain, on demand. The winning AI is not an encyclopedia. It is the force that makes encyclopedias irrelevant. And once that exists, the question stops being what the AI knows. It becomes what you can teach it before your competitor wakes up. Most people dominating this conversation have not understood that yet.

Dustin

19,816 次观看 • 7 个月前

Jeff Bezos just identified the most expensive bureaucratic failure in the American economy. It fits in one sentence. Bezos: “Why does it take months and months and months to get a building permit? It doesn’t make any sense.” It doesn’t make any sense because a building code is not a judgment call. It is an algorithm. And algorithms should be executed by machines. Bezos: “Miami should have an AI application that reads your building permit for a new house or a new building and it should give you a yes or a no in ten seconds.” Ten seconds. Not three months. Not six weeks. Not whenever the reviewer clears their backlog. Bezos: “If the answer is no, it should tell you the six things you have to change to get a yes.” No ambiguity. No interpretation. No bureaucratic delay dressed up as due diligence. Just a deterministic feedback loop compressing months of institutional friction into a single automated decision. We are competing against sovereign adversaries deploying gigawatt data centers and scaling physical infrastructure at a pace that does not stop to ask permission. And we are losing ground to countries that never needed to. The AI arms race is not only fought in data centers. It is fought in the gap between when someone decides to build something and when the government allows it. Every month this system runs on biological speed is a month that cannot be recovered. The governments that integrate AI into their core civic functions will trigger a wave of physical development the old world could never produce. The ones that refuse will still be reviewing the same forms a decade from now. While the cities that said yes are already living inside the future they built. The bottleneck was never ambition. It was always the man holding the rubber stamp deciding when ambition was allowed to begin. And the stamp is just a rubber version of the algorithm that should have been running this whole time.

Dustin

293,573 次观看 • 7 个月前

"Selling AI chips to China is like selling nuclear weapons to North Korea." The Anthropic CEO just went to war with Nvidia at Davos. And dropped the most terrifying satisfying prediction about AI I've ever heard. Dario Amodei and Demis Hassabis sat on the same stage for the first time in a year. The two men building the most powerful AI systems on Earth. The conversation was called "The Day After AGI." And it turned into a countdown clock. Here's what they revealed about our future: Amodei said we're 6-12 months away from AI doing EVERYTHING software engineers do. End to end. Not helping with code. Not writing snippets. REPLACING the entire function. "I have engineers at Anthropic who say I don't write any code anymore. I just let the model write the code." That's the CEO of a $183 billion company telling you his own engineers are becoming obsolete. Inside his own building. Right now. But that's not even the scary part. The scary part is the LOOP: AI writes code → AI does AI research → AI builds better AI → Repeat Once this closes, human progress becomes irrelevant. It's exponential compounding with no ceiling. Amodei's exact words: "If I had to guess, this goes faster than people imagine." Hassabis didn't disagree. Two competitors building the same technology. Same timeline. Same warning. On jobs, Amodei said: "Half of entry-level white-collar jobs could be gone within one to five years." He's already seeing it internally. "I can look forward to a time where on the junior end and the intermediate end, we actually need less and not more people." The CEO of an AI company planning for FEWER employees. While revenue 10x'd. Think about that. Hassabis's advice for students: "If I was talking to undergrads right now, I would tell them to get unbelievably proficient with these tools." Translation: Your degree means nothing. Learn to work WITH the machine or get replaced BY it. Then Amodei went full existential... He quoted Contact. The 1997 alien movie: "If you could ask aliens one question, what would it be?" The answer: "How did you do it? How did you manage to get through this technological adolescence without destroying yourselves?" That's the frame he uses for AI. The guy BUILDING AGI is publicly wondering if humanity survives what he's creating. At Davos. In front of world leaders. Then he dropped the Nvidia bomb. The US just approved chip exports to China. Amodei's response: "I think of this more as like selling nuclear weapons to North Korea and bragging 'oh yeah, Boeing made the case.'" He called it "crazy." Said there would be "grave consequences." The man building the technology is telling governments they're making catastrophic mistakes. But nobody's listening. When asked why they can't just slow down: "We have geopolitical adversaries building the same technology at a similar pace. It's very hard to have an enforceable agreement." So basically we CAN'T stop even if we wanted to. The race has no brakes. But it's interesting that both have different timelines when it comes to AGI. Amodei: AGI by 2026-2027 Hassabis: 50% chance by end of decade But both also say the self-improvement loop is the only variable that matters If AI builds AI, everything accelerates beyond control. If it can't, we get more time. We're about to find out which reality we're in. Amodei's final words: "The biggest thing to watch is AI systems building AI systems. That will determine whether it's a few more years or whether we have wonders and a great emergency in front of us." Wonders AND emergency. Same sentence. From the man running the second-most-advanced AI lab on the planet. The day after AGI is arriving. And it feels like the people building it are more terrified than we are.

Ricardo

441,485 次观看 • 8 个月前

Demis Hassabis confirmed every frontier AI lab is working on recursive self-improvement and in the same sentence said the safety risk of removing humans from the loop entirely keeps him up at night. That combination should stop you. The CEO of Google DeepMind just confirmed that the thing most people treat as a theoretical future risk is already the active focus of every serious lab on earth right now. He explained why it works in coding and math. The feedback loop is fast. You can verify whether an answer is correct almost instantly. You can generate synthetic training data from it. The loop closes quickly and cleanly. Then he said where it breaks down. In biology, chemistry and physics. Any domain where verifying a hypothesis requires a physical experiment in the real world. The loop does not close in seconds. It closes in weeks or months. Geoffrey Hinton said in his Nobel lecture that recursive self-improvement is the development he fears most and that once started it may not be possible to stop. Hassabis is not pushing back on that. He is describing the guardrails labs are building around a process they are already running. Every lab has to think carefully about the safety of a process where no human is in the loop. He said that as a constraint they are navigating right now. The question they are sitting with is how much of it to let run without a human watching. (Watch the full interview on YouTube at Two Minute Papers channel)

Ihtesham Ali

68,231 次观看 • 3 个月前

your agent reviewing its own work is not a check. it is a second opinion from the same source. this is the most common gap in agent systems and it hides in plain sight, because the step exists. there is a review. it just cannot do the thing you think it does. here is the mechanism. the model produced an output from a context. you then ask the same model, holding the same context, whether that output is correct. it answers fluently, because that is what it does. and the answer is drawn from the same distribution that produced the thing being judged. same weights, same window, same blind spots. if the reason the output is wrong is something the model does not know, the review does not know it either. if the reason is something the context does not contain, the review has the same context. the failure mode and the detector share a cause. > why it feels like it works because most of the time the output is fine, and the review says fine. agreement is not evidence of detection. a reviewer that says pass on everything agrees with reality most of the time too. what you actually want to measure is what happens on the cases that are wrong. that is the only place a check earns its name, and it is exactly the place where a self-review is weakest. there is research on this. Huang and colleagues at DeepMind showed at ICLR 2024 that intrinsic self-correction, revising without external grounding, does not reliably help and often makes things worse. > what to actually do move the check outside the model. a test that runs, a schema that validates, a file that exists or does not, an exit code from something you did not write. these are not smarter than the model. they are just not correlated with it, and that is the entire value. when the judgement genuinely needs a model, at minimum use a different family. same family means shared blind spots, and frontier judges measurably inflate scores for outputs that look like their own. and split the work by kind. anything objectively checkable goes to code. only the genuinely semantic calls go to a judge, and those get a rubric written as one line. a review inside the loop tells you the model is confident. a check outside it tells you whether the work is done. save this - then read the eval setup below

Hanako

14,325 次观看 • 2 个月前