Loading video...

Video Failed to Load

Go Home

Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet? Meet TRACES 🧭 — the world's first benchmark for measuring discoverative AI: AI that can work through evidence, test...

827,479 views • 13 hours ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

We've made a breakthrough in self-evolving AI scientists moving from "search" to "principled discovery": Scientific discovery requires that the search space itself changes, and an AI scientist must perceive this shift without intervention. We built an AI that achieves this for the first time with the ability to discover the scientific vocabulary it reasons in. Evidence, tools, artifacts, verifiers, failures & claims become typed provenance. We show three distinct modalities: 1) retrieval, adding known objects; 2) search, exploring a fixed schema; and critically: 3) discovery, a verified regime transition. We solve the open-endedness evaluation problem by lifting agentic workflows into a typed copresheaf and proving, via a Kan obstruction, that true discovery is not unbounded generation but a verifiable schema expansion: old evidence is transported by Left Kan extension, and genuine novelty is mathematically quantified by the pointwise residual beyond the transported image - separating discovery from mere search and making novelty objective and measurable rather than a subjective judgment or benchmark delta. Our AI scientist is built in a way that does not pre-conceive the approach it chooses; instead, we endow the system with formal power to adapt, evolve, and reason from first principles. Case studies include: 1⃣Builder/Breaker model that discovers mode-conditioned compliance in proteins; 2⃣CategoryScienceClaw that finds anisotropic fiber-network stiffness rules. Great work in collaboration with my graduate student Fiona Wang MIT Dept of BE F.Y. Wang & M.J. Buehler, Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence, arXiv:2606.01444, 2026

Markus J. Buehler

789,959 views • 2 months ago

Introducing ALE-Bench, ALE-Agent! Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems Blog: Paper: ALE-Bench is a coding benchmark primarily focused on hard optimization (NP-hard) problems. We developed this benchmark with AtCoder Inc., a leading coding contest platform company. What makes ALE-Bench unique is its focus on hard optimization problems that demand long-horizon and creative reasoning. It’s open-ended, in the sense that true optima are out of reach (NP-hard) and scores can continuously improve. We believe this benchmark has the potential to become one of the key benchmarks for reasoning and coding in the next generation. ALE-Agent is our end-to-end agent that we specifically designed for this challenging domain. In fact, our ALE-Agent has already built an impressive track record in the wild! In May 2025, our agent participated in a live AtCoder Heuristic Competition (AHC), alongside 1,000 other participants in real-time. AHC is considered to be one of the most challenging coding competitions in this domain. Our ALE-Agent achieved an impressive ranking of 21st out of 1,000 human participants in the competition (top 2%), marking a turning point for AI discovery of solutions to hard optimization problems with a wide spectrum of important real world applications such as logistics, routing, packing, factory production planning, power-grid balancing. We look forward to applying this technology to real industrial optimization opportunities. Building on the insights from this study, Sakana AI will continue to tackle the challenge of developing AI with even greater algorithm engineering capabilities. ALE-Bench Dataset: ALE-Bench Code: This research was conducted in collaboration with AtCoder Inc. (AtCoder). We are deeply grateful for their outstanding expertise and contributions in optimization and algorithms, which were invaluable in providing data, analyzing results, and enabling our AI agent’s participation in their contests.

Sakana AI

237,195 views • 1 year ago

The smartest man in AI just exposed the whole AGI narrative as a LIE. And he used a physics problem from 1905 to prove it. His name is Demis Hassabis. He runs Google DeepMind, and won the Nobel Prize for using AI to crack a problem in biology that had stumped scientists for 50 years. Almost nobody in this industry has a track record like his. He went on the NothingButTech podcast and called out the biggest lie in AI right now: Right now the loudest voices in AI are telling you that AGI is basically here. OpenAI has literally defined AGI as a system that can outperform humans at most "economically valuable work." In other words, if it replaces enough jobs, we have arrived. Hassabis thinks that bar is a joke. He said real general intelligence has to do what the human brain can do, because the brain is the only proof we have that this kind of intelligence is even possible. He called that "a higher bar than just being able to do some useful economic work," which is about as close as a polite British Nobel laureate gets to calling his rivals out. Then he gave the actual test: Today's AI has read everything humans have ever written, including the theory of relativity. So when it explains relativity back to you, it's repeating an answer that already exists. That's not intelligence. So Hassabis proposed a test that makes memorization impossible. Train an AI on only what humanity knew in 1901, four years BEFORE Einstein published relativity. Then ask it to come up with relativity on its own. It can't look up the answer, because in 1901 the answer doesn't exist yet. The only way to pass is to do what Einstein actually did: Take the same physics everyone else had and reason its way to an idea no human had ever had. Hassabis says not a single AI today can, no matter how much it has memorized. Which means what we keep calling "almost AGI" is really just the best librarian in history. It can find any answer that already exists but it cannot create one that doesn't. His second version is even sharper: AlphaGo, the system his own team built, famously invented a brand new move that no human had played in 2,000 years of the game. Everyone called it genius but Hassabis says that still is not the bar. The real test is not whether an AI can invent a new move inside Go, it is whether an AI could INVENT a game as deep and as beautiful as Go in the first place. No model that exists today can do it. The people telling you AGI has already arrived are the same people raising hundreds of billions of dollars on that exact promise. The valuations only work if the finish line is right in front of us. So the finish line keeps getting dragged closer, and AGI keeps getting quietly redefined down to "does useful work," until the products they already sell happen to qualify. Hassabis has nothing to prove and nothing to sell you. He already won the Nobel, and he is telling you the machines still cannot do the one thing that would make them genuinely intelligent, which is have a truly original idea. To be fair to him, he is not a pessimist about it. He believes real AGI IS coming, and he is spending his life building it. He just refuses to pretend it is already sitting in your phone. So the next time a founder tells you AGI is months away, remember that the one man in the room with a Nobel Prize built his test around Einstein, and admitted that nothing we have made can pass it. What do you think?

Ricardo

1,286,342 views • 2 months ago

Demis Hassabis just made every AI benchmark on Earth irrelevant. Hassabis: “True creativity, continual learning, long-term planning. They’re not good at those things.” Three capabilities every human develops naturally. No AI system has achieved one. Hassabis: “They can get gold medals in international math olympiad questions, but they can still fall over on relatively simple math problems if you pose it in a certain way.” Gold medals and grade school failures from the same machine. That’s not intelligence with blind spots. That’s mimicry with peaks. Performance without comprehension. Hassabis: “My definition of AGI has never changed. A system that can exhibit all the cognitive capabilities that humans can.” Not most. All. He proposed the only test that can’t be gamed. Hassabis: “Training an AI system with a knowledge cutoff of 1911 and seeing if it could come up with general relativity like Einstein did in 1915. That’s the true test of whether we have a full AGI system.” Give a machine everything humanity knew before Einstein’s breakthrough. Seal it off. See if it reaches what one mind reached alone. No benchmarks. No leaderboards. No curated evals. Just a mind against the unknown. Nothing passes that today. Hassabis: “The brain is the only existence proof we have, maybe in the universe, of a general intelligence.” One structure in the entire known universe has proven general intelligence is physically possible. One. And every lab on Earth is racing to build the second. Hassabis: “I think we’re still a few years away from that.” Not decades. A machine that passes that test doesn’t stop at physics. It runs that same process against every unsolved problem in every discipline. No fatigue. No lifespan. No ceiling. Einstein was one mind. One field. One lifetime. This would be a thousand Einsteins across every domain simultaneously. And it never dies. The greatest intellectual achievement in human history becomes the floor. The thing that made one man singular becomes the entry exam for a machine. That’s not artificial intelligence. That’s a different kind of mind. It’s years away. Not generations.

Dustin

19,931 views • 1 month ago

🧵06/34 Narrow vs General AI --- At first glance, this AGI being generally capable in multiple domains looks like a group of many narrow AIs combined, but that is not a correct way to think about it. It is actually more like… a species, a new life form. To illustrate the point, we’ll compare the general AGI of the near future with a currently existing narrow AI that is optimised at playing chess. Both of them are able to comfortably win a game of chess against any human on earth, every time. And both of them win by making plans and setting goals. The main goal is to achieve checkmate. This is the final destination or otherwise called Terminal Goal. In order to get there though it needs to work on smaller problems, what the AI research geeks call instrumental goals. For example: • attack and capture the opponent’s pieces • defend my pieces • strategically dominate the cetre (etc..) All these instrumental goals have something in common: they only make sense in its narrow world of chess. If you place this Narrow Chess AI behind the wheel of a car, it will simply crash, as it can not work on goals unrelated to chess, like driving. Its model doesn’t have a concept for space, time or movement for that matter. In contrast the AGI by design has no limit on what problems it can work on. So when it tries to figure out a solution to a main problem, the sub-problems it chooses to work on can be anything... literally any path out of the infinite possibilities allowed within the laws of physics and nature.

Lethal Intelligence

570,437 views • 1 year ago