Loading video...
Video Failed to Load
We've made a breakthrough in self-evolving AI scientists moving from "search" to "principled discovery": Scientific discovery requires that the search space itself changes, and an AI scientist must perceive this shift without intervention. We built an AI that achieves this for the first time with the ability to discover... show more
791,554 views • 4 months ago •via X (Twitter)
47 Comments

Paper:

@blader Can we test this on ARC-AGI-3?

very cool paper Category theory fits here because neural networks learn data relationships, not data itself. The paper formalizes those relationships as typed artifacts, morphisms, provenance, and regime transitions. Deep Manifold reads these categorical structures as external boundary conditions that guide numerical manifold traversal. Discovery occurs when a new relation becomes stable enough to form a new fixed-point pathway. #DeepManifoldInterpretation

The jump from searching a fixed schema to actually expanding the schema itself is the part that makes this genuinely different. The fact that novelty is now mathematically quantifiable rather than a benchmark delta is a real step forward.

Thank you @AlphaWireHQ !

I like how this is timed with in Boston! @NoahChrein might find this interesting!

@NoahChrein 😀

the Left Kan extension framing is clean, but the residual's objectivity seems to ride on how you typed the provenance. too coarse and real discoveries collapse into the transported image, too fine and every reformulation reads as novel. feels like the subjectivity moves to the choice of base category rather than disappearing.

Thank you @somi_ai great points. The critique would apply if typing were post hoc but in our framework the schema itself is part of the audited state: typed operations, immutable lineage, gates, rejected alternatives, and explicit regime transitions. Novelty is only the residual beyond the transported image (i.e. what carrying old evidence over by Left Kan extension cannot produce) and only once the new state passes its own gate. In Builder/Breaker mode-conditioned compliance survived MDL on enlarged evidence and arbitrary refinements were rejected/retracted. In other words subjectivity is not hidden in the ontology choice - it is exposed, constrained, and audited.

Excellent work. The key move here is recognizing that discovery is not just better search inside a fixed representation space, but a verified change in the representation space itself. The typed provenance layer is especially important: evidence, tools, artifacts, failures, verifiers, and claims become part of an auditable scientific state, not just loose agent outputs. That makes the distinction between retrieval, search, and discovery much sharper. This resonates with broader categorical views of intelligence: learning as finding invariants and changing the base language in which knowledge is organized. The Kan-extension/residual novelty framework feels like a practical step toward AI systems that can evolve their scientific vocabulary while preserving what was already known. Very exciting direction.

💯

Can you say something about why you think the formalism of Kan extensions is appropriate here? (Also I hope you don’t mind if I ask you not to copy and paste from an LLM in response—I’ve already discussed the paper with a frontier model.)

Great question, thank you. Building a "real" AI scientist means building an AI that can rewrite its own representational regime and here, Kan transport is how we formalize the process of data migration. Specifically we want to preserve what it already knew across the rewrite so that discovery becomes exactly what survives outside the transport.

Thanks! I think I understood that; what I’m asking about is why you think this formalism is a good one. E.g. why was it chosen, what good features does it have, etc.? I’m a bit skeptical that Kan extensions are meaningfully present in the scientific process or agentic workflows.

Incredibly timely work, Prof. Buehler. I recently lived this exact edge case as a human Verifier. When introducing a novel, first principles chemical formulation to a standard LLM 'final check,' the model defaulted to a fixed schema search, confabulating out of date textbook ...

Thank you!

in case it is helpful...

Sheaf Neural Networks on SPD Manifolds: Second-Order Geometric Representation Learning

hello! I’ve been working on the “metatheory of constructive knowledge” and also simultaneously older work from undergrad that tries to speak to equitable partitioning over changing spaces. agent coordination: equitable partition:

still feels like optimizing, just over spaces instead of points. whats the signal it uses to know the current space is the wrong one?

This is really nice I'll add it to my topos theoretic simulator

Funny, I used these exact words on a first date last night! Didn’t go well for some reason… “…lifting agentic workflows into a typed copresheaf, proving via a Kan obstruction that true discovery is not unbounded generation but a verifiable schema expansion…”

This is a much bigger shift than it looks at first glance. Most AI systems optimize within a language of concepts they inherit, your framework asks whether the language itself can evolve. The real breakthrough is not generating more hypotheses, it's making schema expansion something that can be verified rather than narrated after the fact. If this holds up, it pushes AI closer to participating in science instead of just accelerating existing scientific workflows.

Discovery cost = L(I′) − L(im ρ̄)" only subtracts cleanly if description length is additive. It's subadditive. The relative form L(I′ | im ρ̄) is well-defined but it's just a conditional code length, not the before/after gap the figures imply. I’m confused by this

You're right that L is subadditive and we don't assume additivity. Discovery cost is defined as the conditional code length L(I′ | im ρ̄); the subtraction L(I′) − L(im ρ̄) is only its additive special case (a subadditive upper bound otherwise). The figure numbers aren't that cost anyway but instead they are paired MDL acceptance gains at fixed evidence which we flag explicitly as 'not a direct numerical discovery cost'. So the conditional form is the cross-regime primitive; the figures' before/after gaps are a separate fixed-evidence quantity. We never equate them.

Is there a plan to add experiments that directly test the regime-transition formalism itself? For example, a case where the Kan-transport/residual criterion leads to a different conclusion than MDL-guided symbolic search or provenance tracking alone?

Good research

Thank you!

Might be of interest!

principled discovery is the right framing.

Thank you😀

Fascinating work. One aspect that resonates strongly is the distinction between search within a fixed representation and discovery through representational revision. In our Recursive Indexing (RIX) work, we have observed that transfer, compression, and perturbation behavior are governed less by local similarity than by transport topology. This suggests that understanding how representations change may require not only a formal language for regime transitions, but also measurable diagnostics of transport stability, fragmentation, and recursive organization within those regimes. Excited to see more work connecting representation geometry with scientific discovery.

Impressive work!

@waitbutwhy I found our next rabbit hole. This is fascinating stuff!!!

@patience_cave These people are not patient, at all.

I’m an independent researcher building Project Aletheia, an independent research lab for computational phenomenology and adversarial discovery. Your framework for typed provenance, verifiers, and regime transitions feels relevant to what I’m doing. I’d be very open to collaboration. Also, I’d take any help I can get

The Topos of Transformer Networks

Cross pollin-Ai-tion

Amazing !!!

Do you believe an AI agent in the future should be able to win a Nobel/Turing/or Fields medal all on its own or should the prize go to you if the core or framework of that AI scientists comes from some future iteration or evolved version of your work now.

You really didn’t need that extremely overhyped and partially bogus tweet to go along with your paper. It’s a good paper but it’s really not all that you’re saying in the tweet. I don’t think the bots commenting really read the paper.

YES

Another question: The real-data law ends at R²=0.41 and gets less accurate as harder proteins enter. Why is that "evidence widening" rather than the model failing? And what's the R²=0.99999 fit to synthetic data actually demonstrating?

Realy nice work @ProfBuehlerMIT and team 👏 One thing I'm stuck on. The definition is tuned for expansion (new types, new content). But a lot of progress is compressive. Same predictions exactly, smaller ontology. Eliminating an axiom that's derivable from the others. These leave zero residual, yet they clearly change the regime, so they sit in neither "discovery" nor "search." Is compressive progress meant to be discovery, a separate category, or out of scope? Genuinely curious.

An AI that notices when the search space itself changes without being told is a completely different kind of intelligence

Yes!

Most AI optimizes within boundaries humans set. This one questions the boundaries. That is a fundamentally different thing.

