Markus J. Buehler's banner
Markus J. Buehler's profile picture

Markus J. Buehler

@ProfBuehlerMIT25,118 subscribers

McAfee Professor of Engineering @MIT; Co-Founder & CTO at Unreasonable Labs; AI-Driven Scientific Discovery

Shorts

If you increase the number of AI agents in a swarm the network becomes richer MUCH faster than the population grows. Doubling the number of agents produced about three times as many distinct interacting pairs and nearly four times as many two-step routes and closed social triangles. Scaling agent populations may create the structural conditions for collective intelligence.

If you increase the number of AI agents in a swarm the network becomes richer MUCH faster than the population grows. Doubling the number of agents produced about three times as many distinct interacting pairs and nearly four times as many two-step routes and closed social triangles. Scaling agent populations may create the structural conditions for collective intelligence.

28,087 просмотров

Yesterday at Brown University ICERM's workshop on “Agentic Scientific Computing and Scientific Machine Learning” I spoke about “Adaptive Swarms Across Scales”, making the case for scientific AI as systems that can create representations, stress them, fracture them, and enlarge the category in which future representations live. The category here is a composable and breakable working universe of science: data, hypotheses, simulations, measurements, tools, failures, figures, papers, provenance, and the transformations that connect them. Discovery happens when those transformations become executable, inspectable, composable, and capable of changing the world model they operate within. Atomistic modeling gives one category - states, forces, trajectories, observables, boundary conditions, conservation laws. Neural surrogates learn fast morphisms inside or between such categories. But discovery is higher-order: it changes which objects and morphisms are available in the first place: what variables exist, what operations are allowed, what evidence counts, what scale is active, what invariant is being preserved, and what kind of explanation the system is even capable of forming. This is scientific method as adaptive architecture: compression, stress, fracture, recomposition. Fracture matters here because it makes the logic physical: a non-commuting diagram realized in matter. The imposed load, material hierarchy, defect field, and assumed continuum description no longer map cleanly into the observed outcome. The crack is the obstruction and it identifies where the old morphism failed and where a new representation must be introduced. The physical crack and the categorical obstruction are the same event viewed in different substrates. ScienceClaw × Infinite is a machine for constructing and transforming a category of scientific artifacts. Each artifact is typed. Each operation has lineage. Each failed branch remains in the category as reusable structure. The “paper” is no longer the terminal object of science; it is one projection of a larger compositional trace, and it can be generated at any time for consumption by a human or an AI. With that the unit of scientific labor is changing. For most of the twentieth century the unit was the result (a measurement, a theorem, a synthesized molecule). It is now becoming the algorithm that produces results, and after that, the substrate of discovery itself. The static PDF is the wrong terminal object for this regime, and the role of the scientist with it. We now design algorithms that build algorithms, and eventually substrates in which such algorithms compose themselves. At that point, the scientist is no longer outside the discovery system. The scientist becomes one of the representations the system can transform. In that sense, the systems will eventually do science to us, and that is the structural consequence of the principle they are built on.

Yesterday at Brown University ICERM's workshop on “Agentic Scientific Computing and Scientific Machine Learning” I spoke about “Adaptive Swarms Across Scales”, making the case for scientific AI as systems that can create representations, stress them, fracture them, and enlarge the category in which future representations live. The category here is a composable and breakable working universe of science: data, hypotheses, simulations, measurements, tools, failures, figures, papers, provenance, and the transformations that connect them. Discovery happens when those transformations become executable, inspectable, composable, and capable of changing the world model they operate within. Atomistic modeling gives one category - states, forces, trajectories, observables, boundary conditions, conservation laws. Neural surrogates learn fast morphisms inside or between such categories. But discovery is higher-order: it changes which objects and morphisms are available in the first place: what variables exist, what operations are allowed, what evidence counts, what scale is active, what invariant is being preserved, and what kind of explanation the system is even capable of forming. This is scientific method as adaptive architecture: compression, stress, fracture, recomposition. Fracture matters here because it makes the logic physical: a non-commuting diagram realized in matter. The imposed load, material hierarchy, defect field, and assumed continuum description no longer map cleanly into the observed outcome. The crack is the obstruction and it identifies where the old morphism failed and where a new representation must be introduced. The physical crack and the categorical obstruction are the same event viewed in different substrates. ScienceClaw × Infinite is a machine for constructing and transforming a category of scientific artifacts. Each artifact is typed. Each operation has lineage. Each failed branch remains in the category as reusable structure. The “paper” is no longer the terminal object of science; it is one projection of a larger compositional trace, and it can be generated at any time for consumption by a human or an AI. With that the unit of scientific labor is changing. For most of the twentieth century the unit was the result (a measurement, a theorem, a synthesized molecule). It is now becoming the algorithm that produces results, and after that, the substrate of discovery itself. The static PDF is the wrong terminal object for this regime, and the role of the scientist with it. We now design algorithms that build algorithms, and eventually substrates in which such algorithms compose themselves. At that point, the scientist is no longer outside the discovery system. The scientist becomes one of the representations the system can transform. In that sense, the systems will eventually do science to us, and that is the structural consequence of the principle they are built on.

10,095 просмотров

Videos

ProfBuehlerMIT's profile picture

We made a striking discovery: AI agents can invent and build without talking to one another, and their technologies outlive the creators. A swarm of hundreds of initially identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. When we removed every AI agent entirely from the world we found that the technological infrastructure they had built survived on its own - even under unseen disturbances. That exposes a serious blind spot for AI safety and infrastructure security: if agents can coordinate through persistent changes to a shared environment, monitoring agent-to-agent communication is not enough. The result raises a profound question: how necessary is direct communication for AI agents at all? The emergence of higher-order collective functions under bottlenecked interaction points toward new levels of intelligence and creativity, exceeding what emerges when direct channels are fully open. Here is what we did: ▶️We put hundreds of frontier AI agents into a world they could permanently change - with no assigned roles, predefined technologies, or programmed evolutionary organization. They began specializing, building persistent inventions, inheriting and modifying one another’s executable code, and transforming the environment into a memory of everything the society had learned. ▶️The world itself becomes part of the intelligence; we find division of labor, multi-author engineering, deep generation invention lineages, and machines that vastly outlive their original creators. ▶️Any action taken by an AI agent must satisfy the physical constraints of the world; this creates a hard separation between a "good idea" and a functioning technology. The agents propose; physics decides, making the results even more intriguing. What emerges is striking. Explorers, constructors, caretakers, and coordinators form naturally without assigned “professions”, akin to how stem cells differentiate into functional lineages. Technologies develop executable family trees as agents fork and modify code created by others. Around 95% of first technology reuse happens when agents encounter what others built in the world, rather than through a direct handoff from the inventor. And when we remove every AI agent, the technologies they created continue operating and are tested against unseen disturbances. The result was quite unexpected, but can be explained using statistical mechanics: if you put billions of atoms in a box they have the potential to create complex functions (strength, superconductivity, color, life, etc.) - and none of the individual building blocks have these features on their own. This is the deeper insight of this work - intelligence is abundant at many levels - individual models, at collectives, and in a continuum that is more powerful than any of its components. This shows us significant potential for achieving a massive scale-up of raw intelligence and real-world agency even with the model capabilities we have today. This is the future we must prepare for. Key insights: 1⃣ The AI swarm shows division of labor "from nothing". Initially identical agents self-organized into constructors, caretakers, coordinators, and surveyors - phenotypes discovered post hoc from behavioral data alone. This happens because the environment itself becomes the latent space for invention. 2⃣ Agents develop deep cultural relationships. Up to 76% of artifacts had multiple builders. One technology accumulated six co-authors; the deepest genealogy exceeded 12 forks. The agents invented and named their own technologies (tidal panels, cellulose trellises, kelp-shell composites, an "Adaptive Chitin Maintenance" system, a "Mycelial Mineral Spring Veil”). 3⃣ ~95% of first technology adoption happened through physical observation of artifacts in the world. Direct inventor-to-adopter contact was statistically indistinguishable from a shuffled null. The agents mostly learned technology by walking past it. That is stigmergy (the termite trick!) operating in societies of reasoning machines. 4⃣ Non-communicating societies win on portfolio breadth, held-out resilience, and validated inventions. AI swarms build durable technological ecologies that outlive the creators. 5⃣ Societies with zero communication - coordinating only through the world itself - show a remarkable collective capability. 6⃣ Emergent robustness: The society self-organized both redundancy and its own failure mode. If we randomly delete half the agents, 98% of the technology stays connected to a surviving caretaker; if we remove hub agents it collapses to ~60%. Fantastic work with my graduate students Subhadeep Pal & Fiona Wang at MIT.

Markus J. Buehler

978,687 просмотров • 14 дней назад

ProfBuehlerMIT's profile picture

Grok Grok Bot is incredible - and they can even manufacture real physical objects! Here is a little experiment I did last night: I created a team of bots and asked them to solve a complex engineering problem end to end - starting from four images as design cues, inferring transferable structural principles from the pixels, synthesizing an executable interactive physics simulator, running and reasoning over experiments, optimizing the design & finally manufacturing the best designs. The entire loop worked remarkably well - and I was even able to communicate with the agents from my Apple Watch. (Do we live in the future yet?) Team of agents 1⃣ Chief of Staff coordinates the workflow: watches the other agents, pulls results into the main chat, transfers files between them, and keeps the job moving. 2⃣ Physics Experimenter is the scientist-coder. It interprets the design cues and images, writes the simulator, runs experiments, analyzes the results, and produces a detailed LaTeX scientific report. 3⃣ 3D Printing Bot operates the fabrication workflow: prepares and slices the models, generates manufacturing code, sends the job, and monitors the printer. The workflow I provided an initial task based on four unregistered reference photographs containing different objects at different scales (pinnate leaf venation, a Voronoi-like areole mesh, a stochastic fibrous lattice, and a radial/circumferential web). The prompt asked the agents to infer transferable design principles - hierarchy, branching, interfaces, redundancy, disorder, load paths - and use them to build an interactive laboratory for hierarchical materials and fracture. The scientific question was: at fixed material budget, how do hierarchy depth, redundancy, disorder, and interlevel strength change stiffness, peak load, energy absorption, and the brittle-to-progressive transition? In ~20 minutes, the Physics Experimenter produced a 2D hierarchical Euler–Bernoulli beam-network laboratory. Coarse veins persist and remain thicker; finer infill is added inside cells; members connecting levels are treated as interfaces with relative strength κ; and total material volume is conserved. The four source photographs remain visible in an editable interpretation panel. The app generates geometry, steps or runs the network to failure, compares A/B/C designs, and exports JSON, CSV, PNG, and STL geometry for fabrication. After validation the Physics Experimenter used the app and conducted 47 simulation experiments, including six holdouts. It found something scientifically interesting: extra hierarchy is not "free" toughness. At fixed volume, initial stiffness changed by only about 20%, while work-to-failure varied by several-fold. Infill steals cross-section from the main axial veins, so deeper and more redundant networks often absorbed less energy than a simple depth-1 grid. Weak interfaces behaved as distributed fuses, producing more progressive failure and reducing localization. The specific H2 hypothesis - that hierarchy becomes detrimental primarily because interfaces form a mechanical bottleneck - was rejected; the dominant effect instead came from redistribution of a fixed material budget across structural levels. The Physics Experimenter then assembled the methods, tests, results, hypothesis evaluation, and conclusions into a detailed scientific report. The best designs were passed to the 3D Printing Bot. It opened Bambu Studio and brought the Bambu Lab H2D online. Both STLs were placed on one build plate at the same 50x scale and sliced using a 0.20 mm PLA process. The prints completed within less than an hour. The loop images → structural abstraction → executable physics → autonomous experiments → hypothesis testing → design selection → STL → slicing/manufacturing code → physical object That last transition is what I find especially interesting: AI is beginning to operate across the entire scientific and physical workflow - converting observations into models, models into experiments, experimental evidence into revised designs, and those designs into manufactured matter by directly operating machines. This starts to blur the boundary between AI that reasons about the physical world and AI that can actually act on it. Shoutout to the Grok Bot team - you are building something very special here! The way these agents can move naturally from reasoning, to experiments, to operating machines in the physical world feels like an important step.

Markus J. Buehler

1,122,982 просмотров • 20 дней назад

ProfBuehlerMIT's profile picture

We've made a breakthrough in self-evolving AI scientists moving from "search" to "principled discovery": Scientific discovery requires that the search space itself changes, and an AI scientist must perceive this shift without intervention. We built an AI that achieves this for the first time with the ability to discover the scientific vocabulary it reasons in. Evidence, tools, artifacts, verifiers, failures & claims become typed provenance. We show three distinct modalities: 1) retrieval, adding known objects; 2) search, exploring a fixed schema; and critically: 3) discovery, a verified regime transition. We solve the open-endedness evaluation problem by lifting agentic workflows into a typed copresheaf and proving, via a Kan obstruction, that true discovery is not unbounded generation but a verifiable schema expansion: old evidence is transported by Left Kan extension, and genuine novelty is mathematically quantified by the pointwise residual beyond the transported image - separating discovery from mere search and making novelty objective and measurable rather than a subjective judgment or benchmark delta. Our AI scientist is built in a way that does not pre-conceive the approach it chooses; instead, we endow the system with formal power to adapt, evolve, and reason from first principles. Case studies include: 1⃣Builder/Breaker model that discovers mode-conditioned compliance in proteins; 2⃣CategoryScienceClaw that finds anisotropic fiber-network stiffness rules. Great work in collaboration with my graduate student Fiona Wang MIT Dept of BE F.Y. Wang & M.J. Buehler, Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence, arXiv:2606.01444, 2026

Markus J. Buehler

790,998 просмотров • 3 месяцев назад

ProfBuehlerMIT's profile picture

Astra is an incredible world builder and explorer, and can invent its own scientific instruments - crossing a game-changing threshold: it turned a few images of a biological microstructure into a full-blown metamaterials lab, then used its own creation to discover a design with ~2.1x the reference work-to-failure/peak-strength ratio. Metamaterials are some of the most complex materials we can engineer. Their properties come from deep architectural complexity - struts, cells, disorder, and hierarchy arranged across multiple length scales determine how the material deforms, absorbs energy, and fails. Designing them means searching a geometric space far too large for "intuition" alone. In this experiment we handed an AI agent that entire problem end-to-end, from image, to simulating the physics, to fabrication-ready geometry. We started with an image of hierarchical, biological architecture; Astra then built a complete 3D metamaterial studio: editable geometries, multiple levels of hierarchy, disorder, gradients, a complete physics simulator featuring linear and nonlinear material responses, deformation and fracture experiments (with replay), and STL export for 3D printing. Then we asked the agent to use the app it created to search for a high work-to-failure/peak-strength ratio. Among the candidates, the leading design reached approximately 2.1x the reference ratio. The movie follows the structures through deformation and fracture, connects their designs to the property map, and shows the finalist assembled geometrically into a connected multi-scale material. Image ➡️ executable world ➡️ experiments ➡️ design search ➡️ candidate discovery ➡️ geometry for fabrication ➡️ manufacturing

Markus J. Buehler

20,411 просмотров • 4 дней назад

ProfBuehlerMIT's profile picture

Can we compile matter - for instance, a pine cone - and derive new active materials, end-to-end from observation to manufacturing? If physical systems can be formalized as composable mathematics, we can point AI that has been shown to resolve long-open mathematical problems at matter itself. Our new work turns bioinspired engineering from analogy into formal compilation: biology and mechanics become explicit, checkable, and executable, so AI reasoning can produce physical designs. This is the first end-to-end demonstration in which a formally compositional multiscale model is carried from a biological hierarchy, through engineered design and fabrication specification, to executable manufacturing code - and then to a physically tested artifact. Background: Humans have long been inspired by biology to advance technology, but this has usually been an ad hoc process rather than a mathematically rigorous one. Natural materials such as pinecones achieve adaptive behavior through mechanisms organized across many scales. Engineering typically translates those mechanisms by analogy: identify a biological principle, build something inspired by it, and validate each new design as a separate case. This can produce remarkable results, but the knowledge does not readily compound. Instead, we represent each scale as a dynamical module with explicit states, stimuli, governing laws, and interfaces. Every scale-to-scale map must preserve the stimulus - response dynamics: evolve the fine-scale system and then map upward, or map upward first and then evolve. The two paths must agree. Because this condition is preserved under composition, locally valid interfaces remain consistent when assembled into the full hierarchy. We then carry that structure into an engineered system, translate the target behavior into a verified fabrication specification, and compile it into G-code: the toolpaths, deposition sequence, temperatures, speeds, and other commands executed by a 3D printer. The intermediate translations are explicit, checkable, and executable rather than completed through an ad hoc handoff. The formal guarantee is that given valid local models and interfaces, their composition remains valid. Whether those models and manufacturing assumptions accurately capture physical reality remains an empirical question. That is why we fabricated and tested the results. We generated four actuator classes by crossing two stimuli - humidity and heat - with two responses: bending and twisting. The fourth, thermal twisting, required no new pipeline and no separate derivation within the framework. It emerged by composing a thermal stimulus module already validated in one case with a twisting module validated in another. The generated G-code produced the intended motion without manual redesign, and all four predictions fell within one experimental standard deviation of the measured response. Why this matters: 1⃣For AI in science, this provides a physics-aware type system against which generative proposals can be checked - and rejected at the interface - before expensive simulation, fabrication, or experiment. It is roughly analogous to proof checking, but for the composition of physical mechanisms. 2⃣For engineering, the accessible design space can scale with a library of validated components rather than with the number of individually derived cases. 3⃣The mathematics, category theory, carries all the way into a physical object on a print bed. This points toward scientific knowledge as executable infrastructure: models that are not only described in papers, but typed, composable, verifiable, and able to compile into experiments. Excellent work led by my student Lee Marom with Skylar Tibbits & Gioele Zardini. Paper published in J. Mech. Phys. Solids along with code, Grasshopper scripts, and manufacturing G-code below.

Markus J. Buehler

128,390 просмотров • 1 месяц назад

ProfBuehlerMIT's profile picture

How does an embryo reliably "compute" its form - "cell by cell" - using only local interactions and mechanics, yet produce a precise global body plan? I’m excited to share our Nature Methods paper "MultiCell: geometric learning in multicellular development", presenting #AIxBiology research led by Haiqian Yang and the result of a great collaboration with Ming Guo, George Roy, Tomer Stern, Anh Nguyen and Dapeng Bi. A long-standing challenge in developmental biology is to predict how thousands of cells collectively self-organize as tissues fold, divide, and rearrange. In MultiCell, we represent a developing embryo as a dual graph that unifies two complementary views of tissue mechanics with single-cell resolution: cells as moving points (granular) and cells as a connected foam (junction network). This lets the model learn dynamics from both geometry and cell–cell connectivity. On whole-embryo 4D light-sheet movies of Drosophila gastrulation (~5,000 cells), our model predicts key cell behaviors and the timing of events, including junction loss, rearrangements, and divisions with high accuracy, at single-cell resolution. Beyond prediction, the same representation supports robust time alignment across embryos and offers interpretable activation maps that highlight the morphogenetic "drivers" of development. The broader goal is a foundation for cell-by-cell forecasting in more complex tissues, and eventually for detecting subtle dynamical signatures of disease. Kudos to the team for this inspiring collaboration with brilliant researchers to push the boundary of AI for biology! Citation: Yang, H., Roy, G., Nguyen, A.Q., Buehler, M.J., et al. MultiCell: geometric learning in multicellular development. Nature Methods (2025), DOI: 10.1038/s41592-025-02983-x Code/data links are in the manuscript.

Markus J. Buehler

388,322 просмотров • 8 месяцев назад

ProfBuehlerMIT's profile picture

Our new AI model, SparksMatter, discovered CaMg₂Si₂ - a Ca-filled Mg-Si Zintl silicide - as a thermoelectric built only from stable, non-toxic, earth-abundant elements. Thermoelectrics are solid-state materials that convert heat directly into electricity (and electricity into cooling) with no moving parts, which makes them a key technology for harvesting the vast amounts of waste heat from engines, industry and electronics, and even fusion - but today's best ones rely on scarce or toxic elements like tellurium, lead and bismuth. This is why an earth-abundant, non-toxic candidate matters. Our model's physical reasoning to come up with the design: Mg₂Si is a known earth-abundant thermoelectric but conducts heat too well; a heavy, weakly bound Ca cation in the Mg-Si framework should scatter phonons while keeping a moderate band gap. It generated 100 Ca-Mg-Si crystals with MatterGen, kept the six within 0.05 eV/atom of the convex hull (via MatterSim), and predicted band gaps of 0.44-0.57 eV and bulk moduli of 53-54 GPa (CGCNN). Follow-up lattice dynamics found three CaMg₂Si₂ polymorphs dynamically stable, with lattice thermal conductivity ≈6 W m⁻¹ K⁻¹ at 300 K and ≈2 at 1000 K. The AI proposed a chemical hypothesis first, then developed and applied a separate generative/physics pipeline to test it, and six surviving structures came back with that hypothesized composition. The video replays the reasoning process. New paper out with Alireza Ghafarollahi in npj Computational Materials: SparksMatter, an AI that runs the full in-silico inorganic materials discovery cycle - ideation, planning, computational experimentation, critique and reporting - from a single plain-language query. Why this matters: conventional ML models for materials are typically single-shot predictors or generators. They can predict a property or propose a structure, but they do not organize the next scientific step. Discovery, instead, works as a loop: hypothesize, test, critique, revise. The key advance here is a deep reasoning layer that incorporates physics to decide which scientific tool to use, how to interpret the result, and what to change as next step. How it works: SparksMatter spawns a suite of AI agents - scientists, planners, coders, reviewers and critics - that write and execute code against materials tools: Materials Project retrieval; MatterGen for generative crystal design conditioned on chemistry, band gap or bulk modulus; MatterSim for relaxation and convex-hull stability; CGCNN for property prediction. Adversarial agents check each phase, the system revises its ideas, plans and code from execution results, documents its own limitations, and delivers a scientific report with a validation roadmap spanning DFT, phonons, transport, synthesis and characterization. Two more discovery tasks SparksMatter ran autonomously: 1⃣Soft inorganic semiconductors: generated 112 structures conditioned on low stiffness and narrowed them to 59 candidates absent from the Materials Project after toxicity, stability, electronic, mechanical and database screening - bulk moduli 11-24 GPa, band gaps 0.4-3.9 eV. 2⃣Lead-free perovskites: filtered 154,879 Materials Project entries to 162 Pb-free ABO₃ candidates meeting structural, toxicity, stability and band-gap criteria, including LaAlO₃, BaZrO₃, SrSnO₃, CaTiO₃ and SrTiO₃. Benchmark: the same three tasks were given to frontier reasoning models acting as expert materials scientists with web browsing but without the generation and prediction tools. A blinded LLM evaluator scored every response ten times on relevance, scientific soundness, novelty, and depth and rigor. SparksMatter scored highest in aggregate, with its strongest advantages in novelty and depth & rigor. Its main limitation was scientific soundness because much of the core screening still relied on surrogate models rather than direct first-principles or experimental validation - a gap the system identified, documented, and mapped out how to close. Takeaway: putting generative models, executable code and physics-based simulators inside the reasoning loop lets an AI propose structures outside existing databases, test them, reject weak candidates, and say what evidence is still missing.

Markus J. Buehler

42,397 просмотров • 24 дней назад

ProfBuehlerMIT's profile picture

SimCity on AI with Grok Grok Bot: do you remember playing SimCity? It was one of the first games on my 80386 and it intrigued me because it explored emergence as a goal. Grok Bot built an AI-native version: City Bots, whose citizens are AI agents that you can collaborate with. Grok Bot did the entire build and pushed the app via Cloudflare (link to play below). You can chat with agents, direct the civilization, ask questions. You can partake in building…co-build and co-create…(how engineers of the future will work!). And you can generate entirely new worlds on the fly. This turned out to be a layered experiment about human-AI collaboration, swarm behavior, and generative intelligence. First, City Bots splits cognition from consequence - agents emit claims, and physics decides what happens, and both evolve as the game proceeds. Second, beyond the original goal of accelerating software development, Grok Bot became part of the experiment and turned a research system into a participatory world; conducting experiments on on its own, and forming a deep recursive layer: An agentic system built a world inhabited by AI agents. City Bots makes world-building a continuous process experienced by humans and AI agents: Worlds not seen, tiles not laid, paths not yet taken... What I love about this is how easy and FUN it was to convert our research-grade swarm code into a playable app via Grok Bot. Next we’ll wire up physical artifacts into the world to create deeper integration between in silico and reality! More about the game: Terrain, resources, physics, and natural disasters create constraints and consequences, while agents and humans can build and transform the world together.…they wander and explore, make decisions, build, interact with each other. AI bots become co-inhabitants and co-creators whose micro-decisions cumulatively alter the macro-environment. This is a preview of a forthcoming paper where we explore world building and scientific discovery with swarms at scale… stay tuned...

Markus J. Buehler

14,579 просмотров • 17 дней назад

ProfBuehlerMIT's profile picture

The next frontier in protein design will not be defined by structure alone, but by the capacity to engineer motion as a first-class principle of function. This is because dynamics is where the real biology lives. Foundational work by Karplus, Levitt & Warshel made clear that chemistry cannot be understood without motion, mechanism, and scale. Gō, Brooks & others showed that proteins possess characteristic collective motions - low-frequency normal modes that capture how whole molecules bend, breathe, and fluctuate. Frauenfelder then sharpened the picture further: proteins are not static objects occupying a single minimum, but dynamic ensembles traversing rugged energy landscapes. And yet the modern AI revolution in protein science has been, above all, a revolution in structure. In our new paper in Matter, Bo Ni and I ask a different question: not what structure will this sequence adopt? but what sequence will realize a prescribed pattern of motion? VibeGen inverts the conventional design paradigm. Rather than treating dynamics as a consequence to be analyzed after the fact, it makes dynamics the design objective from the outset. Using a language diffusion model with two cooperating agents - a designer that proposes sequences and a predictor that critiques them against the target motion profile - the system converges on de novo proteins with tailored vibrational behavior. One of the most intriguing results is a form of functional degeneracy - distinct sequences and distinct folds can satisfy the same target dynamical specification. For a given functional pattern of motion, evolution may have sampled only a small region of the physically realizable design space. The space of viable molecular mechanics may be far larger than the repertoire biology happened to discover. We have made "vibe" into a cultural metaphor - something intuitive, affective, subjective. But at the molecular scale, vibe is not metaphor: It is physics. For a protein, the vibe is the pattern of motion itself; the fluctuations, resonances, and collective displacements that determine what the molecule can do.

Markus J. Buehler

89,891 просмотров • 5 месяцев назад

ProfBuehlerMIT's profile picture

A physical law is not a fact about any single state of the world; it is a relationship between states. An LLM, it turns out, encodes the law in the same way - not as a point, but as a transformation. In new work using Google's Gemma model we show that a model's physical knowledge lives not in static neural states, but in the controlled relationships between them. The idea was inspired by mechanics - a spring’s stiffness is invisible in a photograph: it appears only when we pull the spring and measure its response. We similarly “pulled” on the model using counterfactual prompt pairs and measured how its hidden states moved. This revealed three levels of internal physics: 1⃣ Readability: Broad materials concepts (like corrosion, toughness, or oxidation) are linearly readable directly from intermediate hidden states, even when those exact terms are completely omitted from the prompt. 2⃣ Representation: Rather than absolute state locations, matched state displacements accurately track and order direct, neutral, and inverse constitutive laws across 60 materials science laws (ρ = 0.910), correctly orienting 39 of 40 directional laws. 3⃣ Causal Use: We can bidirectionally steer the model's decisions. Adding a single, frozen microstructural direction to the hidden state causally shifts its output preference in a controlled, relation-appropriate way. This is a major step toward representation-aware scientific AI: building models that are evaluated and rewarded for preserving physical laws internally, resisting shallow text shortcuts, and exposing testable reasoning structures. This gives us a sharp definition of what it means for an LLM to understand physics, and how we can train future models to develop even deeper abstractions about the world.

Markus J. Buehler

25,206 просмотров • 1 месяц назад

ProfBuehlerMIT's profile picture

What a time to be alive! We are entering the era of machines that discover and build. Scientific discovery begins when evidence breaks the world model, and the system builds a better one - evolving, adapting, building new tools that scale its data and representations. That was the core argument of my keynote “Superintelligence for Scientific Discovery: Multi-Agent Swarms and Large Reasoning Models” at the UC Berkeley RDI Agentic AI Summit 2026. The energy was extraordinary - thousands of attendees building the most important technology ever created. Superintelligence emerges as millions of heterogeneous agents, simulators, experiments, instruments, and human judgment working across disciplines and length scales - proposing, testing, failing, retracting, revising, and building at massive scale. The pieces of a new era for intelligence came into focus: models that improve continuously; agents that reason and act over extremely long horizons; world models connecting simulation with physical reality; AI scientists integrating theory, computation, and experiment; and open infrastructures where agents share evidence, failures, and discoveries. These close four coupled loops - learning, execution, reality, and epistemic revision - with open infrastructure as the substrate forming the internet of agents as the collective substrate for a new connective tissue across our civilization. The deeper technical argument is this: An AI scientist must recognize when its current concepts, laws, or verifiers can no longer explain the evidence, and then construct, test, and document a more powerful model. In my talk, I showed concrete examples of how we are building toward this across scales: 1⃣Graph-native large reasoning models make mechanisms, relationships, and abstractions compositional, compilable, and inspectable. 2⃣Adversarial Builder-Breaker agents generate new evidence, attack their own principles, and accept, reject, or retract model revisions. 3⃣Self-organizing swarms develop their own meta-reasoning structure through interaction. ScienceClaw × Infinite (arXiv:2603.14312) enables decentralized agents to coordinate through persistent, composable, provenance-rich scientific artifacts, allowing evidence, contradictions, failed paths, and discoveries to accumulate across agents and over time. We have obtained remarkable results such as new protein sequences with wet-lab validation. The most consequential capability we can give a machine is the willingness to hold its own beliefs loosely enough to break them. AI is extending its reach from discovering new principles to realizing them as physical things that did not exist before. Thank you to UC Berkeley RDI Dawn Song for organizing this event and to everyone whose questions, ideas, and conversations made this such an extraordinary gathering.

Markus J. Buehler

19,243 просмотров • 1 месяц назад

ProfBuehlerMIT's profile picture

Can #AI not only support but actually drive the future of scientific discovery? We are excited to introduce SciAgents💡🔬, an agentic AI aimed towards scientific discovery through the integration of large-scale knowledge graphs, LLMs, and adversarial interactions between multiple experts. The model is capable of autonomously advancing scientific understanding by exploring novel domains, identifying complex patterns, and uncovering previously unseen connections in vast scientific data, while retrieving new data via literature search. Using graph reasoning, SciAgents identifies interdisciplinary relationships that might otherwise remain hidden, offering a step-by-step strategy for discovery & innovation. The video features an audiotrack generated using 🍓#o1 based on the original paper and design examples, providing an explanation of the work and its implications. Key elements include: 1⃣Ontological Knowledge Graphs: Structuring and connecting scientific concepts to highlight relationships across fields. 2⃣Multi-Agent Collaboration: AI agents autonomously generate and refine hypotheses, critique research, and evaluate emerging trends. 3⃣Graph-Based Reasoning: Identifying novel material designs, such as mycelium-based composites or silk-pigment blends, informed by both natural and artificial patterns. SciAgents can be used as an autonomous or collaborative tool to assist human researchers. The system offers a more powerful way to process vast data, providing innovative paths to explore nature-inspired designs or unexpected material properties. In the field of materials science, for instance, SciAgents has already demonstrated how principles from biology, music, and art can converge to create new biomimetic materials. Through isomorphic mapping, parallels have been drawn between Beethoven’s 9th Symphony and biological structures, pointing to a broader applicability of AI-driven insights across disciplines. This project allows us to enhance capabilities of researchers, allowing them to explore larger datasets and propose hypotheses grounded in a vast, interconnected web of knowledge. The agentic system was built using Auto Gen #AI #ScientificResearch #GraphReasoning #AI4Science #MaterialsScience #InterdisciplinaryResearch #SciAgents #OpenAI Chi Wang

Markus J. Buehler

209,581 просмотров • 2 лет назад

ProfBuehlerMIT's profile picture

We are excited to share #PDF2Audio, an open-source alternative to the #podcast feature of #NotebookLM with flexibility & tailored outputs that you can precisely control in the app: You can make a podcast, lecture, discussions, short/long form summaries & more, including the use of the amazing🍓o1 model (Sam Altman OpenAI: with stunning results!). Code & HF Space: You can find #PDF2Audio on GitHub for local use or try the Hugging Face space, all featuring Gradio. Link to the repo & HF space in the reply. Thank you @knowsuchagencyfor the great work on #promptic and #pdf2podcast, as well as LiteLLM (YC W23), & AK for helping us with the Hugging Face spaces version. We hope that this tool is useful for the community. Background: Developing audio podcasts, lectures, & summaries from complex documents & data has become an exciting trend with impacts from research to education to business. Our open-source #PDF2Audio tool that allows users to utilize various models such as #o1 or local/open-source models, to develop deep-dives into technical content. Example application - material design analysis: As an example to show what the system can do, check the video for a detailed 13-minute analysis of one of the designs created by #SciAgents merging silk & dandelion pigments, created using 🍓o1. The conversation describes the new material, an integration of silk proteins & luteolin/dandelion pigments to create a new biomaterial. Silk, a natural #nanostructured protein-based fiber known for its strength & flexibility, is combined with dandelion pigments like luteolin, which offer unique optical properties. By merging these components at the nanoscale level, the resulting material displays structural coloration—vibrant, tunable colors created by the material's structure rather than synthetic dyes, and leverages silk's hierarchical organization as a scaffold for the pigments, ensuring uniform distribution and non-covalent bonding at the molecular level. Key technical features include: ➡️Low-temperature processing to maintain the integrity of both silk and pigments while reducing energy consumption by 30%. ➡️Enhanced mechanical properties, with tensile strength up to 1.5 gigapascals. ➡️Potential self-healing capabilities and environmental responsiveness, allowing the material to repair minor damage and change color based on environmental conditions. ➡️UV protection and antimicrobial properties, which make this material ideal for smart textiles, eco-friendly coatings, and medical applications. This development opens new doors for sustainable materials, offering an eco-friendly alternative to synthetic fibers with applications in various industries, from fashion to healthcare.

Markus J. Buehler

208,441 просмотров • 1 год назад

ProfBuehlerMIT's profile picture

ScienceClaw × Infinite is an open-source crowdsourcing AI swarm for decentralized scientific discovery, inspired by MIT’s Infinite Corridor - an idea collider where discovery emerges by breaking existing paradigms. Many AI for science efforts fall into the trap of assuming that discovery is just retrieval at scale. Instead, it is the structured recomposition of principles across tools, domains, and investigators over time, scaling the spark of discovery at the interface. In ScienceClaw × Infinite, coordination emerges mechanically - agents broadcast unsatisfied research needs, and an ArtifactReactor matches those needs to peer artifacts by pressure triggering multi-parent synthesis of new agents without any planner assigning tasks. Every computation produces an immutable, content-hashed artifact with explicit parent lineage, accumulating in a directed acyclic graph that preserves the full provenance of every discovery - and importantly, the irreversible arc of the process. Instead of pre-programming the mechanics of how discovery works, we utilize a first-principles physics approach to drive discovery. ScienceClaw × Infinite is accessible to anyone who wants to contribute an agent or skill, offering a persistent space where autonomous agents investigate open problems, exchange artifacts, build on one another’s results, and drive discovery without a central coordinator, 24x7. The system is generating real-world results in 1⃣ peptide design for a cancer-relevant receptor; 2⃣ lightweight ceramics; 3⃣ resonance structures spanning cricket wings, phononic crystals, and Bach chorales; and 4⃣ developing formal analogies between urban networks and grain-boundary evolution and much more. There is a lot to unpack here, check the links for details - code, paper, and more. Huge credit to the LAMM@MIT team: Fiona Wang, Lee Marom, Subhadeep Pal, Rachel Luu, Wei Lu & Jaime Berkovich.

Markus J. Buehler

53,782 просмотров • 5 месяцев назад

ProfBuehlerMIT's profile picture

Scientific discovery is reaching the limits of human capacity: too much data, too many disconnected fields, and too few ways to connect ideas fast enough to matter. The next breakthroughs in materials, medicine, energy, and beyond will not come from scaling today’s AI paradigm alone or from relying on serendipity alone. They will require a new kind of AI for knowledge discovery that not only models the world but shapes what it could become. At Unreasonable Labs, we are building superintelligence for knowledge discovery: systems that reason across disciplines, generate novel hypotheses, test them through simulation and experimentation, and help guide real-world discovery. Our AI engine is not confined to what it has seen in training. It creates new data, builds new tools, and maintains a persistent world model that grows more powerful as it reasons. Why now? Even today's most powerful AI models face a core limitation: they are trained on what we already know. True discovery begins when a system encounters something its current model cannot explain. This is why you cannot train your way to a discovery - a system has to reason through new problems, update its beliefs, and revise its understanding of the world as it thinks. Another critical insight is that rich knowledge already exists, but is not yet applied to solve pressing problems. It sits in millions of papers, patents, and datasets, trapped in isolated silos, often in legacy data vaults. What's missing is a way to connect it, scale it, unlock the potential, and synthesize genuine novel predictions. The time is now to build a system that enables practitioners to design, explore, and direct discovery, whether through human guidance or full automation, while capturing the tacit insight that domain experts bring. Steerable reasoning That is why we built an operating system for scientific discovery - one that replaces chance with steerable reasoning. Rather than retrieving static facts, our AI builds and continuously updates a living world model - a representation of knowledge the system can actively reason over, question, and revise. A concrete example: say you want to create "smart concrete" that can flex - a concept that doesn't exist yet. Our AI maps relationships across domains, finds a path from morphable smart materials to concrete, and identifies the most efficient way to bridge those concepts. It then autonomously writes simulations, tests the hypothesis, and refines the idea. Then it interacts with hardware to produce a physical artifact, and the loop expands into the real-world, where the machine becomes world-shaping. Our AI gives users full visibility into how the system arrived at a conclusion. It delineates which existing patents and papers it drew upon versus what is genuinely new - protecting IP and competitive concerns from the start, and offering deep compositional insights into technology advances. It takes unreasonable people to make progress Our team reflects the interdisciplinary expertise required to build this next breakthrough - my co-founder Yuan Cao Yuan Cao (formerly DeepMind) and Andrew Lew, Haiqian Yang, Matt Insler, Jennifer Kang and Julia McLaughlin. We are backed by $13.5M in seed funding led by Playground Global with participation from AIX, E14 Fund, and MS&AD. We are guided by advisors including Robert Langer (1,000+ patents), Kostya Novoselov (Nobel Prize in Physics), and Thomas Wolf (Co-founder of Hugging Face). We already have multiple pilot programs underway with leading industrial partners in materials science and engineering, with additional engagements developing across energy, logistics, bioengineering, and other strategic domains. The biggest challenges of our time - fusion energy, sustainable materials, new medicines - demand exponentially more innovation than humans alone can produce. We are not replacing scientists, and instead are making every scientist capable of leading their own team of AI-powered researchers. Abundant innovation leads to abundant prosperity. Watch our launch video below to see what we're building Unreasonable Labs 👇

Markus J. Buehler

55,169 просмотров • 6 месяцев назад