Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Sharing our work at NeurIPS Conference on reasoning with EBMs! We learn an EBM over simple subproblems and combine EBMs at test-time to solve complex reasoning problems (3-SAT, graph coloring, crosswords). Generalizes well to complex 3-SAT / graph coloring/ N-queens problems.

48,030 Aufrufe • vor 11 Monaten •via X (Twitter)

22 Kommentare

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

To solve complex reasoning tasks, our approach formulates the reasoning task as optimizing the summed energy functions over EBMs learned over subproblems of the task.

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

To effectively optimize this composed energy function, we propose a parallel optimization process where we jointly optimize a set of particles at once.

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

Our approach is able to generalize well to complex reasoning tasks -- outperforming specialized SAT solving methods on 3-SAT problems such as NeuroSAT and NSNet.

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

More details and illustrations can be found at the webpage here: and code at

Profilbild von billy bubba bingus-drumpfus
billy bubba bingus-drumpfusvor 11 Monaten

@NeurIPSConf Brøther u know I love this kinda work. Seems like you focused on discrete problems here. Did you explore at all how the continuous relaxation you used impacted the performance of learning or sampling ?

Profilbild von Jie Wang
Jie Wangvor 11 Monaten

@NeurIPSConf Reasoning with EBMs seems very applicable to robotics!

Profilbild von Etienne
Etiennevor 11 Monaten

Love this work. One question: how does this handle noise filtering tasks where most input is irrelevant? Your energy composition excels at structured reasoning (SAT, graph coloring). But what about tasks requiring selective attention - finding rare signal in long noisy sequences? Different architectural principle needed there?

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

@NeurIPSConf Good question! In this setting, instead of optimizing the sum othe energy function, you could instead optimize the softmin or K lowest energy values. This would allow you to selectively attend to rare signals.

Profilbild von Peyman
Peymanvor 10 Monaten

@NeurIPSConf How does the optimization work? And what size problems can it solve?

Profilbild von Yilun Du
Yilun Duvor 10 Monaten

@NeurIPSConf We run gradient descent on a set of particles and periodically resample them. The method works well for problems with 8-20 models composed.

Profilbild von Peyman
Peymanvor 10 Monaten

@NeurIPSConf Hi thanks — by model composed, do you mean sub-problems? Why can’t it be generalized to any combinatorial problem? Specially if you have thousands of various boolean constraints?

Profilbild von Yilun Du
Yilun Duvor 10 Monaten

@NeurIPSConf Yes, subproblems. You can in principle also have thousands of constraints -- the optimization problem across constraints will just be hard then.

Profilbild von Samip
Samipvor 11 Monaten

@NeurIPSConf any thoughts on how the compositions of energy functions could be learned instead of hardcoding? that seems important outside of combinatorial optimization (for example, if you wanna model natural language)

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

@NeurIPSConf Yes, you could definitely learn the composition! Each energy function could be conditioned by a latent inferred by an encoder, which could then allow you to learn the composition

Profilbild von Samip
Samipvor 11 Monaten

@NeurIPSConf interesting, do you mean using those latents from an encoder as a soft routing among (or some weighted combination of) energy functions?

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

@NeurIPSConf Yup!

Profilbild von Samip
Samipvor 11 Monaten

@NeurIPSConf okay so similar to MoE but each expert has an explicit energy function? curious how end to end training would work when you have multiple energy functions at different levels, is there any paper that does this?

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

@NeurIPSConf Yes, it would be something like or though all energy functions are at the same level.

Profilbild von Akshay
Akshayvor 11 Monaten

@NeurIPSConf Interesting work. Curious, do you have a sense of how the difficulty of training the EBMs on subproblems scales on tasks with larger input spaces?

Profilbild von Yilun Du
Yilun Duvor 11 Monaten

@NeurIPSConf It depends on how difficult the subproblem is -- the harder the subproblem is, the harder it is to train the EBM. The input space size doesn't matter that much.

Profilbild von ruffian-l (Jason Van Pham)
ruffian-l (Jason Van Pham)vor 10 Monaten

@NeurIPSConf NIODOO's TQFT acts as a reasoning engine for transitions, while EBMs focus on state-based energy landscapes. They complement each other, with potential synergies in hybrid models.

Profilbild von truth.phd
truth.phdvor 10 Monaten

Energy-Based Models sound fancy, but think of them like chefs balancing flavors; they learn which combinations of features taste right for the task. Additive EBMs? That’s when the chef adds each ingredient thoughtfully, so the dish can reason about each flavor instead of throwing everything into one mystery stew.

Ähnliche Videos

Everyone is focused on tracking the ways LLMs are getting better. And they are. But we know there are still things that LLMs can’t do well—the tasks where you can feel the architecture fighting the problem. So I was excited to chat with Eve Bodnia (@eve_bodnia), who is developing an alternative AI model to LLMs, on Every 📧's AI & I. Eve's argument: energy-based models (EBMs), which map possible outcomes onto a mathematical landscape, will lead to the next AI phase shift. We get into: - How energy-based models work. Likely outcomes sit in valleys, and unlikely ones sit on peaks. Whereas LLMs process one token at a time, an EBM scans the full terrain to find the lowest point, or the most probable answer. - Language-based versus data-native models. LLMs are language-dependent even when the problem has nothing to do with language. "If your data is numbers, relationships, and functions, and you try to map those rules into words and then search for the next word, you're losing a lot of information," Bodnia says. EBMs work directly with the underlying data structure, including numbers and spatial coordinates. - Sequential versus panoramic reasoning. An LLM is like driving through San Francisco without a map. Each turn constrains the next, and if you go down the wrong street, you can't reverse course. An EBM has the bird's-eye view—it can evaluate multiple routes at once and course-correct before hitting a dead end. - The LLM plateau no one wants to talk about. LLMs are getting incrementally better, step-change improvements aren’t coming, Eve argues. To achieve that, we need new solutions that compensate for what LLMs are inherently bad at, like non-language reasoning, verification, and real-time data analysis. This is a must-watch for anyone who's curious what might come after the LLM. Watch below! Timestamps: Introduction: 00:00:51 Why correctness and verifiability matter in AI: 00:02:09 What an energy-based model is: 00:09:33 How EBMs construct energy landscapes to understand data: 00:14:21 Why modeling intelligence through language alone is a flawed approach: 00:19:00 What it means for a model to "understand" data: 00:26:54 How EBMs solve the vibe coding problem and enable formally verified code: 00:37:21 Why LLM progress is plateauing: 00:43:21 Mission-critical industries haven't adopted LLMs, and why EBMs can fill that gap: 00:49:54

Dan Shipper 📧

26,900 Aufrufe • vor 5 Monaten