Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Sharing our work at NeurIPS Conference on reasoning with EBMs! We learn an EBM over simple subproblems and combine EBMs at test-time to solve complex reasoning problems (3-SAT, graph coloring, crosswords). Generalizes well to complex 3-SAT / graph coloring/ N-queens problems.

48,030 görüntüleme • 11 ay önce •via X (Twitter)

22 Yorum

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

To solve complex reasoning tasks, our approach formulates the reasoning task as optimizing the summed energy functions over EBMs learned over subproblems of the task.

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

To effectively optimize this composed energy function, we propose a parallel optimization process where we jointly optimize a set of particles at once.

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

Our approach is able to generalize well to complex reasoning tasks -- outperforming specialized SAT solving methods on 3-SAT problems such as NeuroSAT and NSNet.

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

More details and illustrations can be found at the webpage here: and code at

billy bubba bingus-drumpfus profil fotoğrafı
billy bubba bingus-drumpfus11 ay önce

@NeurIPSConf Brøther u know I love this kinda work. Seems like you focused on discrete problems here. Did you explore at all how the continuous relaxation you used impacted the performance of learning or sampling ?

Jie Wang profil fotoğrafı
Jie Wang11 ay önce

@NeurIPSConf Reasoning with EBMs seems very applicable to robotics!

Etienne profil fotoğrafı
Etienne11 ay önce

Love this work. One question: how does this handle noise filtering tasks where most input is irrelevant? Your energy composition excels at structured reasoning (SAT, graph coloring). But what about tasks requiring selective attention - finding rare signal in long noisy sequences? Different architectural principle needed there?

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

@NeurIPSConf Good question! In this setting, instead of optimizing the sum othe energy function, you could instead optimize the softmin or K lowest energy values. This would allow you to selectively attend to rare signals.

Peyman profil fotoğrafı
Peyman10 ay önce

@NeurIPSConf How does the optimization work? And what size problems can it solve?

Yilun Du profil fotoğrafı
Yilun Du10 ay önce

@NeurIPSConf We run gradient descent on a set of particles and periodically resample them. The method works well for problems with 8-20 models composed.

Peyman profil fotoğrafı
Peyman10 ay önce

@NeurIPSConf Hi thanks — by model composed, do you mean sub-problems? Why can’t it be generalized to any combinatorial problem? Specially if you have thousands of various boolean constraints?

Yilun Du profil fotoğrafı
Yilun Du10 ay önce

@NeurIPSConf Yes, subproblems. You can in principle also have thousands of constraints -- the optimization problem across constraints will just be hard then.

Samip profil fotoğrafı
Samip11 ay önce

@NeurIPSConf any thoughts on how the compositions of energy functions could be learned instead of hardcoding? that seems important outside of combinatorial optimization (for example, if you wanna model natural language)

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

@NeurIPSConf Yes, you could definitely learn the composition! Each energy function could be conditioned by a latent inferred by an encoder, which could then allow you to learn the composition

Samip profil fotoğrafı
Samip11 ay önce

@NeurIPSConf interesting, do you mean using those latents from an encoder as a soft routing among (or some weighted combination of) energy functions?

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

@NeurIPSConf Yup!

Samip profil fotoğrafı
Samip11 ay önce

@NeurIPSConf okay so similar to MoE but each expert has an explicit energy function? curious how end to end training would work when you have multiple energy functions at different levels, is there any paper that does this?

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

@NeurIPSConf Yes, it would be something like or though all energy functions are at the same level.

Akshay profil fotoğrafı
Akshay11 ay önce

@NeurIPSConf Interesting work. Curious, do you have a sense of how the difficulty of training the EBMs on subproblems scales on tasks with larger input spaces?

Yilun Du profil fotoğrafı
Yilun Du11 ay önce

@NeurIPSConf It depends on how difficult the subproblem is -- the harder the subproblem is, the harder it is to train the EBM. The input space size doesn't matter that much.

ruffian-l (Jason Van Pham) profil fotoğrafı
ruffian-l (Jason Van Pham)10 ay önce

@NeurIPSConf NIODOO's TQFT acts as a reasoning engine for transitions, while EBMs focus on state-based energy landscapes. They complement each other, with potential synergies in hybrid models.

truth.phd profil fotoğrafı
truth.phd10 ay önce

Energy-Based Models sound fancy, but think of them like chefs balancing flavors; they learn which combinations of features taste right for the task. Additive EBMs? That’s when the chef adds each ingredient thoughtfully, so the dish can reason about each flavor instead of throwing everything into one mystery stew.

Benzer Videolar

Everyone is focused on tracking the ways LLMs are getting better. And they are. But we know there are still things that LLMs can’t do well—the tasks where you can feel the architecture fighting the problem. So I was excited to chat with Eve Bodnia (@eve_bodnia), who is developing an alternative AI model to LLMs, on Every 📧's AI & I. Eve's argument: energy-based models (EBMs), which map possible outcomes onto a mathematical landscape, will lead to the next AI phase shift. We get into: - How energy-based models work. Likely outcomes sit in valleys, and unlikely ones sit on peaks. Whereas LLMs process one token at a time, an EBM scans the full terrain to find the lowest point, or the most probable answer. - Language-based versus data-native models. LLMs are language-dependent even when the problem has nothing to do with language. "If your data is numbers, relationships, and functions, and you try to map those rules into words and then search for the next word, you're losing a lot of information," Bodnia says. EBMs work directly with the underlying data structure, including numbers and spatial coordinates. - Sequential versus panoramic reasoning. An LLM is like driving through San Francisco without a map. Each turn constrains the next, and if you go down the wrong street, you can't reverse course. An EBM has the bird's-eye view—it can evaluate multiple routes at once and course-correct before hitting a dead end. - The LLM plateau no one wants to talk about. LLMs are getting incrementally better, step-change improvements aren’t coming, Eve argues. To achieve that, we need new solutions that compensate for what LLMs are inherently bad at, like non-language reasoning, verification, and real-time data analysis. This is a must-watch for anyone who's curious what might come after the LLM. Watch below! Timestamps: Introduction: 00:00:51 Why correctness and verifiability matter in AI: 00:02:09 What an energy-based model is: 00:09:33 How EBMs construct energy landscapes to understand data: 00:14:21 Why modeling intelligence through language alone is a flawed approach: 00:19:00 What it means for a model to "understand" data: 00:26:54 How EBMs solve the vibe coding problem and enable formally verified code: 00:37:21 Why LLM progress is plateauing: 00:43:21 Mission-critical industries haven't adopted LLMs, and why EBMs can fill that gap: 00:49:54

Dan Shipper 📧

26,900 görüntüleme • 5 ay önce