Video yรผkleniyor...
Video Yรผklenemedi
๐ Introducing ๐๐ช๐ฎ๐ข๐ฅ๐ข๐๐ซ๐ข๐ฎ๐ฆ ๐๐๐๐ฌ๐จ๐ง๐๐ซ๐ฌ (๐๐ช๐) ! Feedforward models and weight-tied models behave very differently on hard reasoning generalization. EqR pushes this difference to the extreme by learning ๐ญ๐๐ฌ๐ค-๐๐จ๐ง๐๐ข๐ญ๐ข๐จ๐ง๐๐ ๐ง๐๐ฎ๐ซ๐๐ฅ ๐๐ญ๐ญ๐ซ๐๐๐ญ๐จ๐ซ๐ฌ . โข Sudoku-Extreme: 99.8% โข Maze: 93% #ICML2026
104,646 gรถrรผntรผleme โข 4 ay รถnce โขvia X (Twitter)
31 Yorum

2/ Standard feedforward transformers barely generalize: 4 / 16 / 64 layers โ 1.8 / 2.1 / 2.6% on Sudoku, 0.0% on Maze. But weight-tied iteration changes the regime. What does the loop buy us?

3/ Our view: weight tying turns a network into a ๐ฅ๐๐ญ๐๐ง๐ญ ๐๐ฒ๐ง๐๐ฆ๐ข๐๐๐ฅ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ. At test time, the model repeatedly updates its hidden state. Scaling works when the modelโs ๐ข๐ง๐ญ๐๐ซ๐ง๐๐ฅ ๐๐ญ๐ญ๐ซ๐๐๐ญ๐จ๐ซ ๐ฅ๐๐ง๐๐ฌ๐๐๐ฉ๐ aligns with the ๐ญ๐๐ฌ๐ค ๐ฆ๐๐ญ๐ซ๐ข๐ ๐ฅ๐๐ง๐๐ฌ๐๐๐ฉ๐: low-residual attractors -> low-error solutions.

4/ When the landscape is aligned, generalization beyond training depth becomes possible. As iterations increase: residual โ accuracy โ Convergence to neural attractors becomes a useful scaling signal.

5/ This gives two axes for inference scaling. ยท ๐๐๐ฉ๐ญ๐ก: run one trajectory longer โ better convergence. ยท ๐๐ซ๐๐๐๐ญ๐ก: random initialization + path noise โ explore more basins. Then select the trajectory that converged best. No external verifier needed. Just the modelโs own attractor dynamics.

6/ But weight tying alone is not enough. Recursive models can learn bad landscapes: 1. no correct attractor 2. spurious attractors 3. correct basin too narrow Then more compute just converges to the wrong place. In EqR, we introduce training interventions that reshape the landscape so correct attractors become reachable.

6.1/ Randomized State Initialization (RI)

6.2/ Path Stochasticity via Noise Injection (NI)

7/ Aligned attractors also make compute adaptive. Easy examples can converge in 1โ5 steps. Hard examples can receive much more depth + breadth.

๐ Paper: โ๏ธ Code: Shout out to my amazing mentors @ZhengyangGeng and @zicokolter for the guidance and support throughout this project!

Bonus 1/ Breadth scaling also becomes more useful after EqR training. Instead of majority voting, we can select the trajectory that actually converged best. For N trajectories, we choose the one with the lowest average residual over the final three iteration steps. This ๐๐จ๐ง๐ฏ๐๐ซ๐ ๐๐ง๐๐-๐๐๐ฌ๐๐ ๐ฌ๐๐ฅ๐๐๐ญ๐ข๐จ๐ง beats majority voting in both performance and efficiency. But importantly: it works for EqR, not for the baseline. That suggests RI + NI do more than improve accuracy. They make residuals meaningful again.

@zicokolter Bonus 2/ For people interested in iterative models with feedback loops, check out this collection! PRs, comments, and suggestions are very welcome!

Side-quest/ A separate question is: ๐๐จ๐ฐ ๐๐จ ๐ฐ๐ ๐ ๐๐ญ ๐๐ซ๐จ๐ฆ ๐ ๐ฌ๐ญ๐๐ง๐๐๐ซ๐ ๐๐๐๐๐๐จ๐ซ๐ฐ๐๐ซ๐ ๐ฆ๐จ๐๐๐ฅ ๐ญ๐จ ๐ ๐๐๐ฉ๐๐๐ฅ๐ ๐ฅ๐จ๐จ๐ฉ ๐ฆ๐จ๐๐๐ฅ ๐ข๐ง ๐ญ๐ก๐ ๐๐ข๐ซ๐ฌ๐ญ ๐ฉ๐ฅ๐๐๐ ? We have also explored this problem in our work. We share our observations and findings in the side-post below:

Btw, the video has sound ๐

Love the video !

Thank you! It really takes me some time hh

Congrats Benhao. Very neat work.

Thank you Hayden! More on the way hh! And I have been trying your model at larger scale, so far so good ๐

Awesome! Excited for whatโs next ๐

Really interesting work.

Thank you! More to release tmrw, stay tuned ๐

This is really cool! Do you think the attractors could eventually be represented implicitly by a verifier? It seems like the flow fields are defined explicitly by a neural network here.

Thanks for your kind words! Yes, this is why we term it as neural attractors. And yeah! Flow field is quite relevant to equilibrium, and you may also intersted in drifting models

Interesting! Thanks for sharing!

๐

๐คฏ bro

congratulations on the icml acceptance ๐ nice paper!

Thank you! ๐

Why not Group-Equivariant Equilibrium Reasoner or Clifford Group Equivariant? Would it make sense as it would need to learn to compose like with monoids?

LLM donโt reason, they compute.

I would say, โThey computeโ isnโt an argument against reasoning. If reasoning emerges from computation in brains, the question is whether LLM computation can instantiate reasoning-like processes.

่ฟ้้ข็ๅจ็ปไนๆฏskill่ชๅทฑ็ๆ็ๅ๏ผ
