Video yükleniyor...
Video Yüklenemedi
Almost all animals sleep. Why don’t LMs? Introducing our new work on language model sleep. tl;dr : A periodic, recurrent “sleep” phase allows LMs to digest their context and transfer it into their weights, improving recall and reasoning on challenging tasks.
126,161 görüntüleme • 4 ay önce •via X (Twitter)
45 Yorum

Prior work on SSM-attention hybrid models introduces a weight-memory that maintains a compressed memory of past tokens, complementing short-term attention memory. This context->weight transfer is done in a single forward pass: learning happens on-the-fly. Good for latency.

But isn’t this too good to be true? Learning a good representation of data is hard and normally requires deep sequential computation–most learning algorithms are iterative (eg gradient descent). For animals, sleep plays a major role in learning. During sleep, most external input is blocked, but the brain is still actively replaying information. We sleep for 4 to 10 hours–why so long, if we could just learn on-the-fly?

On several symbolic and natural language reasoning tasks (cellular automaton, multi-hop graph retrieval, and GSM-Infinite), we show that hybrid models fail as reasoning requirements grow, even when the amount of information to store is held fixed. So, unlike previous belief, this is NOT due to a lack of fast weight capacity!

Our fix is simple: to use N recurrent forward passes for learning fast weights. This gives the model enough time to learn a good representation of context. We call this process “sleep”. This is not the same as looped transformers–our model still uses a single forward pass outside the sleep phase (i.e. when the context window is not full).

For all tasks we evaluate, we see more offline loops -> better task performance, especially the ones that require more reasoning.

Takeaway: Effective learning takes time, so we should give the model enough time to learn. Sufficient sleep duration gives the model enough time to transform tokens into a good weight -> better retention and reasoning. Arxiv:

The ML community has explored the idea of sleep from various angles. Replay buffers in RL, training agents on model-generated experience (eg the Dreamer series), wake-sleep algorithm, contrastive divergence, to name a few. Agentic scaffolds such as OpenClaw and Claude Code call their offline phase “Dreaming”. Tagging a few recent relevant works that explore the sleep idea from different angles: Sleep-time compute paper from Letta: There is also this prior OpenReview submission with the same title as ours (we updated ours to avoid conflict): Let us know if there are more!

Lots of offline memory processing can be called “sleep.” I think ours maps especially directly onto the analogy: recurrent Hebbian-style weight updates, no test-time backprop, no natural-language summarization.

We'd love to see if sleep could replace/complement the natural-language compaction currently used by Codex / Claude Code

Huge thanks to amazing collaborators! @giuliacfanti @SeanMcleish @tomgoldsteincs

And thank you so much to Modal for providing a compute grant to support this work. This work would have been literally impossible without your generous support! They have a really nice CLI interface, so migration to their cluster was almost zero-conflict with the help of coding agents! @charles_irl

Because they aren’t animals?

Now add "emotions" to the context pieces to help the sleep phase "dream" on the most important things that need to be learned. First emotion could be "surprise" whenever something unexpected happens that can't be easily predicted.

So… we finetune the models using context?

I added your paper to my online Arxiv archive, w related research topics, concepts, & questions to your paper. It's been making the rounds!

This is really cool! I wrote something semi-related ( but as a random guy with no background in math/CS I ran out of road. I like your approach more...

this is a fantastic paper! it reminds me of CPU memory design with L1 (Attention KV Cache), L2 (Sliding Window Attention) and L3 (Fast Weights/SSM). If I log the average activation levels of the input gate beta_t, and it is high, should I trigger a sleep to avoid thrashing?

I like this. Simplifying a lot here, but comparing to LLMs, humans are like multi-modal dynamic models. Except we update this dynamic model for the entirety of our consciousness (not just sleep). If you think about it this way biological “compute” is just eons ahead of silicon.

my torment nexus

The app LAYLA is already out and does this same function. While you sleep or have it in background it ingests your information and better stores it in long term memory. It runs most LLMs faster than any other mobile app, it beats anythingLLM. I can load larger models on android

why stop at chunk level. have you tried applying ACT at token level (from second chunk pass onwards). should help with reasoning and also make sleep phase more efficient

This is the way

Very cool work!!! Congrats !

LFG, glad someone figured out a good dream cycle

Is this a LoRA?

But do they dream of electric sheep?

Too bad it’s already been patented. FYI: you’re still missing a critical part to make it work. Sleeping alone doesn’t get you there. Go ahead and ask us how we know.

Interesting concept

眠り続けた人間やLLMはどうなりますか?

LMs should eat too!

@n00rdung

Cool work. Thanks for the pointer Sangyun 🙏

sleep as bulk-boundary compression

I think sleep is no necessary to actually feel alive. I think we are gonna find the same thing true for ai networks.

Is this some advanced method of “dreaming”? How do you prevent weights getting shifted too far in one direction

Im already doing this with local memory for frequently used agents

Very intriguing framing. Treating memory consolidation as an explicit phase rather than a byproduct of training could open up some interesting directions for language model development.

Extremely interesting

그 동안 글로벌 AI 개발사들이 발표해 왔던 내용들은 사실상 거의 전부가 다른 사람들의 통찰과 성과를 재포장해서 발표한 것에 불과했죠. 그런 관점에서 봤을 때 AI의 수면 개념은 정말 획기적인 전환점이 될 수 있는 개념이라고 생각합니다. 생물적 체계에서의 지능과 기계적 체계에서의 지능이 사실은 근본적인 차이가 없을 거라는 가정에 한발 더 다가가는 계기가 될 거라고 생각합니다. AI가 인간과 같다는 말보다는 인간도 AI와 같을 수 있다는 개념에 더 가깝겠죠.

when i sleep i dream in code now

sleep in animals is essentially offline replay for memory consolidation. this is doing something surprisingly similar. curious how sensitive the gains are to sleep frequency and whether there's an optimal wake/sleep ratio like we see in neuroscience

training could be considered sleep

I find 70 to 200ms is very effective which allows applied intelligence to harmonise with governance

这在认知心理学上叫做replay,不是什么新概念。

Basically doesn’t grok build have this

