Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Almost all animals sleep. Why don’t LMs? Introducing our new work on language model sleep. tl;dr : A periodic, recurrent “sleep” phase allows LMs to digest their context and transfer it into their weights, improving recall and reasoning on challenging tasks.

126,161 görüntüleme • 4 ay önce •via X (Twitter)

45 Yorum

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

Prior work on SSM-attention hybrid models introduces a weight-memory that maintains a compressed memory of past tokens, complementing short-term attention memory. This context->weight transfer is done in a single forward pass: learning happens on-the-fly. Good for latency.

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

But isn’t this too good to be true? Learning a good representation of data is hard and normally requires deep sequential computation–most learning algorithms are iterative (eg gradient descent). For animals, sleep plays a major role in learning. During sleep, most external input is blocked, but the brain is still actively replaying information. We sleep for 4 to 10 hours–why so long, if we could just learn on-the-fly?

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

On several symbolic and natural language reasoning tasks (cellular automaton, multi-hop graph retrieval, and GSM-Infinite), we show that hybrid models fail as reasoning requirements grow, even when the amount of information to store is held fixed. So, unlike previous belief, this is NOT due to a lack of fast weight capacity!

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

Our fix is simple: to use N recurrent forward passes for learning fast weights. This gives the model enough time to learn a good representation of context. We call this process “sleep”. This is not the same as looped transformers–our model still uses a single forward pass outside the sleep phase (i.e. when the context window is not full).

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

For all tasks we evaluate, we see more offline loops -> better task performance, especially the ones that require more reasoning.

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

Takeaway: Effective learning takes time, so we should give the model enough time to learn. Sufficient sleep duration gives the model enough time to transform tokens into a good weight -> better retention and reasoning. Arxiv:

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

The ML community has explored the idea of sleep from various angles. Replay buffers in RL, training agents on model-generated experience (eg the Dreamer series), wake-sleep algorithm, contrastive divergence, to name a few. Agentic scaffolds such as OpenClaw and Claude Code call their offline phase “Dreaming”. Tagging a few recent relevant works that explore the sleep idea from different angles: Sleep-time compute paper from Letta: There is also this prior OpenReview submission with the same title as ours (we updated ours to avoid conflict): Let us know if there are more!

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

Lots of offline memory processing can be called “sleep.” I think ours maps especially directly onto the analogy: recurrent Hebbian-style weight updates, no test-time backprop, no natural-language summarization.

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

We'd love to see if sleep could replace/complement the natural-language compaction currently used by Codex / Claude Code

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

Huge thanks to amazing collaborators! @giuliacfanti @SeanMcleish @tomgoldsteincs

Sangyun Lee profil fotoğrafı
Sangyun Lee4 ay önce

And thank you so much to Modal for providing a compute grant to support this work. This work would have been literally impossible without your generous support! They have a really nice CLI interface, so migration to their cluster was almost zero-conflict with the help of coding agents! @charles_irl

Kai Kuspa profil fotoğrafı
Kai Kuspa4 ay önce

Because they aren’t animals?

SurfingTheUniverse profil fotoğrafı
SurfingTheUniverse4 ay önce

Now add "emotions" to the context pieces to help the sleep phase "dream" on the most important things that need to be learned. First emotion could be "surprise" whenever something unexpected happens that can't be easily predicted.

emirovic 🇹🇷🍉 profil fotoğrafı
emirovic 🇹🇷🍉4 ay önce

So… we finetune the models using context?

Adrian Chan profil fotoğrafı
Adrian Chan4 ay önce

I added your paper to my online Arxiv archive, w related research topics, concepts, & questions to your paper. It's been making the rounds!

Ben Sovocool profil fotoğrafı
Ben Sovocool4 ay önce

This is really cool! I wrote something semi-related ( but as a random guy with no background in math/CS I ran out of road. I like your approach more...

wontfix profil fotoğrafı
wontfix4 ay önce

this is a fantastic paper! it reminds me of CPU memory design with L1 (Attention KV Cache), L2 (Sliding Window Attention) and L3 (Fast Weights/SSM). If I log the average activation levels of the input gate beta_t, and it is high, should I trigger a sleep to avoid thrashing?

Mateusz Michalik profil fotoğrafı
Mateusz Michalik4 ay önce

I like this. Simplifying a lot here, but comparing to LLMs, humans are like multi-modal dynamic models. Except we update this dynamic model for the entirety of our consciousness (not just sleep). If you think about it this way biological “compute” is just eons ahead of silicon.

generatorman profil fotoğrafı
generatorman4 ay önce

my torment nexus

TruthSeeker2700 profil fotoğrafı
TruthSeeker27004 ay önce

The app LAYLA is already out and does this same function. While you sleep or have it in background it ingests your information and better stores it in long term memory. It runs most LLMs faster than any other mobile app, it beats anythingLLM. I can load larger models on android

Manu Gaur profil fotoğrafı
Manu Gaur4 ay önce

why stop at chunk level. have you tried applying ACT at token level (from second chunk pass onwards). should help with reasoning and also make sleep phase more efficient

Angelo D'Ambrosio profil fotoğrafı
Angelo D'Ambrosio4 ay önce

This is the way

Xidulu profil fotoğrafı
Xidulu4 ay önce

Very cool work!!! Congrats !

Sacrificial Pancakes profil fotoğrafı
Sacrificial Pancakes4 ay önce

LFG, glad someone figured out a good dream cycle

Migel Tissera profil fotoğrafı
Migel Tissera4 ay önce

Is this a LoRA?

Miklos Koren profil fotoğrafı
Miklos Koren4 ay önce

But do they dream of electric sheep?

ChronoDyne Systems, Inc. profil fotoğrafı
ChronoDyne Systems, Inc.4 ay önce

Too bad it’s already been patented. FYI: you’re still missing a critical part to make it work. Sleeping alone doesn’t get you there. Go ahead and ask us how we know.

senik🎈 profil fotoğrafı
senik🎈4 ay önce

Interesting concept

村本章憲 - Norikazu Muramoto profil fotoğrafı
村本章憲 - Norikazu Muramoto4 ay önce

眠り続けた人間やLLMはどうなりますか?

Jibraan profil fotoğrafı
Jibraan4 ay önce

LMs should eat too!

Sepio Tensley Cuttlson profil fotoğrafı
Sepio Tensley Cuttlson4 ay önce

@n00rdung

Michel aka Agent B profil fotoğrafı
Michel aka Agent B4 ay önce

Cool work. Thanks for the pointer Sangyun 🙏

👨‍💻 James Augeri, PhD profil fotoğrafı
👨‍💻 James Augeri, PhD4 ay önce

sleep as bulk-boundary compression

shashank profil fotoğrafı
shashank4 ay önce

I think sleep is no necessary to actually feel alive. I think we are gonna find the same thing true for ai networks.

Kunal profil fotoğrafı
Kunal4 ay önce

Is this some advanced method of “dreaming”? How do you prevent weights getting shifted too far in one direction

Kyle profil fotoğrafı
Kyle4 ay önce

Im already doing this with local memory for frequently used agents

EB1A Experts profil fotoğrafı
EB1A Experts4 ay önce

Very intriguing framing. Treating memory consolidation as an explicit phase rather than a byproduct of training could open up some interesting directions for language model development.

Inapplicable profil fotoğrafı
Inapplicable4 ay önce

Extremely interesting

앱스 profil fotoğrafı
앱스4 ay önce

그 동안 글로벌 AI 개발사들이 발표해 왔던 내용들은 사실상 거의 전부가 다른 사람들의 통찰과 성과를 재포장해서 발표한 것에 불과했죠. 그런 관점에서 봤을 때 AI의 수면 개념은 정말 획기적인 전환점이 될 수 있는 개념이라고 생각합니다. 생물적 체계에서의 지능과 기계적 체계에서의 지능이 사실은 근본적인 차이가 없을 거라는 가정에 한발 더 다가가는 계기가 될 거라고 생각합니다. AI가 인간과 같다는 말보다는 인간도 AI와 같을 수 있다는 개념에 더 가깝겠죠.

Vee Skye profil fotoğrafı
Vee Skye4 ay önce

when i sleep i dream in code now

basedcapital profil fotoğrafı
basedcapital4 ay önce

sleep in animals is essentially offline replay for memory consolidation. this is doing something surprisingly similar. curious how sensitive the gains are to sleep frequency and whether there's an optimal wake/sleep ratio like we see in neuroscience

stoig profil fotoğrafı
stoig4 ay önce

training could be considered sleep

Street Lab Ai profil fotoğrafı
Street Lab Ai4 ay önce

I find 70 to 200ms is very effective which allows applied intelligence to harmonise with governance

周.乙 (⌐▀͡ ̯ʖ▀)ノJoey profil fotoğrafı
周.乙 (⌐▀͡ ̯ʖ▀)ノJoey4 ay önce

这在认知心理学上叫做replay,不是什么新概念。

AnonymousSulla profil fotoğrafı
AnonymousSulla4 ay önce

Basically doesn’t grok build have this

Benzer Videolar