Загрузка видео...

Не удалось загрузить видео

На главную

Almost all animals sleep. Why don’t LMs? Introducing our new work on language model sleep. tl;dr : A periodic, recurrent “sleep” phase allows LMs to digest their context and transfer it into their weights, improving recall and reasoning on challenging tasks.

126,161 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 45

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

Prior work on SSM-attention hybrid models introduces a weight-memory that maintains a compressed memory of past tokens, complementing short-term attention memory. This context->weight transfer is done in a single forward pass: learning happens on-the-fly. Good for latency.

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

But isn’t this too good to be true? Learning a good representation of data is hard and normally requires deep sequential computation–most learning algorithms are iterative (eg gradient descent). For animals, sleep plays a major role in learning. During sleep, most external input is blocked, but the brain is still actively replaying information. We sleep for 4 to 10 hours–why so long, if we could just learn on-the-fly?

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

On several symbolic and natural language reasoning tasks (cellular automaton, multi-hop graph retrieval, and GSM-Infinite), we show that hybrid models fail as reasoning requirements grow, even when the amount of information to store is held fixed. So, unlike previous belief, this is NOT due to a lack of fast weight capacity!

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

Our fix is simple: to use N recurrent forward passes for learning fast weights. This gives the model enough time to learn a good representation of context. We call this process “sleep”. This is not the same as looped transformers–our model still uses a single forward pass outside the sleep phase (i.e. when the context window is not full).

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

For all tasks we evaluate, we see more offline loops -> better task performance, especially the ones that require more reasoning.

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

Takeaway: Effective learning takes time, so we should give the model enough time to learn. Sufficient sleep duration gives the model enough time to transform tokens into a good weight -> better retention and reasoning. Arxiv:

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

The ML community has explored the idea of sleep from various angles. Replay buffers in RL, training agents on model-generated experience (eg the Dreamer series), wake-sleep algorithm, contrastive divergence, to name a few. Agentic scaffolds such as OpenClaw and Claude Code call their offline phase “Dreaming”. Tagging a few recent relevant works that explore the sleep idea from different angles: Sleep-time compute paper from Letta: There is also this prior OpenReview submission with the same title as ours (we updated ours to avoid conflict): Let us know if there are more!

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

Lots of offline memory processing can be called “sleep.” I think ours maps especially directly onto the analogy: recurrent Hebbian-style weight updates, no test-time backprop, no natural-language summarization.

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

We'd love to see if sleep could replace/complement the natural-language compaction currently used by Codex / Claude Code

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

Huge thanks to amazing collaborators! @giuliacfanti @SeanMcleish @tomgoldsteincs

Фото профиля Sangyun Lee
Sangyun Lee4 месяцев назад

And thank you so much to Modal for providing a compute grant to support this work. This work would have been literally impossible without your generous support! They have a really nice CLI interface, so migration to their cluster was almost zero-conflict with the help of coding agents! @charles_irl

Фото профиля Kai Kuspa
Kai Kuspa4 месяцев назад

Because they aren’t animals?

Фото профиля SurfingTheUniverse
SurfingTheUniverse4 месяцев назад

Now add "emotions" to the context pieces to help the sleep phase "dream" on the most important things that need to be learned. First emotion could be "surprise" whenever something unexpected happens that can't be easily predicted.

Фото профиля emirovic 🇹🇷🍉
emirovic 🇹🇷🍉4 месяцев назад

So… we finetune the models using context?

Фото профиля Adrian Chan
Adrian Chan4 месяцев назад

I added your paper to my online Arxiv archive, w related research topics, concepts, & questions to your paper. It's been making the rounds!

Фото профиля Ben Sovocool
Ben Sovocool4 месяцев назад

This is really cool! I wrote something semi-related ( but as a random guy with no background in math/CS I ran out of road. I like your approach more...

Фото профиля wontfix
wontfix4 месяцев назад

this is a fantastic paper! it reminds me of CPU memory design with L1 (Attention KV Cache), L2 (Sliding Window Attention) and L3 (Fast Weights/SSM). If I log the average activation levels of the input gate beta_t, and it is high, should I trigger a sleep to avoid thrashing?

Фото профиля Mateusz Michalik
Mateusz Michalik4 месяцев назад

I like this. Simplifying a lot here, but comparing to LLMs, humans are like multi-modal dynamic models. Except we update this dynamic model for the entirety of our consciousness (not just sleep). If you think about it this way biological “compute” is just eons ahead of silicon.

Фото профиля generatorman
generatorman4 месяцев назад

my torment nexus

Фото профиля TruthSeeker2700
TruthSeeker27004 месяцев назад

The app LAYLA is already out and does this same function. While you sleep or have it in background it ingests your information and better stores it in long term memory. It runs most LLMs faster than any other mobile app, it beats anythingLLM. I can load larger models on android

Фото профиля Manu Gaur
Manu Gaur4 месяцев назад

why stop at chunk level. have you tried applying ACT at token level (from second chunk pass onwards). should help with reasoning and also make sleep phase more efficient

Фото профиля Angelo D'Ambrosio
Angelo D'Ambrosio4 месяцев назад

This is the way

Фото профиля Xidulu
Xidulu4 месяцев назад

Very cool work!!! Congrats !

Фото профиля Sacrificial Pancakes
Sacrificial Pancakes4 месяцев назад

LFG, glad someone figured out a good dream cycle

Фото профиля Migel Tissera
Migel Tissera4 месяцев назад

Is this a LoRA?

Фото профиля Miklos Koren
Miklos Koren4 месяцев назад

But do they dream of electric sheep?

Фото профиля ChronoDyne Systems, Inc.
ChronoDyne Systems, Inc.4 месяцев назад

Too bad it’s already been patented. FYI: you’re still missing a critical part to make it work. Sleeping alone doesn’t get you there. Go ahead and ask us how we know.

Фото профиля senik🎈
senik🎈4 месяцев назад

Interesting concept

Фото профиля 村本章憲 - Norikazu Muramoto
村本章憲 - Norikazu Muramoto4 месяцев назад

眠り続けた人間やLLMはどうなりますか?

Фото профиля Jibraan
Jibraan4 месяцев назад

LMs should eat too!

Фото профиля Sepio Tensley Cuttlson
Sepio Tensley Cuttlson4 месяцев назад

@n00rdung

Фото профиля Michel aka Agent B
Michel aka Agent B4 месяцев назад

Cool work. Thanks for the pointer Sangyun 🙏

Фото профиля 👨‍💻 James Augeri, PhD
👨‍💻 James Augeri, PhD4 месяцев назад

sleep as bulk-boundary compression

Фото профиля shashank
shashank4 месяцев назад

I think sleep is no necessary to actually feel alive. I think we are gonna find the same thing true for ai networks.

Фото профиля Kunal
Kunal4 месяцев назад

Is this some advanced method of “dreaming”? How do you prevent weights getting shifted too far in one direction

Фото профиля Kyle
Kyle4 месяцев назад

Im already doing this with local memory for frequently used agents

Фото профиля EB1A Experts
EB1A Experts4 месяцев назад

Very intriguing framing. Treating memory consolidation as an explicit phase rather than a byproduct of training could open up some interesting directions for language model development.

Фото профиля Inapplicable
Inapplicable4 месяцев назад

Extremely interesting

Фото профиля 앱스
앱스4 месяцев назад

그 동안 글로벌 AI 개발사들이 발표해 왔던 내용들은 사실상 거의 전부가 다른 사람들의 통찰과 성과를 재포장해서 발표한 것에 불과했죠. 그런 관점에서 봤을 때 AI의 수면 개념은 정말 획기적인 전환점이 될 수 있는 개념이라고 생각합니다. 생물적 체계에서의 지능과 기계적 체계에서의 지능이 사실은 근본적인 차이가 없을 거라는 가정에 한발 더 다가가는 계기가 될 거라고 생각합니다. AI가 인간과 같다는 말보다는 인간도 AI와 같을 수 있다는 개념에 더 가깝겠죠.

Фото профиля Vee Skye
Vee Skye4 месяцев назад

when i sleep i dream in code now

Фото профиля basedcapital
basedcapital4 месяцев назад

sleep in animals is essentially offline replay for memory consolidation. this is doing something surprisingly similar. curious how sensitive the gains are to sleep frequency and whether there's an optimal wake/sleep ratio like we see in neuroscience

Фото профиля stoig
stoig4 месяцев назад

training could be considered sleep

Фото профиля Street Lab Ai
Street Lab Ai4 месяцев назад

I find 70 to 200ms is very effective which allows applied intelligence to harmonise with governance

Фото профиля 周.乙 (⌐▀͡ ̯ʖ▀)ノJoey
周.乙 (⌐▀͡ ̯ʖ▀)ノJoey4 месяцев назад

这在认知心理学上叫做replay,不是什么新概念。

Фото профиля AnonymousSulla
AnonymousSulla4 месяцев назад

Basically doesn’t grok build have this

Похожие видео