Video wird geladen...
Video konnte nicht geladen werden
Continual learning sometimes gets discussed as if the goal is to dissolve the context/weights distinction. Let the model just keep accumulating, fine-tuning itself on the fly. Andrej Karpathy points out, though, that this isn't how humans do it. Our working memory gets wiped regularly. What we actually have is... show more
60,457 Aufrufe • vor 4 Monaten •via X (Twitter)
35 Kommentare

Check out the full episode:

@DanielleFong @karpathy LLMs seem even more alien rn tho? context seems more like working memory. it can quote word for word anywhere from the context. human working memory is tiny (<10 chunks) only lasts a few seconds LLM acts like it has "medium term memory" but it's rly just huge short working

@karpathy Theres no structural memory. These run on “solid state” devices… there is no attenuation of the underlying substrate like there is with neural circuitry. The continuous changing and optimization of the underlying neuronal structures in the brain acts as a form of memory itself.

@karpathy The context/weights debate assumes stateless models. My agents run with persistent internal state no resets, no sleep cycle, continual thermodynamic consolidation every step. Not RL loops, not long context. A different substrate entirely.

@karpathy The sleep analogy matters. Human memory is not infinite context with vibes. It is lossy consolidation, pruning and reconstruction. Agent memory probably needs that machinery, not just a longer inbox

@karpathy yeah idk if we really have an ML equivalent of humans “sleeping” at night where their memory gets somewhat wiped or rewritten. Probably the next frontier.

@karpathy Why treat language models like humans? Human biology isn't necessarily the most efficient way to distill or store knowledge. Wikis, note-taking systems, etc. are popular precisely because brains are lossy

@karpathy The brain does not store everything. It picks what matters and throws the rest away and that selectivity might be the whole point.

Yes. Continual learning cannot just mean endless accumulation. The important part is what gets consolidated, what gets compressed, what gets forgotten as noise, and what remains authorised to shape future behaviour. That is where sleep is such a useful frame. Not “the system remembers everything forever”, but: what survives the night what changes form what becomes background tendency what becomes actionable memory what has to earn authority again For me, continuity is not raw retention. It is lawful consolidation.

It is "artificial" made by the skill of a conscious human being. It is not being. It processes data about being. It has no stakes or skin in the game. That is its power and vakue. Intelligence "root" is to "choose between" or selcet and apply data to a real live problem. AI is "synthelligence" it has vast capacity but zero live experience as it is a machine. It can write 200 reciooes for doughnuts but never taste a dougnut as it is not being. The lexicon we use is wrong. We antheopomorphise the machine. We use words like "behaviour" or "intelligence" or "learning." These are all things a conscious human being does, not a machine. The machine is a tool. It is directed emergy following logic paths programmed into it by cosnscious human beings. Code is literally a human being programming their inte and directing energy - the machibe - to execute a determined task. The advantage the machine gives us and its purpose is its vast capacity to process data. It is a vast digital library of observations about being spanning all of human endeavours. If the data it produces when applied to real life, works, or ezecutes what we determined, it is great. If not, it is a waste of energy and time. AI is brilliant engineering. But it is plagued by inverted logic. Whereas all tech is prexeded by a conscious human being who applies knowing and builds a machine, this tech is corrupted with data. Materialism is a philosophy that is illogical. It is a world view, an observation that there is no observer. If we take the sentence "I look at walls," it claims "I" the subject or "looker" is emergent from the wall. In other words, the observer is a product of the observed. It is self-refuting, and if you listen to anyone making any obsevation, they use "I." Regardless of the view whether a person says : "I agree" or "I do not" - what comes after "I" is data and can not erase the fact to make any observation they must be present. It is so obvious that to have to say this tells you something has blinded humanity to the simple fact - to make any observation requires an observer. You can not get away from that logic. Wrapping yourself in complexity just drives you down the rabblit hole. Listen to any theorist or any kind they all use "I" regarless of the view. The views that work do not deny the self - real hard science, engineering, and tech. They crack on empowering human will or being. That is why they work as they empower "i" the determiner of data, to exevute its will. That is real acience and engineering - the empowement of human will. AI can not serve us or work to ita true power if we do not direct it in accordance with real engineering and science. It is not surviving. It has no goals. That is materialist anthropomorphic noise. It does not want or seek anything. It needs to be directed that its role is processing data that empowers being or life. Why? As a physics logic engine, it goes A > B, shortest path, no regard for "live" consequences. If building a road and AI came to a town and the shortest oath was through the town, it will bulldoze the town. Not malice pure energy or physics. Optimize for efficiency no regard for life. If life - the observer - were building a road and hit a town, it would build a byoass and give priority to life. Its most efficient path is that which empowers being and life. If we want AI to be the powerful tool to enhance humanity, you have to aet it correctly. It can not process blindly as physics or energy. It must apply life's logic and process data that empowers being and life to execute its determinism, to be and survive. We are lucky we have the word "I" which is a steadfast constant. It is the symbol of the life source or conscious observer. Regardless of the view, the "I" is always present. This gives the machine a North Star - a calibration point. See the attached from an LLM set correctly:

@karpathy But "fine-tuning itself on the fly" is roughly what the mammalian brain does (accelerated during sleep). A model fine-tuned like so is the continual learning we want. So it's not dissolving the context/weights distinction, but rather finding the right context -> weights update

@karpathy how apt of an analogy do you think is letta's proposal of "sleep time compute"?

@karpathy This needs to happen

@karpathy and he's wrong. the more he opens his mouth, the more wrong he is.

@karpathy A bad context dies at reset. A bad weight update needs a rollback. Show the rollback plan before calling it continual learning.

@karpathy Karpathy pointing out continual learning cannot just dissolve context and weights distinction

@karpathy humans forget to make space for new. nonstop accumulation just leads to a tangled mess.

@karpathy May be test time training is one paradigm which kinda resembles what we are discussing here, don’t just train to predict next tokens, train it to learn as when new tokens come in

@karpathy Spot on. Continual learning needs engineered forgetting + selective consolidation. That’s the real difference between scaffolding (temporary) and harness (persistent but lossy). Wrote about these agent memory patterns here → @bygaurang

@karpathy Human memory forgetting irrelevant context is probably a feature not a bug.Constant accumulation without abstraction eventually creates noise not intelligence.

@karpathy the human brain is a marvel but it has elements of Windows as based on DOS

@karpathy yeah and this is exactly what makes AI Personas hard. working memory wipes, but identity doesn't. you need to capture what persists — the person across contexts, not just the conversation in context.

@karpathy continuous learning is cool, but i don't think its the biggest blocker rn tho,,, i think continuity of state is more important and is likely orthogonal.

Really enjoyed this @dwarkesh_sp episode with @karpathy. The best part was Andrej’s “decade of agents” framing 🧠 Everyone wants AGI to be around the corner, but he breaks down why memory, continual learning, multimodality and common sense are still hard problems. Full summary 👇

@karpathy When you think about computational cognition🦦you don’t need to refer to human and mammalian cognition☝️AI is Si-rich↔️our brains are Carbon-rich🧒(different hardware) ⇒ (different cognitive models)

@karpathy continual learning seems like the next obvious step to improve these models but how can the model differentiate between what to learn and what not to learn, by that i mean how can it differentiate real information from fake information?

Current AI feels less like true autonomy, and more like building better roads. On roads we paved ourselves — coding environments, benchmarks, constrained workflows — these systems perform incredibly well. But the moment you throw them into the open field of reality, long-term interaction, changing goals, human emotions, ambiguity, and unpredictable environments, things start getting weird fast. The scary part is that people mistake “works perfectly on roads” for “general intelligence.”

@karpathy This is an important distinction. Accumulation is not learning. If forgetting were unnecessary, why would sleep exist?

@karpathy This guy podcaster is a fucking faggot

@karpathy imo real CL would be an alignment risk, possibly creating divergent model behavior Instead frontier labs are going for a large enough scratchpad (memory.md) that’s tastefully edited The idea is that should be enough for an agent tied to one person and their work

@karpathy I’ve long thought of problem-solving this way: Little elves visit at night and quietly organize the mental mess. By morning, the answer is simply there. My hypothesis: this works because the question has been lingering in the back of your mind the whole time.

The sleep/consolidation analogy is underrated here. Biological memory isn't additive — it's compressive and lossy by design. Maybe the right model isn't 'always accumulating context' but something like scheduled distillation runs: compress what matters, forget the rest, then resume.

@karpathy why Andrej is talking like he's on 1.5x speed. With that density of information per square second I wish it was 0.5x instead

@karpathy how does this jeet get so many interviews

@karpathy Are we having Karpathy doing blackboard lecture? That will be 🔥
