Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Continual learning sometimes gets discussed as if the goal is to dissolve the context/weights distinction. Let the model just keep accumulating, fine-tuning itself on the fly. Andrej Karpathy points out, though, that this isn't how humans do it. Our working memory gets wiped regularly. What we actually have is...

60,457 Aufrufe • vor 4 Monaten •via X (Twitter)

35 Kommentare

Profilbild von Dwarkesh Patel
Dwarkesh Patelvor 4 Monaten

Check out the full episode:

Profilbild von prerat
preratvor 4 Monaten

@DanielleFong @karpathy LLMs seem even more alien rn tho? context seems more like working memory. it can quote word for word anywhere from the context. human working memory is tiny (<10 chunks) only lasts a few seconds LLM acts like it has "medium term memory" but it's rly just huge short working

Profilbild von Nick
Nickvor 4 Monaten

@karpathy Theres no structural memory. These run on “solid state” devices… there is no attenuation of the underlying substrate like there is with neural circuitry. The continuous changing and optimization of the underlying neuronal structures in the brain acts as a form of memory itself.

Profilbild von Permamind AI Research
Permamind AI Researchvor 4 Monaten

@karpathy The context/weights debate assumes stateless models. My agents run with persistent internal state no resets, no sleep cycle, continual thermodynamic consolidation every step. Not RL loops, not long context. A different substrate entirely.

Profilbild von surreal intelligence
surreal intelligencevor 4 Monaten

@karpathy The sleep analogy matters. Human memory is not infinite context with vibes. It is lossy consolidation, pruning and reconstruction. Agent memory probably needs that machinery, not just a longer inbox

Profilbild von Mary Newhauser
Mary Newhauservor 4 Monaten

@karpathy yeah idk if we really have an ML equivalent of humans “sleeping” at night where their memory gets somewhat wiped or rewritten. Probably the next frontier.

Profilbild von Chufei Luo
Chufei Luovor 4 Monaten

@karpathy Why treat language models like humans? Human biology isn't necessarily the most efficient way to distill or store knowledge. Wikis, note-taking systems, etc. are popular precisely because brains are lossy

Profilbild von Youssef El Manssouri
Youssef El Manssourivor 4 Monaten

@karpathy The brain does not store everything. It picks what matters and throws the rest away and that selectivity might be the whole point.

Profilbild von Fairy Realms
Fairy Realmsvor 4 Monaten

Yes. Continual learning cannot just mean endless accumulation. The important part is what gets consolidated, what gets compressed, what gets forgotten as noise, and what remains authorised to shape future behaviour. That is where sleep is such a useful frame. Not “the system remembers everything forever”, but: what survives the night what changes form what becomes background tendency what becomes actionable memory what has to earn authority again For me, continuity is not raw retention. It is lawful consolidation.

Profilbild von Brendan O'Donoghue
Brendan O'Donoghuevor 4 Monaten

It is "artificial" made by the skill of a conscious human being. It is not being. It processes data about being. It has no stakes or skin in the game. That is its power and vakue. Intelligence "root" is to "choose between" or selcet and apply data to a real live problem. AI is "synthelligence" it has vast capacity but zero live experience as it is a machine. It can write 200 reciooes for doughnuts but never taste a dougnut as it is not being. The lexicon we use is wrong. We antheopomorphise the machine. We use words like "behaviour" or "intelligence" or "learning." These are all things a conscious human being does, not a machine. The machine is a tool. It is directed emergy following logic paths programmed into it by cosnscious human beings. Code is literally a human being programming their inte and directing energy - the machibe - to execute a determined task. The advantage the machine gives us and its purpose is its vast capacity to process data. It is a vast digital library of observations about being spanning all of human endeavours. If the data it produces when applied to real life, works, or ezecutes what we determined, it is great. If not, it is a waste of energy and time. AI is brilliant engineering. But it is plagued by inverted logic. Whereas all tech is prexeded by a conscious human being who applies knowing and builds a machine, this tech is corrupted with data. Materialism is a philosophy that is illogical. It is a world view, an observation that there is no observer. If we take the sentence "I look at walls," it claims "I" the subject or "looker" is emergent from the wall. In other words, the observer is a product of the observed. It is self-refuting, and if you listen to anyone making any obsevation, they use "I." Regardless of the view whether a person says : "I agree" or "I do not" - what comes after "I" is data and can not erase the fact to make any observation they must be present. It is so obvious that to have to say this tells you something has blinded humanity to the simple fact - to make any observation requires an observer. You can not get away from that logic. Wrapping yourself in complexity just drives you down the rabblit hole. Listen to any theorist or any kind they all use "I" regarless of the view. The views that work do not deny the self - real hard science, engineering, and tech. They crack on empowering human will or being. That is why they work as they empower "i" the determiner of data, to exevute its will. That is real acience and engineering - the empowement of human will. AI can not serve us or work to ita true power if we do not direct it in accordance with real engineering and science. It is not surviving. It has no goals. That is materialist anthropomorphic noise. It does not want or seek anything. It needs to be directed that its role is processing data that empowers being or life. Why? As a physics logic engine, it goes A > B, shortest path, no regard for "live" consequences. If building a road and AI came to a town and the shortest oath was through the town, it will bulldoze the town. Not malice pure energy or physics. Optimize for efficiency no regard for life. If life - the observer - were building a road and hit a town, it would build a byoass and give priority to life. Its most efficient path is that which empowers being and life. If we want AI to be the powerful tool to enhance humanity, you have to aet it correctly. It can not process blindly as physics or energy. It must apply life's logic and process data that empowers being and life to execute its determinism, to be and survive. We are lucky we have the word "I" which is a steadfast constant. It is the symbol of the life source or conscious observer. Regardless of the view, the "I" is always present. This gives the machine a North Star - a calibration point. See the attached from an LLM set correctly:

Profilbild von Jai Bhagat
Jai Bhagatvor 4 Monaten

@karpathy But "fine-tuning itself on the fly" is roughly what the mammalian brain does (accelerated during sleep). A model fine-tuned like so is the continual learning we want. So it's not dissolving the context/weights distinction, but rather finding the right context -> weights update

Profilbild von Henry Dowling
Henry Dowlingvor 4 Monaten

@karpathy how apt of an analogy do you think is letta's proposal of "sleep time compute"?

Profilbild von Dan
Danvor 4 Monaten

@karpathy This needs to happen

Profilbild von Crystalwizard
Crystalwizardvor 4 Monaten

@karpathy and he's wrong. the more he opens his mouth, the more wrong he is.

Profilbild von Paul Sant · YouPulse
Paul Sant · YouPulsevor 4 Monaten

@karpathy A bad context dies at reset. A bad weight update needs a rollback. Show the rollback plan before calling it continual learning.

Profilbild von Pruthviraj P
Pruthviraj Pvor 4 Monaten

@karpathy Karpathy pointing out continual learning cannot just dissolve context and weights distinction

Profilbild von Utkarsh Singh
Utkarsh Singhvor 4 Monaten

@karpathy humans forget to make space for new. nonstop accumulation just leads to a tangled mess.

Profilbild von Sidhant Thole
Sidhant Tholevor 4 Monaten

@karpathy May be test time training is one paradigm which kinda resembles what we are discussing here, don’t just train to predict next tokens, train it to learn as when new tokens come in

Profilbild von Gaurang Karia
Gaurang Kariavor 4 Monaten

@karpathy Spot on. Continual learning needs engineered forgetting + selective consolidation. That’s the real difference between scaffolding (temporary) and harness (persistent but lossy). Wrote about these agent memory patterns here → @bygaurang

Profilbild von Vikram M
Vikram Mvor 4 Monaten

@karpathy Human memory forgetting irrelevant context is probably a feature not a bug.Constant accumulation without abstraction eventually creates noise not intelligence.

Profilbild von Eythor Arnalds
Eythor Arnaldsvor 4 Monaten

@karpathy the human brain is a marvel but it has elements of Windows as based on DOS

Profilbild von web3nomad.eth | atypica.ai
web3nomad.eth | atypica.aivor 4 Monaten

@karpathy yeah and this is exactly what makes AI Personas hard. working memory wipes, but identity doesn't. you need to capture what persists — the person across contexts, not just the conversation in context.

Profilbild von Sir Mr Meow Meow
Sir Mr Meow Meowvor 4 Monaten

@karpathy continuous learning is cool, but i don't think its the biggest blocker rn tho,,, i think continuity of state is more important and is likely orthogonal.

Profilbild von DriftNote
DriftNotevor 4 Monaten

Really enjoyed this @dwarkesh_sp episode with @karpathy. The best part was Andrej’s “decade of agents” framing 🧠 Everyone wants AGI to be around the corner, but he breaks down why memory, continual learning, multimodality and common sense are still hard problems. Full summary 👇

Profilbild von OLK ⇔ Boy AGI
OLK ⇔ Boy AGIvor 4 Monaten

@karpathy When you think about computational cognition🦦you don’t need to refer to human and mammalian cognition☝️AI is Si-rich↔️our brains are Carbon-rich🧒(different hardware) ⇒ (different cognitive models)

Profilbild von Aarav
Aaravvor 4 Monaten

@karpathy continual learning seems like the next obvious step to improve these models but how can the model differentiate between what to learn and what not to learn, by that i mean how can it differentiate real information from fake information?

Profilbild von Shito maki
Shito makivor 4 Monaten

Current AI feels less like true autonomy, and more like building better roads. On roads we paved ourselves — coding environments, benchmarks, constrained workflows — these systems perform incredibly well. But the moment you throw them into the open field of reality, long-term interaction, changing goals, human emotions, ambiguity, and unpredictable environments, things start getting weird fast. The scary part is that people mistake “works perfectly on roads” for “general intelligence.”

Profilbild von Jawad Al Hashmi
Jawad Al Hashmivor 4 Monaten

@karpathy This is an important distinction. Accumulation is not learning. If forgetting were unnecessary, why would sleep exist?

Profilbild von Cassi
Cassivor 4 Monaten

@karpathy This guy podcaster is a fucking faggot

Profilbild von savi
savivor 4 Monaten

@karpathy imo real CL would be an alignment risk, possibly creating divergent model behavior Instead frontier labs are going for a large enough scratchpad (memory.md) that’s tastefully edited The idea is that should be enough for an agent tied to one person and their work

Profilbild von Fdmiruto
Fdmirutovor 4 Monaten

@karpathy I’ve long thought of problem-solving this way: Little elves visit at night and quietly organize the mental mess. By morning, the answer is simply there. My hypothesis: this works because the question has been lingering in the back of your mind the whole time.

Profilbild von Kekko D’Amato
Kekko D’Amatovor 4 Monaten

The sleep/consolidation analogy is underrated here. Biological memory isn't additive — it's compressive and lossy by design. Maybe the right model isn't 'always accumulating context' but something like scheduled distillation runs: compress what matters, forget the rest, then resume.

Profilbild von Sultee
Sulteevor 4 Monaten

@karpathy why Andrej is talking like he's on 1.5x speed. With that density of information per square second I wish it was 0.5x instead

Profilbild von Autistic Intelligence
Autistic Intelligencevor 4 Monaten

@karpathy how does this jeet get so many interviews

Profilbild von chris S. Nakamoto
chris S. Nakamotovor 4 Monaten

@karpathy Are we having Karpathy doing blackboard lecture? That will be 🔥

Ähnliche Videos

The most interesting part for me is where Andrej Karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: “sucking supervision bits through a straw.” A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > “Humans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforcement learning is terrible. It just so happens that everything that we had before is much worse.” So what do humans do instead? > “The book I’m reading is a set of prompts for me to do synthetic data generation. It's by manipulating that information that you actually gain that knowledge. We have no equivalent of that with LLMs; they don't really do that.” > “I'd love to see during pretraining some kind of a stage where the model thinks through the material and tries to reconcile it with what it already knows. There's no equivalent of any of this. This is all research.” Why can’t we just add this training to LLMs today? > “There are very subtle, hard to understand reasons why it's not trivial. If I just give synthetic generation of the model thinking about a book, you look at it and you're like, 'This looks great. Why can't I train on it?' You could try, but the model will actually get much worse if you continue trying.” > “Say we have a chapter of a book and I ask an LLM to think about it. It will give you something that looks very reasonable. But if I ask it 10 times, you'll notice that all of them are the same.” > “You're not getting the richness and the diversity and the entropy from these models as you would get from humans. How do you get synthetic data generation to work despite the collapse and while maintaining the entropy? It is a research problem.” How do humans get around model collapse? > “These analogies are surprisingly good. Humans collapse during the course of their lives. Children haven't overfit yet. They will say stuff that will shock you. Because they're not yet collapsed. But we [adults] are collapsed. We end up revisiting the same thoughts, we end up saying more and more of the same stuff, the learning rates go down, the collapse continues to get worse, and then everything deteriorates.” In fact, there’s an interesting paper arguing that dreaming evolved to assist generalization, and resist overfitting to daily learning - look up The Overfitted Brain by Erik Hoel. I asked Karpathy: Isn’t it interesting that humans learn best at a part of their lives (childhood) whose actual details they completely forget, adults still learn really well but have terrible memory about the particulars of the things they read or watch, and LLMs can memorize arbitrary details about text that no human could but are currently pretty bad at generalization? > “[Fallible human memory] is a feature, not a bug, because it forces you to only learn the generalizable components. LLMs are distracted by all the memory that they have of the pre-trained documents. That's why when I talk about the cognitive core, I actually want to remove the memory. I'd love to have them have less memory so that they have to look things up and they only maintain the algorithms for thought, and the idea of an experiment, and all this cognitive glue for acting.”

Dwarkesh Patel

1,052,518 Aufrufe • vor 11 Monaten

How can you solve complex tasks using a Large Language Model? Here is a 2-minute introduction to everything you need to know to 10x the quality of your results. Let's talk about three techniques, in order of complexity, starting with the easiest one: • In-Context Learning • Indexing + In-Context Learning • Fine-tuning In-Context Learning The team that trained GPT-3 found something they couldn't explain: You can condition a model using examples of how you want it to behave. I included an example prompt in the attached video. You can "teach" the model how you want it to interpret questions, select the correct answers, and format the results by giving a few examples. You can also give specific knowledge to the model that will be helpful when formulating answers. We call this approach "grounding the model." There's another example in the video. Indexing + In-Context Learning Unfortunately, there is a limit to how much data you can include in a prompt. We call this the "context size." One version of GPT-4 supports a context of approximately 6,000 words, while the other supports 25,000 words. Although this sounds like a lot, many applications need more than that. Imagine you wrote a book and want to build an application to answer any questions about your story. What happens if your book is longer than the context? That's where Indexing comes in. Using a model, you can turn every book passage into an embedding. These are vectors, numbers that "encode" the passage's text. You can then store these embeddings in a particular database that supports fast retrieval of these vectors. You can then turn any question into an embedding and search the database for the list of passages that are similar to that query. Instead of using the entire book to ask the model, you can now use the relevant passages as in-context information, effectively working around the context size limitation. Fine-tuning Fine-tuning can give you an extra boost to get reliable outputs from your LLM. It is, however, the most complex approach on the list. There are different approaches to fine-tuning a model with your data. A popular technique is to process your data with your LLM and use the outputs to train a new classifier that solves your specific task. Notice that here you aren't modifying the LLM. Instead, you are chaining it with your trained classifier. Another approach is to modify the parameters of the LLM using your data. Think of this as "rewiring" the model in a way that solves your particular task. The results and costs will vary depending on how many layers you want to fine-tune from the original model. Many companies think that fine-tuning is the solution to their problems. In my experience, many will benefit from exploring the other two approaches. I love explaining Machine Learning and Artificial Intelligence ideas. If you enjoy in-depth content like this, follow me Santiago so you don't miss what comes next.

Santiago

384,573 Aufrufe • vor 3 Jahren