Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚨 Grant Sanderson DID IT AGAIN Language compressibility is not just a neat math trick: it is the core engine of modern LLMs. Grant's latest video boils Shannon's entropy down to a single, powerful idea: Prediction IS compression. → Predict the next word better, use fewer bits to store...

160,113 görüntüleme • 1 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,110 görüntüleme • 3 ay önce

Yann LeCun just told the most well-funded industry in human history it is solving the wrong problem. LeCun: “Babies learn this around the age of eight or nine months, that objects don’t float, they fall.” No dataset. No labels. No reward signal. A nine month old drops a spoon and builds a physics engine no machine can match. LeCun: “Most of us can learn to drive in about 20 or 30 hours of training without ever crashing, causing any accident.” Twenty hours. Tesla has built the most capable driving system on the road. It took billions of miles of data to get there. A sixteen year old gets there over a long weekend. Not because the teenager is the better driver. Because the teenager is not learning to drive. They are deploying a model of reality they have been building since birth. LeCun: “If we drive next to a cliff, we know that if we turn the wheel to the right, the car is going to run off the cliff and nothing good is going to come out of this.” You simulate the crash. You see the wreckage. You feel the fall. You turn the wheel. None of it was real. All of it was intelligence. Every AI has to crash a thousand times to learn what you imagined once and never did. That is not a performance gap. That is an architecture gap. LeCun: “The main problem we need to solve is how do we learn models of the world.” Not bigger models. Not more compute. Not another trillion tokens. World models. A machine that can run reality forward before it acts. The industry is scaling language. LeCun says language is a compression of thought. Not thought itself. You understood gravity before you could say the word. You grasped cause and effect before your first sentence. The deepest intelligence you will ever possess was built in total silence. And every lab on Earth is trying to reconstruct the mind from words alone. Physics does not care about your context window. A baby who learns that cups fall in a kitchen already knows that rocks fall off cliffs. No retraining. No fine-tuning. One model. Every environment. That is what intelligence actually is. Not prediction. Not pattern matching. Not scale. A simulation of reality so precise you rehearse the future before it exists. Every infant on Earth builds one. No machine ever has.

Dustin

121,544 görüntüleme • 16 gün önce