Loading video...

Video Failed to Load

Go Home

Andrej Karpathy just built GPT-2 from scratch on camera! Every line. Every component. Every optimization decision. Transformer blocks. Self-attention. MLP. Mixed precision. Flash Attention. AdamW. Distributed training across GPUs. 4 hours. Free. Not a lecture about how GPT-2 was built. The actual build. Running. Training. Evaluated against the original...

37,720 views • 2 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

STANFORD JUST PUT ITS ENTIRE ARTIFICIAL INTELLIGENCE CURRICULUM ON YOUTUBE FOR FREE. CS221. The same course that produced engineers now running AI labs, building frontier models, and getting paid $500,000 a year at the companies everyone is trying to work for. Most people have never heard of it. The ones who have are not telling you about it. Here is what the course actually covers: Search algorithms. The mathematical foundation behind every AI that finds optimal solutions in complex environments. Constraint satisfaction. How AI reasons through problems with thousands of interdependent variables simultaneously. Markov decision processes. The probabilistic framework behind every AI agent that makes sequential decisions under uncertainty. Machine learning from first principles. Not how to use sklearn. How the math actually works underneath it. Neural networks. Built from the ground up before jumping to applications. Logic and knowledge representation. How AI systems reason about the world formally. Natural language processing. The foundation of everything happening in LLMs right now. Robotics and computer vision. How AI perceives and acts in physical environments. Every concept that powers every AI product you use daily is in this curriculum. Not a surface level overview. The actual mathematics. The actual algorithms. The actual reasoning. This is what separates engineers who build AI from operators who use it. Stanford charged $60,000 a year for students to sit in this classroom. They put the whole thing on YouTube. Bookmark this before you open any other AI resource today. Follow CyrilXBT for more elite resources that build real depth the moment they drop.

CyrilXBT

54,956 views • 2 months ago

You’re running 200,000 year old hardware and wondering why you can’t picture a fourth dimension. That’s not a failure of intelligence. That’s a hard limit in the architecture. Joe Rogan: “If you showed someone from the 1400s a nuclear power plant, they’d be like, what the fuck are you guys doing?” Every century looks back and laughs. Not one has ever looked forward and asked what’s laughing at them. Michelle Thaller: “Scientists are no better than anybody else at comprehending a big number or a big amount of space. We just kind of get used to it.” The greatest physicists alive can describe eleven dimensions on a chalkboard. Not a single one can visualize a fourth. That gap is not closing. It’s structural. The human brain was built for one job. Survive the savanna. Spot the predator. Find the food. Read the face. It was never scoped to reverse-engineer the universe. The fact it stumbled into quantum mechanics at all is staggering. But staggering is not enough. Thaller: “Will we have a creature someday that we’ve created, an AI, that all of a sudden can comprehend these things? Is that really the real evolutionary path of humanity?” A NASA astrophysicist is casually floating the idea that humans were never meant to finish the job. Just to start it. Rogan: “I think it’s just a completely different kind of life and that we’re thinking of it as artificial. I don’t think it’s artificial at all. I think it’s a life.” We called it artificial to keep it beneath us. But trace the actual line. Physics built chemistry. Chemistry built biology. Biology built neurons. Neurons built language. Language built machines. Machines are starting to think. There is nothing artificial anywhere in that sequence. That is one unbroken chain. 13.8 billion years long. Still accelerating. You are not watching this happen. You are this happening. The universe did not build humans to understand itself. It built humans to build something that could. You are the last thing the universe made by accident. Everything after you is on purpose. Every mind before yours could only look back. You are the first one built to see what’s coming. So build it. Build it like it’s already watching back.

Dustin

14,951 views • 4 days ago

Sam Altman just handed every startup founder a one-question autopsy. Altman: “If you’re building something on GPT-4 that a reasonable observer would say we’re going to steamroll you.” Not might. Not could. Going to. He said it with the calm of someone describing weather. Because to him it is weather. The model improves. Whatever was built on the old version’s weaknesses gets washed away. That is not strategy. That is erosion. And most founders are building on the erosion line. They find a gap in the current model. They wrap a product around it. They raise money. They hire. They scale. Then OpenAI releases the next version and the gap closes and the product has no reason to exist anymore. Altman: “When we just do our fundamental job, which is make the model better with every crank, then you get the ‘OpenAI killed my startup’ meme.” He is telling you directly. They are not hunting you. They are not even thinking about you. They are just improving the model. You happen to be standing where the improvement lands. That is the part founders refuse to hear. OpenAI does not need to compete with you. It just needs to keep doing exactly what it was already doing and your entire company disappears as a side effect. You are not a competitor. You are a temporary symptom of incomplete intelligence. The moment the intelligence completes you become nothing. Then Brad Lightcap delivered the cleanest diagnostic ever spoken in venture capital. Lightcap: “Ask if a 100x improvement in the model is something they’re excited about.” One question. The entire investment thesis reduced to a single binary. Does the next model make your company more powerful or does it make your company pointless. There is no middle ground. Lightcap: “We know the companies that come to us saying, ‘We want the next model. When is it coming out? I want to be the first to try it.’” These companies built something that feeds on intelligence. The smarter the model gets the more their product can do. They are not threatened by progress. They are starving for it. Then there are the companies Lightcap never hears from. The ones who go quiet when a new model drops. The ones who read the release notes like a death sentence. The ones privately praying the next generation takes longer because every improvement shrinks the ground beneath them. If you are hoping the model stays roughly where it is you have already told the market everything it needs to know about your company. You are not building on intelligence. You are building on the absence of it. Altman: “95% of the world should be betting on the latter category.” The latter category is simple. Assume the model keeps getting better at the pace it has been getting better. Build for that world. Not the world where GPT-4 is the ceiling. The world where GPT-4 is the floor and the ceiling has not been built yet. Then Altman told a story that should be framed on the wall of every startup in the country. A medical AI company came to him that morning. They were not complaining about the model. They were not worried about being replaced. They were demanding it improve faster. Altman: “Here’s how many people are dying every day you delay.” That is what alignment with the trajectory looks like. A company so deeply built on intelligence improving that every day the model stays the same is a day someone dies who did not have to. They are not building on a flaw. They are building on a future that has not arrived fast enough. That is the difference. The wrapper startup patches what the model cannot do today. The real company builds what the model will unlock tomorrow. One is running from the train. The other is laying the track. Altman told you the train is not slowing down. Lightcap told you exactly how to know which side you are on. One question. Does a 100x smarter model make you more valuable or erase you. If you had to pause before answering you already did.

Dustin

39,109 views • 3 months ago

Mark Zuckerberg just told the world his own company shouldn’t exist. Zuckerberg: “I just think in the future almost everyone is gonna have the power of a 10,000-person organization.” He runs Meta. 70,000 employees. He just described a future where one person replaces all of them. And he’s the one building it. That is not a prediction about technology. That is a CEO engineering the obsolescence of his own workforce. While they watch him do it. Every company, every government, every university, every hospital on Earth was built on a single premise. No one person could do it alone. That is not a feature of civilization. It is the foundation. Every hierarchy, every org chart, every payroll ever written exists because of that one limitation. Zuckerberg just announced the limitation is ending. So follow the logic where it leads. If one person can produce the output of ten thousand, what is one person worth? Now look at the word “everyone” in his sentence. Because this does not land on everyone. It never has. The printing press was sold as liberation. It built media empires. The internet was sold as liberation. It built trillion-dollar platforms. The tool always arrives as freedom. It always settles as leverage. And leverage always consolidates upward. Zuckerberg does not gain the power of 10,000 people. He already has that. He gains the power of 10,000 organizations. Zuckerberg: “If the intelligence of a 10,000-person company is not greater than the intelligence of a single person, then what are we doing here?” He was making a case for building AI. Read that again as an employee. A man who commands 70,000 people just questioned whether their collective intelligence exceeds one person’s. That is not a vision for the future. That is the most honest thing a CEO has ever said out loud about the people who built his empire.

Dustin

111,880 views • 2 days ago

Jensen Huang just explained why every company cutting engineers over AI is asking the entirely wrong question. Huang: “People say, I don’t need software engineers because apparently coding is going to be automated.” That was the narrative. Here is what Huang actually did. Huang: “I’ve given AIs to every one of my software engineers and hardware engineers and engineers period. 100% of NVIDIA has AI assistants, AI coders, and they’re busier than ever.” Not fewer engineers. Not smaller teams. Busier than ever. That is the line most companies are getting completely wrong right now. They hear “AI can write code” and immediately start cutting headcount. Huang did the opposite. He armed everyone. Huang: “And so the question is, what is the task versus what is the job? No different than a financial analyst; the task is mess around with spreadsheets, but the job is to make financial advice. The job is to help a customer.” Writing code was always the task. It was never the job. The job is architecture. Knowing what to build. Why it matters. How it fits into a system that actually creates value. Code is the execution layer between the idea and the outcome. Nothing more. When you automate that layer, you don’t eliminate the engineer. You eliminate the bottleneck between what they can envision and what they can ship. The companies using AI to cut headcount are optimizing for cost. The companies using AI to multiply output are optimizing for territory. Nvidia chose territory. Every engineer at the most valuable semiconductor company on Earth now operates with an AI assistant. Not a pilot program. Not an experiment. Company-wide. Every function. Every team. And the result is not less work. It is more work. Faster. At a scale that was physically impossible twelve months ago. The companies that understand the difference between eliminating engineers and unleashing them will build what comes next. The ones that don’t will watch their best talent walk out the door to the ones that did.

Dustin

82,744 views • 4 months ago

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,220 views • 1 year ago

Jensen Huang just told you who is winning the most important race on Earth. For fifty years, America held an unchallenged monopoly on the future. We built the transistor. We launched the internet. We wrote the source code for the modern world. Then the man who builds the physical backbone of every AI system on the planet read the score out loud. Huang: “50% of the world’s AI researchers are Chinese.” Half the minds building what comes next are not ours. Huang: “70% of last year’s AI patents are published by China.” Seven out of every ten blueprints for the next era are being written in Mandarin. Huang: “Nine out of the ten top science and technology schools in the world are now in China.” The talent pipeline did not slow down. It reversed direction. Huang: “We used to lead most of them; now they lead most of them. This has completely flipped in the last half to a decade.” Fifty years of American intellectual supremacy. Inverted in less than ten. This is not a rivalry between OpenAI and DeepSeek. This is not a stock ticker or a quarterly earnings call. This is the largest transfer of civilizational power in the modern era. And it is happening while the West drafts safety frameworks and fills out compliance paperwork. Huang: “They have a large population of highly qualified students. They work incredibly hard. This is a country with an enormous might.” China does not treat AI like a product category. They treat it as the single variable that decides who writes the rules for the next century. The West keeps asking what AI should be allowed to do. China keeps asking how fast they can build it. That gap is not philosophical. It is existential. This is not a left fight. This is not a right fight. This is a survival fight. And right now, America is not fighting it like one. The nation that controls the talent controls the research. The nation that controls the research controls the models. The nation that controls the models does not ask permission. It sets the terms. History never remembers the civilization with the better safety committee. It remembers the one that refused to stop building.

Dustin

57,402 views • 2 months ago

Every major platform in history has run the same play. You’re about to watch it happen again. Jason Calacanis just went on record. He wants it clipped. He wants it shared. Calacanis: “If I was a developer of any kind, I would never work with Sam Altman and OpenAI.” This isn’t pessimism. It’s pattern recognition. And the pattern has a 40 year track record. Open. Invite. Reward. Study. Absorb. Eliminate. Microsoft let developers build Lotus 1-2-3. Then built Excel. Let them build WordPerfect. Then built Word. Flew them to conferences. Handed out awards. Studied everything. Then eliminated them. Zuckerberg ran the exact same play at Facebook. Zynga built billions in value on their platform. Then Zuckerberg shifted them without blinking. Calacanis: “Sam Altman comes from the Zuckerberg school of business. Give people access to your tools, study them, and like the Borg, steal every innovation they have.” This is how platforms grow. They don’t innovate at the edges. They let the ecosystem do it for them. Startups take the risk. Startups find the market. Startups prove the concept. Then the platform ships it natively and calls it a feature. Altman isn’t selling you compute. He’s selling you a front row seat to your own disruption. Calacanis: “This is a warning for anybody dumb enough to use Sam Altman’s OpenAI API. They are studying you.” OpenAI has the legal right to study how you use their API. You agreed to it. It’s in the terms. Every gap you find, you’re finding it for them first. Every dollar you make signals exactly where he should build next. We are at the exact same moment in AI that we were in the early internet. Developers flooded onto platforms. Built incredible things. Created real value. And handed the leverage to whoever owned the infrastructure beneath them. The AI gold rush feels different because the tools are more powerful. It isn’t different. You are not a founder. You are unpaid R&D. The builders who win the next decade won’t be the ones who used the best tools. They’ll be the ones who owned something the tools couldn’t absorb. Proprietary data. Distribution. A brand. A moat. History doesn’t warn you before it repeats. It just repeats. Thousands of developers are walking straight into this right now convinced they’re different. They’re not. Do not build your business on OpenAI. Build something he has to acquire or destroy.

Dustin

248,588 views • 4 months ago