Loading video...
Video Failed to Load
Another quick lecture -- I've been asked many times for prereq's to my book and what you should know, so built a little lecture (with GLM 5.2) to cover some more basics. Topics include: 00:00 Introduction & Course Prerequisites 01:37 Language Models Overview 02:47 The LM Head 04:29 Softmax... show more
52,668 views • 3 months ago •via X (Twitter)
15 Comments

Lecture 0:

(video is ft. phoebe, second half with her on my lap)

I’ve really gotta sit down and spend some time with these. Btw is there a space you’ve made yet for students taking the videos to yap? (Discord etc)

Yeah there’s a book discord, link on site / description (don’t past discord links on x or get got)

Bootiful

5.2 has been blowing my socks off. I've never considered selling a kidney for a Mac studio to run it locally. Until now.

thank you so much. learning A LOT from everything you’re sharing: from to rlhf book.

That sounds super helpful! It’s great to see you breaking down the essentials for everyone. Can’t wait to dive into it!

This is amazing , thanks 🤩

KL divergence ve entropy başlığında bir detay: ikisi de bilgi teorisinin temeli ama çoğu kişi entropy'yi "belirsizlik" olarak, KL divergence'ı ise "mesafe" olarak düşünüyor. Aslında ikisi de olasılık dağılımlarının "ekonomisi" hakkında konuşuyor — entropy kaynağın ne kadar "pahalı" olduğunu, KL divergence ise yanlış modelle kodlama yaparken ne kadar "harcadığını" ölçüyor. Post-training'de bu fark kritik: cross-entropy sadece loss, KL divergence ise modelin orijinal bilgisinden ne kadar sapmaya izin verdiğini kontrol ediyor.

Great Thanks for your effort and posting this prerequisite, Nathan!

If it does one annoying thing for me every day, I am listening.

The timestamped prereq lecture is a good filter. If someone cannot point to the LM head screen in the video, I trust their agent product claim a little less.

Interesting

Is the LM head really the biggest bottleneck for training efficiency now? I am reading that the softmax bottleneck in that last layer can suppress 95% of the gradient norm during pretraining. It is good you covered this because it makes some patterns almost impossible to learn.
