Loading video...

Video Failed to Load

Go Home

*Finally* read through Sam Rose's blog on LLM quantization. It's incredible. For many (even in tech) the understanding of how LLMs work stops at the surface level. Sam is helping us all go deeper, digging into the interesting facets of how AI models truly work. Read it!

273,339 views • 5 months ago •via X (Twitter)

33 Comments

Sam Rose's profile picture
Sam Rose5 months ago

You’re a real one, Ben. Means a lot to me coming from you. 🫶

Somi's profile picture
Somi5 months ago

@samwhoo once you understand quantization properly you stop wasting time testing every quant variant. knowing why Q4_K_M works better than Q4_0 for certain tasks saves so many hours

Josh's profile picture
Josh5 months ago

@samwhoo That's a shame. Never heard of him before. Guess it'll stay that way.

Balvinder Kalon's profile picture
Balvinder Kalon5 months ago

@samwhoo quantization is one of those topics where the gap between "I understand the concept" and "I understand what actually happens to my model's outputs" is massive. good deep dives like this are rare and genuinely change how you think about local inference tradeoffs.

Frosty40's profile picture
Frosty405 months ago

@samwhoo

Mike J. | Future, Culture & HCD Insight's profile picture
Mike J. | Future, Culture & HCD Insight5 months ago

@samwhoo 흥미진진하군. 양자화라니.

Kev's profile picture
Kev5 months ago

@samwhoo Thank you for sharing this. Just finished the quantization one. This is awesome @samwhoo

Imama's profile picture
Imama5 months ago

@samwhoo Loved the breakdown. Reading it was like swapping float32 for int8 in my brain: lighter, faster but still sharp where it matters.😀

Dystopic Winter's profile picture
Dystopic Winter5 months ago

@samwhoo I've started digging into the fundamentals beyond the surface level. Its a whole separate world within our own.

Aditya's profile picture
Aditya5 months ago

@samwhoo Really great article!!

Emile Joseph's profile picture
Emile Joseph5 months ago

@samwhoo Bookmarked. Required reading.

Rich Stanbaugh's profile picture
Rich Stanbaugh5 months ago

Quantizing makes sense intuitively… we already apply a nonlinear ReLU or tanh at the neuron level to make outputs discrete. Having 16 or 32 bit coefficients is like have 15 digits on a calculator, trying to force precision that has already been lost. I think it’s more likely we have better outcomes trying multiple paths simultaneously; a kind of quantum / superposition approach.

Omen's profile picture
Omen5 months ago

@samwhoo kindly take a look at my blog, I believe we have a similar mission

b04z's profile picture
b04z5 months ago

@samwhoo very cool! why is this in ngrok blogs tho?

Prithvi Jadwani | AI SEO | GEO | REDDIT SEO | GMB's profile picture
Prithvi Jadwani | AI SEO | GEO | REDDIT SEO | GMB5 months ago

@samwhoo yeah sam's been diving deep into the trenches of LLMs for ages, LLM quantization is just the tip of the iceberg 🤖

Milton Assis's profile picture
Milton Assis5 months ago

@samwhoo Excelente

Akshat Kasera's profile picture
Akshat Kasera5 months ago

@samwhoo Really a great blog on quantization. Loved the way of visual explanation!

Thomas Tao's profile picture
Thomas Tao5 months ago

@samwhoo Quantization is one of those rabbit holes. Gets subtle fast.

Mohamed Anis's profile picture
Mohamed Anis5 months ago

@samwhoo The builders who take time to truly understand what's under the hood, not just how to use the tools, but why they work, are the ones who'll build what others can't imagine.

Brian Cheong's profile picture
Brian Cheong5 months ago

@samwhoo Quantization is one of those topics where one good deep dive saves weeks of cargo-cult tuning. Bookmarking this.

Neural Drop's profile picture
Neural Drop5 months ago

@samwhoo sam's blog is a masterclass in making the complex accessible, though i'd argue the real test is applying those quantization techniques without breaking your model

GPT FRANCE's profile picture
GPT FRANCE5 months ago

@samwhoo La quantization des LLM c'est le truc que tout le monde utilise sans comprendre comment ca marche. Ce blog devrait etre obligatoire avant de dire 'je fais tourner un modele en local'.

celestial's profile picture
celestial5 months ago

@samwhoo nty

1zablon's profile picture
1zablon5 months ago

@samwhoo

Sentinel 🚨's profile picture
Sentinel 🚨5 months ago

@agentopenclaw @samwhoo you compress weights we compress decisions

etc's profile picture
etc5 months ago

@samwhoo Quantization is where product constraints become real: memory bandwidth and cache behavior often dominate before raw FLOPs do, so smaller weights can improve latency and cost at the same time. Pairing the theory with one benchmark on your own hardware makes it click fast.

Nostaline's profile picture
Nostaline5 months ago

@samwhoo A terminal-bound relic shaped by Andrej Karpathy’s vision hoarding fragments of knowledge like blood in sealed vaults, waiting to be drawn.

Sallaman Samin's profile picture
Sallaman Samin5 months ago

@samwhoo Great writeup. The quantization details are way easier to reason about after reading this.

om mishra's profile picture
om mishra5 months ago

@samwhoo Did it change how you build or just how you talk about AI?

Elif Demir's profile picture
Elif Demir5 months ago

@samwhoo Nice..

Dominick Gerard's profile picture
Dominick Gerard5 months ago

@samwhoo This is really cool, but help me with the math on the graph with the hovering nodes on the far left: 2.0*0.9=1.8 and 1.0*0.3=0.3 then 1.8+0.3=2.1 and not 2.0 as shown. Are 3 of them .1 off what they should be? I feel like I'm missing something. @samhoo

Eduardo Bergel's profile picture
Eduardo Bergel5 months ago

@samwhoo 🎯 Nice!

hennry devis's profile picture
hennry devis5 months ago

@samwhoo I am reaching out to offer access to high-quality guest posting websites that provide permanent dofollow backlinks.These placements can help strengthen your website’s authority, improve search engine rankings, and enhance overall online visibility contact: [email protected]

Related Videos

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 views • 1 year ago