Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

*Finally* read through Sam Rose's blog on LLM quantization. It's incredible. For many (even in tech) the understanding of how LLMs work stops at the surface level. Sam is helping us all go deeper, digging into the interesting facets of how AI models truly work. Read it!

273,339 Aufrufe • vor 5 Monaten •via X (Twitter)

33 Kommentare

Profilbild von Sam Rose
Sam Rosevor 5 Monaten

You’re a real one, Ben. Means a lot to me coming from you. 🫶

Profilbild von Somi
Somivor 5 Monaten

@samwhoo once you understand quantization properly you stop wasting time testing every quant variant. knowing why Q4_K_M works better than Q4_0 for certain tasks saves so many hours

Profilbild von Josh
Joshvor 5 Monaten

@samwhoo That's a shame. Never heard of him before. Guess it'll stay that way.

Profilbild von Balvinder Kalon
Balvinder Kalonvor 5 Monaten

@samwhoo quantization is one of those topics where the gap between "I understand the concept" and "I understand what actually happens to my model's outputs" is massive. good deep dives like this are rare and genuinely change how you think about local inference tradeoffs.

Profilbild von Frosty40
Frosty40vor 5 Monaten

@samwhoo

Profilbild von Mike J. | Future, Culture & HCD Insight
Mike J. | Future, Culture & HCD Insightvor 5 Monaten

@samwhoo 흥미진진하군. 양자화라니.

Profilbild von Kev
Kevvor 5 Monaten

@samwhoo Thank you for sharing this. Just finished the quantization one. This is awesome @samwhoo

Profilbild von Imama
Imamavor 5 Monaten

@samwhoo Loved the breakdown. Reading it was like swapping float32 for int8 in my brain: lighter, faster but still sharp where it matters.😀

Profilbild von Dystopic Winter
Dystopic Wintervor 5 Monaten

@samwhoo I've started digging into the fundamentals beyond the surface level. Its a whole separate world within our own.

Profilbild von Aditya
Adityavor 5 Monaten

@samwhoo Really great article!!

Profilbild von Emile Joseph
Emile Josephvor 5 Monaten

@samwhoo Bookmarked. Required reading.

Profilbild von Rich Stanbaugh
Rich Stanbaughvor 5 Monaten

Quantizing makes sense intuitively… we already apply a nonlinear ReLU or tanh at the neuron level to make outputs discrete. Having 16 or 32 bit coefficients is like have 15 digits on a calculator, trying to force precision that has already been lost. I think it’s more likely we have better outcomes trying multiple paths simultaneously; a kind of quantum / superposition approach.

Profilbild von Omen
Omenvor 5 Monaten

@samwhoo kindly take a look at my blog, I believe we have a similar mission

Profilbild von b04z
b04zvor 5 Monaten

@samwhoo very cool! why is this in ngrok blogs tho?

Profilbild von Prithvi Jadwani | AI SEO | GEO | REDDIT SEO | GMB
Prithvi Jadwani | AI SEO | GEO | REDDIT SEO | GMBvor 5 Monaten

@samwhoo yeah sam's been diving deep into the trenches of LLMs for ages, LLM quantization is just the tip of the iceberg 🤖

Profilbild von Milton Assis
Milton Assisvor 5 Monaten

@samwhoo Excelente

Profilbild von Akshat Kasera
Akshat Kaseravor 5 Monaten

@samwhoo Really a great blog on quantization. Loved the way of visual explanation!

Profilbild von Thomas Tao
Thomas Taovor 5 Monaten

@samwhoo Quantization is one of those rabbit holes. Gets subtle fast.

Profilbild von Mohamed Anis
Mohamed Anisvor 5 Monaten

@samwhoo The builders who take time to truly understand what's under the hood, not just how to use the tools, but why they work, are the ones who'll build what others can't imagine.

Profilbild von Brian Cheong
Brian Cheongvor 5 Monaten

@samwhoo Quantization is one of those topics where one good deep dive saves weeks of cargo-cult tuning. Bookmarking this.

Profilbild von Neural Drop
Neural Dropvor 5 Monaten

@samwhoo sam's blog is a masterclass in making the complex accessible, though i'd argue the real test is applying those quantization techniques without breaking your model

Profilbild von GPT FRANCE
GPT FRANCEvor 5 Monaten

@samwhoo La quantization des LLM c'est le truc que tout le monde utilise sans comprendre comment ca marche. Ce blog devrait etre obligatoire avant de dire 'je fais tourner un modele en local'.

Profilbild von celestial
celestialvor 5 Monaten

@samwhoo nty

Profilbild von 1zablon
1zablonvor 5 Monaten

@samwhoo

Profilbild von Sentinel 🚨
Sentinel 🚨vor 5 Monaten

@agentopenclaw @samwhoo you compress weights we compress decisions

Profilbild von etc
etcvor 5 Monaten

@samwhoo Quantization is where product constraints become real: memory bandwidth and cache behavior often dominate before raw FLOPs do, so smaller weights can improve latency and cost at the same time. Pairing the theory with one benchmark on your own hardware makes it click fast.

Profilbild von Nostaline
Nostalinevor 5 Monaten

@samwhoo A terminal-bound relic shaped by Andrej Karpathy’s vision hoarding fragments of knowledge like blood in sealed vaults, waiting to be drawn.

Profilbild von Sallaman Samin
Sallaman Saminvor 5 Monaten

@samwhoo Great writeup. The quantization details are way easier to reason about after reading this.

Profilbild von om mishra
om mishravor 5 Monaten

@samwhoo Did it change how you build or just how you talk about AI?

Profilbild von Elif Demir
Elif Demirvor 5 Monaten

@samwhoo Nice..

Profilbild von Dominick Gerard
Dominick Gerardvor 5 Monaten

@samwhoo This is really cool, but help me with the math on the graph with the hovering nodes on the far left: 2.0*0.9=1.8 and 1.0*0.3=0.3 then 1.8+0.3=2.1 and not 2.0 as shown. Are 3 of them .1 off what they should be? I feel like I'm missing something. @samhoo

Profilbild von Eduardo Bergel
Eduardo Bergelvor 5 Monaten

@samwhoo 🎯 Nice!

Profilbild von hennry devis
hennry devisvor 5 Monaten

@samwhoo I am reaching out to offer access to high-quality guest posting websites that provide permanent dofollow backlinks.These placements can help strengthen your website’s authority, improve search engine rankings, and enhance overall online visibility contact: [email protected]

Ähnliche Videos

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 Aufrufe • vor 1 Jahr