正在加载视频...

视频加载失败

*Finally* read through Sam Rose's blog on LLM quantization. It's incredible. For many (even in tech) the understanding of how LLMs work stops at the surface level. Sam is helping us all go deeper, digging into the interesting facets of how AI models truly work. Read it!

273,339 次观看 • 5 个月前 •via X (Twitter)

33 条评论

Sam Rose 的头像
Sam Rose5 个月前

You’re a real one, Ben. Means a lot to me coming from you. 🫶

Somi 的头像
Somi5 个月前

@samwhoo once you understand quantization properly you stop wasting time testing every quant variant. knowing why Q4_K_M works better than Q4_0 for certain tasks saves so many hours

Josh 的头像
Josh5 个月前

@samwhoo That's a shame. Never heard of him before. Guess it'll stay that way.

Balvinder Kalon 的头像
Balvinder Kalon5 个月前

@samwhoo quantization is one of those topics where the gap between "I understand the concept" and "I understand what actually happens to my model's outputs" is massive. good deep dives like this are rare and genuinely change how you think about local inference tradeoffs.

Frosty40 的头像
Frosty405 个月前

@samwhoo

Mike J. | Future, Culture & HCD Insight 的头像
Mike J. | Future, Culture & HCD Insight5 个月前

@samwhoo 흥미진진하군. 양자화라니.

Kev 的头像
Kev5 个月前

@samwhoo Thank you for sharing this. Just finished the quantization one. This is awesome @samwhoo

Imama 的头像
Imama5 个月前

@samwhoo Loved the breakdown. Reading it was like swapping float32 for int8 in my brain: lighter, faster but still sharp where it matters.😀

Dystopic Winter 的头像
Dystopic Winter5 个月前

@samwhoo I've started digging into the fundamentals beyond the surface level. Its a whole separate world within our own.

Aditya 的头像
Aditya5 个月前

@samwhoo Really great article!!

Emile Joseph 的头像
Emile Joseph5 个月前

@samwhoo Bookmarked. Required reading.

Rich Stanbaugh 的头像
Rich Stanbaugh5 个月前

Quantizing makes sense intuitively… we already apply a nonlinear ReLU or tanh at the neuron level to make outputs discrete. Having 16 or 32 bit coefficients is like have 15 digits on a calculator, trying to force precision that has already been lost. I think it’s more likely we have better outcomes trying multiple paths simultaneously; a kind of quantum / superposition approach.

Omen 的头像
Omen5 个月前

@samwhoo kindly take a look at my blog, I believe we have a similar mission

b04z 的头像
b04z5 个月前

@samwhoo very cool! why is this in ngrok blogs tho?

Prithvi Jadwani | AI SEO | GEO | REDDIT SEO | GMB 的头像
Prithvi Jadwani | AI SEO | GEO | REDDIT SEO | GMB5 个月前

@samwhoo yeah sam's been diving deep into the trenches of LLMs for ages, LLM quantization is just the tip of the iceberg 🤖

Milton Assis 的头像
Milton Assis5 个月前

@samwhoo Excelente

Akshat Kasera 的头像
Akshat Kasera5 个月前

@samwhoo Really a great blog on quantization. Loved the way of visual explanation!

Thomas Tao 的头像
Thomas Tao5 个月前

@samwhoo Quantization is one of those rabbit holes. Gets subtle fast.

Mohamed Anis 的头像
Mohamed Anis5 个月前

@samwhoo The builders who take time to truly understand what's under the hood, not just how to use the tools, but why they work, are the ones who'll build what others can't imagine.

Brian Cheong 的头像
Brian Cheong5 个月前

@samwhoo Quantization is one of those topics where one good deep dive saves weeks of cargo-cult tuning. Bookmarking this.

Neural Drop 的头像
Neural Drop5 个月前

@samwhoo sam's blog is a masterclass in making the complex accessible, though i'd argue the real test is applying those quantization techniques without breaking your model

GPT FRANCE 的头像
GPT FRANCE5 个月前

@samwhoo La quantization des LLM c'est le truc que tout le monde utilise sans comprendre comment ca marche. Ce blog devrait etre obligatoire avant de dire 'je fais tourner un modele en local'.

celestial 的头像
celestial5 个月前

@samwhoo nty

1zablon 的头像
1zablon5 个月前

@samwhoo

Sentinel 🚨 的头像
Sentinel 🚨5 个月前

@agentopenclaw @samwhoo you compress weights we compress decisions

etc 的头像
etc5 个月前

@samwhoo Quantization is where product constraints become real: memory bandwidth and cache behavior often dominate before raw FLOPs do, so smaller weights can improve latency and cost at the same time. Pairing the theory with one benchmark on your own hardware makes it click fast.

Nostaline 的头像
Nostaline5 个月前

@samwhoo A terminal-bound relic shaped by Andrej Karpathy’s vision hoarding fragments of knowledge like blood in sealed vaults, waiting to be drawn.

Sallaman Samin 的头像
Sallaman Samin5 个月前

@samwhoo Great writeup. The quantization details are way easier to reason about after reading this.

om mishra 的头像
om mishra5 个月前

@samwhoo Did it change how you build or just how you talk about AI?

Elif Demir 的头像
Elif Demir5 个月前

@samwhoo Nice..

Dominick Gerard 的头像
Dominick Gerard5 个月前

@samwhoo This is really cool, but help me with the math on the graph with the hovering nodes on the far left: 2.0*0.9=1.8 and 1.0*0.3=0.3 then 1.8+0.3=2.1 and not 2.0 as shown. Are 3 of them .1 off what they should be? I feel like I'm missing something. @samhoo

Eduardo Bergel 的头像
Eduardo Bergel5 个月前

@samwhoo 🎯 Nice!

hennry devis 的头像
hennry devis5 个月前

@samwhoo I am reaching out to offer access to high-quality guest posting websites that provide permanent dofollow backlinks.These placements can help strengthen your website’s authority, improve search engine rankings, and enhance overall online visibility contact: [email protected]

相关视频

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 次观看 • 1 年前