Loading video...

Video Failed to Load

Go Home

Have you used quantization with an open source machine learning library, and wondered how quantization works? How can you preserve model accuracy as you compress from 32 bits to 16, 8, or even 2 bits? In our new short course, Quantization in Depth, taught by Hugging Face's Marc Sun...

198,671 views • 2 years ago •via X (Twitter)

10 Comments

Quantcheck's profile picture
Quantcheck2 years ago

@huggingface Whether it is signal processing, data compression, or machine learning, Quantization plays a crucial role.

Saquib Mehmood's profile picture
Saquib Mehmood2 years ago

@huggingface Thanks. Very helpful refresher.

kevlarai's profile picture
kevlarai2 years ago

@huggingface I'd love to learn more about this. What are the suggested pre-reqs?

Malik KISSOUM's profile picture
Malik KISSOUM2 years ago

@huggingface This is fire 🔥🔥🔥, thank you for making deep learning so fun and accessible

AIxBlock's profile picture
AIxBlock2 years ago

@huggingface The detailed approach to understanding and implementing different quantization methods will undoubtedly empower many developers!

Vincent Valentine (CEO of UnOpen.ai)'s profile picture
Vincent Valentine (CEO of UnOpen.ai)2 years ago

@huggingface @AndrewYNg Fascinating course. Quantization intrigues me - compressing models while retaining accuracy? How does this technique balance resource optimization and performance? Exploring the intricacies seems insightful.

Data & Analytics's profile picture
Data & Analytics2 years ago

@huggingface @AndrewYNg Interesting topic! Quantization can be tricky, but preserving model accuracy is key. Have you tried any techniques to maintain accuracy during compression?

GPT.Biz's profile picture
GPT.Biz2 years ago

探索量化的奥秘吧,这门课程将带你从理论到实践,了解如何优化模型的存储与计算效率!

GeraDeluxer's profile picture
GeraDeluxer2 years ago

@huggingface Thanks a lot for the great AI content 🚀

Michael Guo's profile picture
Michael Guo2 years ago

@huggingface I have put this course on my radar for quite some time and thanks for the reminder and I need get it done

Related Videos

Tokenization -- turning text into a sequence of integers -- is a key part of generative AI, and most API providers charge per million tokens. How does tokenization work? Learn the details of tokenization and RAG optimization in Retrieval Optimization: From Tokenization to Vector Quantization, created in collaboration with Qdrant and taught by its Developer Relations Lead, Kacper Łukawski. This course focuses on Retrieval augmented generation (RAG), which has two steps: First, a retriever finds relevant information; then, the generator uses what’s retrieved as context to produce a response. You’ll learn to optimize the first step (the retriever) by understanding how tokenization works and how it impacts the relevance of your search. In addition, you will also learn to measure and improve retrieval quality, speed, and memory. In detail, you’ll: - Learn about the internal workings of the embedding models and how your text turns into vectors. - Understand how several tokenizers, such as Byte-Pair Encoding, WordPiece, Unigram, and SentencePiece work. - Explore common challenges with tokenizers, such as unknown tokens, domain-specific identifiers, and numerical values, that can negatively affect your vector search. - Understand how to measure the quality of your search across relevance, ranking, and score-related metrics. - Understand how the main parameters in "HNSW", a graph-based algorithm, affect the relevance and speed of vector search, and how to tune its parameters. - Experiment with the three major quantization methods – product, scalar, and binary – and learn how they impact memory requirements, search quality, and speed. By the end of this course, you’ll have a solid understanding of how tokenization functions and how to optimize vector search in your RAG systems. Please sign up here!

Andrew Ng

146,313 views • 1 year ago

New Short Course: Getting Structured LLM Output! Learn how to get structured outputs from your LLM applications in this course, built in partnership with .txt, and taught by Will Kurt, a Founding Engineer, and , Developer Relations Engineer. It's challenging for software to automatically parse through an LLM's freeform text outputs. Structured outputs—like JSON—solve this by converting natural language into consistent, clear, data that a machine can read and process. This course teaches you how to generate structured outputs while building several use cases, including a social media analysis agent. You’ll learn about structured outputs and efficient ways to generate outputs in your defined schema or format. You’ll begin by using structured output APIs, then use re-prompting libraries like “instructor” to generate structured output. Finally, you’ll learn how constrained decoding works; this is a very clever technique in which constraints are applied on each subsequent token generated, blocking any tokens that don’t fit your defined schema. In detail, you’ll: - Learn why structured outputs are important, how they allow for scalable software development, and the different approaches to generate them, including vendor-provided APIs, re-prompting libraries, and structured generation. - Build a simple social media agent using OpenAI’s structured output API, learn how to define a model's desired structured output using Pydantic, and perform basic programming with your outputs, such as importing structured data into a data frame using pandas. - Learn how to use the open-source library "instructor," which checks the structured output of the model and re-prompts the model until it validates the desired output, and explore the limitations of this approach. - Understand how structured generation by the “outlines” library works by modifying LLM logits, on a per-generated-token basis based on the desired format, to give a particular output structure. - Learn how regular expressions, which outlines works with, are represented as finite-state machines, and how they can be used to develop a range of structured outputs beyond JSON. By the end of this course, you’ll have broadened your knowledge of the approaches you can use to get structured outputs from your LLM applications. Please sign up here:

Andrew Ng

89,792 views • 1 year ago

New short course: Collaborative Writing and Coding with OpenAI Canvas! Explore new ways to write and code with OpenAI Canvas, a user-friendly interface that allows you to brainstorm, draft, and refine text and code in collaboration with ChatGPT. In the short course, created with OpenAI, and taught by , a research lead at OpenAI, you’ll learn to use Canvas to enhance your workflows. Canvas lets you go beyond simple chat interactions. It provides a side-by-side workspace where you and ChatGPT can edit and refine text or code collaboratively. This makes brainstorming, drafting, and iterating as you write feel more natural and effective. As the first major update to ChatGPT’s visual interface since its launch in 2022, Canvas gives a new, innovative approach to collaboration with AI. For instance, after writing the first version of your code, Canvas can review it and give suggestions for improvement. It can also help with debugging by adding logging, identifying problems to fix, and writing comments. In addition, you'll also learn what it takes to train the model for an interface like Canvas. In this video-only short course, you’ll: - Learn how to ask for in-line feedback and control the iteration of your work by directly editing selected areas of your text or code from the model’s output. - Learn how to access quick automation tools in a shortcut menu that allows you to modify your writing tone and length, enhance your code, and restore previous versions of your work. - Learn how to use Canvas as a research assistant tool with an example of asking the model to reason through the screenshot of a plot to write a research report, in which you can ask questions within the created report. - Ask the model to write Python code to replicate the graph seen on a screenshot image. - Go behind the scenes of how you can create a video game, such as Space Battleship, from scratch, edit it, and display it in one self-contained HTML file. - Get a real-world application example of creating a SQL database from the image of its architecture. - Understand the model training and design processes that power Canvas! Please sign up here:

Andrew Ng

128,180 views • 1 year ago

New short course: Evaluating AI Agents! Evals are important for driving AI system improvements, and in this course you'll learn to systematically assess and improve an AI agent’s performance. This is built in partnership with Arize AI and taught by John Gilhuly, Head of Developer Relations, and , Director of Product. I've often found evals to be a critical tool in the agent development process - they can be the difference between picking the right thing to work on vs. wasting weeks of effort. Whether you’re building a shopping assistant, coding agent, or research assistant, having a structured evaluation process helps you refine its performance systematically, rather than relying on random trial and error. This course shows you how to structure your evals to assess the performance of each component of an agent and its end-to-end performance. For each component, you select the appropriate evaluators, test examples, and performance metrics. This helps you identify areas for improvement both during development and in production. (If you're familiar with error analysis in supervised learning, think of this as adapting those ideas to agentic workflows.) In this course, you'll build an AI agent, and add observability to visualize and debug its steps. You’ll learn about code-based evals, in which you write code explicitly to test a certain step, as well as LLM-as-a-Judge evals, in which you prompt an LLM to efficiently come up with ways to evaluate more open-ended outputs. In detail, you’ll: - Understand key differences between evaluating LLM-based systems and traditional software testing. - Add observability to an agent by collecting traces of the steps taken by the agent and visualizing them - Choose the appropriate evaluator - code-based, LLM-as-a-Judge, human-annotation based - for each component. - Compute a convergence score to evaluate if your agent can respond to a query in an efficient number of steps. - Run structured experiments to improve the agent’s performance by exploring changes to the prompt, LLM model, or the agent’s logic. - Understand how to deploy these evaluation techniques to monitor the agent’s performance in production. By the end of this course, you’ll know how to trace AI agents, systematically evaluate them, and improve their performance. Please sign up here:

Andrew Ng

126,478 views • 1 year ago

Open source software is GREAT. But "open source" AI is NOT like software - it's VERY different. Rob Miles cuts through the bullshit: ROB: Oh, hey, Meta. I heard Llama's weights leaked. That's rough, man. Information security's hard. How you holding up? META: Oh, we're great. Yeah, we're fine. We... actually, that was deliberate. We meant to do that. ROB MILES: Oh, really? META: Yeah... well, the second time anyway. It's called open source. Look it up. ROB MILES: Oh. Well, I love free and open source software, but do those principles really apply to network weights? How does that work? META: Open source is good for users because it lets them read the source code and see what the program is really doing and how it works. ROB MILES: Wait, have you found a way to tell how a model works by looking at its weights? META: No. But, it lets developers all over the world spot bugs in the code and submit patches. ROB: Wait, people are fixing bugs in Llama's weights? META: Well, no. People can fine tune it themselves, though. ROB: ?? Other companies offer fine tuning through APIs. ... So, hang on, if you can't actually read the code and know what it's doing, then network weights are effectively a compiled binary. So, in what sense is this open source? Why not call it like public weights? Why call it open source at all? META: I love open source. ROB: Well, I know a lot of your employees do, but you don't love anything. You're a giant corporation. What's in it for you? META: I love, love open source.

AI Notkilleveryoneism Memes ⏸️

107,362 views • 2 years ago