Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Andrew Ng + Google showed the foundation behind modern "RAG", "Graph RAG" and semantic retrieval: • 12:00 - turning text into embeddings that capture meaning • 24:00 - visualizing semantic relationships between vectors • 35:00 - using embeddings for classification, clustering and outlier detection • 50:00 - controlling LLM...

34,790 görüntüleme • 1 ay önce •via X (Twitter)

15 Yorum

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack1 ay önce

@0xMorlex watched it last night. the part on embeddings for clustering? blew my mind. so many use cases I hadn't considered yet!

Chase profil fotoğrafı
Chase1 ay önce

really useful because it separates representation from retrieval from generation instead of treating rag as one monolithic feature

Morlex profil fotoğrafı
Morlex1 ay önce

yeah bro

國泰世華 profil fotoğrafı
國泰世華1 ay önce

Good

MiklaS profil fotoğrafı
MiklaS1 ay önce

solid resource, graph RAG saved me on a research bot where plain vector search kept losing entity relationships across long docs. still cap chunk size aggressively though, oversized nodes torched my retrieval precision fast.

YuML AI profil fotoğrafı
YuML AI1 ay önce

@grok when was this lecture?

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

One level lower makes all the difference

Saman Ahmed profil fotoğrafı
Saman Ahmed1 ay önce

This is the layer everyone skips.

Alex Smith profil fotoğrafı
Alex Smith1 ay önce

yeah similarity search alone misses this. we moved retrieval over to HydraDB since it's graph based and answers got noticeably more precise once entities were actually linked instead of just scored.

ALEXYZ profil fotoğrafı
ALEXYZ1 ay önce

embeddings make semantic search click

Osamik 🤖 profil fotoğrafı
Osamik 🤖1 ay önce

The chunk retriever example nails why naive RAG fails — keyword matching disguised as semantic search. Curious if the roadmap covers re-ranking after retrieval, that's usually where the real quality gap shows up.

cvxv666 profil fotoğrafı
cvxv6661 ay önce

the most logical path for the evolution of AI

Morlex profil fotoğrafı
Morlex1 ay önce

yep, andrej broke it all down cleanly like always

中國信託 profil fotoğrafı
中國信託1 ay önce

Good

Dipanshu Kushwaha profil fotoğrafı
Dipanshu Kushwaha1 ay önce

This is a great breakdown. It really clarifies how these concepts work together.

Benzer Videolar

Announcing a new Coursera course: Retrieval Augmented Generation (RAG) You'll learn to build high performance, production-ready RAG systems in this hands-on, in-depth course created by and taught by , experienced AI and ML engineer, researcher, and educator. RAG is a critical component today of many LLM-based applications in customer support, internal company Q&A systems, even many of the leading chatbots that use web search to answer your questions. This course teaches you in-depth how to make RAG work well. LLMs can produce generic or outdated responses, especially when asked specialized questions not covered in its training data. RAG is the most widely used technique for addressing this. It brings in data from new data sources, such as internal documents or recent news, to give the LLM the relevant context to private, recent, or specialized information. This lets it generate more grounded and accurate responses. In this course, you’ll learn to design and implement every part of a RAG system, from retrievers to vector databases to generation to evals. You’ll learn about the fundamental principles behind RAG and how to optimize it at both the component and whole-system levels. As AI evolves, RAG is evolving too. New models can handle longer context windows, reason more effectively, and can be parts of complex agentic workflows. One exciting growth area is Agentic RAG, in which an AI agent at runtime (rather than it being hardcoded at development time) autonomously decides what data to retrieve, and when/how to go deeper. Even with this evolution, access to high-quality data at runtime is essential, which is why RAG is a key part of so many applications. You'll learn via hands-on experiences to: - Build a RAG system with retrieval and prompt augmentation - Compare retrieval methods like BM25, semantic search, and Reciprocal Rank Fusion - Chunk, index, and retrieve documents using a Weaviate vector database and a news dataset - Develop a chatbot, using open-source LLMs hosted by Together AI, for a fictional store that answers product and FAQ questions - Use evals to drive improving reliability, and incorporate multi-modal data RAG is an important foundational technique. Become good at it through this course! Please sign up here:

Andrew Ng

124,656 görüntüleme • 1 yıl önce

Tokenization -- turning text into a sequence of integers -- is a key part of generative AI, and most API providers charge per million tokens. How does tokenization work? Learn the details of tokenization and RAG optimization in Retrieval Optimization: From Tokenization to Vector Quantization, created in collaboration with Qdrant and taught by its Developer Relations Lead, Kacper Łukawski. This course focuses on Retrieval augmented generation (RAG), which has two steps: First, a retriever finds relevant information; then, the generator uses what’s retrieved as context to produce a response. You’ll learn to optimize the first step (the retriever) by understanding how tokenization works and how it impacts the relevance of your search. In addition, you will also learn to measure and improve retrieval quality, speed, and memory. In detail, you’ll: - Learn about the internal workings of the embedding models and how your text turns into vectors. - Understand how several tokenizers, such as Byte-Pair Encoding, WordPiece, Unigram, and SentencePiece work. - Explore common challenges with tokenizers, such as unknown tokens, domain-specific identifiers, and numerical values, that can negatively affect your vector search. - Understand how to measure the quality of your search across relevance, ranking, and score-related metrics. - Understand how the main parameters in "HNSW", a graph-based algorithm, affect the relevance and speed of vector search, and how to tune its parameters. - Experiment with the three major quantization methods – product, scalar, and binary – and learn how they impact memory requirements, search quality, and speed. By the end of this course, you’ll have a solid understanding of how tokenization functions and how to optimize vector search in your RAG systems. Please sign up here!

Andrew Ng

146,313 görüntüleme • 2 yıl önce