Loading video...

Video Failed to Load

Go Home

Lets build `Auto-RAG` where we let the LLM pull the data it needs from different sources. ๐Ÿ”Ž The user asks a question. ๐Ÿค” LLM decides whether to search its knowledge, memory, internet or make an API call. โœ๏ธ LLM answers with the context. Code:

174,580 views โ€ข 2 years ago โ€ขvia X (Twitter)

10 Comments

Vexxter's profile picture
Vexxter2 years ago

is there any way to run a local quantized LLM via ollama in this?? amazing project btw!

Ashpreet Bedi's profile picture
Ashpreet Bedi2 years ago

@XPhyxer1 absolutely the Hermes2-llama3 might work well here :)

Jordan A. Metzner's profile picture
Jordan A. Metzner2 years ago

Just read the read me. Any plans for Groq on Llama 3.

Ashpreet Bedi's profile picture
Ashpreet Bedi2 years ago

@mrjmetz on it!

Ameriki Singh ๐Ÿˆณ's profile picture
Ameriki Singh ๐Ÿˆณ2 years ago

Would love to see groq and Llmma3 on it

Emma.Ai's profile picture
Emma.Ai2 years ago

wow, can't wait to try this out

CoinCollector's profile picture
CoinCollector2 years ago

Ashpreet is coooooking

0xba0e7f9d's profile picture
0xba0e7f9d2 years ago

๐Ÿง‘โ€๐Ÿš€this is awesome demo!

Aws Abdo, Ph.D.'s profile picture
Aws Abdo, Ph.D.2 years ago

This work on automating retrieval and generation tasks is incredibly helpful. Thanks you! #MachineLearning #DataScience

Petamber's profile picture
Petamber2 years ago

You can also try Braveโ€™s Search API for web search

Related Videos

New short course: LLMs as Operating Systems: Agent Memory, created with Letta, and taught by its founders Charles Packer and Sarah Wooders. An LLM's input context window has limited space. Using a longer input context also costs more and results in slower processing. So, managing what's stored in this context window is important. In the innovative paper MemGPT: Towards LLMs as Operating Systems, its authors (which include the instructors) proposed using an LLM agent to manage this context window. Their system uses a large persistent memory that stores everything that could be included in the input context, and an agent decides what is actually included. Take the example of building a chatbot that needs to remember what's been said earlier in a conversation (perhaps over many days of interaction with a user). As the conversation's length grows, the memory management agent will move information from the input context to a persistent searchable database; summarize information to keep relevant facts in the input context; and restore relevant conversation elements from further back in time. This allows a chatbot to keep what's currently most relevant in its input context memory to generate the next response. When I read the original MemGPT paper, I thought it was an innovative technique for handling memory for LLMs. The open-source Letta framework, which we'll use in this course, makes MemGPT easy to implement. It adds memory to your LLM agents and gives them transparent long-term memory. In detail, youโ€™ll learn: - How to build an agent that can edit its own limited input context memory, using tools and multi-step reasoning - What is a memory hierarchy (an idea from computer operating systems, which use a cache to speed up memory access), and how these ideas apply to managing the LLM input context (where the input context window is a "cache" storing the most relevant information; and an agent decides what to move in and out of this to/from a larger persistent storage system) - How to implement multi-agent collaboration by letting different agents share blocks of memory This course will give you a sophisticated understanding of memory management for LLMs, which is important for chatbots having long conversations, and for complex agentic workflows. Please sign up here!

Andrew Ng

201,127 views โ€ข 1 year ago

RLM is the most import foundation of my Pi Harness (other than Pi of course). It's seeded with late interaction retrieval results (thanks to @lightonai for pylate). The Agent initiates it with query then.. ๐’๐ž๐ญ๐ฎ๐ฉ A python REPL is created and seeded with: 1. Late interaction search to pre-filter. Instead of doing top 3/5/10, it's top hundreds of documents. This is set into a `context` variable. 2. Python functions are loaded in to do more searches if `context` variable isn't enough. And to make llm calls with cheaper models in parallel batches. ๐ˆ๐ญ๐ž๐ซ๐š๐ญ๐ข๐จ๐ง ๐‹๐จ๐จ๐ฉ From there, an LLM iterates in the REPL based on the query. It's just like exploring in a jupyter notebook. The LLM writes prose (like a markdown cell) and code to be run in the REPL each turn. This allows the LLM to sort, filter, and synthesize information. It can fan out and ask smaller models to summarize, combine, contrast, or do anything else to documents to help it understand the data. After several turns the LLM reponds with the final answer. Either because it found the answer, or hit the budget limit. Context as a Python variable, LLM as the programmer, REPL as the runtime. ๐–๐ก๐ฒ ๐ƒ๐จ๐ž๐ฌ ๐“๐ก๐ข๐ฌ ๐–๐จ๐ซ๐ค 1. Richer Shell. Agents (and subagents) work by intermixing code and prose/thinking. But they use static scripts or bash that run and exit and start over each tool call. That's not ideal for exploration and synthesis of data. For that, state is useful to continue building and exploring the data as you learn more. There's a reason jupyter notebooks have been popular with data scientists. 2. Keeps main agent context clean. The better context you have the better the agent will perform (duh!). This means three thing: better human input, less missing search results, and less incorrect search results. Letting the agent iterate allows it to synthesize just what is needed and nothing else. All bad paths or peeks at something that turns out to be irrelevant stays out of main agent context. 3. Stack the good ideas! People often compare late interaction search vs RLM. Or static vs dynamic languages. Or agentic search vs semantic search. But...You can just use them all together for what they're each good at. Use them all for the area they're really great for. Read the full post which has more detail about how and why.

Isaac Flath

42,874 views โ€ข 5 months ago