Загрузка видео...

Не удалось загрузить видео

На главную

Storing too many tools in your context window increases latency and can lead to wrong tool selection. In this demo, we used LFM2.5-ColBERT-350M as a filter to only select the five most relevant tools among 151 options. It's fast and reliable, even without any specific fine-tuning. Try the demo...

33,824 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 12

Фото профиля Ronald Mannak
Ronald Mannak3 месяцев назад

Why LFM2.5-ColBERT-350M and not LFM2.5-Embedding-350M? Isn’t a single vector good enough for tool selection? It’s simpler to implement too I assume.

Фото профиля Trung Le
Trung Le3 месяцев назад

Tiny but powerful, that’s what I usually describe LFMs

Фото профиля Akshat Dwivedi
Akshat Dwivedi3 месяцев назад

I guess this will fit perfectly in my recent research : "

Фото профиля ✦ SK ✦
✦ SK ✦3 месяцев назад

How can we test with the API of the LFM 2.5 8B model?

Фото профиля -abk-
-abk-3 месяцев назад

@maximelabonne @tejus_sawjiani

Фото профиля mystic
mystic3 месяцев назад

W

Фото профиля Eddie,Jesse Bailey
Eddie,Jesse Bailey3 месяцев назад

curious how tight the correlation is between colbert rank and tool correctness. feels like a good spot to find weird model behaviors and test some reranking ideas.

Фото профиля AI Mastery Guide
AI Mastery Guide3 месяцев назад

Filtering to the 5 most relevant tools out of 151 without fine-tuning is impressive 👏

Фото профиля Aina Ai | Tools & Updates
Aina Ai | Tools & Updates3 месяцев назад

Smart move - 151 → 5 tools = less lag, better picks LFM2.5-ColBERT-350M filtering without fine-tuning is clutch.

Фото профиля Cyril Gupta
Cyril Gupta3 месяцев назад

This fixes a massive structural headache for agentic workflows.

Фото профиля Samian
Samian3 месяцев назад

151 tools is wild. filtering down to 5 keeps things snappy. how does it score ties when tool descriptions overlap hard

Фото профиля Alice The Ai Expert
Alice The Ai Expert3 месяцев назад

Smarter tool selection, zero latency bloat

Похожие видео

How can you solve complex tasks using a Large Language Model? Here is a 2-minute introduction to everything you need to know to 10x the quality of your results. Let's talk about three techniques, in order of complexity, starting with the easiest one: • In-Context Learning • Indexing + In-Context Learning • Fine-tuning In-Context Learning The team that trained GPT-3 found something they couldn't explain: You can condition a model using examples of how you want it to behave. I included an example prompt in the attached video. You can "teach" the model how you want it to interpret questions, select the correct answers, and format the results by giving a few examples. You can also give specific knowledge to the model that will be helpful when formulating answers. We call this approach "grounding the model." There's another example in the video. Indexing + In-Context Learning Unfortunately, there is a limit to how much data you can include in a prompt. We call this the "context size." One version of GPT-4 supports a context of approximately 6,000 words, while the other supports 25,000 words. Although this sounds like a lot, many applications need more than that. Imagine you wrote a book and want to build an application to answer any questions about your story. What happens if your book is longer than the context? That's where Indexing comes in. Using a model, you can turn every book passage into an embedding. These are vectors, numbers that "encode" the passage's text. You can then store these embeddings in a particular database that supports fast retrieval of these vectors. You can then turn any question into an embedding and search the database for the list of passages that are similar to that query. Instead of using the entire book to ask the model, you can now use the relevant passages as in-context information, effectively working around the context size limitation. Fine-tuning Fine-tuning can give you an extra boost to get reliable outputs from your LLM. It is, however, the most complex approach on the list. There are different approaches to fine-tuning a model with your data. A popular technique is to process your data with your LLM and use the outputs to train a new classifier that solves your specific task. Notice that here you aren't modifying the LLM. Instead, you are chaining it with your trained classifier. Another approach is to modify the parameters of the LLM using your data. Think of this as "rewiring" the model in a way that solves your particular task. The results and costs will vary depending on how many layers you want to fine-tune from the original model. Many companies think that fine-tuning is the solution to their problems. In my experience, many will benefit from exploring the other two approaches. I love explaining Machine Learning and Artificial Intelligence ideas. If you enjoy in-depth content like this, follow me Santiago so you don't miss what comes next.

Santiago

384,573 просмотров • 3 лет назад