Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

In depth explanation + examples of why sometimes you don't need rag (especially with gemini 2.0)

34,228 Aufrufe • vor 1 Jahr •via X (Twitter)

11 Kommentare

Profilbild von Sully
Sullyvor 1 Jahr

video up on youtube too

Profilbild von Aly Khairy
Aly Khairyvor 1 Jahr

It's simpler, but still costs a lot more. Maybe break down 1 query into several prompts like "is management optimistic about iphone sales next Q" and feed that to a "risks"prmpr and an "mda" prmpt, retrieve more chunks then feed those chunks to Gemini to get a more rounded answer

Profilbild von Sully
Sullyvor 1 Jahr

Would you rather pay more and have a accurate answer ? I think the cost equation starts to matter less and less

Profilbild von Sohail Hosseini
Sohail Hosseinivor 1 Jahr

Thanks for sharing!

Profilbild von Raduan Al-Shedivat
Raduan Al-Shedivatvor 1 Jahr

amazing explainer video, Sully. one thing that you've missed, which is HUGE: **context caching**[1]. basically, now you can send this whole 50k transcript, cache it for subsequent calls, and never ever bother breaking this down for RAG, even for long conversations, because this whole thing will get cached. I am switching most of my LLM use to Gemini after this. [1]:

Profilbild von Sully
Sullyvor 1 Jahr

Agreed I didn’t even mentioned which makes it way cheaper I will say googles caching needs work

Profilbild von D33PS33K
D33PS33Kvor 1 Jahr

KAG vs RAG

Profilbild von David
Davidvor 1 Jahr

great video, i just don't understand how google was able to achieve this without crazy hallucinations

Profilbild von Derek Cheung
Derek Cheungvor 1 Jahr

Thanks Sully! Great work

Profilbild von Nate Pratt
Nate Prattvor 1 Jahr

im extracting information from bank statments and need it REALLY accurate what would you recommend?

Profilbild von Sully
Sullyvor 1 Jahr

Gemini

Ähnliche Videos