Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

How do we cache prompts with LLMs? Interview with Tanishq Singh. This is a real interview question from a big tech company, asked to a candidate in their technical interview round. The video explains the answer in roughly 6 minutes. 00:00 Question - Prompt Caching with LLMs 00:39 Exact...

94,937 görüntüleme • 1 ay önce •via X (Twitter)

21 Yorum

kirankoukuntla profil fotoğrafı
kirankoukuntla1 ay önce

Is it right answer .. for the question .. question was how are you caching.. ?? .. But answer was how caching works.. only extra point added was about age out and LRU..

Dhruba Goswami profil fotoğrafı
Dhruba Goswami1 ay önce

hi gaurav, can you please update your ai engineering Course on interview ready also, so far loving your system design simplified course will go for ai course next

tfi profil fotoğrafı
tfi1 ay önce

have you developed any production grade agent?

Aditya Raj profil fotoğrafı
Aditya Raj1 ay önce

In most of the cases no one stores the prompt embedding to do a redis data send, it happens with pre-fix caching(KV) which are stored near to the GPU on premise, such that no two people get same responses even at enterprise level, and whatever places has cache, they use raw high-speed memory and exact-prefix hashing right on their GPU clusters, they don't use vector similarity to give the exact same response, as on enterprise level vector similarity can't be trusted. @gkcs_

Aman Madan profil fotoğrafı
Aman Madan1 ay önce

Hey gaurav, I really enjoy your informative videos. Was going thru above link, What is the fees for the program?

Chen Avnery profil fotoğrafı
Chen Avnery1 ay önce

semantic caching looks so clean on a whiteboard. in prod it's the thing that quietly serves the wrong cached answer and nobody notices for a week.

Manav Shah profil fotoğrafı
Manav Shah1 ay önce

Yes caching great for repeated tasks wanting the same information

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan1 ay önce

prompt caching breaks fast when user-specific context changes inside the supposedly stable prefix

Akash B A profil fotoğrafı
Akash B A1 ay önce

Wish you taught me the whole 4 year engineering course. 🥹💯🙏

vivek varikuti profil fotoğrafı
vivek varikuti1 ay önce

One nuance worth making explicit: cache hit rate is a product metric, not just an infra metric. Stable system-prompt prefixes, deterministic tool schemas, and avoiding volatile fields can matter as much as cache size.

Satgur Rao profil fotoğrafı
Satgur Rao1 ay önce

What a neat way of answering @gkcs_ .

ChengyouXin profil fotoğrafı
ChengyouXin1 ay önce

One subtle issue with semantic caching is that people often optimize the similarity threshold instead of validating correctness. Threshold tuning changes the precision recall tradeoff, but it doesn’t solve incorrect cache hits.

Maduxoo profil fotoğrafı
Maduxoo1 ay önce

Wonderful explanation. My question is at what point does having a cache becomes costlier than an LLM? What if I have SLMs in this case? At what point of LLM size, can we avoid cache completely?

Polyglot profil fotoğrafı
Polyglot1 ay önce

Naive

Ankush Singal profil fotoğrafı
Ankush Singal1 ay önce

why not use llm cache?

Proud Indian American profil fotoğrafı
Proud Indian American1 ay önce

KV cache

Dipanshu Kushwaha profil fotoğrafı
Dipanshu Kushwaha1 ay önce

This is a fantastic question to ask! It's great that they're focusing on practical aspects like prompt caching. Thanks for sharing!

Saeed Anwar profil fotoğrafı
Saeed Anwar1 ay önce

Prompt caching is one of those topics that sounds trivial until you realize cache invalidation across distributed inference servers is its own nightmare. Semantic similarity based caching sounds great until two nearly identical prompts need completely different outputs.

未知 profil fotoğrafı
未知1 ay önce

现象背后有更深的逻辑。高手和普通人的区别,就在于愿不愿意多问几个为什么。大多数人停在第一个答案,真正的高手会一直追问到根因。

Ashutosh Sharma profil fotoğrafı
Ashutosh Sharma1 ay önce

Really useful

Fatah profil fotoğrafı
Fatah1 ay önce

prompt caching dropped our API costs 40% almost overnight. if you have a heavy system prompt you're repeating every call and you're not caching, that's just money you're leaving on the table.

Benzer Videolar

AI has a trust problem. Verifiability is the solution. Our GM of AI Nima Vaziri sat down with a16z’s Ali Yahya and Dan Boneh of Stanford University to map the deepest fault lines in AI today. ☁️ Models we can’t trust ☁️ Current providers can censor, shut down, or shift rules overnight. Outsourced training hides backdoors. Even “open” weights don’t prove what’s actually running. Trust. Backdoors. Black boxes. The path forward is clear: 🔥 Verifiable evals 🔥 Verifiable inference 🔥 TEEs for hardware-backed integrity 🔥 Infra beyond single points of control 🔥 Blockchains as coordination layers for AI From “trust us” to “verify yourself.” That’s the shift. That’s the unlock. The frontier is here. The builders decide what comes next. Create and use AI that’s incentive aligned with you. Timestamps: 00:00:00 Introduction: AI & Crypto Intersection Overview 00:01:58 Four Major AI-Crypto Trends 00:02:44 AI Agents Need Financial Infrastructure 00:04:03 Proof of Humanity: Fighting AI-Generated Content 00:04:17 Decentralizing AI Infrastructure Networks 00:04:44 Synthetic Life: Autonomous AI Agents 00:06:20 Verifiable AI 00:10:16 Current Performance Numbers for AI Proofs 00:13:18 The Era of Experience in AI Learning 00:14:56 AI Agents Having Life of its Own 00:18:21 Algorithmic Fairness & Verifiable Models 00:23:18 Privacy in AI: Trusted Execution Environments 00:25:47 Economic Incentive for Open Weight Models 00:31:39 Attribution Problem: Who Gets Paid for AI Training? 00:35:52 Content Provenance & Authentication (C2PA) 00:48:03 AI Security: Finding Exploits & Vulnerabilities 00:54:53 Educational Applications: LLMs as Learning Partner 00:58:29 Reliance on LLMs and Cognitive Abilities 01:03:57 Content Providers’ Fear of LLM Training

EigenCloud

62,099 görüntüleme • 1 yıl önce