
Gaurav Sen
@gkcs_ • 73,634 subscribers
CEO of AIEngg
Videos

How do we cache prompts with LLMs? Interview with Tanishq Singh. This is a real interview question from a big tech company, asked to a candidate in their technical interview round. The video explains the answer in roughly 6 minutes. 00:00 Question - Prompt Caching with LLMs 00:39 Exact Query Caching 02:04 Semantic Caching 04:14 Tradeoffs 05:52 Blooper :p AI engineering program for software engineers: #LLMs #AI #InterviewReady
Gaurav Sen94,211 views • 29 days ago

RAG vs. Agents Interview with Tanishq Singh. This is a real interview question from a big tech company, asked to a candidate in their technical interview round. The video explains the answer in roughly 4 minutes. 00:00 Question - RAG vs Agents 00:35 Difference between RAG and Agents 01:45 What can RAG make tool calls 02:54 Agentic RAG 03:47 Bloopers :p AI engineering program for software engineers:
Gaurav Sen28,949 views • 11 days ago

How do we run evals with LLMs? Interview with Tanishq Singh. This is a real interview question from a big tech company, asked to a candidate in their technical interview round. The video explains the answer in roughly 11 minutes. 00:00 Question - Reliability with LLMs 01:00 Observability & Evals 02:50 What are evals? 06:41 Who does evals? 08:11 Answer relevancy, groundedness, tool sequence 10:39 Bloopers :p AI engineering program for software engineers:
Gaurav Sen35,079 views • 20 days ago

What's better? LLMs or Agents? This is a real interview question from a big tech company. For more questions, subscribe to the channel! AI engineering program for software engineers: 00:00 Question - LLMs vs. Agents 01:02 Basic Answer 02:16 Follow-up question 03:47 Context Management 05:08 Bloopers #LLMs #AI #InterviewReady
Gaurav Sen18,889 views • 1 month ago

Engineers need to communicate effectively when building AI Systems. These terms will help you use a shared vocabulary. This is useful when discussing concepts, reading papers, or collaborating with teammates. Listed in order. 00:00 Agenda 00:28 1. Large Language Model 01:28 2. Tokenization 02:53 3. Vectorization 04:15 4. Attention 07:22 5. Self-Supervised Learning 12:07 6. Transformer 14:32 7. Fine-tuning 17:05 8. Few-shot Prompting 18:11 9. Retrieval Augmented Generation 20:33 10. Vector Database 23:03 11. Model Context Protocol 25:43 12. Context Engineering 28:17 13. Agents 29:19 14. Reinforcement Learning 34:42 15. Chain of Thought 35:55 16. Reasoning Models 36:36 17. Multi-modal Models 38:21 18. Small Language Models 40:24 19. Distillation 41:47 20. Quantization If you are a software engineer looking to transition to AI, click the link below. AI Engineering Course: #AI #SoftwareEngineering #Agents
Gaurav Sen80,282 views • 11 months ago

DeepSeek's R1 Model has shocked the World. Here is how it works. #DeepSeek #R1 #AI
Gaurav Sen87,427 views • 1 year ago

WhatsApp serves 100 Million international calls everyday. This is how.
Gaurav Sen74,008 views • 1 year ago

This video is a response to the recent IT rules. These rules have been suggested by the Ministry of Electronics and Information Technology. They raise serious concerns on free speech in India. Thanks to Internet Freedom Foundation (IFF) for flagging this. I have shared a template that we can use to connect with Meity regarding this. Link: Please raise your voice, while you still can. Cheers :)
Gaurav Sen20,947 views • 4 months ago

This video explains how diffusion models are overtaking Large Language Models for generation tasks like: 1. Code Generation 2. Image Generation 3. Video Generation 00:00 Agenda 00:20 How are they different from LLMs? 05:09 Internal Mechanism 10:09 How are vectors generated? 12:08 Conclusion 13:02 Opinion Piece AI Engineering Course: #Diffusion #AI #LLMs
Gaurav Sen28,027 views • 10 months ago

In this video, we break down the Model Context Protocol — a massively underrated concept in AI that's quietly redefining what large language models can do. While most people are distracted by the hype, MCP is solving a core problem: LLMs that can perform actions. We dive into: 1. The limitations of current AI systems 2. What makes MCP different 3. How MCP can reduce engineering overhead 4. Real-world scenarios where MCP shines 00:00 What is MCP? 01:48 Usecase - SEO 03:50 Usecase - RAG 04:50 Usecase - Apps 06:30 Conclusion 07:00 The future? #MCP #AI #LLM
Gaurav Sen41,154 views • 1 year ago

Amazon S3 stores exabytes of data across millions of disks. This video explains it's internal architecture. In this video, we break down: 1. How S3 uploads files using parallel multipart uploads 2. How checksums and monitoring prevent data corruption 3. How S3 survives disk, shard, and even region failures 4. How shuffle sharding and request cancellation keep latency low 5. How auto-scaling avoids hot prefixes 6. How error-correcting codes deliver 11 nines of durability with only ~1.8× storage overhead If you want to understand real-world distributed storage design, this is how S3 does it at scale. Links: #AWS #S3 #SystemDesign
Gaurav Sen22,111 views • 8 months ago