Загрузка видео...
Не удалось загрузить видео
Part 1: CUDA Graphs - Intro & Basics - how normal CPU ←→ GPU execution works - what a kernel launch is and where CPU launch overhead comes from - why launching many small kernels individually can become a bottleneck - what CUDA Graphs are and how they reduce... show more
29,135 просмотров • 2 месяцев назад •via X (Twitter)
Комментарии: 13

So if a beginner like me can follow u from day 1 Any prerequisites for this?

i would say you would still need some inference knowledge or context of if u get what I mean because I'm doing topics that I haven't deeply done or missing in upcoming days, so following my exact path as beginner won't work but if u happen to learn about cuda graphs or torch.compile, later ptx etc, you can easily watch and follow these topics

my goat

Bro one shottted ? Sick man keep it coming

nah lol, ~60 tries for 2 clips - its 2 seperate clips that i merged

Start a youtube

Please help me with DMing you.

sure

still can’t

dmed you

Great breakdown. Where this really bites is LLM inference decoding: you launch a flurry of tiny kernels per token and CPU launch overhead dominates at small batch sizes. Capturing the decode step as a graph and replaying it is one of the biggest easy wins for tokens per second.

🗿🗿

second follow
