Loading video...
Video Failed to Load
Two days ago, Deepseek surprised everyone with an "undefined-behavior" PTX optimization speeding up particular ML workloads on a Hopper NVIDIA GPU Kernel. Let's reverse engineer the hack, implement it ourselves, and benchmark the speedup on an H100.
228,591 views • 1 year ago •via X (Twitter)
11 Comments

LaurieWired1 year ago
Full Video:

LaurieWired1 year ago
My test code:

NetMind.AI2 years ago
Get access to a wide range of GPUs like H100, A100, 4090, 3090 and save over 90% at NetMind Power. Rent Now!

numanumabruh1 year ago
You'd never know she's 6'5"

Jason Ho1 year ago
laurie supremacy

Bob (Moderna #7) Kerns1 year ago
Until recently, I'd only seen your tweets; the first video I encountered was the 2025 prediction ones. Assumptions violated: higher voice, younger. Always good to have one's assumptions flagged, but especially the age. I was struck by the maturity of your analysis!

Dave 🚀1 year ago
LFG!

KnowledgeisMostValuable1 year ago
I'd fight off a bear for you

nisten - e/acc1 year ago
lfg

🥀shiVam🥀1 year ago
wait, did you film this at Google HQ? (must appreciate the audio recording and editing)

Calcs1 year ago
Fantastic video, more please, lol 😂

