Video yükleniyor...
Video Yüklenemedi
Two days ago, Deepseek surprised everyone with an "undefined-behavior" PTX optimization speeding up particular ML workloads on a Hopper NVIDIA GPU Kernel. Let's reverse engineer the hack, implement it ourselves, and benchmark the speedup on an H100.
228,591 görüntüleme • 1 yıl önce •via X (Twitter)
11 Yorum

LaurieWired1 yıl önce
Full Video:

LaurieWired1 yıl önce
My test code:

NetMind.AI2 yıl önce
Get access to a wide range of GPUs like H100, A100, 4090, 3090 and save over 90% at NetMind Power. Rent Now!

numanumabruh1 yıl önce
You'd never know she's 6'5"

Jason Ho1 yıl önce
laurie supremacy

Bob (Moderna #7) Kerns1 yıl önce
Until recently, I'd only seen your tweets; the first video I encountered was the 2025 prediction ones. Assumptions violated: higher voice, younger. Always good to have one's assumptions flagged, but especially the age. I was struck by the maturity of your analysis!

Dave 🚀1 yıl önce
LFG!

KnowledgeisMostValuable1 yıl önce
I'd fight off a bear for you

nisten - e/acc1 yıl önce
lfg

🥀shiVam🥀1 yıl önce
wait, did you film this at Google HQ? (must appreciate the audio recording and editing)

Calcs1 yıl önce
Fantastic video, more please, lol 😂

