Загрузка видео...
Не удалось загрузить видео
Two days ago, Deepseek surprised everyone with an "undefined-behavior" PTX optimization speeding up particular ML workloads on a Hopper NVIDIA GPU Kernel. Let's reverse engineer the hack, implement it ourselves, and benchmark the speedup on an H100.
228,591 просмотров • 1 год назад •via X (Twitter)
Комментарии: 11

LaurieWired1 год назад
Full Video:

LaurieWired1 год назад
My test code:

NetMind.AI2 лет назад
Get access to a wide range of GPUs like H100, A100, 4090, 3090 and save over 90% at NetMind Power. Rent Now!

numanumabruh1 год назад
You'd never know she's 6'5"

Jason Ho1 год назад
laurie supremacy

Bob (Moderna #7) Kerns1 год назад
Until recently, I'd only seen your tweets; the first video I encountered was the 2025 prediction ones. Assumptions violated: higher voice, younger. Always good to have one's assumptions flagged, but especially the age. I was struck by the maturity of your analysis!

Dave 🚀1 год назад
LFG!

KnowledgeisMostValuable1 год назад
I'd fight off a bear for you

nisten - e/acc1 год назад
lfg

🥀shiVam🥀1 год назад
wait, did you film this at Google HQ? (must appreciate the audio recording and editing)

Calcs1 год назад
Fantastic video, more please, lol 😂

