Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Two days ago, Deepseek surprised everyone with an "undefined-behavior" PTX optimization speeding up particular ML workloads on a Hopper NVIDIA GPU Kernel. Let's reverse engineer the hack, implement it ourselves, and benchmark the speedup on an H100.

228,591 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

LaurieWired profil fotoğrafı
LaurieWired1 yıl önce

Full Video:

LaurieWired profil fotoğrafı
LaurieWired1 yıl önce

My test code:

NetMind.AI profil fotoğrafı
NetMind.AI2 yıl önce

Get access to a wide range of GPUs like H100, A100, 4090, 3090 and save over 90% at NetMind Power. Rent Now!

numanumabruh profil fotoğrafı
numanumabruh1 yıl önce

You'd never know she's 6'5"

Jason Ho profil fotoğrafı
Jason Ho1 yıl önce

laurie supremacy

Bob (Moderna #7) Kerns profil fotoğrafı
Bob (Moderna #7) Kerns1 yıl önce

Until recently, I'd only seen your tweets; the first video I encountered was the 2025 prediction ones. Assumptions violated: higher voice, younger. Always good to have one's assumptions flagged, but especially the age. I was struck by the maturity of your analysis!

Dave 🚀 profil fotoğrafı
Dave 🚀1 yıl önce

LFG!

KnowledgeisMostValuable profil fotoğrafı
KnowledgeisMostValuable1 yıl önce

I'd fight off a bear for you

nisten - e/acc profil fotoğrafı
nisten - e/acc1 yıl önce

lfg

🥀shiVam🥀 profil fotoğrafı
🥀shiVam🥀1 yıl önce

wait, did you film this at Google HQ? (must appreciate the audio recording and editing)

Calcs profil fotoğrafı
Calcs1 yıl önce

Fantastic video, more please, lol 😂

Benzer Videolar

Dylan Patel on the importance of memory and storage Two key quotes: "An $NVDA GPU is faster than an $AMD GPU in most cases, but because AMD GPUs have more memory, they can outperform Nvidia in certain workloads." “It is a difficult, multivariable problem. Generally, you need the best GPU, such as a GB300, but you also need the best storage solutions. I will not spoil who comes out on top, but storage solutions matter a lot, memory solutions matter a lot, and frontend networking also matters significantly" Full Quote: “We have over $80 million of compute: GPUs from $NVDA and $AMD, TPUs from Google, and Trainium from Amazon. We constantly run this benchmark using the newest inference engines, drivers, PyTorch versions, and other software. It runs every day through automated CI across the latest Chinese models from GLM, Zhipu, Moonshot, Kimi, Alibaba, and others. Initially, when we were benchmarking the differences between these chips, inference engines, and parallelism schemes, we used fixed context lengths. But with Agent X, we have now analyzed more than $5 million worth of Claude Code traces. This is real production traffic that users have donated to us, combined with internally generated data, so we now understand what an actual agent workload looks like. When we implement those workloads and run the benchmarks, it turns out that the chip you are using is very important, but how you handle memory offload can be even more important. An Nvidia GPU is faster than an AMD GPU in most cases, but because AMD GPUs have more memory, they can outperform Nvidia in certain workloads. Similarly, you can use a less powerful GPU with a much better storage solution and outperform the best GPU when it lacks those solutions. Simply buying the newest GPU does not necessarily give you the best inference economics. You need to layer in other innovations, including storage and memory.” Interviewer: “Who is the top player on your chart? Can you tell us?” Dylan Patel: “It is a difficult, multivariable problem. Generally, you need the best GPU, such as a GB300, but you also need the best storage solutions. I will not spoil who comes out on top, but storage solutions matter a lot, memory solutions matter a lot, and frontend networking also matters significantly.”

Daniel Romero

38,220 görüntüleme • 1 ay önce