memer of technical staff at @modal. he/him.
ex @full_stack_dl, @weights_biases (acq. @CoreWeave), phd Berkeley @Redwood_Neuro.
try https://t.co/SYWVMCb7OB
Shorts
Low-precision floats are weird. I have been building up my intuition by playing with them outside of inference/training. Adam Azzam and I cooked up this visualizer for micro-scaling/block quant formats like NVFP4, MXFP4, and friends. Try it:
13,028 次观看
Added a fun lil widget to the LLM Engineer's Almanac -- a "Token Timing Simulator" so you can get a visceral feel for what a benchmark perf number means. Here's David Wang's latest work with Zhijian Liu's DFlash technique in SGLang -- ~1k TPS!
18,764 次观看
Step 4 to achieve truly serverless GPUs for AI inference: skip over unserializable inference engine setup steps like CUDA graph capture and Torch compilation by stacking GPU snapshots and CPU snapshots.