Charles 🎉 Frye's banner
Charles 🎉 Frye's profile picture

Charles 🎉 Frye

@charles_irl19,273 subscribers

memer of technical staff at @modal. he/him. ex @full_stack_dl, @weights_biases (acq. @CoreWeave), phd Berkeley @Redwood_Neuro.

Shorts

Added a fun lil widget to the LLM Engineer's Almanac -- a "Token Timing Simulator" so you can get a visceral feel for what a benchmark perf number means. Here's David Wang's latest work with Zhijian Liu's DFlash technique in SGLang -- ~1k TPS!

Added a fun lil widget to the LLM Engineer's Almanac -- a "Token Timing Simulator" so you can get a visceral feel for what a benchmark perf number means. Here's David Wang's latest work with Zhijian Liu's DFlash technique in SGLang -- ~1k TPS!

18,764 Aufrufe

Step 4 to achieve truly serverless GPUs for AI inference: skip over unserializable inference engine setup steps like CUDA graph capture and Torch compilation by stacking GPU snapshots and CPU snapshots.

Step 4 to achieve truly serverless GPUs for AI inference: skip over unserializable inference engine setup steps like CUDA graph capture and Torch compilation by stacking GPU snapshots and CPU snapshots.

17,452 Aufrufe

Low-precision floats are weird. I have been building up my intuition by playing with them outside of inference/training. Adam Azzam and I cooked up this visualizer for micro-scaling/block quant formats like NVFP4, MXFP4, and friends. Try it:

Low-precision floats are weird. I have been building up my intuition by playing with them outside of inference/training. Adam Azzam and I cooked up this visualizer for micro-scaling/block quant formats like NVFP4, MXFP4, and friends. Try it:

13,136 Aufrufe

nother banger in the pipeline btw

nother banger in the pipeline btw

12,600 Aufrufe

Videos

Keine weiteren Inhalte verfügbar