Loading video...
Video Failed to Load
I trained a 12M parameter LLM on my own ML framework using a Rust backend and CUDA kernels for flash attention, AdamW, and more. Wrote the full transformer architecture, and BPE tokenizer from scratch. The framework features: - Custom CUDA kernels (Flash Attention, fused LayerNorm, fused GELU) for 3x... show more
820,861 views • 4 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

