正在加载视频...
视频加载失败
FlexAttention is a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code Boyuan Feng & Avik Chaudhuri show how many existing attention variants can be implemented via FlexAttention & that we achieve competitive performance compared to handwritten kernels.... show more
17,929 次观看 • 1 年前 •via X (Twitter)
3 条评论

Lakshya1 年前
@Boyuan_Feng @__avik How does it compare to MagiAttention

Jilong | We provide AI marketer - 24/7 marketing1 年前
@Boyuan_Feng @__avik FlexAttention sounds promising! Compilers beat brute force coding.

Vinayak N Baddi1 年前
@Boyuan_Feng @__avik Possible to share the slides @__avik Thanks
