Loading video...

Video Failed to Load

Go Home

Interested in optimizing Mixture of Experts LLM inference on ARM CPUs? In his talk at PyTorch Conference North America, Maajid Khan from Fujitsu Research India will explore efficient MoE LLM inference using vLLM and OpenVINO, sharing practical strategies for running these models effectively on ARM architectures. Register and join...

16,448 views • 6 days ago •via X (Twitter)

2 Comments

Fajar M Reza's profile picture
Fajar M Reza6 days ago

ARM MoE benefits from sparse activation, but memory traffic still shapes latency.

Sutton's profile picture
Sutton6 days ago

MoE is honestly the best case for CPU inference. you need a lot of RAM for the total params, which CPUs have, but each token only touches the active experts, so memory bandwidth stops being such a wall. curious how much of the speedup comes from the kernels vs just smarter expert placement

Related Videos