Loading video...
Video Failed to Load
Interested in optimizing Mixture of Experts LLM inference on ARM CPUs? In his talk at PyTorch Conference North America, Maajid Khan from Fujitsu Research India will explore efficient MoE LLM inference using vLLM and OpenVINO, sharing practical strategies for running these models effectively on ARM architectures. Register and join... show more
16,448 views • 6 days ago •via X (Twitter)
2 Comments

Fajar M Reza6 days ago
ARM MoE benefits from sparse activation, but memory traffic still shapes latency.

Sutton6 days ago
MoE is honestly the best case for CPU inference. you need a lot of RAM for the total params, which CPUs have, but each token only touches the active experts, so memory bandwidth stops being such a wall. curious how much of the speedup comes from the kernels vs just smarter expert placement





