Loading video...

Video Failed to Load

Go Home

Apple’s “LLM in a Flash” is definitely worth checking out. Going to 2-bit for the shared-expert MLP means disk I/O is no longer dominant. 14–15 tok/s from SSD is still wild for a ~400B MoE model streamed from storage. Qwen3.5-397B-A17B Credit: Dan Woods

35,664 views • 3 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos