Loading video...
Video Failed to Load
Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.📜 Apache 2.0. 🤖 ⚡ Reaches 23.9–25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark. 🏆 Retains most of the FP16 base model’s quality, with an average gap of... show more
16,923 views • 14 days ago •via X (Twitter)
6 Comments

Very cool! Decode speed under memory pressure would help show how much performance depends on the OS keeping streamed experts cached.

Sparse activation and low memory could make local MoE inference surprisingly practical.

interesting to see edge0-8b-a1b-preview's performance on short-context benchmark. do you have any insights on the memory efficiency of the moe mechanism in this setup, considering the 1.0 gib peak active memory?

1.2B active params is a smart move. Balancing scale and memory efficiency is key for real...

Is vision or tool supported?

Worth adding the HF receipts: sibling Edge0-35B-A3B-preview hit 17.8K dl / 2.4K likes in its first week, this 8B took 3.2K dl by day 4 — Apache-2.0 weights are getting real adoption, not just benchmark praise. (HF API, checked today)
