Загрузка видео...
Не удалось загрузить видео
DFlash speculative decoding on Apple Silicon Qwen3.5-9B bf16 · M5 Max · greedy exact match ▸ 85 tok/s, 3.3× at 1024 tokens (runtime) ▸ ~70 tok/s, 2.6× in the video (terminal I/O overhead) ▸ 80 tok/s, 3.1× at 2048 tokens (runtime) Currently working on: → Long context (speedup degrades... show more
37,049 просмотров • 5 месяцев назад •via X (Twitter)
Комментарии: 18

@alexocheema @ivanfioravanti

@awnihannun This rules, please continue, and let us know if you need any support!

@awnihannun Thanks, will reach out if needed!

nice numbers but remember that 4k context is where the kv cache actually starts eating your ram on m-series chips, so keep an eye on those limits before you ship it

dogggggg you beat me to it lol, have been up for two days working on this

check this @Prince_Canuma bstn is implementing DFlash speculative decoding on Apple Silicon in Qwen3.5-9B 85 tok/s, 3.3× at 1024 tokens

@danveloper Looking fwd

@awnihannun How much ram did the Mac have?

@awnihannun M5 Max, 64GB unified

@awnihannun That’s awesome, I can’t wait till they will sell bigger ram models

Great to see replication of the speed increase!

holy shit, goat

Using a diffusion model for the draft phase is a ridiculously clean way to dodge the VRAM tax of normal speculative decoding. Does the acceptance rate tank past 2k tokens?

Lossless?

*drools* 🤤

ok wow but what do you want to achieve with potato-class llm?

Running a qwen 9b on m5 is like taking a bus when you have a chauffer

what test is this?
