
Youssof Al Toukhi
@Youssofal_ • 10,854 subscribers
I’m a 21 y/o hyperfixated dabbler and Forbes featured viral meme building the next biggest mobile game based on memes.
Videos

20 hours later and I’m done! Qwen 3.8 27B @ 73 TPS (peak) on a Macbook pro M5 max. MTPLX V2.7 out now with 3 new models. - Bare Speed: short burst decode, good for chat. - Optimized Speed: higher quality and faster on long coding. - Optimized Quality: 30% slower but super!
Youssof Al Toukhi96,285 Aufrufe • vor 29 Tagen

MTPLX V2.10 is out! I put a lot of work into this one. Highlights: - Qwen 3.8 Next Support Added: Bare Speed (4 bit) and Optimised Speed (4 bit/8 bit) models added running at 50 - 80 TPS sustained at high contexts. (expect speed improvements in the coming days) - Qwen 3.8 Next mmap: The 32 GB n-gram table streams from SSD via mmap instead of RAM. Both models fit a 96 GB Mac with 0 speed impact. - Qwen 3.8 27B Decode & Prefill boost EVERYWHERE: +15% TPS at 3k context, +29% at 88k, +54% at 147k. Prefill +41%. - 48 GB Macs went from 3 to 4 TPS in swap @ 30k context to 33 TPS. - KV Cache has been solved!: 8 bit and 4 bit cache now barely impacts speed instead of a 50% decline. - Coding agents wall time optimisations: file task dropped from 150s to 44s, and mid-session first token went from ~2s to 0.11s. - CLI UX improvements + 34 bug fixes. My thoughts on Qwen Next? Incredible model. Enormously better at design, and creative tasks with scary good vision. It created what I think is the best flappy bird. maybe even better than the Fable version. It had various difficulties and skins, moving pillars, and upgrades. It also had the most memeified look. Incredibly impressed. However, it is slightly worse than 27B at having code just work straight away. But it's personality is also far better than 27B. Here is the Flappy bird it made below! Made with Qwen Next Optimized Speed one shot@ 57 TPS:
Youssof Al Toukhi20,975 Aufrufe • vor 15 Tagen

I put Qwen 3.6 27B on a B200 for the fun of it. I told Fable to apply my past optimisation principles. Result? 1000 TPS Qwen 3.6 27B NVFP4 (550 TPS average long decode) I used the same principles that got me 250 TPS on 2x 3090s. Now I need to buy a DGX Station 🤪.
Youssof Al Toukhi46,035 Aufrufe • vor 1 Monat

MTPLX V2.9 is OUT! - 30% faster TPS in CLI & 60% in app. - All MTPLX models retuned to be smaller AND faster at the same quality (re-download!) - 80% lower CPU utilisation in app. - 95% less streaming freezes. Here is our bare speed model using session bank to reach 175 TPS:
Youssof Al Toukhi24,745 Aufrufe • vor 24 Tagen

Introducing MTPLX V1: The fastest and simplest way to run MTP compatible models on your Mac. - New Swift based app. 2x speed increase without bloat - Easy OpenCode, Hermes & Pi integration - Convert your own MTP models with forge And more try now at:
Youssof Al Toukhi31,741 Aufrufe • vor 3 Monaten

Managed to get QWEN 3.6 27B running in Cursor with localhost (no Ngrok) at 120 - 140 TPS on 2x 3090s.
Youssof Al Toukhi21,008 Aufrufe • vor 3 Monaten
Keine weiteren Inhalte verfügbar