Video yükleniyor...
Video Yüklenemedi
First results are in. Llama 4 Maverick 17B active / 400B total is blazing fast with MLX on an M3 Ultra. Here is the 4-bit model generating 1100 tokens at 50 tok/sec:
149,855 görüntüleme • 1 yıl önce •via X (Twitter)
11 Yorum

Awni Hannun1 yıl önce
PR here:

Rainmaker2 yıl önce
Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

Sharan Narang1 yıl önce
That’s fast @awnihannun !

Zack Angelo1 yıl önce
what throughput are you seeing during the prefill phase?

moskstraumen1 yıl önce
Thanks! Would it be possible to test the 8-bit quant?

Daniel Byalsky 🇺🇦🎗️1 yıl önce
Whoa, think my M1 Max could pull off 5-10t/s (if it even loads into memory, lol)?

Tom Jeans1 yıl önce
you’re getting 50 tps on a single M3 Ultra?! 🤯

Paul Marin1 yıl önce
I am very curious about q6 or q8 performance on 512gb as well as q4 with very long context. If it’s good, i will probably buy one.

SuperBadGPT1 yıl önce
Is this M3Ultra-512GB ?

Awni Hannun1 yıl önce
Yes. It should fit with 256GB as well.

ROBERT D3REZZ1 yıl önce
Looks amazing 👏
