Loading video...
Video Failed to Load
First results are in. Llama 4 Maverick 17B active / 400B total is blazing fast with MLX on an M3 Ultra. Here is the 4-bit model generating 1100 tokens at 50 tok/sec:
149,855 views • 1 year ago •via X (Twitter)
11 Comments

Awni Hannun1 year ago
PR here:

Rainmaker2 years ago
Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

Sharan Narang1 year ago
That’s fast @awnihannun !

Zack Angelo1 year ago
what throughput are you seeing during the prefill phase?

moskstraumen1 year ago
Thanks! Would it be possible to test the 8-bit quant?

Daniel Byalsky 🇺🇦🎗️1 year ago
Whoa, think my M1 Max could pull off 5-10t/s (if it even loads into memory, lol)?

Tom Jeans1 year ago
you’re getting 50 tps on a single M3 Ultra?! 🤯

Paul Marin1 year ago
I am very curious about q6 or q8 performance on 512gb as well as q4 with very long context. If it’s good, i will probably buy one.

SuperBadGPT1 year ago
Is this M3Ultra-512GB ?

Awni Hannun1 year ago
Yes. It should fit with 256GB as well.

ROBERT D3REZZ1 year ago
Looks amazing 👏
