正在加载视频...
视频加载失败
First results are in. Llama 4 Maverick 17B active / 400B total is blazing fast with MLX on an M3 Ultra. Here is the 4-bit model generating 1100 tokens at 50 tok/sec:
11 条评论

Awni Hannun1 年前
PR here:

Rainmaker2 年前
Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

Sharan Narang1 年前
That’s fast @awnihannun !

Zack Angelo1 年前
what throughput are you seeing during the prefill phase?

moskstraumen1 年前
Thanks! Would it be possible to test the 8-bit quant?

Daniel Byalsky 🇺🇦🎗️1 年前
Whoa, think my M1 Max could pull off 5-10t/s (if it even loads into memory, lol)?

Tom Jeans1 年前
you’re getting 50 tps on a single M3 Ultra?! 🤯

Paul Marin1 年前
I am very curious about q6 or q8 performance on 512gb as well as q4 with very long context. If it’s good, i will probably buy one.

SuperBadGPT1 年前
Is this M3Ultra-512GB ?

Awni Hannun1 年前
Yes. It should fit with 256GB as well.

ROBERT D3REZZ1 年前
Looks amazing 👏
