Загрузка видео...
Не удалось загрузить видео
First results are in. Llama 4 Maverick 17B active / 400B total is blazing fast with MLX on an M3 Ultra. Here is the 4-bit model generating 1100 tokens at 50 tok/sec:
149,855 просмотров • 1 год назад •via X (Twitter)
Комментарии: 11

Awni Hannun1 год назад
PR here:

Rainmaker2 лет назад
Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

Sharan Narang1 год назад
That’s fast @awnihannun !

Zack Angelo1 год назад
what throughput are you seeing during the prefill phase?

moskstraumen1 год назад
Thanks! Would it be possible to test the 8-bit quant?

Daniel Byalsky 🇺🇦🎗️1 год назад
Whoa, think my M1 Max could pull off 5-10t/s (if it even loads into memory, lol)?

Tom Jeans1 год назад
you’re getting 50 tps on a single M3 Ultra?! 🤯

Paul Marin1 год назад
I am very curious about q6 or q8 performance on 512gb as well as q4 with very long context. If it’s good, i will probably buy one.

SuperBadGPT1 год назад
Is this M3Ultra-512GB ?

Awni Hannun1 год назад
Yes. It should fit with 256GB as well.

ROBERT D3REZZ1 год назад
Looks amazing 👏
