Video wird geladen...
Video konnte nicht geladen werden
First results are in. Llama 4 Maverick 17B active / 400B total is blazing fast with MLX on an M3 Ultra. Here is the 4-bit model generating 1100 tokens at 50 tok/sec:
149,855 Aufrufe • vor 1 Jahr •via X (Twitter)
11 Kommentare

Awni Hannunvor 1 Jahr
PR here:

Rainmakervor 2 Jahren
Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

Sharan Narangvor 1 Jahr
That’s fast @awnihannun !

Zack Angelovor 1 Jahr
what throughput are you seeing during the prefill phase?

moskstraumenvor 1 Jahr
Thanks! Would it be possible to test the 8-bit quant?

Daniel Byalsky 🇺🇦🎗️vor 1 Jahr
Whoa, think my M1 Max could pull off 5-10t/s (if it even loads into memory, lol)?

Tom Jeansvor 1 Jahr
you’re getting 50 tps on a single M3 Ultra?! 🤯

Paul Marinvor 1 Jahr
I am very curious about q6 or q8 performance on 512gb as well as q4 with very long context. If it’s good, i will probably buy one.

SuperBadGPTvor 1 Jahr
Is this M3Ultra-512GB ?

Awni Hannunvor 1 Jahr
Yes. It should fit with 256GB as well.

ROBERT D3REZZvor 1 Jahr
Looks amazing 👏
