Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

First results are in. Llama 4 Maverick 17B active / 400B total is blazing fast with MLX on an M3 Ultra. Here is the 4-bit model generating 1100 tokens at 50 tok/sec:

149,855 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

Awni Hannun profil fotoğrafı
Awni Hannun1 yıl önce

PR here:

Rainmaker profil fotoğrafı
Rainmaker2 yıl önce

Can Machine Learning beat the market? Check out this post on my free Substack where I share code and commentary for an XGBoost model and a Random Forest model that both deliver powerful performances.

Sharan Narang profil fotoğrafı
Sharan Narang1 yıl önce

That’s fast @awnihannun !

Zack Angelo profil fotoğrafı
Zack Angelo1 yıl önce

what throughput are you seeing during the prefill phase?

moskstraumen profil fotoğrafı
moskstraumen1 yıl önce

Thanks! Would it be possible to test the 8-bit quant?

Daniel Byalsky 🇺🇦🎗️ profil fotoğrafı
Daniel Byalsky 🇺🇦🎗️1 yıl önce

Whoa, think my M1 Max could pull off 5-10t/s (if it even loads into memory, lol)?

Tom Jeans profil fotoğrafı
Tom Jeans1 yıl önce

you’re getting 50 tps on a single M3 Ultra?! 🤯

Paul Marin profil fotoğrafı
Paul Marin1 yıl önce

I am very curious about q6 or q8 performance on 512gb as well as q4 with very long context. If it’s good, i will probably buy one.

SuperBadGPT profil fotoğrafı
SuperBadGPT1 yıl önce

Is this M3Ultra-512GB ?

Awni Hannun profil fotoğrafı
Awni Hannun1 yıl önce

Yes. It should fit with 256GB as well.

ROBERT D3REZZ profil fotoğrafı
ROBERT D3REZZ1 yıl önce

Looks amazing 👏

Benzer Videolar

Introducing "Building with Llama 4." This short course is created with Meta AI at Meta, and taught by Amit Sangani, Director of Partner Engineering for Meta’s AI team. Meta’s new Llama 4 has added three new models and introduced the Mixture-of-Experts (MoE) architecture to its family of open-weight models, making them more efficient to serve. In this course, you’ll work with two of the three new models introduced in Llama 4. First is Maverick, a 400B parameter model, with 128 experts and 17B active parameters. Second is Scout, a 109B parameter model with 16 experts and 17B active parameters. Maverick and Scout support long context windows of up to a million tokens and 10M tokens, respectively. The latter is enough to support directly inputting even fairly large GitHub repos for analysis! In hands-on lessons, you’ll build apps using Llama 4’s new multimodal capabilities including reasoning across multiple images and image grounding, in which you can identify elements in images. You’ll also use the official Llama API, work with Llama 4’s long-context abilities, and learn about Llama’s newest open-source tools: its prompt optimization tool that automatically improves system prompts and synthetic data kit that generates high-quality datasets for fine-tuning. If you need an open model, Llama is a great option, and the Llama 4 family is an important part of any GenAI developer's toolkit. Through this course, you’ll learn to call Llama 4 via API, use its optimization tools, and build features that span text, images, and large context. Please sign up here:

Andrew Ng

67,846 görüntüleme • 1 yıl önce