Loading video...

Video Failed to Load

Go Home

DeepSeek R1 (the full 680B model) runs nicely in higher quality 4-bit on 3 M2 Ultras with MLX. Asked it a coding question and it thought for ~2k tokens and generated 3500 tokens overall:

997,284 views • 1 year ago •via X (Twitter)

11 Comments

Awni Hannun's profile picture
Awni Hannun1 year ago

Prompt (h/t @ivanfioravanti): "write a python script for a bouncing yellow ball within a square, make sure to handle collision detection properly. make the square slowly rotate. implement it in python. make sure ball stays within the square" Code ran no edits and the result is actually pretty decent:

Lab4crypto's profile picture
Lab4crypto1 year ago

🚀 Don't gamble with your portfolio! Use our advanced hybrid quant risk tool using on/off-chain data daily and make informed decisions. 📈 Acess to 1000+ charts for your crypto journey. 📚Receive free weekly quant analysis. 📊+21 projects supported. 🏗️ Beginners and experts.

Clem's profile picture
Clem1 year ago

And drawing only 60watts is insane. If we assume scaling of the last 20 years, the amount of AI resources available to us in our pockets or watches is gonna be insane

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ1 year ago

WOW! 4bit and the Big 680B model! AMAZING!

Kim Noël ⚡ 📖's profile picture
Kim Noël ⚡ 📖1 year ago

Soon in 8-bit on only 2 M4 Ultra 👀

Mark Lord's profile picture
Mark Lord1 year ago

This is so cool 🥹 Still got my fingers crossed a new M4 Ultra comes out at some point so I can hopefully buy some second hand M2 Ultras for cheap 😂 Can't wait to one day run these mega models at home!

Lewis N Watson's profile picture
Lewis N Watson1 year ago

Have you noticed any trade offs between high parameter but heavily/quantised vs lower parameter but full precision or even half?

Ronald Mannak's profile picture
Ronald Mannak1 year ago

How is it possible that 4-bit models are useful and reasonably accurate now? Just a few months ago 4-bit models were close to useless

sour coach sauers's profile picture
sour coach sauers1 year ago

Don’t look Jensen Huang!!!

allinasecond's profile picture
allinasecond1 year ago

apple is toast anyway, making money or having a moat on inference will not be a thing

X Æ A-12's profile picture
X Æ A-121 year ago

Do you do some test with concurrent prompt to see how many prompt can run in the same time for acceptable speed ?

Related Videos