Loading video...
Video Failed to Load
DeepSeek R1 (the full 680B model) runs nicely in higher quality 4-bit on 3 M2 Ultras with MLX. Asked it a coding question and it thought for ~2k tokens and generated 3500 tokens overall:
997,284 views • 1 year ago •via X (Twitter)
11 Comments

Prompt (h/t @ivanfioravanti): "write a python script for a bouncing yellow ball within a square, make sure to handle collision detection properly. make the square slowly rotate. implement it in python. make sure ball stays within the square" Code ran no edits and the result is actually pretty decent:

🚀 Don't gamble with your portfolio! Use our advanced hybrid quant risk tool using on/off-chain data daily and make informed decisions. 📈 Acess to 1000+ charts for your crypto journey. 📚Receive free weekly quant analysis. 📊+21 projects supported. 🏗️ Beginners and experts.

And drawing only 60watts is insane. If we assume scaling of the last 20 years, the amount of AI resources available to us in our pockets or watches is gonna be insane

WOW! 4bit and the Big 680B model! AMAZING!

Soon in 8-bit on only 2 M4 Ultra 👀

This is so cool 🥹 Still got my fingers crossed a new M4 Ultra comes out at some point so I can hopefully buy some second hand M2 Ultras for cheap 😂 Can't wait to one day run these mega models at home!

Have you noticed any trade offs between high parameter but heavily/quantised vs lower parameter but full precision or even half?

How is it possible that 4-bit models are useful and reasonably accurate now? Just a few months ago 4-bit models were close to useless

Don’t look Jensen Huang!!!

apple is toast anyway, making money or having a moat on inference will not be a thing

Do you do some test with concurrent prompt to see how many prompt can run in the same time for acceptable speed ?
