正在加载视频...

视频加载失败

DeepSeek R1 (the full 680B model) runs nicely in higher quality 4-bit on 3 M2 Ultras with MLX. Asked it a coding question and it thought for ~2k tokens and generated 3500 tokens overall:

997,284 次观看 • 1 年前 •via X (Twitter)

11 条评论

Awni Hannun 的头像
Awni Hannun1 年前

Prompt (h/t @ivanfioravanti): "write a python script for a bouncing yellow ball within a square, make sure to handle collision detection properly. make the square slowly rotate. implement it in python. make sure ball stays within the square" Code ran no edits and the result is actually pretty decent:

Lab4crypto 的头像
Lab4crypto1 年前

🚀 Don't gamble with your portfolio! Use our advanced hybrid quant risk tool using on/off-chain data daily and make informed decisions. 📈 Acess to 1000+ charts for your crypto journey. 📚Receive free weekly quant analysis. 📊+21 projects supported. 🏗️ Beginners and experts.

Clem 的头像
Clem1 年前

And drawing only 60watts is insane. If we assume scaling of the last 20 years, the amount of AI resources available to us in our pockets or watches is gonna be insane

Ivan Fioravanti ᯅ 的头像
Ivan Fioravanti ᯅ1 年前

WOW! 4bit and the Big 680B model! AMAZING!

Kim Noël ⚡ 📖 的头像
Kim Noël ⚡ 📖1 年前

Soon in 8-bit on only 2 M4 Ultra 👀

Mark Lord 的头像
Mark Lord1 年前

This is so cool 🥹 Still got my fingers crossed a new M4 Ultra comes out at some point so I can hopefully buy some second hand M2 Ultras for cheap 😂 Can't wait to one day run these mega models at home!

Lewis N Watson 的头像
Lewis N Watson1 年前

Have you noticed any trade offs between high parameter but heavily/quantised vs lower parameter but full precision or even half?

Ronald Mannak 的头像
Ronald Mannak1 年前

How is it possible that 4-bit models are useful and reasonably accurate now? Just a few months ago 4-bit models were close to useless

sour coach sauers 的头像
sour coach sauers1 年前

Don’t look Jensen Huang!!!

allinasecond 的头像
allinasecond1 年前

apple is toast anyway, making money or having a moat on inference will not be a thing

X Æ A-12 的头像
X Æ A-121 年前

Do you do some test with concurrent prompt to see how many prompt can run in the same time for acceptable speed ?

相关视频