
Ash Hart
@ashxhart • 3,085 subscribers
Doing my bit to bring local AI to everyone | https://t.co/BWp6YHfYSd | https://t.co/XfmTXso9Us | https://t.co/naxpf8lpU8 | Thoughts are my own.
Shorts
Videos

Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and there’s still more on the table. Running Qwen3.8-Flash-Next @ 4-bit locally. Now for the crazy bit: Artificial Analysis Intelligence Index v4.3: GPT-5.6 Sol 39 (Medium) GPT-5.6 Sol 42 (High) Qwen3.8-Flash-Next 42 GPT-5.6 Sol 44 (XHigh) GPT-5.6 Sol 47 (Max) Grok 4.5 39 (High) Claude Sonnet 5 38 (Max) *Qwen score currently estimated by Artificial Analysis. A 4-bit open-weights model running at 73 tok/s on a Mac under my desk... with an estimated intelligence score matching GPT-5.6 Sol at High reasoning. Local AI is getting ridiculous. 🔥 Plus, this model does not say no to a bit of red teaming. Once I am happy with it, I will do a pr for Jun Kim's oMLX :)
Ash Hart18,827 görüntüleme • 28 gün önce

I put MCDMA through its paces today and tested Qwen 3.8 Flash Next on my dual Spark / Mac Studio cluster. My first test was disaggregated prefill across my two Sparks, then decode onto my Studio. PP 2,100 tok/s Decode 80 tok/s at 20k context, higher on short replies. That's faster than either machine manages on its own: the Studio alone prefills at about 1,200 tok/s, and the Sparks alone decode at 35 to 65. 🤯 Mac prefill has always been way behind the Sparks; now I get the best of both worlds with the hardware I currently have.
Ash Hart13,121 görüntüleme • 21 gün önce
Daha fazla içerik yok.