100 tok/s on GLM-5.3-Flash, GLM-5.3 runs on the 6000s @ 80-180 tok/s I have not tried the Studio, I’m sure it’s an amazing machine especially if you are in the Mac world already. From a pure logistics POV the sparks use little power, make less noise, serve good model & tokens
67,414 次观看
Tibo, why? Why must the performance suck so bad.
121,863 次观看
Deepseek-v4-Flash running locally. Makes me dizzy, this is one of 6 sessions running rn
37,143 次观看
Codex nooooooo I can't close this popup
135,878 次观看
Kitty litter, GLM-5.2 running locally. Have you seen 180 tok/s? Do you see how much room we have to push things forward? Most owners don’t know how to squeeze every token out of their GPUs
22,222 次观看
GLM-5.2-REAP-NVFP4 This experiment I calibrated the model on 30,000+ samples of my agent sessions, tweets, writing, and codebases. Pretty charming, every dataset produced something different.
22,535 次观看
If you've done work in OSS AI, or Local AI and want to have you and your company on State of Local AI - 2026 Get in touch, you'll be on as a coauthor with the companies and people who make this all possible. We moved release date to August 7th, we have 6 contributors so far
13,786 次观看
The fact that a 284B parameter model is able to in a single 400K token session - ssh into my Homelab to find docs - ssh into lambda / prime intellect - rewrite reap to support the DS4 attention - download models & datasets - run tests to check it works
22,762 次观看
What 4 concurrency on GLM sounds like. So gosh damn loud, open air frames are easier to build but so gosh darn loud
11,809 次观看
Asking Codex (GPU) to find and kill Codex (Cerebras)