0xSero's banner
0xSero's profile picture

0xSero

@0xSero65,414 subscribers

Do you see what I see?

Shorts

100 tok/s on GLM-5.3-Flash, GLM-5.3 runs on the 6000s @ 80-180 tok/s I have not tried the Studio, I’m sure it’s an amazing machine especially if you are in the Mac world already. From a pure logistics POV the sparks use little power, make less noise, serve good model & tokens

100 tok/s on GLM-5.3-Flash, GLM-5.3 runs on the 6000s @ 80-180 tok/s I have not tried the Studio, I’m sure it’s an amazing machine especially if you are in the Mac world already. From a pure logistics POV the sparks use little power, make less noise, serve good model & tokens

67,414 görüntüleme

Tibo, why? Why must the performance suck so bad.

Tibo, why? Why must the performance suck so bad.

121,863 görüntüleme

Deepseek-v4-Flash running locally. Makes me dizzy, this is one of 6 sessions running rn

Deepseek-v4-Flash running locally. Makes me dizzy, this is one of 6 sessions running rn

37,143 görüntüleme

Codex nooooooo I can't close this popup

Codex nooooooo I can't close this popup

135,878 görüntüleme

Kitty litter, GLM-5.2 running locally. Have you seen 180 tok/s? Do you see how much room we have to push things forward? Most owners don’t know how to squeeze every token out of their GPUs

Kitty litter, GLM-5.2 running locally. Have you seen 180 tok/s? Do you see how much room we have to push things forward? Most owners don’t know how to squeeze every token out of their GPUs

22,222 görüntüleme

GLM-5.2-REAP-NVFP4 This experiment I calibrated the model on 30,000+ samples of my agent sessions, tweets, writing, and codebases. Pretty charming, every dataset produced something different.

GLM-5.2-REAP-NVFP4 This experiment I calibrated the model on 30,000+ samples of my agent sessions, tweets, writing, and codebases. Pretty charming, every dataset produced something different.

22,535 görüntüleme

If you've done work in OSS AI, or Local AI and want to have you and your company on State of Local AI - 2026 Get in touch, you'll be on as a coauthor with the companies and people who make this all possible. We moved release date to August 7th, we have 6 contributors so far

If you've done work in OSS AI, or Local AI and want to have you and your company on State of Local AI - 2026 Get in touch, you'll be on as a coauthor with the companies and people who make this all possible. We moved release date to August 7th, we have 6 contributors so far

13,786 görüntüleme

The fact that a 284B parameter model is able to in a single 400K token session - ssh into my Homelab to find docs - ssh into lambda / prime intellect - rewrite reap to support the DS4 attention - download models & datasets - run tests to check it works

The fact that a 284B parameter model is able to in a single 400K token session - ssh into my Homelab to find docs - ssh into lambda / prime intellect - rewrite reap to support the DS4 attention - download models & datasets - run tests to check it works

22,762 görüntüleme

What 4 concurrency on GLM sounds like. So gosh damn loud, open air frames are easier to build but so gosh darn loud

What 4 concurrency on GLM sounds like. So gosh damn loud, open air frames are easier to build but so gosh darn loud

11,809 görüntüleme

Asking Codex (GPU) to find and kill Codex (Cerebras)

Asking Codex (GPU) to find and kill Codex (Cerebras)

24,131 görüntüleme

Videos

0xSero's profile picture

Opus at home, at 200+ tok/s I love ZAI

0xSero

200,857 görüntüleme • 7 gün önce

0xSero's profile picture

Yours to keep.

0xSero

144,267 görüntüleme • 29 gün önce

0xSero's profile picture

An autist in its natural habitat

0xSero

13,038 görüntüleme • 16 gün önce