
antirez
@antirez • 73,895 subscribers
Reproducible bugs are candies. I like programming too much for not liking automatic programming.
Shorts
Videos

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.
antirez317,088 Aufrufe • vor 27 Tagen

DeepSeek v4 PRO running via SSD streaming on my 128GB MacBook m5 max. 1.6 trillion parameters.
antirez271,840 Aufrufe • vor 1 Monat

If you need AI to do a search for you in the real world, ds4-agent is basically SOTA, because it can access the web sites without any limitations given that it uses your local Chrome browser (no, not in headless mode, that's the trick...), and DeepSeek v4 is great at search.
antirez148,707 Aufrufe • vor 1 Monat

I didn't expect DeepSeek v4 PRO (not Flash) to run well on the Mac Studio M3 Ultra with 512GB of RAM. This is 2 bit quantized with the same DwarfStar recipe used for Flash. 433GB GGUF file. 130 t/s prefill, 13 t/s generation. Prefill in the video is low because small prompt.
antirez169,897 Aufrufe • vor 2 Monaten

DS4 running on DGX Spark (GB10 / CUDA), private branch for now. 12 tokens/sec, the memory bandwidth is limited in this system, at 270GB/sec. But prefill is ways more alighed to M3 Max at ~200 t/s. I'll release when more mature, but it is almost sure that it will get merged.
antirez84,069 Aufrufe • vor 2 Monaten

Much better at 17 tokens/sec in M3 Max, fusing a few operations. 128gb m3 max.
antirez25,147 Aufrufe • vor 3 Monaten
Keine weiteren Inhalte verfügbar