
antirez
@antirez • 73,895 subscribers
Reproducible bugs are candies. I like programming too much for not liking automatic programming.
Shorts
Videos

GLM 5.2, Q2_K routed experts (effectively ~2.6 bits) running with SSD streaming on an M5 Max 128GB computer.
antirez317,088 просмотров • 27 дней назад

DeepSeek v4 PRO running via SSD streaming on my 128GB MacBook m5 max. 1.6 trillion parameters.
antirez271,840 просмотров • 1 месяц назад

If you need AI to do a search for you in the real world, ds4-agent is basically SOTA, because it can access the web sites without any limitations given that it uses your local Chrome browser (no, not in headless mode, that's the trick...), and DeepSeek v4 is great at search.
antirez148,707 просмотров • 1 месяц назад

I didn't expect DeepSeek v4 PRO (not Flash) to run well on the Mac Studio M3 Ultra with 512GB of RAM. This is 2 bit quantized with the same DwarfStar recipe used for Flash. 433GB GGUF file. 130 t/s prefill, 13 t/s generation. Prefill in the video is low because small prompt.
antirez169,897 просмотров • 2 месяцев назад

DS4 running on DGX Spark (GB10 / CUDA), private branch for now. 12 tokens/sec, the memory bandwidth is limited in this system, at 270GB/sec. But prefill is ways more alighed to M3 Max at ~200 t/s. I'll release when more mature, but it is almost sure that it will get merged.
antirez84,069 просмотров • 2 месяцев назад

The first M5 max arrived! Many many thanks to our sponsors ⿻ Audrey Tang 唐鳳 and Niels Gron, the next will arrive on Monday.
antirez27,840 просмотров • 2 месяцев назад

Much better at 17 tokens/sec in M3 Max, fusing a few operations. 128gb m3 max.
antirez25,147 просмотров • 3 месяцев назад
Больше нет контента для загрузки