Загрузка видео...
Не удалось загрузить видео
Deepseek V4.1 Flash Max Optimized 365 code / 163 prose / 333 json 3081 toks 32 concurrency 131k context, 3.7mil KV, 16k prefill at 60k 203GB engram on RAM Tuned block-FP8 GEMM table Annihilated all my agentic tasks
19,593 просмотров • 18 дней назад •via X (Twitter)
Комментарии: 21

60% of the model is experts and attention, 307 GB, resident across 4× 96 GB. The other 40% is Engram: n-gram lookup tables read a few KB per token, so they sit in host RAM (639 GB here) at zero measured cost. V4.1 was built for exactly this split.

Ok Killy! What should I buy if I was a mad lad, but not an absolute mad lad

4 x RTX 6000 if you want any chance of running these large models, with at least 256GB DDR5 RAM. CPU doesn't have to the best, something like TR Pro 9965wx Most importantly though, WATERCOOL IT. +6000mem OC, 600w full throttle all day, everyday and not see 60c is where its at

@chalkers Which waterblocks, rads and pump did you use? Thanks.

@chalkers Optimus, only 1 slot waterblock available and quality is outstanding. 2 560x60mm, 560x86 and 360x86 rads, two D5 pumps. Loop being so large with lots of choke points, flow was on the low side so bought four more that will be installed next week

oeeehhhhhh the speed is getting me hyped like I wanna cheer for it 😁

that netterm harness/terminal is still looking really cool btw

try with engram q4 for less ram needed

how many model calls in the workflow, or is it single shot per task

my apologizes, having trouble understanding the ask. Which workflow specifically?

My wet dream? 😅🤣😂

did u setup and assemble the hardware yourself ?

Sure did

do you have anything written or filmed about the setup process from scratch, and things like tradeoffs/decisions you had to make?

Nothing like that, just few picture of the process here and there. If you are interested, I’d be happy to put together something for Watercooling Blackwell GPUs and server parts

yes please It would be much appreciated by me and many others I think

the dgx station is inferior?

Its two different beast really. DGX Station is superior for LLM specific tasks, setup like mine is more flexible

@softpoo "Superior"? You have 384 worth of VRAM, DGX station only 252Gb (ok, higher bw, but still..)

@softpoo Yeah I’d rather have 7.1TB bandwidth compared to 1.9TB on 6000s. Real talk, I’ve found running multiple small/mid size models superior to running one large one like this

Bro thats ridiculous lol. I want 80k in gpus 😆 the amount of crazy shit u can do with that kind of throughput 🤔
