Loading video...

Video Failed to Load

Go Home

Deepseek V4.1 Flash Max Optimized 365 code / 163 prose / 333 json 3081 toks 32 concurrency 131k context, 3.7mil KV, 16k prefill at 60k 203GB engram on RAM Tuned block-FP8 GEMM table Annihilated all my agentic tasks

19,593 views • 18 days ago •via X (Twitter)

21 Comments

Killy's profile picture
Killy18 days ago

60% of the model is experts and attention, 307 GB, resident across 4× 96 GB. The other 40% is Engram: n-gram lookup tables read a few KB per token, so they sit in host RAM (639 GB here) at zero measured cost. V4.1 was built for exactly this split.

Chalkers's profile picture
Chalkers18 days ago

Ok Killy! What should I buy if I was a mad lad, but not an absolute mad lad

Killy's profile picture
Killy18 days ago

4 x RTX 6000 if you want any chance of running these large models, with at least 256GB DDR5 RAM. CPU doesn't have to the best, something like TR Pro 9965wx Most importantly though, WATERCOOL IT. +6000mem OC, 600w full throttle all day, everyday and not see 60c is where its at

moskstraumen's profile picture
moskstraumen18 days ago

@chalkers Which waterblocks, rads and pump did you use? Thanks.

Killy's profile picture
Killy18 days ago

@chalkers Optimus, only 1 slot waterblock available and quality is outstanding. 2 560x60mm, 560x86 and 360x86 rads, two D5 pumps. Loop being so large with lots of choke points, flow was on the low side so bought four more that will be installed next week

A.K.A CS's profile picture
A.K.A CS18 days ago

oeeehhhhhh the speed is getting me hyped like I wanna cheer for it 😁

A.K.A CS's profile picture
A.K.A CS18 days ago

that netterm harness/terminal is still looking really cool btw

CV.YH's profile picture
CV.YH18 days ago

try with engram q4 for less ram needed

Aleksandar Janca's profile picture
Aleksandar Janca18 days ago

how many model calls in the workflow, or is it single shot per task

Killy's profile picture
Killy18 days ago

my apologizes, having trouble understanding the ask. Which workflow specifically?

don lee's profile picture
don lee18 days ago

My wet dream? 😅🤣😂

bashar's profile picture
bashar18 days ago

did u setup and assemble the hardware yourself ?

Killy's profile picture
Killy18 days ago

Sure did

bashar's profile picture
bashar18 days ago

do you have anything written or filmed about the setup process from scratch, and things like tradeoffs/decisions you had to make?

Killy's profile picture
Killy18 days ago

Nothing like that, just few picture of the process here and there. If you are interested, I’d be happy to put together something for Watercooling Blackwell GPUs and server parts

bashar's profile picture
bashar18 days ago

yes please It would be much appreciated by me and many others I think

Neko Legends's profile picture
Neko Legends18 days ago

the dgx station is inferior?

Killy's profile picture
Killy18 days ago

Its two different beast really. DGX Station is superior for LLM specific tasks, setup like mine is more flexible

moskstraumen's profile picture
moskstraumen18 days ago

@softpoo "Superior"? You have 384 worth of VRAM, DGX station only 252Gb (ok, higher bw, but still..)

Killy's profile picture
Killy18 days ago

@softpoo Yeah I’d rather have 7.1TB bandwidth compared to 1.9TB on 6000s. Real talk, I’ve found running multiple small/mid size models superior to running one large one like this

DaayTerkErJerbs's profile picture
DaayTerkErJerbs18 days ago

Bro thats ridiculous lol. I want 80k in gpus 😆 the amount of crazy shit u can do with that kind of throughput 🤔

Related Videos