Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Deepseek V4.1 Flash Max Optimized 365 code / 163 prose / 333 json 3081 toks 32 concurrency 131k context, 3.7mil KV, 16k prefill at 60k 203GB engram on RAM Tuned block-FP8 GEMM table Annihilated all my agentic tasks

19,593 Aufrufe • vor 18 Tagen •via X (Twitter)

21 Kommentare

Profilbild von Killy
Killyvor 18 Tagen

60% of the model is experts and attention, 307 GB, resident across 4× 96 GB. The other 40% is Engram: n-gram lookup tables read a few KB per token, so they sit in host RAM (639 GB here) at zero measured cost. V4.1 was built for exactly this split.

Profilbild von Chalkers
Chalkersvor 18 Tagen

Ok Killy! What should I buy if I was a mad lad, but not an absolute mad lad

Profilbild von Killy
Killyvor 18 Tagen

4 x RTX 6000 if you want any chance of running these large models, with at least 256GB DDR5 RAM. CPU doesn't have to the best, something like TR Pro 9965wx Most importantly though, WATERCOOL IT. +6000mem OC, 600w full throttle all day, everyday and not see 60c is where its at

Profilbild von moskstraumen
moskstraumenvor 18 Tagen

@chalkers Which waterblocks, rads and pump did you use? Thanks.

Profilbild von Killy
Killyvor 18 Tagen

@chalkers Optimus, only 1 slot waterblock available and quality is outstanding. 2 560x60mm, 560x86 and 360x86 rads, two D5 pumps. Loop being so large with lots of choke points, flow was on the low side so bought four more that will be installed next week

Profilbild von A.K.A CS
A.K.A CSvor 18 Tagen

oeeehhhhhh the speed is getting me hyped like I wanna cheer for it 😁

Profilbild von A.K.A CS
A.K.A CSvor 18 Tagen

that netterm harness/terminal is still looking really cool btw

Profilbild von CV.YH
CV.YHvor 18 Tagen

try with engram q4 for less ram needed

Profilbild von Aleksandar Janca
Aleksandar Jancavor 18 Tagen

how many model calls in the workflow, or is it single shot per task

Profilbild von Killy
Killyvor 18 Tagen

my apologizes, having trouble understanding the ask. Which workflow specifically?

Profilbild von don lee
don leevor 18 Tagen

My wet dream? 😅🤣😂

Profilbild von bashar
basharvor 18 Tagen

did u setup and assemble the hardware yourself ?

Profilbild von Killy
Killyvor 18 Tagen

Sure did

Profilbild von bashar
basharvor 18 Tagen

do you have anything written or filmed about the setup process from scratch, and things like tradeoffs/decisions you had to make?

Profilbild von Killy
Killyvor 18 Tagen

Nothing like that, just few picture of the process here and there. If you are interested, I’d be happy to put together something for Watercooling Blackwell GPUs and server parts

Profilbild von bashar
basharvor 18 Tagen

yes please It would be much appreciated by me and many others I think

Profilbild von Neko Legends
Neko Legendsvor 18 Tagen

the dgx station is inferior?

Profilbild von Killy
Killyvor 18 Tagen

Its two different beast really. DGX Station is superior for LLM specific tasks, setup like mine is more flexible

Profilbild von moskstraumen
moskstraumenvor 18 Tagen

@softpoo "Superior"? You have 384 worth of VRAM, DGX station only 252Gb (ok, higher bw, but still..)

Profilbild von Killy
Killyvor 18 Tagen

@softpoo Yeah I’d rather have 7.1TB bandwidth compared to 1.9TB on 6000s. Real talk, I’ve found running multiple small/mid size models superior to running one large one like this

Profilbild von DaayTerkErJerbs
DaayTerkErJerbsvor 18 Tagen

Bro thats ridiculous lol. I want 80k in gpus 😆 the amount of crazy shit u can do with that kind of throughput 🤔

Ähnliche Videos