Загрузка видео...

Не удалось загрузить видео

На главную

Deepseek V4.1 Flash Max Optimized 365 code / 163 prose / 333 json 3081 toks 32 concurrency 131k context, 3.7mil KV, 16k prefill at 60k 203GB engram on RAM Tuned block-FP8 GEMM table Annihilated all my agentic tasks

19,593 просмотров • 18 дней назад •via X (Twitter)

Комментарии: 21

Фото профиля Killy
Killy18 дней назад

60% of the model is experts and attention, 307 GB, resident across 4× 96 GB. The other 40% is Engram: n-gram lookup tables read a few KB per token, so they sit in host RAM (639 GB here) at zero measured cost. V4.1 was built for exactly this split.

Фото профиля Chalkers
Chalkers18 дней назад

Ok Killy! What should I buy if I was a mad lad, but not an absolute mad lad

Фото профиля Killy
Killy18 дней назад

4 x RTX 6000 if you want any chance of running these large models, with at least 256GB DDR5 RAM. CPU doesn't have to the best, something like TR Pro 9965wx Most importantly though, WATERCOOL IT. +6000mem OC, 600w full throttle all day, everyday and not see 60c is where its at

Фото профиля moskstraumen
moskstraumen18 дней назад

@chalkers Which waterblocks, rads and pump did you use? Thanks.

Фото профиля Killy
Killy18 дней назад

@chalkers Optimus, only 1 slot waterblock available and quality is outstanding. 2 560x60mm, 560x86 and 360x86 rads, two D5 pumps. Loop being so large with lots of choke points, flow was on the low side so bought four more that will be installed next week

Фото профиля A.K.A CS
A.K.A CS18 дней назад

oeeehhhhhh the speed is getting me hyped like I wanna cheer for it 😁

Фото профиля A.K.A CS
A.K.A CS18 дней назад

that netterm harness/terminal is still looking really cool btw

Фото профиля CV.YH
CV.YH18 дней назад

try with engram q4 for less ram needed

Фото профиля Aleksandar Janca
Aleksandar Janca18 дней назад

how many model calls in the workflow, or is it single shot per task

Фото профиля Killy
Killy18 дней назад

my apologizes, having trouble understanding the ask. Which workflow specifically?

Фото профиля don lee
don lee18 дней назад

My wet dream? 😅🤣😂

Фото профиля bashar
bashar18 дней назад

did u setup and assemble the hardware yourself ?

Фото профиля Killy
Killy18 дней назад

Sure did

Фото профиля bashar
bashar18 дней назад

do you have anything written or filmed about the setup process from scratch, and things like tradeoffs/decisions you had to make?

Фото профиля Killy
Killy18 дней назад

Nothing like that, just few picture of the process here and there. If you are interested, I’d be happy to put together something for Watercooling Blackwell GPUs and server parts

Фото профиля bashar
bashar18 дней назад

yes please It would be much appreciated by me and many others I think

Фото профиля Neko Legends
Neko Legends18 дней назад

the dgx station is inferior?

Фото профиля Killy
Killy18 дней назад

Its two different beast really. DGX Station is superior for LLM specific tasks, setup like mine is more flexible

Фото профиля moskstraumen
moskstraumen18 дней назад

@softpoo "Superior"? You have 384 worth of VRAM, DGX station only 252Gb (ok, higher bw, but still..)

Фото профиля Killy
Killy18 дней назад

@softpoo Yeah I’d rather have 7.1TB bandwidth compared to 1.9TB on 6000s. Real talk, I’ve found running multiple small/mid size models superior to running one large one like this

Фото профиля DaayTerkErJerbs
DaayTerkErJerbs18 дней назад

Bro thats ridiculous lol. I want 80k in gpus 😆 the amount of crazy shit u can do with that kind of throughput 🤔

Похожие видео