Alexey Fateev's banner
Alexey Fateev's profile picture

Alexey Fateev

@superalesha • 7,227 subscribers

how far can 4x RTX 3090s go?

Shorts

1-bit Qwen3.8-27B just more than DOUBLED its MMLU-Pro score. They didn’t change the decoder. 29.04% -> 61.54%, according to a new paper from ISTA-DASLab. Same lab behind GSQ-RCO, some of the best tiny Qwen3.8-27B quants in my book. Their method is called Disaggregated Quantization. You use different weights to read the prompt and generate the answer. A trained NVFP4 prefiller processes your input and builds the cache. Then the tiny GGUF decoder takes over. They train the prefiller specifically for that decoder, so it learns to produce representations the heavily compressed model can use. The extra checkpoint is 12.8 GiB, but you don’t need all of it in VRAM. Stream it from SSD layer by layer, reusing GPU memory. With a long enough prompt, loading can overlap computation. That’s a pretty fucking good reason to pay attention if your GPU is short on memory. I want to test this on my 4x3090 rig with NVMe. The released implementation targets Blackwell, so this needs adaptation. First I’ll check whether the quality gains survive on Ampere, then measure whether offloading actually helps. Paper:

1-bit Qwen3.8-27B just more than DOUBLED its MMLU-Pro score. They didn’t change the decoder. 29.04% -> 61.54%, according to a new paper from ISTA-DASLab. Same lab behind GSQ-RCO, some of the best tiny Qwen3.8-27B quants in my book. Their method is called Disaggregated Quantization. You use different weights to read the prompt and generate the answer. A trained NVFP4 prefiller processes your input and builds the cache. Then the tiny GGUF decoder takes over. They train the prefiller specifically for that decoder, so it learns to produce representations the heavily compressed model can use. The extra checkpoint is 12.8 GiB, but you don’t need all of it in VRAM. Stream it from SSD layer by layer, reusing GPU memory. With a long enough prompt, loading can overlap computation. That’s a pretty fucking good reason to pay attention if your GPU is short on memory. I want to test this on my 4x3090 rig with NVMe. The released implementation targets Blackwell, so this needs adaptation. First I’ll check whether the quality gains survive on Ampere, then measure whether offloading actually helps. Paper:

19,210 görüntüleme

I think I need professional medical help. I cant stop. Told myself just one more and that was 40 videos ago. Until Qwen3.8 27B drops my 4x3090 will keep cooking these until they die. More videos and the prompts are in the replies.

I think I need professional medical help. I cant stop. Told myself just one more and that was 40 videos ago. Until Qwen3.8 27B drops my 4x3090 will keep cooking these until they die. More videos and the prompts are in the replies.

90,314 görüntüleme

Minimax H3, same 4x3090 rig, same clip. Yesterday it rendered in 11:21. Today it renders in 3:45 🤯 The hardware did not change. Two things happened overnight. An AI agent rewrote the attention CUDA kernel and hit a half speed trap inside GeForce tensor cores that most people never heard of. Then a stranger on HuggingFace dropped a LoRA that cuts 20 sampling steps down to 4. One of these two mattered way more than the other. Breakdown below. I packed all of it into ready ComfyUI workflows for my rig. If you want them, ask in the replies and I'll share.

Minimax H3, same 4x3090 rig, same clip. Yesterday it rendered in 11:21. Today it renders in 3:45 🤯 The hardware did not change. Two things happened overnight. An AI agent rewrote the attention CUDA kernel and hit a half speed trap inside GeForce tensor cores that most people never heard of. Then a stranger on HuggingFace dropped a LoRA that cuts 20 sampling steps down to 4. One of these two mattered way more than the other. Breakdown below. I packed all of it into ready ComfyUI workflows for my rig. If you want them, ask in the replies and I'll share.

70,024 görüntüleme

Videos

superalesha's profile picture

pretty sure Mia blocks me after this one

Alexey Fateev

30,394 görüntüleme • 1 gün önce

superalesha's profile picture

Opus 5.5 made this in Blender This is AGI?

Alexey Fateev

133,333 görüntüleme • 12 gün önce

superalesha's profile picture

YES, FUCKING YES! I DID IT !!! In my last post, I heavily trashed Qwen3.8 Flash Next, claiming it couldn't do anything at all. I managed to get awesome speed out of it, but in nearly 3 hours it couldn't even generate a simple Pagoda Garden. I had no clue what was going on. Turns out the issue was with my custom quant, which I built myself to fit 256k context + vision into 4x3090 with good throughput. Ton Cao released a properly calibrated AWQ INT4 of this model, but it was larger than mine and didn't fit into my 96GB as-is. CUDA OOM - sounds familiar to everyone, right? I decided to tackle it and make something that actually works. A full day of work, interrupted only by 5 hours of sleep. To fit 262K context, I had to un-fuse the fused layers and quantize all non-sensitive projections to FP8. vLLM allows only 1 quantization scheme per fused layer, which turned into writing a custom loader and 13 launch attempts 😭 Attempt 12 finally loaded, and I was so hyped. But the model's output was pure gibberish, total mush. Turns out TP4 was slicing the fused QKV matrix in the wrong places, for fuck's sake. Attempt 13 fixed the slicing. 262K fit. The model kept spitting garbage on any complex task. Short answers were fine, a 40K needle test was fine, but ask for a three.js scene - and you get 16,384 tokens of complete shit. I was devastated. The root cause of all this was just mind-bogglingly fucked up. 48 tiny tensors, shared_expert_gate.weight, ~5 KB each. My build script quantized them to FP8, but the runtime read them as BF16 without applying the scale 🤡 The gate expects values around 0.077, but was receiving 448 🤡🤡 Eat up, my dear friend. And this gate determines how much the shared expert contributes to every single token. I restored all 48 from the original checkpoint - and the pagoda scene, which I had spent hours grinding on and ran dozens of times - just passed twice in a row. Wanna know how mindblown I was? Exactly. Then the model swallowed 258,008 real prompt tokens with an image and nailed the needle test. Decode at 63-68 tok/s at 220 W per GPU - identical to what the broken build pushed. Video is the proof. One prompt, two sets of weights: reference BF16 vs. my Franken-quant running on 4 gaming GPUs from 2020. One-shot, zero fixes. Mine finished first, and the scene came out tighter. It's just 1 prompt, so I won't call it a benchmark. 24 hours of my life for 240 KB of broken tensors 🤡🤡🤡

Alexey Fateev

129,122 görüntüleme • 1 ay önce

superalesha's profile picture

I asked Deepseek V4.1 Flash to build an FPS for me, just like I usually do with every new model. I was literally fucking blown away. To be honest, when I launched it, I wasn't expecting to see anything that would actually pleasantly surprise me. Man, was I wrong. Deepseek did what no other model has done before. Lately, I've been using the exact same prompt, no clear or strict rules, just something along the lines of "Bro, make it good." And it fucking delivered. Every model before this just made some local, small-scale arenas, roughly 5x5 meters, and spawned waves. No distinct style, nothing. They just made okay, generic shooters in different settings. Deepseek went a completely different route. It created a massive, dark, almost horror-like map where enemies could be waiting around every corner. And the atmosphere isn't one where you're the hero about to smash a hundred heads or so - it's more like they're the hunters and you're just trying to survive. No other model before this, not even Fable 5.1, has produced such a cool, realistic, and vibrant visual style. When I fired this submachine gun, at first I didn't realize what that was underneath the barrel - turned out to be a heat haze effect from the red-hot barrel. Link to the game will be in the replies so you can try it yourself. Don't take my word for it, see for yourself Yeah, there are some collision issues and probably optimization problems, but this is pretty much a one-shot without any polish.

Alexey Fateev

81,822 görüntüleme • 25 gün önce