
Alexey Fateev
@superalesha • 7,123 subscribers
how far can 4x RTX 3090s go?
Shorts
Videos

98.2% of Qwen3.8 27B my ass. I got hyped and gave it a real agent job right away. Build me an FPS in three.js, 6 hours on a 3090. My most standard and default prompt that I always use. It spent the first 32K tokens on a plan without writing a single file, then shipped a black screen, 2 shaders that dont compile and a player who spawns dead. And wrote "verified" in the final report. Ok, too hard. I gave it the easiest thing I have, a voxel pagoda garden in one html file. Video attached. 3 hours for THIS. On the way it deleted its own file and spent an hour debugging a raycaster nobody asked for. Where the 98.2% is in all this I have no fucking idea. Same Qwen3.8 27B, same pagoda task, ISTA-DASLab GSQ-RCO IQ2_XS on a 12 GB 3080 Ti gave me day and night, real shadows and koi fish in the pond. 8.4 GB on disk, real 2.50 bpw, 131072 ctx with q4_0 KV, 47 tok/s at 128K. 2.5 bits beats 2.13 bits by a lot when the 2.13 is this shit. Post with the video and the exact command in the replies. In the replies, I'll attach what a proper Qwen3.8 27B created in my hands. Bottom line: if you have at least 12GB VRAM, use ISTA-DASLab GSQ-RCO of all THIS.
Alexey Fateev371,144 views • 16 days ago

I put the fly brain on a skateboard and turned on dubstep..... It rides in rhythm for a while, even pushes with its back leg on the beat. Then the bass comes in and it starts shaking, falls off and lies next to the board. I did this to it again and I'll do it again.....
Alexey Fateev235,849 views • 25 days ago

YES, FUCKING YES! I DID IT !!! In my last post, I heavily trashed Qwen3.8 Flash Next, claiming it couldn't do anything at all. I managed to get awesome speed out of it, but in nearly 3 hours it couldn't even generate a simple Pagoda Garden. I had no clue what was going on. Turns out the issue was with my custom quant, which I built myself to fit 256k context + vision into 4x3090 with good throughput. Ton Cao released a properly calibrated AWQ INT4 of this model, but it was larger than mine and didn't fit into my 96GB as-is. CUDA OOM - sounds familiar to everyone, right? I decided to tackle it and make something that actually works. A full day of work, interrupted only by 5 hours of sleep. To fit 262K context, I had to un-fuse the fused layers and quantize all non-sensitive projections to FP8. vLLM allows only 1 quantization scheme per fused layer, which turned into writing a custom loader and 13 launch attempts 😭 Attempt 12 finally loaded, and I was so hyped. But the model's output was pure gibberish, total mush. Turns out TP4 was slicing the fused QKV matrix in the wrong places, for fuck's sake. Attempt 13 fixed the slicing. 262K fit. The model kept spitting garbage on any complex task. Short answers were fine, a 40K needle test was fine, but ask for a three.js scene - and you get 16,384 tokens of complete shit. I was devastated. The root cause of all this was just mind-bogglingly fucked up. 48 tiny tensors, shared_expert_gate.weight, ~5 KB each. My build script quantized them to FP8, but the runtime read them as BF16 without applying the scale 🤡 The gate expects values around 0.077, but was receiving 448 🤡🤡 Eat up, my dear friend. And this gate determines how much the shared expert contributes to every single token. I restored all 48 from the original checkpoint - and the pagoda scene, which I had spent hours grinding on and ran dozens of times - just passed twice in a row. Wanna know how mindblown I was? Exactly. Then the model swallowed 258,008 real prompt tokens with an image and nailed the needle test. Decode at 63-68 tok/s at 220 W per GPU - identical to what the broken build pushed. Video is the proof. One prompt, two sets of weights: reference BF16 vs. my Franken-quant running on 4 gaming GPUs from 2020. One-shot, zero fixes. Mine finished first, and the scene came out tighter. It's just 1 prompt, so I won't call it a benchmark. 24 hours of my life for 240 KB of broken tensors 🤡🤡🤡
Alexey Fateev129,122 views • 1 month ago

I asked Deepseek V4.1 Flash to build an FPS for me, just like I usually do with every new model. I was literally fucking blown away. To be honest, when I launched it, I wasn't expecting to see anything that would actually pleasantly surprise me. Man, was I wrong. Deepseek did what no other model has done before. Lately, I've been using the exact same prompt, no clear or strict rules, just something along the lines of "Bro, make it good." And it fucking delivered. Every model before this just made some local, small-scale arenas, roughly 5x5 meters, and spawned waves. No distinct style, nothing. They just made okay, generic shooters in different settings. Deepseek went a completely different route. It created a massive, dark, almost horror-like map where enemies could be waiting around every corner. And the atmosphere isn't one where you're the hero about to smash a hundred heads or so - it's more like they're the hunters and you're just trying to survive. No other model before this, not even Fable 5.1, has produced such a cool, realistic, and vibrant visual style. When I fired this submachine gun, at first I didn't realize what that was underneath the barrel - turned out to be a heat haze effect from the red-hot barrel. Link to the game will be in the replies so you can try it yourself. Don't take my word for it, see for yourself Yeah, there are some collision issues and probably optimization problems, but this is pretty much a one-shot without any polish.
Alexey Fateev81,822 views • 25 days ago

Qwen3.8 27B wrote a playable first person shooter start to finish on 2 used 3090s in my apartment. No cloud. No API key. No subscription. 687 steps. 87.9M tokens in, 822k out. 5 hours 11 minutes of model time plus 58 minutes of tool calls. The agent loop never broke once. And it plays. Enemies spawn and push you, the gun kicks, shadows stretch across the whole block while the sun goes down behind the towers. I sat there clearing waves instead of grading the output. Same shooter prompt I threw at the frontier models a few weeks ago. That time the tokens went to somebody elses datacenter. This time nothing left the flat. 87.9 million tokens through my own cards. On an API that run has a price tag. Here it has an electricity bill. 60 tok/s all the way through. Slower than frontier, and it stops mattering when the thing works through the night while you sleep. Local models were a toy 18 months ago. This one finished a game.
Alexey Fateev150,780 views • 1 month ago

27B model in 8.45 GB. 60 tok/s on a single 3090. And it built this voxel pagoda from a single prompt. Qwen3.8-27B in 2.4 bits from Mirai Labs is on a completely different level compared to the mess from PrismML. Yeah, it's slightly larger in size, but these guys don't claim 98% accuracy either, unlike some others. Pagoda: 1 prompt, medium reasoning, pi agent, 10 turns, 32.5K output tokens, 12 minutes. A single HTML file on three.js. Speed on my rig, GPU at 290 W, vLLM: 1x 3090: 60 tok/s with MTP, 44 without, prefill ~1190 tok/s on a 7K prompt, 159K context 2x 3090 (pipeline parallel): 66 tok/s, prefill ~1890 tok/s, full 262K context Huge thanks to Artur Chakhvadze and ryan mathieu for the opportunity to test this model before release and help the team with testing.
Alexey Fateev29,548 views • 10 days ago

Fable 5.1 made this shooter with just this short prompt. Honestly, I'm completely fucking blown away. Prompt: "Make me the most insane and blast of a shooter you can possibly build on ThreeJS + Web shaders bro! The most important things are shooting, dynamics, recoil, graphics, hit impact, gunplay. It's a beautifully designed arena where enemies swarm in hordes. Must have ADS aiming, weapon inertia/weight. 3 weapon types: assault rifle / submachine gun, shotgun, and a heavy-hitting marksman rifle. Red dot sight only on the rifle, iron sights on the rest. This is a Call Of Duty Style AAA game in the browser. Do it right bro, I believe in you!" Link on game in reply
Alexey Fateev75,649 views • 1 month ago

Im done with GLM-5.3 Flash 🤯 I went deep into the guts of GLM-5.3 Flash. Ripped out every expert, every memory head, every quantization scale, looked at all of it up close, then put the model back together so you can see the inside too. 🎉 GLM-5.3 Flash NVFP4 is now in Weight Atlas. 320B total, 18B active, 45 language layers and a 24 block Vision tower. And none of it is read from the config. I ran the deployed NVFP4 checkpoint itself on 4x RTX PRO 6000 and measured what it does. - 12,096 expert cells, 42 layers x 288 experts, each with exact REAP importance from 12.59M tokens per layer, route share and output contribution. Slice it by 14 domains, russian and vision included - one chart invalidated an assumption I had. Route count does not track expert importance, Spearman -0.222 against exact REAP. The ranking itself is stable, 0.994 split half - 2,176 KDA memory heads, each with its own half life, from 0.25 tokens to 2.7M - 19 billion NVFP4 block scales, plus the deployed FC2 QDQ error: median 21.05 dB and 21.7% of values turn into exact zeros - a causal check of REAP at inference, weights untouched. Dropping the top 2% experts changes 79% of sequences, dropping the bottom 2% changes 64% Every number comes from a capture I checksummed, 15 artifacts, and the limits are written next to each chart. Sound on I will record a video walk through every chart later. For now go poke around inside it:
Alexey Fateev56,181 views • 1 month ago

I made FlappyFly. I put a fly's brain into FlappyBird. It plays terribly. As a pipe approaches, LC4 neurons detect it getting closer. That input goes into the 166,700-neuron network, and when DNp01 fires, the wings beat. DNp01 is the giant fibre neuron that triggers a fly's escape when you try to swat it. I used that same escape circuit to get it past the pipes.
Alexey Fateev41,149 views • 24 days ago

I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins. Now it takes 53 secs. 284b in 2bit, 87gb, barely squeezes into 96gb. Me and Fable 5 spent 4 days in llama.cpp. Fixed everything that was broken
Alexey Fateev121,757 views • 2 months ago

Im done with Qwen3.8 27B 😎 Atlas covers it completely now: all 1,199 tensors, the full architecture, and what happens inside the model while it runs 🚀 The final part is the Living Model. I ran English, code and agent traces through Qwen and mapped signal flow across all 64 layers, activation quantization, every full attention layer, linear memory, 1.1M neurons, Vision, and the damage from swapping each layer to INT4. Everything comes from the actual weights or real runs. No placeholder data anywhere. Dark mode and a proper mobile view are in too. It finally feels like a product.
Alexey Fateev42,238 views • 1 month ago