Loading video...

Video Failed to Load

Go Home

Qwen3.8-Flash-Next just got a massive speed boost on my old hardware. This guy started with roughly 11 tok/s. After tuning the inference stack, I’m now getting around 24–28 tok/s. That’s roughly a 2× speedup without changing the GPU. The setup: → RTX 3090 24GB → 64GB DDR4 system RAM...

23,243 views • 13 days ago •via X (Twitter)

13 Comments

Peasant Smith's profile picture
Peasant Smith13 days ago

I tried that fork on my hardware and -unfortunately- it didn't move a needle from what I was getting before. My hardware is probably the limiting factor here. I was already getting 14.1t/s decode on 3x 3060 12gb and ddr4 2400 quad channel. With this fork I got pretty much 14t/s

Stanislaw Weston's profile picture
Stanislaw Weston13 days ago

This is a very good fork of llama.cpp. It has many useful features for MoE models, especially Qwen 3.8 Flash Next. BUT MTP actually slows things down by 1.5x. So while I use this fork, I keep MTP turned off.

Mohammad Zoheb's profile picture
Mohammad Zoheb13 days ago

I am using on server, with 158GB RAM + RTX3090 but I only allow 12GB VRAM since I have other apps to use it. Getting 17 to 18 token per second with original llama.cpp. I tested the cafe llama.cpp and made no difference. It's just basic optimizations, It is just another llama.cpp

Abdarahmane Traore's profile picture
Abdarahmane Traore13 days ago

2x from tuning alone is the real story with this model, the kernels are young: rebuilding llama.cpp at the merged PR head nearly doubled decode for us too. Next big one for your card: MTP spec decode, once GGUFs get re-exported with the head (PR #27836). Recipes:

LocalAIMaxxer // Niko's profile picture
LocalAIMaxxer // Niko13 days ago

Could I do this with 5080 + 96gb ram and ssd? How much sits on the gpu?

Diego Garcia's profile picture
Diego Garcia13 days ago

Hey mate Mind to share the recipe you used? Thank you!

Alois Vaclav's profile picture
Alois Vaclav13 days ago

Stop this bullshit, 30tps cunt be your daily driver...nice experiments, but thats all. Cunt call it "locall AI" when it can be used as Codex / Claude..makes no sense.

why's profile picture
why13 days ago

Which part of the stack made the biggest difference: quantization, batching, kernel settings, or CPU offload?

LinuxVersion's profile picture
LinuxVersion13 days ago

any way to quantize for use with 8GB vram?

De doncker kenny's profile picture
De doncker kenny13 days ago

weird, i am on 17 tok/s on 8 gb vram now.

Zima's profile picture
Zima13 days ago

If I want to build this system. What’s a brief shopping list?

Binsen Huang's profile picture
Binsen Huang13 days ago

我的用2x5090 32GB,drr5 96g出现了预充填时间过长和解码速度过于缓慢,模型无法正常使用的情况,

Idriss's profile picture
Idriss13 days ago

2x on the same 3090 usually means you stopped spilling into system RAM, not that the kernels got faster. Worth reporting tok/s at 8k and 32k context too, since that gap is where the tuning actually shows up.

Related Videos

🤯 A localmaxxer hit ~381 tok/s on a SINGLE RTX 3090 with Qwen3.8-27B. This developer has turned a 24GB RTX 3090 into a monster Qwen inference box - w/some creativity. Four days ago 👉 ⚡ ~82 tok/s single-user Then 👉 ⚡ ~114 tok/s with optimized MTP ⚡ ~138 tok/s with DFlash2 + lookup drafting Now 👉 🔥 ~381 tok/s on ONE request How? The recipe combines ... 🧠 Qwen3.8-27B 🎮 1× RTX 3090 24GB @ 250W ⚙️ heavily optimized vLLM ⚡ DFlash2 speculative decoding 🔎 lookup-augmented drafting 📚 prefix caching 🧮 16-token verification blocks 💾 quantized KV / heads / activations DFlash2 normally proposes 7 tokens. The developer realized the verification block doesn't have to stop there. If Qwen is answering from a document already sitting in the prompt, the system can fill the remaining draft positions using tokens found directly in that context. 🎯 So the target model can verify 16 tokens at once. On a ~25K-token document reproduction task: Previous DFlash2 👉 ~260 tok/s Longer verification + context lookup 👉 🔥 ~382 tok/s Acceptance: 🤯 15 of 16 tokens per verification step That is where the crazy number comes from. ⚠️ On ordinary real-world chat prompts, the same setup is around ~133 tok/s Still extremely fast for a dense 27B model on an RTX 3090. The 381 tok/s mode shines when the answer largely comes from material already in context so these are best use cases 📚 RAG / document Q&A 💻 Coding assistants applying edits 📝 Quoting or rewriting documents 🔎 Extracting information from long prompts And another optimization 👉 With prefix caching, a second question against the same 25K-token document reportedly goes from: 🐌 22.4 sec TTFT → ⚡ 0.56 sec TTFT Because the model doesn't need to process the whole document again. 🎯 It's specifically a mode for RAG front ends and coding agents. Follow iamMess on Reddit or syv-ai on GitHub 🔗 Reddit: r/LocalLLaMA/comments/1vtup5s/ 🔗 GitHub: /syv-ai/qwen38-27b-rtx3090

David Hendrickson

113,342 views • 24 days ago

i spent 3 hours finding the sweet spot for hermes 4.3 36B on a single RTX 3090. saving you the trouble, anon. the model is 21.8GB at Q4_K_M. that leaves 2.2GB free on 24GB VRAM. not much room for KV cache. here's what actually happened: 4K: 35.3 tok/s 8K: 35.2 tok/s 16K: 34.8 tok/s 32K: 34.6 tok/s 64K: 6.4 tok/s 128K: 1.9 tok/s flat from 4K to 32K. then it falls off a cliff at 64K. the trick is quantized KV cache. without it you OOM at 16K. with quantized KV cache you get 32K at full speed. all 64 layers on GPU. at 64K something weird happens. ngl 99 (all layers on GPU) = 3.96 tok/s. the KV cache silently spills to CPU. drop to ngl 55 and speed jumps to 6.37. drop to ngl 48 and it gets worse again (3.46). there's an offload sweet spot where you free just enough VRAM for the cache without losing too much compute to PCIe transfers. 128K works at ngl 32 but you're at 1.95 tok/s. half the model on CPU. usable for batch work, not for interactive. the sweet spot command: llama-server -m hermes-4.3-36b-Q4_K_M.gguf -ngl 99 -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 32K context. 34.6 tok/s. all on GPU. this is where dense 36B lives on 24GB. for comparison, qwen 3.5 (35B MoE, 3B active) holds 112 tok/s from 4K all the way to 262K on the same GPU. no speed drop. same total params, completely different architecture. hybrid linear attention means flat context scaling. dense pays for every token in the KV cache. code and quality comparison coming next. fast vs slow generation side by side in the videos below.

Sudo su

52,268 views • 6 months ago