Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Qwen3.8-Flash-Next just got a massive speed boost on my old hardware. This guy started with roughly 11 tok/s. After tuning the inference stack, I’m now getting around 24–28 tok/s. That’s roughly a 2× speedup without changing the GPU. The setup: → RTX 3090 24GB → 64GB DDR4 system RAM...

23,243 Aufrufe • vor 14 Tagen •via X (Twitter)

13 Kommentare

Profilbild von Peasant Smith
Peasant Smithvor 13 Tagen

I tried that fork on my hardware and -unfortunately- it didn't move a needle from what I was getting before. My hardware is probably the limiting factor here. I was already getting 14.1t/s decode on 3x 3060 12gb and ddr4 2400 quad channel. With this fork I got pretty much 14t/s

Profilbild von Stanislaw Weston
Stanislaw Westonvor 13 Tagen

This is a very good fork of llama.cpp. It has many useful features for MoE models, especially Qwen 3.8 Flash Next. BUT MTP actually slows things down by 1.5x. So while I use this fork, I keep MTP turned off.

Profilbild von Mohammad Zoheb
Mohammad Zohebvor 13 Tagen

I am using on server, with 158GB RAM + RTX3090 but I only allow 12GB VRAM since I have other apps to use it. Getting 17 to 18 token per second with original llama.cpp. I tested the cafe llama.cpp and made no difference. It's just basic optimizations, It is just another llama.cpp

Profilbild von Abdarahmane Traore
Abdarahmane Traorevor 13 Tagen

2x from tuning alone is the real story with this model, the kernels are young: rebuilding llama.cpp at the merged PR head nearly doubled decode for us too. Next big one for your card: MTP spec decode, once GGUFs get re-exported with the head (PR #27836). Recipes:

Profilbild von LocalAIMaxxer // Niko
LocalAIMaxxer // Nikovor 13 Tagen

Could I do this with 5080 + 96gb ram and ssd? How much sits on the gpu?

Profilbild von Diego Garcia
Diego Garciavor 14 Tagen

Hey mate Mind to share the recipe you used? Thank you!

Profilbild von Alois Vaclav
Alois Vaclavvor 13 Tagen

Stop this bullshit, 30tps cunt be your daily driver...nice experiments, but thats all. Cunt call it "locall AI" when it can be used as Codex / Claude..makes no sense.

Profilbild von why
whyvor 14 Tagen

Which part of the stack made the biggest difference: quantization, batching, kernel settings, or CPU offload?

Profilbild von LinuxVersion
LinuxVersionvor 13 Tagen

any way to quantize for use with 8GB vram?

Profilbild von De doncker kenny
De doncker kennyvor 13 Tagen

weird, i am on 17 tok/s on 8 gb vram now.

Profilbild von Zima
Zimavor 13 Tagen

If I want to build this system. What’s a brief shopping list?

Profilbild von Binsen Huang
Binsen Huangvor 13 Tagen

我的用2x5090 32GB,drr5 96g出现了预充填时间过长和解码速度过于缓慢,模型无法正常使用的情况,

Profilbild von Idriss
Idrissvor 13 Tagen

2x on the same 3090 usually means you stopped spilling into system RAM, not that the kernels got faster. Worth reporting tok/s at 8k and 32k context too, since that gap is where the tuning actually shows up.

Ähnliche Videos

🤯 A localmaxxer hit ~381 tok/s on a SINGLE RTX 3090 with Qwen3.8-27B. This developer has turned a 24GB RTX 3090 into a monster Qwen inference box - w/some creativity. Four days ago 👉 ⚡ ~82 tok/s single-user Then 👉 ⚡ ~114 tok/s with optimized MTP ⚡ ~138 tok/s with DFlash2 + lookup drafting Now 👉 🔥 ~381 tok/s on ONE request How? The recipe combines ... 🧠 Qwen3.8-27B 🎮 1× RTX 3090 24GB @ 250W ⚙️ heavily optimized vLLM ⚡ DFlash2 speculative decoding 🔎 lookup-augmented drafting 📚 prefix caching 🧮 16-token verification blocks 💾 quantized KV / heads / activations DFlash2 normally proposes 7 tokens. The developer realized the verification block doesn't have to stop there. If Qwen is answering from a document already sitting in the prompt, the system can fill the remaining draft positions using tokens found directly in that context. 🎯 So the target model can verify 16 tokens at once. On a ~25K-token document reproduction task: Previous DFlash2 👉 ~260 tok/s Longer verification + context lookup 👉 🔥 ~382 tok/s Acceptance: 🤯 15 of 16 tokens per verification step That is where the crazy number comes from. ⚠️ On ordinary real-world chat prompts, the same setup is around ~133 tok/s Still extremely fast for a dense 27B model on an RTX 3090. The 381 tok/s mode shines when the answer largely comes from material already in context so these are best use cases 📚 RAG / document Q&A 💻 Coding assistants applying edits 📝 Quoting or rewriting documents 🔎 Extracting information from long prompts And another optimization 👉 With prefix caching, a second question against the same 25K-token document reportedly goes from: 🐌 22.4 sec TTFT → ⚡ 0.56 sec TTFT Because the model doesn't need to process the whole document again. 🎯 It's specifically a mode for RAG front ends and coding agents. Follow iamMess on Reddit or syv-ai on GitHub 🔗 Reddit: r/LocalLLaMA/comments/1vtup5s/ 🔗 GitHub: /syv-ai/qwen38-27b-rtx3090

David Hendrickson

113,342 Aufrufe • vor 24 Tagen

i spent 3 hours finding the sweet spot for hermes 4.3 36B on a single RTX 3090. saving you the trouble, anon. the model is 21.8GB at Q4_K_M. that leaves 2.2GB free on 24GB VRAM. not much room for KV cache. here's what actually happened: 4K: 35.3 tok/s 8K: 35.2 tok/s 16K: 34.8 tok/s 32K: 34.6 tok/s 64K: 6.4 tok/s 128K: 1.9 tok/s flat from 4K to 32K. then it falls off a cliff at 64K. the trick is quantized KV cache. without it you OOM at 16K. with quantized KV cache you get 32K at full speed. all 64 layers on GPU. at 64K something weird happens. ngl 99 (all layers on GPU) = 3.96 tok/s. the KV cache silently spills to CPU. drop to ngl 55 and speed jumps to 6.37. drop to ngl 48 and it gets worse again (3.46). there's an offload sweet spot where you free just enough VRAM for the cache without losing too much compute to PCIe transfers. 128K works at ngl 32 but you're at 1.95 tok/s. half the model on CPU. usable for batch work, not for interactive. the sweet spot command: llama-server -m hermes-4.3-36b-Q4_K_M.gguf -ngl 99 -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 32K context. 34.6 tok/s. all on GPU. this is where dense 36B lives on 24GB. for comparison, qwen 3.5 (35B MoE, 3B active) holds 112 tok/s from 4K all the way to 262K on the same GPU. no speed drop. same total params, completely different architecture. hybrid linear attention means flat context scaling. dense pays for every token in the KV cache. code and quality comparison coming next. fast vs slow generation side by side in the videos below.

Sudo su

52,268 Aufrufe • vor 6 Monaten