Loading video...
Video Failed to Load
Qwen3.6 27B Q6_K (Unsloth) 199 tok/s average throughput on code + text combo on RTX 5090. Power capped 400W and clock 2200 MHz. Keeping --spec-draft-n-max 12 gives nice bump for 125k context with symmetric q8_0 KV with vision. High ceiling compression q8_0/q5_1 still possible but noticeable loss above 128k.... show more
13,972 views • 2 months ago •via X (Twitter)
22 Comments

Adjust to cater. ``` command: /opt/llama/bin/llama-server -m /models/Qwen3.6-27B-Q6_K.gguf --mmproj /models/mmproj-Qwable-5-27B-Coder-f16.gguf --image-min-tokens 1024 --host 0.0.0.0 --port 8081 -t 10 -tb 16 --jinja --chat-template-kwargs {preserve_thinking:true} --reasoning-preserve --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --repeat-last-n 512 --seed 42 --flash-attn on -b 1024 -ub 512 --no-mmap --kv-unified --parallel 1 -ngl 999 -ctk q8_0 -ctv q8_0 --metrics --spec-type draft-dflash --spec-draft-model /models/qwen3.6-27b-dflash-IQ4_XS.gguf --spec-draft-n-max 12 --spec-draft-p-min 0.0 -c 125072 ```

This is nice. But then the question is what's better quality q6 or nvfp4?

Tough trades. Code fidelity, Q6. NVFP4 for faster checkpoints. Reloading on llama.cpp can be painful on long horizon contexts >150k.

I thought the fp4 math made it closer to q8 in quality ?

It doesn't intrinsically make it better than a good Q6. Good is keyword here. fp4 partially compensates quantization and uses a dynamic range vs int4 and ModelOpt optimisation. Against a Q4 I'd bet NVFP4 any day.

Bom post! Obrigado. Estou pensando em limitar a energia a 400w na minha 5090 também. Você notou alguma queda de qualidade fazendo isso, ou há apenas potencial para afetar a velocidade?

Sim tem uma pequena queda, mas com o desenvolvimento parece que não perdeu. O que ganha em eficiência compensa. Tem povo que acha 450W é o ideal, tem de ser escolha própria. Tenho rodado 400W mais de um ano tranquilo. Só recentemente diminui o clock para 2200 MHz. Tranquilo

@3682539376x Did you do any under volt + OC ? I've to 5090s (MSI gaming trio) i might try it on

@3682539376x Just power cap and lower clock. Less fiddly.

Finally someone doing this right! I like your methods man, but i think you need a better tool for Llama, than CLI. Thanks me later - you will be the first one ever test what i have been cooking for 6 months of sleepless nights ;-) This is the absolute cutting edge and power users paradise. You will iterate those flag combinations an order of magnitude faster and get the full UX comfort.

Cool, are you using a model that you quantized yourself from Z Lab's official Qwen3.6 27B dflash? I tried quantizing it myself and found the acceptance rate to be really low. 🤔

yes, Z Lab's

Thanks for the reply. On my 4090, I'm only getting 0.08 acceptance rate with --spec-draft-n-max 8, so it seems like MTP would be a better fit for me.🥲

Depends other flags you might have. Repetition penalty usually is a gain stomper

I'm using the official recommended parameters for Qwen3.6 27B, here's my full command: ``` llama-server.exe ` --model "{MODEL_PATH}/Qwen3.6-27B-Q6_K.gguf" ` --mmproj "{MODEL_PATH}/mmproj-BF16.gguf" ` --model-draft "{DRAFT_MODEL_PATH}/Qwen3.6-27B-DFlash-Q5_K_M.gguf" ` --alias "qwen3.6-27b" ` --spec-type draft-dflash ` --spec-draft-n-max 2 ` --spec-draft-ngl all ` -ngl all ` --flash-attn on ` -b 2048 ` -ub 1024 ` -np 2 ` -c 200000 ` --kv-unified ` --spec-draft-type-k bf16 ` --spec-draft-type-v bf16 ` --cache-type-k bf16 --cache-type-v bf16 ` --ctx-checkpoints 32 ` --checkpoint-min-step 8192 ` --cache-ram 20480 ` --image-min-tokens 1024 ` --image-max-tokens 16384 ` --temp 0.6 ` --top-p 0.95 ` --top-k 20 ` --min-p 0.0 ` --repeat-penalty 1.0 ` --presence-penalty 0.0 ` --chat-template-kwargs '{\"preserve_thinking\":true}' ` --reasoning on ` --host 0.0.0.0 ` --port 8001 ` --metrics ` --log-prefix ` --log-timestamps ` --verbose ` --log-verbosity 3 ``` Note that `--repeat-penalty` is set to 1.0 (no penalty), so that shouldn't be the issue. If everything is working as intended, the acceptance rate shouldn't be as low as 0.08 🤔I'll look into other possible causes as well.

unless it's OS related as I'm on Fedora.

Wow that's fast!

2-3 tokens of prediction is fastest for me but I've never gone anywhere near 12.

dont be shy

I limited power to 408 and saw only 3% decline in speed with temps sitting low 60s instead of high 70s

Have you noticed any accuracy shift at higher draft? I haven’t ventured past 2

Any concurrency?
