
wd 🇵🇹🔺
@populartourist • 1,231 subscribers
RUSH | Jack of all trades | Trigger warning
Videos

Qwen3.8 Next Flash UD-IQ3_XXS ~1 hour + 117k total tokens RTX5090 + R9 9950X on 96GB DDR5 Decode falls from 50 tps to low 20's above 40k ctx window whilst prefill holds around 400-700 tps - no drafters. Never got to hit reasoning budget (max set 107k). Flags below.
wd 🇵🇹🔺62,911 görüntüleme • 1 ay önce

Qwen3.6 27B Q6_K (Unsloth) 199 tok/s average throughput on code + text combo on RTX 5090. Power capped 400W and clock 2200 MHz. Keeping --spec-draft-n-max 12 gives nice bump for 125k context with symmetric q8_0 KV with vision. High ceiling compression q8_0/q5_1 still possible but noticeable loss above 128k. Better lower quant at this rate or reduce --spec-draft-n-max 9 to bump closer to 140k context whilst maintaining full q8_0 KV. Mainline llama.cpp
wd 🇵🇹🔺13,972 görüntüleme • 2 ay önce
Daha fazla içerik yok.