Загрузка видео...

Не удалось загрузить видео

На главную

Qwen3.6 35B-A3B MTP is a devil performing around 1,000 tok/s on code with RTX 5090 once the prediction gets easier. Been a while I've tried 35B-A3B. Seems I might be sticking around for the fun.

49,211 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 28

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

Mainline llama.cpp ``` command: - /opt/llama/bin/llama-server - -m - /models/Huihui-Qwen3.6-35B-A3B-abliterated-ggml-model-Q4_K.gguf - --host - 0.0.0.0 - --port - 8081 - -t - 10 - --chat-template-file - /models/qwen36_chat_template.jinja - --temp - 0.05 - --top-p - 0.75 - --top-k - 10 - --min-p - 0.0 - --presence-penalty - 0.0 - --repeat-penalty - 1.00 - --repeat-last-n - 256 - --seed - 42 - --flash-attn - on - -b - 256 - -ub - 256 - --no-mmap - --parallel - 4 - -ngl - 999 - --kv-unified - --spec-type - draft-mtp,ngram-mod,ngram-map-k4v - --spec-draft-n-max - 3 - --spec-draft-p-min - 0.0 - --spec-ngram-mod-n-match - 24 - --spec-ngram-mod-n-min - 48 - --spec-ngram-mod-n-max - 64 - --spec-ngram-map-k4v-size-n - 16 - --spec-ngram-map-k4v-size-m - 96 - --spec-ngram-map-k4v-min-hits - 1 - --no-cache-idle-slots - --cache-ram - 0 - --ctx-checkpoints - 32 - --cache-reuse - 0 ```

Фото профиля Gabu
Gabu2 месяцев назад

Cmon man ;) It's perrforming copy-paste @ 1000t/s ;) Real decode was ~280t/s

Фото профиля luke
luke2 месяцев назад

wow 1k tok/s is impressive! i also just did a benchmark for qwen3.6 27b nvfp4 by @UnslothAI on a 5090, it was ~120 tok/s with nspec = 3 plan to do another benchmark for qwen3.6 35b nvfp4 today.

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

Haven't had the patience to make another docker to try Unsloth's NVFP4 so I got stuck with Nvidia's. Pretty awesome PP - upwards 15k tok/s throughput. Decode is pretty stable, but I think circa 100 tok/s. I got DFlash working above 200 tok/s on up to 50/70k context fill. Need to see if it drops above 100k context. Seems good. Might miss vLLMs PP but might be changing flavours with llama.cpp for now. gg

Фото профиля Paul
Paul2 месяцев назад

The speed is only part of the picture. For example, I would not use 35b for long chain agentic coding. It starts to hallucinate when the context gets large, like over 120k, and then makes too many mistakes, keeps doubting itself, etc. On a 5090 I would go for qwen3.6 27b nvfp4 with built in mtp and fp8 kv cache or better. I love the speed of 35b, but it’s just not good enough for a lot of real coding work in large repos.

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

indeed 35B is not great for "complex" agentic coding, but still has utility. however, you're meant to choose the right tool for the job. also 27B on NVFP4 is not that superb either, when a Q5 or Q6 offers better fidelity. altough the 27B still beats 35B-A3B at 4-bit in that case👍

Фото профиля Phillip Hamilton
Phillip Hamilton2 месяцев назад

3.6 35B-A3B MTP has been my favorite so far, especially for gaming GPUs, im switching around desktop configuration at the moment, and adding a second dusty RTX3060 gpu to run another model, trying to get everything up and running this week, waiting on shipment of a m.2 -> pcie adapter. then use a shared context pool/coordinator so they can share context quickly. hopefully cuda isnt going to complain and Ampere + Blackwell wont cause problems. So ill probably be on Open Kernel Modules package

Фото профиля azzurro
azzurro2 месяцев назад

just switched to it from 27B dense and I'm not regretting it. ~2k tps prefill and up to 100tps decode at Q5_K_XL 200k Q8_0 KV cache on two 5060 TIs. And that's on a Dual E5-2640 v4 Xeon from the stone-age.

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

I still use both. The MoE is useful on some personal workflows and these bursts help a lot.

Фото профиля Space Oddity🇺🇸⚓⚡🪶⚛️
Space Oddity🇺🇸⚓⚡🪶⚛️2 месяцев назад

Its not a good model for code. Crap tokens are still crap tokens.

Фото профиля Henry Mascot
Henry Mascot2 месяцев назад

What is it best at ? Coding or agent or chat ?

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

use it for all when tasks are light enough when isn't gaslighting me short of info - otherwise 27B

Фото профиля Voluntaryist Minister
Voluntaryist Minister2 месяцев назад

If you like that try Ornith, though I don’t know offhand if a MTP version is available.

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

originally it didn't. I know some folks grafted MTP.

Фото профиля Gg
Gg2 месяцев назад

Running local or through virtual environment?

Фото профиля feat7
feat72 месяцев назад

Context size?

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

iirc 262k as I switched later to Q5 to get about circa 180k

Фото профиля feat7
feat72 месяцев назад

Interesting, will try this out.

Фото профиля Stephen Reed 🤖
Stephen Reed 🤖2 месяцев назад

Runs batches of requests too - e.g. 24 concurrently, at much higher throughput decode toks/s.

Фото профиля Mat
Mat2 месяцев назад

That only leaves around 12GB for rag and context

Фото профиля LordOfTheYaml
LordOfTheYaml2 месяцев назад

Yeah.. But what's the point ? The quality is substandard.

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

if you know better, enlighten, or stay quiet forever.

Фото профиля LordOfTheYaml
LordOfTheYaml2 месяцев назад

Sorry, I didn't mean to be rude. The 5090 runs Qwen 27b, which offers better token quality and is more interesting in my opinion. Sacrificing quality for speed is not worth it to me.

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

That's cool. I'm just fatigued by the 27B cult. Different tools for different ends. My daily driver IS a 27B and I occasionally pair the MoE variant for fast loops - which is where this drafting is useful as I can loop super fast where it matters. I don't disagree 27B is overall better, but I have private repo forks for benchmarks and 35B-A3B does come out on top of 27B for my usecases - I have still some old Rust/Next.JS benchmarks shared on my profile as proofs of receipts. I can do more work per unit of time, with less stress on the GPU for equivalent results where I need.

Фото профиля Lyes Sbahi
Lyes Sbahi2 месяцев назад

Avec une 5090 tu ne peux pas rester sur le 35b-A3B… il faut passer sur le 27B

Фото профиля Michael Baumert
Michael Baumert2 месяцев назад

Do you train the weights with qlora? Or out of the box?

Фото профиля Robert
Robert2 месяцев назад

I like the strive to optimize performance for local AI for sports. But why would you want a 1000 tok/s slop machine? I dont see the use case, please enlighten me! Its not like we are gonna trust the knowledge unless it has 4 digit b parameters

Фото профиля wd 🇵🇹🔺
wd 🇵🇹🔺2 месяцев назад

I'm an iterative coder. Function over form here. Coding loops exposes agent to repeated patterns, since code repeats itself, it makes it easier for drafter to predict next tokens. This means during loops, the agent will have these bursts. Faster inference, fewer heat cycles, more work done.

Похожие видео

here's what i vibecoded today: punchingface 🥊 an app to make Hugging Face models fight each other on coding and canvas challenges, built with qwen3.6 35b a3b in 24 hours! benchmarks numbers don't mean anything anymore, we need a way to visualize what the models are actually capable of, and canvas are one the best way to showcase it imo. why? because one single error in the code and everything breaks it shows the differences in a matter of seconds, way easier than manually reviewing the quality of the code produced for a complex project; much needed in the space with all the new finetunes dropping everyday! i did a little demo here with qwen3.6 vs qwopus glm 18b merged, the frankenstein model from Kyle Hessling the winner is clear here, qwen is crazy good and has nothing to prove. that said, qwopus 18b isn't terrible at all; the result isn’t the prettiest to the eye, but hey… it works! i've seen so many models just output a blank page (completely non-working code) so this is already a win frankenstein talks and thinks but he needs some extra brain surgery 🧠 results were expected (it's a very experimental model) but love the effort in the 18b direction from jackrong and kyle! the app was entirely vibecoded with qwen3.6, i didn't edit a single file manually. i can say with confidence that it really has the intelligence of claude sonnet 4.5 at a speed of 125tok/s on an rtx 5080 which models should i make fight next?

left curve dev

19,229 просмотров • 5 месяцев назад