Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I put MCDMA through its paces today and tested Qwen 3.8 Flash Next on my dual Spark / Mac Studio cluster. My first test was disaggregated prefill across my two Sparks, then decode onto my Studio. PP 2,100 tok/s Decode 80 tok/s at 20k context, higher on short replies....

13,121 görüntüleme • 21 gün önce •via X (Twitter)

44 Yorum

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Unexpected side effect of splitting inference across the Mac Studio and the DGX Sparks over RDMA: it all runs cooler, by a big margin for the Sparks. Prefill on the Sparks alone pinned them at 80-85 °C. With decode moved to the Studio, 45 min into a task, the Sparks sit at 44 °C and the Studio at 50 °C.

Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹 profil fotoğrafı
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹21 gün önce

Have you measured this but just with Ethernet? I got pretty good results just with that

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

I did it with USBC, and it was slower, but not via Ethernet. I will try that, as I do need to see where the benefits are.

Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹 profil fotoğrafı
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹21 gün önce

Also I tried with Muse. Models have diff kv cache patterns.

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

What was the outcome?

Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹 profil fotoğrafı
Alvaro Videla - 🇺🇾🇨🇳🇨🇭🇮🇹21 gün önce

You can see in the repos I shared. Muser is the sever implementation for muse. 4x ttft, prefil etc. Also if you want secure kv cache you should check our kvpack project.

ecohash.co profil fotoğrafı
ecohash.co21 gün önce

Bro, we need to connect on TP=4 Sparks to 2x M3 Ultras. I want to see DSv4.1 Flash or GLM 5.3 move with pace in that config.

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Currently trying to get DSv4.1 to pipeline across MCDMA into the two sparks and studio. Two studios would be naughty 😈

ecohash.co profil fotoğrafı
ecohash.co21 gün önce

Question do I need two Mellanox cards/enclosures or can I run TB to each studio from one? I think I am going to need a bigger QSPFP Fabric either way, right now I only have a Microtik 504…

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Yeah you will need two enclosures and Mellanox cards one per Mac. You could hang more than one enclosure off a Mac though and that’s what I’m going to test later down the line.

pdp profil fotoğrafı
pdp21 gün önce

Damn.. good progress.

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Honestly did not expect it to be this good.

Faiz profil fotoğrafı
Faiz21 gün önce

cuda to metal sits between PP and decode, so ttft is one that tells you if the handoff paid for itself.

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Here is my first test, dude!

J. Gamboa profil fotoğrafı
J. Gamboa21 gün önce

@volatilemarkts nice one Ash, did you always expect to have to use the OWC and additional hardware or was this a decision that happened along the way? I ask because the set up is somewhat reminiscent of what Alex Ziskind showed once in one of his videos on disagg set ups.

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

@volatilemarkts I tried wirh usbc but couldn’t get the sparks to enter TB4 mode so this was the next logical step to make rdma possible.

J. Gamboa profil fotoğrafı
J. Gamboa21 gün önce

@volatilemarkts have been resisting for the longest time but maybe just have to accept this is what it takes with the hardware we are given. I’ll most def keep an eye on your updates. Super effort mate.

Tyler Folkman profil fotoğrafı
Tyler Folkman21 gün önce

How are you liking this dual machine setup?

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

I’m genuinely impressed. Only downside I can see is I would normally run qwen3.8 flash next on my studio and glm5.3 on the sparks simultaneously.

Tyler Folkman profil fotoğrafı
Tyler Folkman20 gün önce

Good to know. I've got a 256gb Mac on pre order and have been wondering whether to keep it or instead move towards 4 sparks.

Steve Darlow profil fotoğrafı
Steve Darlow21 gün önce

Casually solving the GPU vs RAM speed gap 🤯

Yuku2SPL profil fotoğrafı
Yuku2SPL21 gün önce

AFD community seems hard working now, maybe 5090 prefill for mac studio and run glm-5.3-flash

Sonypig profil fotoğrafı
Sonypig21 gün önce

the Studio alone prefills at about 1,200 tok/s——Can the M3 Ultra really achieve such high prefill throughput? Also, what if two M3 Ultra Mac Studios are connected via RDMA?

Ash Hart profil fotoğrafı
Ash Hart20 gün önce

Measured, yes. 6B active of 125B = ~14 TFLOP/s of matmul. Works because 3 of 4 layers are DeltaNet with fixed size state, no quadratic attention. Two Studios is wrong for prefill/decode; both are bandwidth-rich and compute-poor. Right for running a model that won't fit on one.

RefurbSitter profil fotoğrafı
RefurbSitter21 gün önce

Spark prefill into Studio decode is exactly why people keep unified memory on the desk.

Jonathan Spangler profil fotoğrafı
Jonathan Spangler21 gün önce

Can you pin/post concurrency runs for prefill speeds . That’s the only main painpoint I foresee on M5U , testing next week though .

kkiran profil fotoğrafı
kkiran21 gün önce

This is cool! ConnectX to TB5 cable would be great for this - hope we find a supplier soon

hehe profil fotoğrafı
hehe21 gün önce

disaggregated prefill on sparks then decode on studio used to be a paper abstract. now it is a tuesday.

Super Orc Trader profil fotoğrafı
Super Orc Trader20 gün önce

I am so so so excited man!! So impressed with momentum and progress. How can I try this with my Macs and Sparks ASAP?

Ash Hart profil fotoğrafı
Ash Hart20 gün önce

Point your agent at the repo dude!

Super Orc Trader profil fotoğrafı
Super Orc Trader20 gün önce

Just did! I need to get the hardware first right? My agent tells me this: First get Ash’s runnable inference code and confirm whether it can also use ordinary networking.

Ash Hart profil fotoğrafı
Ash Hart20 gün önce

RDMA uses verbs so you need to ask your agent to set up your inference engine. I’m going to do a full pr on oMLX for this part.

Dejan Marjanović profil fotoğrafı
Dejan Marjanović21 gün önce

M5 Ultra will erase the prefill advantage of Spark, what's a good combo there?

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Maybe but what about all the people who have sparks already and M3U studios, etc? What happens when Nvidia release a new Spark that slaps the M5U's prefill again?

Dejan Marjanović profil fotoğrafı
Dejan Marjanović21 gün önce

I'm the people 😂 Just that I have 512GB so it's a bit of a mismatch, but 256GB or even 96GB is a bingo.

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Do you have any Sparks as well? I have 2 Sparks connected via CX7, feeding my 256GB Studio.

Dejan Marjanović profil fotoğrafı
Dejan Marjanović21 gün önce

I have 5, but 4 in a cluster, so it's possible to combine? 4x will match or exceed M5U 512GB according to my calculations. I would probably buy now 2x 256GB M5U (more bandwidth/concurrency?)

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Yeah, we can combine that. Do you have a MikroTik switch handling all the comms for the Sparks? You will need a TB5 enclosure and a Mellanox card but it should work.

Dejan Marjanović profil fotoğrafı
Dejan Marjanović21 gün önce

Yes, CRS804. Can you tell which models exactly? (I'm in Europe not sure if there's a lot of choices)

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

I’m in the UK. I ordered these; Mellanox ConnectX®-5 Ex EN NIC, 100GbE dual-port QSFP28, PCIe MCX516A-CDAT Mellanox Passive Copper Cable 100GbE QSFP28 to QSFP28 1M MCP1600-C001E30N x2 OWC Mercury Helios 5S Thunderbolt 5 (80Gb/s) Single Slot PCIe Card Expansion Solution

pidgraph profil fotoğrafı
pidgraph21 gün önce

Would it work for Macbook M5 Max and Spark?

Ash Hart profil fotoğrafı
Ash Hart21 gün önce

Yes.

Derek profil fotoğrafı
Derek21 gün önce

@volatilemarkts Amazing!

draslan.eth profil fotoğrafı
draslan.eth20 gün önce

Prefill on mac mini m4 Pro is a pain... :D

Benzer Videolar

bonsai 2 27b on an rtx 3060 12gb, the full receipt sheet. save this one, the 12gb row of the small gpu guide is built from it. speed by depth, then what context costs, live server, thinking on, real sessions > 7k deep: 24.4 tok/s > 12k deep: 21.9 tok/s > 35k deep: 17.8 tok/s > 77k deep: 13.0 tok/s > 64k window: 7.3gb resident > 128k window: 8.8gb resident > 192k window: 10.2gb resident > 262k window: 11.7gb resident, 0.6gb to spare, the whole native window on a 12gb card > every 64k of context costs 1.47gb, so 327k would not fit > a 41,312 token build session from 35k to 77k of context averaged 15.0 tok/s across 46 minutes > prefill 295 tok/s at 2k of context, 243 tok/s at 35k, first token in 0.6 seconds, fresh decode 26.1 tok/s > the card pinned 149 of 150 w the entire time, 78c, fan at 80%, power bound, not heat bound > 0.158 tok/s per watt at the fresh end the setup > model: ternary bonsai 2 27b, PTQ1_0, 1.75 bits per weight, 5.95gb on disk, base qwen 3.8 27b, apache 2.0 > runtime: prismml llama.cpp fork, prebuilt cuda 12.4 binary, no compile > serve: full 262k native window resident, q4 kv cache, flash attention, one slot, 11.7 of 12gb in use for anyone who followed bonsai 1 in july, that was the 3.9gb 1bit file at 42 tok/s on a 3060 ti, a faster card and a smaller file, so the same card comparison is not on the table yet, it comes with the 8gb test. what changed is the base, qwen 3.8 instead of 3.6, and the retention, 98.2% on their suite instead of 95%, and the whole 262k window fitting on 12gb.

Sudo su

31,174 görüntüleme • 18 gün önce

Google's Gemma 4 26B A4B QAT hits 25+ tokens/sec and 320+ tokens/sec prefill on 8 GB VRAM (RTX 4060) + 16 GB RAM using TurboQuant Prefill just went from 200 → 320+ tok/s on the same 8GB card. 1.6x, no new hardware, no new quant, just a KV cache trick stacked on top of the Gemma 4 26B MoE setup from a few days ago. A few days ago I posted Gemma 4 26B A4B hitting 28 tok/s decode on 8GB VRAM using native MTP. prefill was stuck around 200 tok/s. fair callout by the community. So today I tested something I'd already been meaning to try: TheTom/llama-cpp-turboquant, the TurboQuant KV cache fork by Tom Turney (Tom Turney). (github link in the comments) thanks to him, the fork just got resynced to mainline, so MTP + TurboQuant now run together cleanly (I didnt see any meaningful gains by using MTP with this setup though but you can try). The flags (No MTP): -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -cnv -c 64000 --cache-type-k q8_0 --cache-type-v turbo3 Results on the same RTX 4060 8GB, tested with a 27k token prompt at 64k context loaded: Prefill: 200 tok/s → 320+ tok/s Decode: stayed above 25 tok/s (without MTP) Why it works: TurboQuant uses walsh hadamard rotation + polar quantization on the KV cache. keys are sensitive to compression, values aren't much, so it splits the difference: K stays at q8_0, V drops to turbo3 (~3 bits). bonus from the memory savings: same 8GB card can now stretch to 100-120k context with minimal decode penalty. It should now be snappier with any agent harness such as hermes agent without compromise on intelligence. If you're already running Gemma 4 on a small card, this stacks on top for free. Try --cache-type-k q8_0 --cache-type-v turbo3 on your setup and report back what your prefill/decode split looks like. unsloth model gguf and llama.cpp turboquant fork links in the comments. what's your prefill number before vs after?

Alok

119,821 görüntüleme • 3 ay önce

The VRAM barrier is officially dead. I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer hardware. Tested on Ubuntu 22 | CUDA 13.0 | PCIe 4.0 x16 | 110 GB DDR4 System RAM with a continuous 28k prompt across all runs. ### The Benchmarks & Scaling # 1. Hybrid Offload (-ncmoe 40 @ 80k Context) Offloaded 40 expert layers to the GPU, pushing VRAM to the ceiling. ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 80000 --port 8080 -v --fit off -b 4096 -ub 4096 -ncmoe 40 Prefill: 383.85 t/s | Decode: 22.52 t/s Footprint: 23.85 GB VRAM | 97 GB RAM # 2. Full CPU MoE Offload (-cmoe @ 80k Context) Pinned all 512 expert layers to DDR4 RAM (-cmoe), keeping attention on the 4090. llama.cpp flags: (Same as above, replace -ncmoe 40 with -cmoe) Prefill: 355.72 t/s | Decode: 20.84 t/s Footprint: 11.66 GB VRAM (12GB+ VRAM freed up!) | 110 GB RAM # 3. The 180,000 Context Run Prefill: 357.75 t/s | Decode: 20.98 t/s | VRAM: 15.6 GB | RAM: 110 GB # 4. The 250,000 Context Absolute Ceiling ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 250000 --port 8080 -v --fit off -b 4096 -ub 4096 -cmoe Prefill: 364.29 t/s | Decode: 20.97 t/s Footprint: 18.3 GB VRAM (Still ~5.7 GB of VRAM headroom!) | 110 GB RAM ### Key Insights: -b 4096 -ub 4096: doubles the prompt ingestion from ~150 to 364+ t/s. -cmoe Free Lunch: Shifting expert layers to DDR4 RAM slashes VRAM from 24GB to 11.6GB with virtually zero decode penalty (22.5 -> 20.9 t/s), enabling the 250k context ceiling. Qwen 3.8 Flash-Next (UD-Q4_K_XL) is a massive 111.4 GB model split across 4 shards. To run this architecture, you must build from the experimental PR branch (#27742) by Daniel Han: git clone && cd llama.cpp git fetch origin pull/27742/head:qwen-next && git checkout qwen-next cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DBUILD_SHARED_LIBS=OFF cmake --build build --config Release -j $(nproc) --target llama-server A single 4090 paired with 100 GB of cheap DDR4 RAM will comfortably serve production grade 125B inference. While Qwen 3.8 27B (dense) still holds the crown for single 3090/4090 rigs, Flash Next proves 125B hybrid models are officially viable on consumer hardware. Hugging Face GGUF link and complete performance telemetry graphs are dropped in the replies below. GLM 5.3 Flash VS Qwen 3.8 Flash Next, which one takes the open weights crown this week?

Alok

1,031,691 görüntüleme • 1 ay önce

"which quant should I download?" is a question you may never have to answer again the team Hamster Labs has figured out how to kill it with pMLX. download once at full precision (bf16) and the engine re-fits it to your machine on the fly, based on the job you give it tell it two things: how much context you need, and the slowest speed you'll accept. it reads your mac and picks the quantization plus how many experts stay in ram vs. stream from disk. no need to download a smaller quant or figure out which quant fits your hardware. same model for every use case and the config changes based on what you need I ran this on my own M4 Max and Qwen3.8-Flash-Next splits into two configs. short context, under 32k: full bf16, ~85% of experts in ram, fast at full precision since it fits my ram long context, 64k to 256k: keep ~70% bf16 and stream the rest from SSD at 22 tok/s, or drop to q8 and get 40 tok/s. I pick per task and the model itself never changes. there are many possibilities since I have the ram to spar - if i need speed, Q4 100% resident (74gb ram @ 62 tok/s) - if i need balance, Q8 95% resident (76gb @ 38 tok/s) - if i need accuracy, bf16 70% resident (96gb @ 16 tok/s) given whatever RAM you have (16/32/64/128/256 GB) + your context + your min speed, the engine picks the precision (bf16→q8→q4) and the expert-residency/paging split that fits your box and maximizes quality & speed last thing to optimize is speed. there are so many things we want to power with open models at Hamster and these 180b-300b class models have the potential to play a big role in that

Eyal Toledano

10,590 görüntüleme • 1 ay önce

Deepseek V4 Flash 0731 (Q2) - 12 tokens/sec - Single RTX 4090 - 650+ tokens/sec prefill - 250k context - no kv cache quantization! DeepSeek just dropped the official V4 Flash 0731 two days ago with a massive agent capabilities upgrade. The official benchmarks are literally crushing their own V4-Pro-Preview on agentic tasks like Terminal Bench 2.1 and DeepSWE. Unsloth AI said they couldn't wait to bring it to local devices, and they delivered. If you thought my 118B Poolside Laguna S 2.1 MoE run last week on a single GPU was wild, hold onto your hardware. I just successfully ran Unsloth’s brand new 91GB DeepSeek-V4-Flash-0731 (UD-IQ2_M) GGUF entirely locally. And I pushed it to a mind-bending 250,000 context window. The VRAM ceiling is an illusion if you know how to optimize llama.cpp. Here are the benchmarks and the cheat codes to run a local frontier class model yourself. For the hardware and setup, I used a single NVIDIA RTX 4090 (24GB VRAM) hooked up via a PCIe 4 bus, running Ubuntu 22.04 LTS and CUDA 13.0. You don't need a massive enterprise server for this, if you have more than 80 GB of standard DDR4 RAM and a 24GB card like an RTX 3090 or 4090, you can run this exact stack yourself. All benchmarks were run using a massive 28k token prompt to truly stress test the prefill limits. no kv cache quantization THE BENCHMARKS (Scaling Context): # 80k Context (Baseline: -b 2048 -ub 2048): Prefill: 465.43 t/s | Decode: 13.00 t/s | VRAM: 22.87 GB # 80k Context (Optimized: -b 4096 -ub 4096): Prefill: 643.15 t/s | Decode: 12.20 t/s | VRAM: 23.00 GB (Notice how doubling the batch flags spiked my prefill throughput by nearly 200 t/s with almost zero VRAM penalty) # 180k Context (-b 4096 -ub 4096): Prefill: 629.18 t/s | Decode: 11.92 t/s | VRAM: 23.40 GB # 250k Context MAXIMUM (-b 4096 -ub 4096): Prefill: 619.02 t/s | Decode: 11.54 t/s | VRAM: 23.40 GB # THE SECRET SAUCE (Why this works): Unsloth’s UD-IQ2_M quant is ~91GB across 3 files. Since I only have 24GB of VRAM, the PCIe 4 bus and system RAM have to do the heavy lifting. The magic bullet is the --no-mmap flag. By completely bypassing OS disk paging, I forced llama.cpp to load the massive model weights directly into the system RAM upfront. Combined with Flash Attention (-fa on) and exactly 12 CPU threads (--threads 12), I maintained an incredibly stable 11.5+ tokens/sec decode speed even at a quarter million token context. # THE EXACT COMMAND: ./build/bin/llama-server -m /workspace/models/DeepSeek-V4-Flash-0731-UD-IQ2_M-00001-of-00003.gguf -c 250000 -fa on --port 8080 --threads 12 -b 4096 -ub 4096 --no-mmap -v Local conversational and agentic coding AI is fully here. You don’t need an API or an H100 cluster. Qwen 3.8 27b drops next week making the 24GB VRAM tier even more worthwhile. What does your current local AI rig look like, and what's the craziest model you've managed to squeeze into it? Official huggingface GGUF links from Unsloth and performance graphs are dropped in the replies below!

Alok

46,100 görüntüleme • 2 ay önce

I just crammed the updated Gemma 4 26B A4B QAT (MoE) with 180k context into an 8GB RTX 4060 (8 GB VRAM + 16 GB RAM only!!) and optimized the batch size. 23 tokens/sec decode, 300 tokens/sec prefill Yesterday I showed you a Gemma 4 31B dense model running flawlessly on an RTX 4090. Today, we're breaking the VRAM bank on a budget card using Unsloth’s new Gemma 4 26B (A4B) QAT quants. Following Google’s chat template update that boosted agentic benchmarks by +10%, I pushed this model to its absolute limits. Here is how you squeeze 250k context out of 8GB of VRAM. # The Setup & The Optimization - Hardware: Nvidia RTX 4060 (8GB VRAM) + 16GB System RAM - Environment: CUDA 13.0 build of llama.cpp - Model: gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf - Prompt: 28,000 tokens of prompt for each run If you read my L2 cache breakdown (attached in replies), you know the 4060’s 24MB cache maxes out at `-b 1024 -ub 1024`. Push past that, and prefill crashes. I locked those flags in for every test below to ensure maximum GEMM throughput. # 1. The Raw Context Push (Unquantized KV Cache) First, I wanted to see how far pure 8GB VRAM + 16GB RAM could stretch without touching the KV cache: - 80k Context: Prefill 385 t/s | Decode 25.5 t/s - 120k Context: Prefill 270 t/s | Decode 24 t/s llama.cpp flags: .\llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 120000 --port 8080 -ub 1024 -b 1024 Without KV quantization, 120k is your hard ceiling. push past that prefill throughput drops off a cliff, making the model practically unusable for large agentic workloads. # 2. The Q8 KV Cache Lifeline To survive 250k context on a budget card, you have to quantize the KV cache. I enabled 8 bit KV cache (`-ctk q8_0 -ctv q8_0`) and re ran: - 180k Context: Prefill 280 t/s | Decode 22.8 t/s - 250k Context: Prefill 115 t/s | Decode 20 t/s llama.cpp flags: .\llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 180000 --port 8080 -b 1024 -ub 1024 -ctk q8_0 -ctv q8_0 Result: Q8 KV cache brings 250k context back from the dead. Decode speed stabilizes at a highly usable 20 t/s. You are trading a very small bit amount of reasoning precision for an extra 130,000 tokens of context window. if you own a single rtx 3050, 3060, 3070, 4050, 4060, 5050 or 5060, you must try this model and optimize your batch size for higher prefill. Hugging Face links to the updated Unsloth's QAT quants and performance graph are in the replies below. What model are you running on your 6GB, 8GB or 12GB cards right now? Let's see your setups.

Alok

36,617 görüntüleme • 2 ay önce

Qwen3.8 Flash Next is starting to look ridiculous on Apple Silicon. I’m running the 4-bit MTP build locally on an M3 Ultra Studio, and the latest OMP run hit: 97.1 tok/s decode 23K context ~1,131 tok/s uncached prompt processing That first number is the one that caught my attention. Nearly 100 tokens per second from a local Qwen3.8 Flash Next setup is already fast enough that the usual “local models are slow” argument starts feeling pretty outdated. And the prompt processing speed is even crazier. Over 1,100 tok/s on an uncached prompt means the model can chew through a large amount of context before generation even starts. The setup matters here. This isn’t just Qwen3.8 Flash Next running untouched. It’s a 4-bit quantized build with MTP, and the inference stack is clearly doing a lot of work behind the scenes to make the hardware perform like this. There’s already a PR open for the implementation on oMLX, so this isn’t just a one-off local experiment either. If these optimizations make their way into the broader MLX ecosystem, running large models locally on Apple Silicon gets even more interesting. The other thing I like about numbers like this is that they put the focus back on the entire inference stack. Model size is one variable. Quantization is another. Then you have MTP, KV cache configuration, runtime optimizations and the hardware itself. Change the recipe and the same model can feel completely different. Qwen3.8 Flash Next at ~97 tok/s on an M3 Ultra is a pretty good demonstration of that.

FHILY👑

15,378 görüntüleme • 23 gün önce

Run Updated Gemma 4 26B A4B QAT (MoE) with Vision at 25 tokens/sec and massive 120k context window on a single RTX 4060 (8 GB VRAM + 16 GB RAM Only!!) Yesterday I pushed Gemma 4 26B A4B QAT to 250k context on a single RTX 4060 using nothing but Q8 KV cache and optimized -b and -ub flags for higher prefill throughput. Today I stacked Multi Token Prediction (MTP) self speculative decoding AND the vision projector (mmproj) on top of that same card, same batch size optimization, same $250 GPU and pushed it until it broke, then found the fix. All text only runs consist of a 28k prompt. vision runs consist of 28k text prompt + an image. # 1. MTP alone. near free decode speed, no catch MTP draft assistant is a separate small model (MTP heads are backed into the main model itself for the qwen 3.5+ models but its a separate small model for gemma 4 series), 240 MB gguf 80k ctx: Prefill 510 t/s | Decode 29.5 t/s 120k ctx: Prefill 433 t/s | Decode 29 t/s 180k ctx: Prefill 240 t/s | Decode 24.9 t/s 250k ctx: Prefill 63 t/s | Decode 13 t/s llama.cpp flags: m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp -md mtp-gemma-4-26B-A4B-it.gguf-c 180000 -b 1024 -ub 1024 --spec-draft-n-max 6 --spec-draft-p-min 0.7 -ctk q8_0 -ctv q8_0 # 2. Add vision on top. the tax you actually pay the vision projector gguf is about 1.1 GBs 80k ctx: Prefill 360 t/s | Decode 25.4 t/s 120k ctx: Prefill 230 t/s | Decode 23.8 t/s 180k ctx (Q8 KV): Prefill 75 t/s | Decode 12.5 t/s - cliff flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp -md mtp-gemma-4-26B-A4B-it.gguf -c 80000 --port 8080 -b 1024 -ub 1024 --spec-draft-n-max 6 --spec-draft-p-min 0.7 -ctk q8_0 -ctv q8_0 --mmproj mmproj-F16.gguf # 3. The fix if you want to run vision over 120k context: swap Q8 KV for Q4 KV past 120k Stack MTP + vision + Q8 KV past 120k context and you hit a wall. draft model overhead plus KV pressure tanks everything. Drop to Q4 KV and the wall disappears: 180k ctx (Q4 KV): Prefill 220 t/s | Decode 25.5 t/s -ctk q4_0 -ctv q4_0 --mmproj mmproj-F16.gguf (rest same as above) Bottom line: MTP gives you a near free +20-30% decode boost up to 120k context. Past that, it's fighting your VRAM, not helping and if vision is loaded too, Q4 KV isn't optional past 120k, it's mandatory. 30% boost is model and card specific, MTP boosted decode 2x for gemma 4 31b on a single rtx 4090. Same 8GB card. Same $250 GPU. Multimodal, speculative decoding, 180k usable context, zero upgrades. You gotta try this if you have a single NVIDIA RTX 3050, 3060, 3070, 4050, 4060, 5050 or 5060. You can try it with a 6 GB VRAM card as well but you will have to lower the context window. Hugging Face links to the updated Unsloth's QAT quants and performance graph are in the replies below. Which models are you running on your 6/8/12GB cards with MTP?

Alok

16,434 görüntüleme • 2 ay önce

dflash-mlx v0.1.7 is out. Big adaptive-runtime update, still focused mostly on Qwen3.6 27B 4-bit. @ 2048 tokens, M5 Max, stock mlx_lm baseline: ► 1024: 33.26 → 98.05 tok/s (x2.95) ► 2048: 32.34 → 90.67 tok/s (x2.81) ► 4096: 30.58 → 93.55 tok/s (x3.06) ► 8192: 26.03 → 79.12 tok/s (x3.04) ► 16384: 21.50 → 60.77 tok/s (x2.78) Main change: adaptive verify got a lot smarter. Instead of blindly trying to verify large 16-token blocks all the time, DFlash now watches acceptance + tokens/cycle + real cycle cost. When the draft gets weaker, it drops to smaller 4-token blocks, then probes back up only when the recent cycles make sense. In practice: less wasted verify work, better long-context behavior, and much more useful metrics to understand what is happening. ► retuned adaptive verify for long-context / agentic decode ► richer metrics: tokens/cycle, adaptive block state, CopySpec counters ► /metrics now has real decode avg + logical/real/restored prefill rates ► AIME25 benchmark suite with exact integer scoring ► Qwen thinking default now follows tokenizer/request behavior ► GDN recurrent exactness fixes I also started running AIME25-style long generations. Even around 45k generated tokens, I was still seeing ~40 tok/s on 27B 4-bit. Over the next few days I’ll share more demos: AIME runs, real OpenCode game/project sessions, and full metrics along the way. Still optimizing hard for 27B 4-bit first, while working on custom kernels per Apple GPU generation so more machines can benefit.

bstn 👁️

16,464 görüntüleme • 4 ay önce