
Yume_X
@yume_arasaki • 1,333 subscribers
夢 . Chasing free AI. Mapping the AI era. Your hardware. Your Intelligence. Local Models + Hermes Fleet RTX 4090 | Mac Mini M4 64GB | 2x DGX Spark online
Shorts
They told me DGX Sparks are useless. Doesn't seem so when I'm running GLM 5.3 Flash, a Opus-level AI on my desk. A $100 to $300 a month subscription, now 24/7 at home, completely yours. No rate limits. I ran the numbers. Here's what to know, and here are the caveats. The model is GLM-5.3-Flash EXL3 on two Sparks, DFlash2 speculative decode on. Same rig that ran Qwen 3.8 Flash-Next last week. My one line review "GLM 5.3 Flash can see things, other models simply cannot" Qwen 3.8 Flash one shotted my most one of complex agent workflows, without human intervention What is EXL3 and DFLASH 2? What EXL3 is: your GPU reads weights in chunks. FP16 is 2 bytes each. NVFP4 is half a byte. EXL3 packs them to about 4 bits and stores the rounding error in a second, smaller matrix. The engine decompresses on the fly, layer by layer. That's how a 164 GB model fits in 2×128 GB unified memory. What DFlash2 is: a small draft model guesses the next 7 tokens. The big model checks them in one pass. Agree, and you got 8 tokens for the price of 1. Disagree, and you paid for nothing. Speed is not "is draft on." Speed is accept rate. HTTP benches, thinking off. Throughput is the sum of each stream's decode rate, not tokens divided by the slowest wall clock. - Single stream prose: 25.6 tok/s - Two streams: 38.4 combined - Four streams: 73.7 combined - Qwen 3.8 Flash-Next, same protocol: 45.5 / 69.7 / 104.5. Qwen wins speed, no contest - Structured count 1→200: 57.0. Mia's README claims 62.9. We did not beat it - Needle found at 4.7K, 18.8K, 76K, 152K, 305K. KV pool held 972K tokens in FP8 Caveats, from hours of Oh My Pi on the same stack: Same model, same two Sparks, same DFlash2. 58 tok/s on one prompt, 16 on the next. The screenshot is not the model. The accept rate is. - Prose / hard tools: 12–20 tok/s at 10–20% draft accept - Mixed: 20–30 at 25–40% - Predictable stretches: 42–58 at 64–95% - Peak: 58.2 tok/s at 95% accept, mean length 7.66 out of 8 When accept is 95%, DFlash2 is a 3x over the miss lane. When it's under 20%, the draft is a tax. I would rather run without it than sit there. Concurrency is weaker in this setup. Recommend NVFP4 if speed and concurrency is a concern. Slower than Qwen 3.8 Flash, worth it. Less warm, less human-sounding, more raw intelligence. If you're using your local model to discover things you didn't know about your own work, this is the model. Genuine frontier. Full numbers in the video. Which part of the accept curve are you actually living on? Used recipe and attribution in reply 👇
12,092 görüntüleme