Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Two models. One DGX Spark One identical voxel Eiffel Tower prompt. Qwen3.8-27B xhigh: 83K tokens, 41m29s, 33.4 tok/s Ornith 1.5 35B-A3B: 47K tokens, 10m35s, 74.0 tok/s Thinking on, no token ceiling, matched camera. Ornith finished much faster; Qwen spent far more tokens.

51,311 Aufrufe • vor 21 Tagen •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt on ONE DGX Spark. ⚡🚀 MTP k=2 won my draft-depth sweep, with +42.5% mean decode over no draft. The PLE table stays fully on-device. I promised the deeper MTP tests. Here are the results, and now you can explore them in an interactive benchmark page too. 𝗧𝗪𝗢 𝗗𝗥𝗔𝗙𝗧 𝗧𝗢𝗞𝗘𝗡𝗦 𝗪𝗢𝗡 Mean single-request decode with 32K context configured: MTP k=2: 39.21 tok/s MTP k=3: 36.42 tok/s MTP k=1: 35.18 tok/s No draft: 27.51 tok/s k=2 also produced the fastest individual sweep run: 41.34 tok/s. Four runs each for no draft, k=1 and k=2. Seven for k=3. Decode excludes time to first token. Here, k means speculative draft depth, not quantization bits. k=3 produced more tokens per step, but the extra drafting work did not pay off in throughput. k=2 is my current pick for this setup. 𝗧𝗛𝗘 𝟭𝟮𝟯𝗞-𝗧𝗢𝗞𝗘𝗡 𝗣𝗥𝗢𝗠𝗣𝗧 𝗧𝗘𝗦𝗧 I then ran a separate long-prompt comparison: Actual input: 122,902 tokens Configured context: 262,144 Requested output: 128 tokens One request at a time MTP k=2: ~43 tok/s No draft: 26.2 tok/s Time to first token: 110.6 seconds with MTP 107.0 seconds without it The win here is generation speed, not faster prefill. To keep the scope clear: 256K was the configured limit. This was a real ~123K input, not a completely filled 256K window or a full k sweep at that depth. 𝗣𝗟𝗘 𝗦𝗧𝗔𝗬𝗦 𝗢𝗡 𝗧𝗛𝗘 𝗦𝗣𝗔𝗥𝗞 Whole model on-device: 78.57 GiB Packed 5-bit PLE table: 30.4 GiB, included in that total No NVMe PLE offload in this build. This is still turboderp’s 3.05bpw_h5_ng5 EXL3 pack, served through my vllm-exl3 integration. My work here is the serving integration and testing. These are preliminary performance measurements, not a quality evaluation or a claim of bit-exact full-output parity. 𝗘𝗫𝗣𝗟𝗢𝗥𝗘 𝗧𝗛𝗘 𝗥𝗘𝗦𝗨𝗟𝗧𝗦 The benchmark page has the individual sweep values, long-prompt comparison, and measurement scope. You can play the animation, export the charts, or download the HTML and data to render them yourself. No Spark needed to view the results. Credit to turboderp / ExLlamaV3 for the pack and kernels, vLLM for the serving engine, and Qwen Qwen Developers for the model. Recipe + reproduction: Interactive benchmark:

Cruz

12,175 Aufrufe • vor 1 Tag