Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐—™๐—ข๐—ฅ๐—ง ๐—ช๐—ข๐—ฅ๐—ง๐—› ๐—ข๐—ฅ ๐—ก๐—ข๐—ช๐—›๐—˜๐—ฅ๐—˜.

11,461 Aufrufe โ€ข vor 7 Monaten โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels for 2-bit EXL3 on NVIDIA DGX Spark GB10. GLM-5.3-Flash-EXL3-K2 jumped from 16.9 โ†’ 24.6 tok/s average single-stream decode, a +45.6% gain. Coding hit 27.6 tok/s, +85.6%. ๐Ÿš€ The previous ExLlamaV3-backed path inside vLLM was leaving a lot of GB10 bandwidth on the table. So I rewrote the hot path specifically for EXL3 on Blackwell sm_121: โ†’ in-register Trellis dequantization โ†’ native fused MoE decode โ†’ power-of-two chunked prefill GEMM โ†’ parallel NVMe pre-warm Then I tested it side-by-side on physical DGX Spark hardware using my GLM-5.3-Flash-EXL3-K2 pack and live vLLM HTTP streaming. ๐Ÿš€ ๐——๐—˜๐—–๐—ข๐——๐—˜ ๐—ง๐—›๐—ฅ๐—ข๐—จ๐—š๐—›๐—ฃ๐—จ๐—ง Single-stream C1: Coding 14.9 โ†’ 27.6 tok/s +85.6% Prose 13.7 โ†’ 24.6 tok/s +79.3% Reasoning 18.9 โ†’ 25.1 tok/s +32.7% Summary 17.1 โ†’ 25.6 tok/s +50.0% Format 16.3 โ†’ 24.0 tok/s +47.7% Average: 16.9 โ†’ 24.6 tok/s ๐—ก๐—˜๐—ง ๐—š๐—”๐—œ๐—ก: +45.6% โฑ๏ธ ๐—™๐—œ๐—ฅ๐—ฆ๐—ง-๐—ง๐—ข๐—ž๐—˜๐—ก ๐—ฅ๐—˜๐—ฆ๐—ฃ๐—ข๐—ก๐—ฆ๐—œ๐—ฉ๐—˜๐—ก๐—˜๐—ฆ๐—ฆ Coding TTFT: 2,344 ms โ†’ 859 ms That is a 63.3% reduction, or about 2.7ร— faster to first token. Follow-up turn with prefix cache hit: 5,608 ms โ†’ 3,588 ms 1.56ร— faster. โšก ๐—ช๐—›๐—”๐—ง ๐—–๐—›๐—”๐—ก๐—š๐—˜๐—— ๐—ข๐—ก ๐—ง๐—›๐—˜ ๐—š๐—ฃ๐—จ 40 routed-MoE layers: 19.9 ms โ†’ 11.5 ms per token Per-layer MoE compute: 497 ฮผs โ†’ 287.8 ฮผs That removes 8.4 ms of MoE compute from every generated token. Total per-step wall time: 59.2 ms โ†’ 40.6 ms -31.4% The key is `p2b_fused_moe`. Instead of expanding EXL3 weights through a traditional intermediate path, the new kernel performs Trellis dequantization in-register while executing the routed expert computation. The weights stay compressed until the GPU actually needs them. ๐Ÿ”ฅ ๐—ฃ๐—ฅ๐—˜๐—™๐—œ๐—Ÿ๐—Ÿ ๐—š๐—ข๐—ง ๐—” ๐—ก๐—”๐—ง๐—œ๐—ฉ๐—˜ ๐—ฃ๐—”๐—ง๐—› ๐—ง๐—ข๐—ข The new `exl3_gemm` uses power-of-two chunked prefill GEMM. Measured: 7.85 TFLOPS 13.0ร— faster than the legacy prefill kernel 1,875 tok/s cold prefill sustained across 65K context ๐Ÿ’พ ๐—ง๐—›๐—˜ ๐—•๐—ข๐—ข๐—ง ๐—ฃ๐—”๐—ง๐—› ๐—ก๐—˜๐—˜๐——๐—˜๐—— ๐—ช๐—ข๐—ฅ๐—ž ๐—ง๐—ข๐—ข Loading a ~91 GiB model is part of the user experience. Standard shard loading is mostly serial. The updated recipe parallelizes NVMe pre-warm across 8 workers so the storage controller gets used properly instead of feeding a ~100 GiB model one shard at a time. That turns boot-time storage into another optimization target instead of something we simply accept. ๐Ÿ’ก ๐—ง๐—ช๐—ข ๐—ฆ๐—˜๐—ฅ๐—ฉ๐—œ๐—ก๐—š ๐—™๐—Ÿ๐—”๐—š๐—ฆ ๐—ช๐—ข๐—ฅ๐—ง๐—› ๐—ž๐—ก๐—ข๐—ช๐—œ๐—ก๐—š `--long-prefill-token-threshold 1024` Prevents giant prefill chunks from monopolizing step budgets and starving parallel decode sessions. `--enable-prefix-caching` Avoids paying for the same conversational prefix again on follow-up turns. ๐Ÿ“ฆ ๐—˜๐—ฉ๐—˜๐—ฅ๐—ฌ๐—ง๐—›๐—œ๐—ก๐—š ๐—œ๐—ฆ ๐—ข๐—ฃ๐—˜๐—ก vllm-exl3: GLM-5.3-Flash one-Spark recipe: Model: This is why I like working at the kernel level. The model did not change. The quant did not change. The hardware did not change. The execution path did. 16.9 โ†’ 24.6 tok/s. ๐Ÿ› ๏ธ vLLM turboderp

Cruz

21,173 Aufrufe โ€ข vor 23 Tagen

โ€œCum inside you๐Ÿ˜‚โ€
0:14

Sensitive content

โ€œCum inside you๐Ÿ˜‚โ€

ANIME X

205,673 Aufrufe โ€ข vor 1 Tag