Video yükleniyor...
Video Yüklenemedi
vLLM fast inference running on monte carlo synthetic data generation with: > peak generation throughput of ~ 23k token/s & avg of 20k token/s > ~200 reqs/s > Qwen/Qwen2.5-0.5B-Instruct > on 1x Grace Hopper 200 - 480 GB(~96GB HBM3) vllm config: --max-num-seqs 512 --chunked-prefill-enabled (for better throughput) --dtype float16:... show more
26,130 görüntüleme • 7 ay önce •via X (Twitter)
0 Yorum
Yorum bulunmuyor
Orijinal gönderinin yorumları burada görünecek
