vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels...

Cruz's profile picture

Cruz

20,834 Aufrufe • vor 9 Tagen

I think the CMP 170HX just became one of...

Own Your Compute's profile picture

Own Your Compute

18,858 Aufrufe • vor 19 Tagen

If you have an RTX 3090 or 4090, Mia...

Yume_X's profile picture

Yume_X

38,511 Aufrufe • vor 9 Tagen

Qwen3.8-Flash-Next now reaches ~43 tok/s after a 122,902-token prompt...

Cruz's profile picture

Cruz

12,258 Aufrufe • vor 4 Tagen

One task, two models. GLM 5.3 Flash running slower...

Mia's profile picture

Mia

27,202 Aufrufe • vor 5 Tagen

100 tok/s on GLM-5.3-Flash, GLM-5.3 runs on the 6000s...

0xSero's profile picture

0xSero

68,904 Aufrufe • vor 12 Tagen

167 tok/s on a single RTX 4090. FreeToken just...

FHILY👑's profile picture

FHILY👑

33,058 Aufrufe • vor 18 Tagen

Day-0 Qwen3.8-27B vs Qwen3.6-27B: the voxel pagoda test. Same...

Wësche's profile picture

Wësche

97,988 Aufrufe • vor 29 Tagen

Gemma 4 12B QAT (dense) achieves 1000+ tokens/sec prefill...

Alok's profile picture

Alok

34,500 Aufrufe • vor 2 Monaten

Day 12/90 of Inference Engineering What is chunked prefill...

max fu's profile picture

max fu

29,556 Aufrufe • vor 1 Monat

Qwen3.8-Flash-Next is starting to feel like the local model...

FHILY👑's profile picture

FHILY👑

20,217 Aufrufe • vor 15 Tagen

🎉 Congrats to Thinking Machines on TML Inkling—a 1T-parameter...

vLLM's profile picture

vLLM

49,455 Aufrufe • vor 1 Monat

5 days ago it took 2 GPUs to build...

Sudo su's profile picture

Sudo su

34,624 Aufrufe • vor 6 Monaten

NVIDIA just dropped Nemotron-3-Nano:4b — a tiny 2.8GB model....

stevibe's profile picture

stevibe

127,570 Aufrufe • vor 5 Monaten

nvidia/Qwen3.6-35B-A3B-NVFP4 running in vLLM nightly on my Nvidia GB10...

Onur Solmaz's profile picture

Onur Solmaz

27,887 Aufrufe • vor 2 Monaten

This is awesome. Thanks Mia for this Was able...

Melvin Vivas's profile picture

Melvin Vivas

97,808 Aufrufe • vor 11 Tagen

Three Cline agents. Same 120B model. One rule: terminate...

Cline's profile picture

Cline

14,782 Aufrufe • vor 5 Monaten

SmallThinker 3B is one of the AI models I’d...

Aaron Ng's profile picture

Aaron Ng

35,858 Aufrufe • vor 1 Jahr

Laguna XS 2.1 just matched Qwen 3.6 35B on...

0xMarioNawfal's profile picture

0xMarioNawfal

63,839 Aufrufe • vor 1 Monat

Gemma 4 26B A4B MoE - 500+ t/s decode...

Alok's profile picture

Alok

17,465 Aufrufe • vor 1 Monat

now it's live🔥 GLM-5.2 on RTX 5090s: 80-110 tok/s...

Ning's profile picture

Ning

157,301 Aufrufe • vor 1 Monat