vllm-exl3 v0.3.0 is LIVE with custom native CUDA kernels...

Cruz's profile picture

Cruz

20,016 просмотров • 4 дней назад

I think the CMP 170HX just became one of...

Own Your Compute's profile picture

Own Your Compute

18,740 просмотров • 15 дней назад

If you have an RTX 3090 or 4090, Mia...

Yume_X's profile picture

Yume_X

37,007 просмотров • 4 дней назад

One task, two models. GLM 5.3 Flash running slower...

Mia's profile picture

Mia

23,986 просмотров • 1 день назад

100 tok/s on GLM-5.3-Flash, GLM-5.3 runs on the 6000s...

0xSero's profile picture

0xSero

68,500 просмотров • 8 дней назад

167 tok/s on a single RTX 4090. FreeToken just...

FHILY👑's profile picture

FHILY👑

33,058 просмотров • 13 дней назад

Day-0 Qwen3.8-27B vs Qwen3.6-27B: the voxel pagoda test. Same...

Wësche's profile picture

Wësche

97,988 просмотров • 24 дней назад

Gemma 4 12B QAT (dense) achieves 1000+ tokens/sec prefill...

Alok's profile picture

Alok

34,500 просмотров • 2 месяцев назад

Day 12/90 of Inference Engineering What is chunked prefill...

max fu's profile picture

max fu

29,449 просмотров • 1 месяц назад

Qwen3.8-Flash-Next is starting to feel like the local model...

FHILY👑's profile picture

FHILY👑

19,992 просмотров • 11 дней назад

🎉 Congrats to Thinking Machines on TML Inkling—a 1T-parameter...

vLLM's profile picture

vLLM

49,419 просмотров • 1 месяц назад

5 days ago it took 2 GPUs to build...

Sudo su's profile picture

Sudo su

34,624 просмотров • 6 месяцев назад

NVIDIA just dropped Nemotron-3-Nano:4b — a tiny 2.8GB model....

stevibe's profile picture

stevibe

127,570 просмотров • 5 месяцев назад

nvidia/Qwen3.6-35B-A3B-NVFP4 running in vLLM nightly on my Nvidia GB10...

Onur Solmaz's profile picture

Onur Solmaz

27,887 просмотров • 2 месяцев назад

This is awesome. Thanks Mia for this Was able...

Melvin Vivas's profile picture

Melvin Vivas

87,900 просмотров • 7 дней назад

Three Cline agents. Same 120B model. One rule: terminate...

Cline's profile picture

Cline

14,782 просмотров • 4 месяцев назад

SmallThinker 3B is one of the AI models I’d...

Aaron Ng's profile picture

Aaron Ng

35,858 просмотров • 1 год назад

Laguna XS 2.1 just matched Qwen 3.6 35B on...

0xMarioNawfal's profile picture

0xMarioNawfal

63,839 просмотров • 1 месяц назад

Gemma 4 26B A4B MoE - 500+ t/s decode...

Alok's profile picture

Alok

17,465 просмотров • 1 месяц назад

now it's live🔥 GLM-5.2 on RTX 5090s: 80-110 tok/s...

Ning's profile picture

Ning

157,244 просмотров • 1 месяц назад

Got DeepSeek V4 Flash running on 2x H200s at...

Atharva Ingle's profile picture

Atharva Ingle

21,913 просмотров • 1 месяц назад