Wide Expert Parallelism increases the total memory bandwidth available...

SemiAnalysis's profile picture

SemiAnalysis

30,675 views • 2 months ago

MTP speedup Qwen by 2.5x in Atomic Chat Dense...

atomic.chat's profile picture

atomic.chat

171,139 views • 3 months ago

167 tok/s on a single RTX 4090. FreeToken just...

FHILY👑's profile picture

FHILY👑

33,058 views • 17 days ago

GLM 5.2 INPUT FELL FROM $1.40 TO SEVEN CENTS...

slash1s's profile picture

slash1s

30,231 views • 1 month ago

Native time-tracking in Notion This setup lets you track...

Thomas Frank's profile picture

Thomas Frank

12,480 views • 1 year ago

This is what happens when 48GB of VRAM isn't...

KyzoroX's profile picture

KyzoroX

81,125 views • 6 days ago

AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME...

starmex's profile picture

starmex

193,226 views • 2 months ago

this is the worst local AI will ever be....

Sudo su's profile picture

Sudo su

106,822 views • 6 months ago

Newborns are so tiny. Most full term infants weigh...

Dan Wuori's profile picture

Dan Wuori

203,718 views • 2 years ago

Run Gemma 4 26B MoE on 8GB VRAM with...

Alok's profile picture

Alok

292,770 views • 3 months ago

"Pakistan is the only country that has good relations...

BALA's profile picture

BALA

301,541 views • 5 months ago

5 days ago it took 2 GPUs to build...

Sudo su's profile picture

Sudo su

34,624 views • 6 months ago

I can’t believe I got to design this dream...

Joey Chou's profile picture

Joey Chou

61,678 views • 1 year ago

Gemma 4 26B A4B MoE - 500+ t/s decode...

Alok's profile picture

Alok

17,465 views • 1 month ago

THE CLOUD BILL WAS $14,000 A MONTH. EVERY MONTH....

Framez's profile picture

Framez

18,004 views • 1 month ago

The bulk/cut thing is literally so simple: Spend as...

Dean Turner's profile picture

Dean Turner

144,011 views • 18 days ago

$IREN "we haven't disclosed the specific amount of GPUs"...

Frans Bakker's profile picture

Frans Bakker

148,167 views • 4 months ago