Wide Expert Parallelism increases the total memory bandwidth available...

SemiAnalysis's profile picture

SemiAnalysis

30,675 次观看 • 2 个月前

MTP speedup Qwen by 2.5x in Atomic Chat Dense...

atomic.chat's profile picture

atomic.chat

171,139 次观看 • 3 个月前

167 tok/s on a single RTX 4090. FreeToken just...

FHILY👑's profile picture

FHILY👑

33,058 次观看 • 13 天前

GLM 5.2 INPUT FELL FROM $1.40 TO SEVEN CENTS...

slash1s's profile picture

slash1s

30,117 次观看 • 27 天前

Native time-tracking in Notion This setup lets you track...

Thomas Frank's profile picture

Thomas Frank

12,480 次观看 • 1 年前

This is what happens when 48GB of VRAM isn't...

KyzoroX's profile picture

KyzoroX

80,805 次观看 • 3 天前

AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME...

starmex's profile picture

starmex

193,226 次观看 • 2 个月前

this is the worst local AI will ever be....

Sudo su's profile picture

Sudo su

106,822 次观看 • 6 个月前

Newborns are so tiny. Most full term infants weigh...

Dan Wuori's profile picture

Dan Wuori

203,718 次观看 • 2 年前

Run Gemma 4 26B MoE on 8GB VRAM with...

Alok's profile picture

Alok

292,770 次观看 • 3 个月前

"Pakistan is the only country that has good relations...

BALA's profile picture

BALA

301,541 次观看 • 5 个月前

5 days ago it took 2 GPUs to build...

Sudo su's profile picture

Sudo su

34,624 次观看 • 6 个月前

Gemma 4 26B A4B MoE - 500+ t/s decode...

Alok's profile picture

Alok

17,465 次观看 • 1 个月前

I can’t believe I got to design this dream...

Joey Chou's profile picture

Joey Chou

61,678 次观看 • 1 年前

THE CLOUD BILL WAS $14,000 A MONTH. EVERY MONTH....

Framez's profile picture

Framez

18,004 次观看 • 1 个月前

The bulk/cut thing is literally so simple: Spend as...

Dean Turner's profile picture

Dean Turner

144,011 次观看 • 14 天前

$IREN "we haven't disclosed the specific amount of GPUs"...

Frans Bakker's profile picture

Frans Bakker

148,167 次观看 • 4 个月前