Wide Expert Parallelism increases the total memory bandwidth available...

SemiAnalysis's profile picture

SemiAnalysis

30,675 次观看 • 1 个月前

MTP speedup Qwen by 2.5x in Atomic Chat Dense...

atomic.chat's profile picture

atomic.chat

170,704 次观看 • 2 个月前

Native time-tracking in Notion This setup lets you track...

Thomas Frank's profile picture

Thomas Frank

12,475 次观看 • 1 年前

AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME...

starmex's profile picture

starmex

192,758 次观看 • 1 个月前

this is the worst local AI will ever be....

Sudo su's profile picture

Sudo su

106,710 次观看 • 5 个月前

Newborns are so tiny. Most full term infants weigh...

Dan Wuori's profile picture

Dan Wuori

42,919 次观看 • 10 个月前

Run Gemma 4 26B MoE on 8GB VRAM with...

Alok's profile picture

Alok

292,770 次观看 • 1 个月前

"Pakistan is the only country that has good relations...

BALA's profile picture

BALA

301,425 次观看 • 4 个月前

5 days ago it took 2 GPUs to build...

Sudo su's profile picture

Sudo su

34,569 次观看 • 5 个月前

I can’t believe I got to design this dream...

Joey Chou's profile picture

Joey Chou

61,678 次观看 • 1 年前

$IREN "we haven't disclosed the specific amount of GPUs"...

Frans Bakker's profile picture

Frans Bakker

148,167 次观看 • 2 个月前

Amazon’s machine learning model collects 300 million data points...

Joe Pompliano's profile picture

Joe Pompliano

688,463 次观看 • 2 年前

Day 12/90 of Inference Engineering What is chunked prefill...

max fu's profile picture

max fu

29,111 次观看 • 14 天前

K-Means is simple. Making it fast on GPU isn't....

Daily Dose of Data Science's profile picture

Daily Dose of Data Science

23,748 次观看 • 3 个月前

K-Means is simple. Making it fast on GPU isn't....

Akshay 🚀's profile picture

Akshay 🚀

36,317 次观看 • 4 个月前

i found a way to make UNCENSORED AI AGENT...

chiefofautism's profile picture

chiefofautism

341,710 次观看 • 5 个月前