Wide Expert Parallelism increases the total memory bandwidth available...

SemiAnalysis's profile picture

SemiAnalysis

30,205 views • 1 month ago

MTP speedup Qwen by 2.5x in Atomic Chat Dense...

atomic.chat's profile picture

atomic.chat

170,338 views • 2 months ago

Native time-tracking in Notion This setup lets you track...

Thomas Frank's profile picture

Thomas Frank

12,475 views • 1 year ago

AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME...

starmex's profile picture

starmex

192,758 views • 1 month ago

this is the worst local AI will ever be....

Sudo su's profile picture

Sudo su

106,710 views • 4 months ago

Newborns are so tiny. Most full term infants weigh...

Dan Wuori's profile picture

Dan Wuori

42,919 views • 10 months ago

Run Gemma 4 26B MoE on 8GB VRAM with...

Alok's profile picture

Alok

292,096 views • 1 month ago

"Pakistan is the only country that has good relations...

BALA's profile picture

BALA

301,425 views • 3 months ago

5 days ago it took 2 GPUs to build...

Sudo su's profile picture

Sudo su

34,569 views • 4 months ago

I can’t believe I got to design this dream...

Joey Chou's profile picture

Joey Chou

61,678 views • 1 year ago

$IREN "we haven't disclosed the specific amount of GPUs"...

Frans Bakker's profile picture

Frans Bakker

146,717 views • 2 months ago

Amazon’s machine learning model collects 300 million data points...

Joe Pompliano's profile picture

Joe Pompliano

688,463 views • 2 years ago

Day 12/90 of Inference Engineering What is chunked prefill...

max fu's profile picture

max fu

28,891 views • 7 days ago

K-Means is simple. Making it fast on GPU isn't....

Daily Dose of Data Science's profile picture

Daily Dose of Data Science

23,748 views • 2 months ago

K-Means is simple. Making it fast on GPU isn't....

Akshay 🚀's profile picture

Akshay 🚀

36,317 views • 4 months ago

i found a way to make UNCENSORED AI AGENT...

chiefofautism's profile picture

chiefofautism

341,589 views • 5 months ago