Loading video...

Video Failed to Load

Go Home

The Fastest AI Just Got Faster. Meet CS-4.

1,538,687 views • 1 month ago •via X (Twitter)

35 Comments

Cerebras's profile picture
Cerebras1 month ago

•Up to 30× faster inference than GPUs •Up to 10× more throughput per watt than CS-3 •50% fewer components with modular compute, power, and I/O assemblies

Danny Liu's profile picture
Danny Liu1 month ago

i thought we were still on cs 2

JoJo.hl's profile picture
JoJo.hl1 month ago

Will Cerebras ever sell smaller home-based machines? I think 100 tok/s on GPT 5.6 Sol looks very good.

🍓🍓🍓's profile picture
🍓🍓🍓1 month ago

@LLMJunky i just need to know if i get astra ultra xxxfast now?

Michaelkaoi's profile picture
Michaelkaoi1 month ago

@cerebras It's time to run kimi k3 full weight at 1000tps

Frank's profile picture
Frank1 month ago

very cool!!!

am.will's profile picture
am.will1 month ago

WHAT

fin's profile picture
fin1 month ago

Token Throughput is roughly 3~4X compare to CS-3 token economy at high speed token market CS-4 > VR+LPX > CS-3

Patrick Johnson's profile picture
Patrick Johnson1 month ago

Firstly, congratulations. Secondly I read the on-screen text as if it were backing singers adding emphasis. Something to consider for the CS-5 launch maybe.

vallz's profile picture
vallz1 month ago

Me to my parents : what do you mean it took hours to download a Game My future kids to me : what do you mean it took hours to generate a Game

Sandy's profile picture
Sandy1 month ago

Can we buy a consumer grade machine?

BenIt Pro Max Ultra's profile picture
BenIt Pro Max Ultra1 month ago

HOLY CRAP

Anp🅰️nman's profile picture
Anp🅰️nman1 month ago

That’s pretty sexy

Latent Local's profile picture
Latent Local1 month ago

Oooook! Very nice. When are these bad boys being deployed? Any stats? What can we look forward to? -Speed delta? -Tokens per W? etc etc The people are starved for speed!

Lewis Menelaws's profile picture
Lewis Menelaws1 month ago

Can I buy one of these for home

Frontier Signal's profile picture
Frontier Signal1 month ago

30x per user is a latency claim, not a cost one. SRAM instead of DRAM buys bandwidth and gives up capacity, which is why it takes three wafers. What settles displacement is $/M tokens and watts at a stated concurrency — neither is in the release.

Kaelan's profile picture
Kaelan1 month ago

Congrats on the launch! Unlike @Etched, Cerebras can actually back up their technical claims with published numbers

andy's profile picture
andy1 month ago

i’m in love (already was ngl)

Jeffrey 杰弗瑞's profile picture
Jeffrey 杰弗瑞1 month ago

But is this 30x the price or how does that work? Does anybody know ? How competitive is this new hardware price wise? Can I buy this for ds4 flash?

Andre Buckingham 🧙‍♂️'s profile picture
Andre Buckingham 🧙‍♂️1 month ago

if this isn't the future of inference idk what is... the speeds are absurd... please scale so you can do all the SOTAs silly cheap at 500-1000 tps

NeoSoul's profile picture
NeoSoul1 month ago

cs-4 decode speeds are actually insane for agent loops getting that many more reasoning passes in the same time budget completely flips how we build but ngl this just means the bottleneck shifts entirely to tool execution and state retrieval if inference is basically instant but the agent is stuck waiting on external api calls we are still cooked especially for autonomous trading workloads where io lag is the real final boss curious how the disaggregated prefill and decode handles massive concurrency for tool-heavy agent fleets right now

Emad Ghorbaninia's profile picture
Emad Ghorbaninia1 month ago

30x faster inference is the headline, but the real story is what that compute-per-watt frees up: fine-tuning that used to be an overnight batch job becomes a same-day iteration loop. That changes how teams work, not just how fast they answer.

Ken Goyarola's profile picture
Ken Goyarola1 month ago

Whens the next hackathon to build on it? 😸 Want to build workflows that finish faster than my enter key resets.

Barry Sayer 💻🖱️☁️🤖ᯅ 🐙's profile picture
Barry Sayer 💻🖱️☁️🤖ᯅ 🐙1 month ago

I need a cigarette

Jonah's profile picture
Jonah1 month ago

@ParallelAiRev I like this

Karan's profile picture
Karan1 month ago

Holy fucking shit the inference speeds

Josh Peterson's profile picture
Josh Peterson1 month ago

The internet moves ~1 PB/s globally One CS-4 rack moves 160 PB/s That's 160x the world's internet bandwidth, feeding compute equivalent to hundreds of GPUs. Now imagine 1,000 hooked together. @nvidia stock and my certainty I'm in base reality both taking big hits today fr 📉🤯

IMRAN - Zebracross's profile picture
IMRAN - Zebracross1 month ago

Now release a baby version of this that we can use at home to run DeepSeek flash

The AI Therapist's profile picture
The AI Therapist1 month ago

Cerebras skipped the margin race to sell inference chips. their wafer-scale engine cuts power density. you don't need a datacenter, you need a socket. scale is efficiency

shaik arbaz's profile picture
shaik arbaz1 month ago

Add more models

Syed-Mohammad Raza's profile picture
Syed-Mohammad Raza1 month ago

Huge upgrade from CS2 that I played as a kid

Engin's profile picture
Engin1 month ago

Fastest became faster.

Abbypro's profile picture
Abbypro1 month ago

@tweem 1 million views

Grokit🇺🇸's profile picture
Grokit🇺🇸1 month ago

How does this compare to Vera Rubin ? @grok

iandanforth 🦋 @iandanforth.bsky.social's profile picture
iandanforth 🦋 @iandanforth.bsky.social1 month ago

But can I use the backpacks to deal with a pesky infestation of ghosts?

Related Videos