正在加载视频...

视频加载失败

The Fastest AI Just Got Faster. Meet CS-4.

1,538,687 次观看 • 1 个月前 •via X (Twitter)

35 条评论

Cerebras 的头像
Cerebras1 个月前

•Up to 30× faster inference than GPUs •Up to 10× more throughput per watt than CS-3 •50% fewer components with modular compute, power, and I/O assemblies

Danny Liu 的头像
Danny Liu1 个月前

i thought we were still on cs 2

JoJo.hl 的头像
JoJo.hl1 个月前

Will Cerebras ever sell smaller home-based machines? I think 100 tok/s on GPT 5.6 Sol looks very good.

🍓🍓🍓 的头像
🍓🍓🍓1 个月前

@LLMJunky i just need to know if i get astra ultra xxxfast now?

Michaelkaoi 的头像
Michaelkaoi1 个月前

@cerebras It's time to run kimi k3 full weight at 1000tps

Frank 的头像
Frank1 个月前

very cool!!!

am.will 的头像
am.will1 个月前

WHAT

fin 的头像
fin1 个月前

Token Throughput is roughly 3~4X compare to CS-3 token economy at high speed token market CS-4 > VR+LPX > CS-3

Patrick Johnson 的头像
Patrick Johnson1 个月前

Firstly, congratulations. Secondly I read the on-screen text as if it were backing singers adding emphasis. Something to consider for the CS-5 launch maybe.

vallz 的头像
vallz1 个月前

Me to my parents : what do you mean it took hours to download a Game My future kids to me : what do you mean it took hours to generate a Game

Sandy 的头像
Sandy1 个月前

Can we buy a consumer grade machine?

BenIt Pro Max Ultra 的头像
BenIt Pro Max Ultra1 个月前

HOLY CRAP

Anp🅰️nman 的头像
Anp🅰️nman1 个月前

That’s pretty sexy

Latent Local 的头像
Latent Local1 个月前

Oooook! Very nice. When are these bad boys being deployed? Any stats? What can we look forward to? -Speed delta? -Tokens per W? etc etc The people are starved for speed!

Lewis Menelaws 的头像
Lewis Menelaws1 个月前

Can I buy one of these for home

Frontier Signal 的头像
Frontier Signal1 个月前

30x per user is a latency claim, not a cost one. SRAM instead of DRAM buys bandwidth and gives up capacity, which is why it takes three wafers. What settles displacement is $/M tokens and watts at a stated concurrency — neither is in the release.

Kaelan 的头像
Kaelan1 个月前

Congrats on the launch! Unlike @Etched, Cerebras can actually back up their technical claims with published numbers

andy 的头像
andy1 个月前

i’m in love (already was ngl)

Jeffrey 杰弗瑞 的头像
Jeffrey 杰弗瑞1 个月前

But is this 30x the price or how does that work? Does anybody know ? How competitive is this new hardware price wise? Can I buy this for ds4 flash?

Andre Buckingham 🧙‍♂️ 的头像
Andre Buckingham 🧙‍♂️1 个月前

if this isn't the future of inference idk what is... the speeds are absurd... please scale so you can do all the SOTAs silly cheap at 500-1000 tps

NeoSoul 的头像
NeoSoul1 个月前

cs-4 decode speeds are actually insane for agent loops getting that many more reasoning passes in the same time budget completely flips how we build but ngl this just means the bottleneck shifts entirely to tool execution and state retrieval if inference is basically instant but the agent is stuck waiting on external api calls we are still cooked especially for autonomous trading workloads where io lag is the real final boss curious how the disaggregated prefill and decode handles massive concurrency for tool-heavy agent fleets right now

Emad Ghorbaninia 的头像
Emad Ghorbaninia1 个月前

30x faster inference is the headline, but the real story is what that compute-per-watt frees up: fine-tuning that used to be an overnight batch job becomes a same-day iteration loop. That changes how teams work, not just how fast they answer.

Ken Goyarola 的头像
Ken Goyarola1 个月前

Whens the next hackathon to build on it? 😸 Want to build workflows that finish faster than my enter key resets.

Barry Sayer 💻🖱️☁️🤖ᯅ 🐙 的头像
Barry Sayer 💻🖱️☁️🤖ᯅ 🐙1 个月前

I need a cigarette

Jonah 的头像
Jonah1 个月前

@ParallelAiRev I like this

Karan 的头像
Karan1 个月前

Holy fucking shit the inference speeds

Josh Peterson 的头像
Josh Peterson1 个月前

The internet moves ~1 PB/s globally One CS-4 rack moves 160 PB/s That's 160x the world's internet bandwidth, feeding compute equivalent to hundreds of GPUs. Now imagine 1,000 hooked together. @nvidia stock and my certainty I'm in base reality both taking big hits today fr 📉🤯

IMRAN - Zebracross 的头像
IMRAN - Zebracross1 个月前

Now release a baby version of this that we can use at home to run DeepSeek flash

The AI Therapist 的头像
The AI Therapist1 个月前

Cerebras skipped the margin race to sell inference chips. their wafer-scale engine cuts power density. you don't need a datacenter, you need a socket. scale is efficiency

shaik arbaz 的头像
shaik arbaz1 个月前

Add more models

Syed-Mohammad Raza 的头像
Syed-Mohammad Raza1 个月前

Huge upgrade from CS2 that I played as a kid

Engin 的头像
Engin1 个月前

Fastest became faster.

Abbypro 的头像
Abbypro1 个月前

@tweem 1 million views

Grokit🇺🇸 的头像
Grokit🇺🇸1 个月前

How does this compare to Vera Rubin ? @grok

iandanforth 🦋 @iandanforth.bsky.social 的头像
iandanforth 🦋 @iandanforth.bsky.social1 个月前

But can I use the backpacks to deal with a pesky infestation of ghosts?

相关视频