正在加载视频...
视频加载失败
The Fastest AI Just Got Faster. Meet CS-4.
1,538,687 次观看 • 1 个月前 •via X (Twitter)
35 条评论

•Up to 30× faster inference than GPUs •Up to 10× more throughput per watt than CS-3 •50% fewer components with modular compute, power, and I/O assemblies

i thought we were still on cs 2

Will Cerebras ever sell smaller home-based machines? I think 100 tok/s on GPT 5.6 Sol looks very good.

@LLMJunky i just need to know if i get astra ultra xxxfast now?

@cerebras It's time to run kimi k3 full weight at 1000tps

very cool!!!

WHAT

Token Throughput is roughly 3~4X compare to CS-3 token economy at high speed token market CS-4 > VR+LPX > CS-3

Firstly, congratulations. Secondly I read the on-screen text as if it were backing singers adding emphasis. Something to consider for the CS-5 launch maybe.

Me to my parents : what do you mean it took hours to download a Game My future kids to me : what do you mean it took hours to generate a Game

Can we buy a consumer grade machine?

HOLY CRAP

That’s pretty sexy

Oooook! Very nice. When are these bad boys being deployed? Any stats? What can we look forward to? -Speed delta? -Tokens per W? etc etc The people are starved for speed!

Can I buy one of these for home

30x per user is a latency claim, not a cost one. SRAM instead of DRAM buys bandwidth and gives up capacity, which is why it takes three wafers. What settles displacement is $/M tokens and watts at a stated concurrency — neither is in the release.

Congrats on the launch! Unlike @Etched, Cerebras can actually back up their technical claims with published numbers

i’m in love (already was ngl)

But is this 30x the price or how does that work? Does anybody know ? How competitive is this new hardware price wise? Can I buy this for ds4 flash?

if this isn't the future of inference idk what is... the speeds are absurd... please scale so you can do all the SOTAs silly cheap at 500-1000 tps

cs-4 decode speeds are actually insane for agent loops getting that many more reasoning passes in the same time budget completely flips how we build but ngl this just means the bottleneck shifts entirely to tool execution and state retrieval if inference is basically instant but the agent is stuck waiting on external api calls we are still cooked especially for autonomous trading workloads where io lag is the real final boss curious how the disaggregated prefill and decode handles massive concurrency for tool-heavy agent fleets right now

30x faster inference is the headline, but the real story is what that compute-per-watt frees up: fine-tuning that used to be an overnight batch job becomes a same-day iteration loop. That changes how teams work, not just how fast they answer.

Whens the next hackathon to build on it? 😸 Want to build workflows that finish faster than my enter key resets.

I need a cigarette

@ParallelAiRev I like this

Holy fucking shit the inference speeds

The internet moves ~1 PB/s globally One CS-4 rack moves 160 PB/s That's 160x the world's internet bandwidth, feeding compute equivalent to hundreds of GPUs. Now imagine 1,000 hooked together. @nvidia stock and my certainty I'm in base reality both taking big hits today fr 📉🤯

Now release a baby version of this that we can use at home to run DeepSeek flash

Cerebras skipped the margin race to sell inference chips. their wafer-scale engine cuts power density. you don't need a datacenter, you need a socket. scale is efficiency

Add more models

Huge upgrade from CS2 that I played as a kid

Fastest became faster.

@tweem 1 million views

How does this compare to Vera Rubin ? @grok

But can I use the backpacks to deal with a pesky infestation of ghosts?
