Loading video...

Video Failed to Load

Go Home

When will AI personal assistants be fast enough to be useful? Using Qwen 3.8 27B on Cerebras at ~1,500 tokens/sec, we made an AI personal assistant 19x faster than a suite of other AI personal assistants - Grok Bot, Meta Muse, and Claude Cowork - on the same dinner...

27,910 views • 8 days ago •via X (Twitter)

15 Comments

Cerebras's profile picture
Cerebras8 days ago

Read the full blog by @MilksandMatcha:

Tapa Ghosh's profile picture
Tapa Ghosh8 days ago

Seems slower than TPU?

Dominik Gstöhl's profile picture
Dominik Gstöhl8 days ago

using it with @trycua - booooom

Talk AI Today's profile picture
Talk AI Today8 days ago

Always old model 😂 cerebras you suck now for developers … continue to sell to OpenAI your hardware

ShadowAguy's profile picture
ShadowAguy8 days ago

The useful number is probably reservations confirmed per minute, not tokens per second. How much of the run was model time versus browser and booking-tool latency?

selene's profile picture
selene8 days ago

Suck

EDDY VU's profile picture
EDDY VU8 days ago

At 1,500 tokens per second, does the bottleneck completely shift over to external API calls now?

Mega Kilo's profile picture
Mega Kilo8 days ago

Yes but 128K context. A few rounds of agentic actions would use it up. When 256K context?

무자본진리's profile picture
무자본진리8 days ago

내부자 매도가 너무 많다

Clint Fix's profile picture
Clint Fix8 days ago

Would be awesome if you brought Gemma back

Claudio Aldana's profile picture
Claudio Aldana8 days ago

yeah but is cerebras used in Ultrafast?, that would be a start at least

Yashas's profile picture
Yashas8 days ago

@MilksandMatcha Hey, I'm working on something in personal AI... Wanna partner again?👀

selene's profile picture
selene8 days ago

Bro spends 24/7 hyping his own stock on here, yet OpenAI didn’t even bother using your chips for their latest model. Truly embarrassing."

Sathya's profile picture
Sathya8 days ago

Please bring more models.

Matt Albrizio's profile picture
Matt Albrizio8 days ago

Ok but release it so we can use it.

Related Videos