Loading video...
Video Failed to Load
When will AI personal assistants be fast enough to be useful? Using Qwen 3.8 27B on Cerebras at ~1,500 tokens/sec, we made an AI personal assistant 19x faster than a suite of other AI personal assistants - Grok Bot, Meta Muse, and Claude Cowork - on the same dinner... show more
27,910 views • 8 days ago •via X (Twitter)
15 Comments

Read the full blog by @MilksandMatcha:

Seems slower than TPU?

using it with @trycua - booooom

Always old model 😂 cerebras you suck now for developers … continue to sell to OpenAI your hardware

The useful number is probably reservations confirmed per minute, not tokens per second. How much of the run was model time versus browser and booking-tool latency?

Suck

At 1,500 tokens per second, does the bottleneck completely shift over to external API calls now?

Yes but 128K context. A few rounds of agentic actions would use it up. When 256K context?

내부자 매도가 너무 많다

Would be awesome if you brought Gemma back

yeah but is cerebras used in Ultrafast?, that would be a start at least

@MilksandMatcha Hey, I'm working on something in personal AI... Wanna partner again?👀

Bro spends 24/7 hyping his own stock on here, yet OpenAI didn’t even bother using your chips for their latest model. Truly embarrassing."

Please bring more models.

Ok but release it so we can use it.
