Loading video...

Video Failed to Load

Go Home

17,255 views • 3 days ago •via X (Twitter)

36 Comments

Sufyan's profile picture
Sufyan3 days ago

500

Miyuru's profile picture
Miyuru3 days ago

locking in 900. now tell me the model and how many cards are coughing up smoke behind it

Meet Limbani's profile picture
Meet Limbani3 days ago

i know that scrolling speed in PI, it's around 350

Gil's profile picture
Gil3 days ago

750

Neo's profile picture
Neo3 days ago

420

0xSero's profile picture
0xSero3 days ago

420.69

Decatalyst 🌱's profile picture
Decatalyst 🌱3 days ago

OH MY TOMATOES I FEEL LIKE A CAVEMAN USING OPUS 5 NOW 😭

ruleryak's profile picture
ruleryak3 days ago

went from 24k output at 7s to 44k output at 34s so end to end throughput including tool calls, prefill, and decode is around 750 tps there. So the actual decode is probably 1k+

Maciej's profile picture
Maciej3 days ago

No tool call streaming? They spend bilions on hardware and lack money for API software.

Valentin Yanakiev's profile picture
Valentin Yanakiev3 days ago

384

D's profile picture
D3 days ago

450

pythongiant's profile picture
pythongiant3 days ago

a majillion

PaulNL's profile picture
PaulNL3 days ago

400ish ?

Mike Reese's profile picture
Mike Reese3 days ago

How does cerebras compare to your local setup with deepseek 4.1 flash in terms of tok/s?

AArchimedes64's profile picture
AArchimedes643 days ago

1024

cherki's profile picture
cherki3 days ago

How?

Arman's profile picture
Arman3 days ago

1337

Fraser Price's profile picture
Fraser Price3 days ago

Bout tree fiddy

Crispybits || Parroty account's profile picture
Crispybits || Parroty account3 days ago

3 - you've sped up the AI writing sections of the video 1000x 😹

Redmix's profile picture
Redmix3 days ago

~1850

Wayne's profile picture
Wayne3 days ago

I bet that's at least 1t/s

Serhii Y.'s profile picture
Serhii Y.3 days ago

you're typing 1-2tts, model 100+

mo's profile picture
mo3 days ago

320

Kutluk's profile picture
Kutluk3 days ago

~300-400?

Jason Frisch's profile picture
Jason Frisch3 days ago

What do I need to run GLM-5.3-flash? Can a Mac Studio manage it for regular coding tasks?

Loco Legend's profile picture
Loco Legend3 days ago

170-250

Sektur Toilet's profile picture
Sektur Toilet3 days ago

3-400?

Jeff Steve's profile picture
Jeff Steve3 days ago

700

Simao Soares 🇪🇺🇵🇹🇳🇱's profile picture
Simao Soares 🇪🇺🇵🇹🇳🇱3 days ago

I've seen a tok/s, is there a metric for token quality? I'm honestly curious. So many tools to reduce token usage like context compression, context understanding, prompt engineering for summarised high quality replies, MTP and non-MTP...

saltyclaw707's profile picture
saltyclaw7073 days ago

300 ish

vveerrgg's profile picture
vveerrgg3 days ago

280 T/Sec. … how are you configuring your cards … that’s the question I have

Petey struggles's profile picture
Petey struggles3 days ago

1500

Matt's profile picture
Matt3 days ago

700+

Asif Sheikh 💸's profile picture
Asif Sheikh 💸3 days ago

~1,850 tokens/s you're welcome.

byovd's profile picture
byovd3 days ago

300

Kachow Towmater's profile picture
Kachow Towmater3 days ago

280

Related Videos