Загрузка видео...
Не удалось загрузить видео
Guess the tok/s
17,255 просмотров • 3 дней назад •via X (Twitter)
Комментарии: 36

500

locking in 900. now tell me the model and how many cards are coughing up smoke behind it

i know that scrolling speed in PI, it's around 350

750

420

420.69

OH MY TOMATOES I FEEL LIKE A CAVEMAN USING OPUS 5 NOW 😭

went from 24k output at 7s to 44k output at 34s so end to end throughput including tool calls, prefill, and decode is around 750 tps there. So the actual decode is probably 1k+

No tool call streaming? They spend bilions on hardware and lack money for API software.

384

450

a majillion

400ish ?

How does cerebras compare to your local setup with deepseek 4.1 flash in terms of tok/s?

1024

How?

1337

Bout tree fiddy

3 - you've sped up the AI writing sections of the video 1000x 😹

~1850

I bet that's at least 1t/s

you're typing 1-2tts, model 100+

320

~300-400?

What do I need to run GLM-5.3-flash? Can a Mac Studio manage it for regular coding tasks?

170-250

3-400?

700

I've seen a tok/s, is there a metric for token quality? I'm honestly curious. So many tools to reduce token usage like context compression, context understanding, prompt engineering for summarised high quality replies, MTP and non-MTP...

300 ish

280 T/Sec. … how are you configuring your cards … that’s the question I have

1500

700+

~1,850 tokens/s you're welcome.

300

280
