Загрузка видео...
Не удалось загрузить видео
Same Mac. Same model. Same words. My 512 GB M3 Ultra ran GLM-5.3-Flash on MLX last week: 27 tok/s, first token in 0.46 s. Today it runs TensorFold 0.6.2: 61 tok/s, first token in 0.20 s. In the clip it laps the reply twice while the MLX side is... show more
26,010 просмотров • 4 дней назад •via X (Twitter)
Комментарии: 11

Let’s get that to 100tps 👀

61 tok/s on the same box is a nice jump. does it hold once the context gets ugly?

Like duct tape.

Yep, it almost doubled the speed on my sparks with the same model. And I have already seen improvements with the new updates and recipes. Excited to see where it goes with so many people being able to find these new techniques.

The eight-request result is the real systems datapoint: concurrency exposes memory-bandwidth headroom that single-stream tok/s hides.

🔥🔥

2.25x on the same box, same model, measured with your own probe, is hard to argue with. The first token drop, 0.46 to 0.20, is the part that gets me. That's where interactive use lives or dies.

Love it.

With vision?

Crushing

mlx 27 then tf 61 and eight concurrent still matches solo byte for byte
