
Volatile Markets
@volatilemarkts • 2,384 subscribers
Primary work in Health Care. Passion for Markets, Technicals, and AI Modeling. Acquire the compute while you can…
Videos

Same Mac. Same model. Same words. My 512 GB M3 Ultra ran GLM-5.3-Flash on MLX last week: 27 tok/s, first token in 0.46 s. Today it runs TensorFold 0.6.2: 61 tok/s, first token in 0.20 s. In the clip it laps the reply twice while the MLX side is still typing. Same reply, byte for byte. Then eight requests at once, 78 tok/s combined, every answer identical to its solo run. The old engine took one at a time. Context stays at 1,048,576 on the Mac. Every number measured on my own machine with the same probe, one week apart. Nothing estimated. Ash Hart built it. Apache 2.0.
Volatile Markets26,010 görüntüleme • 4 gün önce
Daha fazla içerik yok.