
Ronald Mannak
@ronaldmannak • 6,489 subscribers
Local AI. Building @picogpt Download: https://t.co/o0PKhVM691 Discord: https://t.co/KNirjAzl32
Videos

Apple Silicon + Gemma 4 fans: this is for you. Pico AI Server now supports continuous batching with MLX-Swift. 43 tok/s on 1 stream. 26 tok/s per stream on 2 concurrent streams. That’s 52 tok/s total. a 21% throughput gain on a six-year-old MacBook Pro M1 Max!
Ronald Mannak50,180 Aufrufe • vor 3 Monaten

Your Mac is about to run inference like a datacenter. Coming soon to MLX-Swift: Continuous batching: the fastest way to handle multiple inference streams locally. It starts with regular inference and seamlessly upgrades to batched mode when new requests arrive. The best of both worlds. Based on the work of and Awni Hannun
Ronald Mannak60,570 Aufrufe • vor 8 Monaten
Keine weiteren Inhalte verfügbar