
Ronald Mannak
@ronaldmannak • 6,489 subscribers
Local AI. Building @picogpt Download: https://t.co/o0PKhVM691 Discord: https://t.co/KNirjAzl32
Videos

Apple Silicon + Gemma 4 fans: this is for you. Pico AI Server now supports continuous batching with MLX-Swift. 43 tok/s on 1 stream. 26 tok/s per stream on 2 concurrent streams. That’s 52 tok/s total. a 21% throughput gain on a six-year-old MacBook Pro M1 Max!
Ronald Mannak50,180 görüntüleme • 3 ay önce

Your Mac is about to run inference like a datacenter. Coming soon to MLX-Swift: Continuous batching: the fastest way to handle multiple inference streams locally. It starts with regular inference and seamlessly upgrades to batched mode when new requests arrive. The best of both worlds. Based on the work of and Awni Hannun
Ronald Mannak60,570 görüntüleme • 8 ay önce
Daha fazla içerik yok.