Loading video...
Video Failed to Load
Next mlx-vlm release will ship with continuous batching support on the server ๐ What's coming: โ Continuous batching โ new requests join the active batch immediately, no waiting. Mixed image + text batches supported โ OpenAI-compatible API โ field-for-field match with mlx-lm, reasoning/content split for thinking models, tag-aware streaming... show more
82,349 views โข 4 months ago โขvia X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
