正在加载视频...
视频加载失败
Got continuous batching working with SSMs in mlx-lm. Here's four OpenCode agents simultaneously running Nvidia's Nemotron Nano on 64GB M4 Max. This is a nice model for smaller machines since it's MoE + hybrid attention (small cache).
35,078 次观看 • 5 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
