正在加载视频...
视频加载失败
397 BILLION PARAMETER MODEL RUNS LOCALLY ON A $32,000 DESKTOP CLUSTER An 8-node cluster of Nvidia GB10 mini-supercomputers pooled 1TB of memory to run Qwen 3.5 offline. Thanks to Mixture-of-Experts sparsity (activating only 17B parameters per token) the $32,000 setup delivers private 24 token/sec inference without data leaving the room.
54,183 次观看 • 10 天前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
