Loading video...
Video Failed to Load
397 BILLION PARAMETER MODEL RUNS LOCALLY ON A $32,000 DESKTOP CLUSTER An 8-node cluster of Nvidia GB10 mini-supercomputers pooled 1TB of memory to run Qwen 3.5 offline. Thanks to Mixture-of-Experts sparsity (activating only 17B parameters per token) the $32,000 setup delivers private 24 token/sec inference without data leaving the room.
54,183 views • 10 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
