Video yükleniyor...
Video Yüklenemedi
Successfully deployed Deepseek R1 Distilled 70B (AWQ) across 8x NVIDIA RTX 3080 10G GPUs, achieving 60 tokens/s with full tensor parallelism via PCIe. Total hardware cost: $6,400 This demonstrates that consumer GPUs can deliver substantial ML inference capabilities at a fraction of the cost of datacenter hardware. For perspective,... show more
16,219 görüntüleme • 1 yıl önce •via X (Twitter)
6 Yorum

TensorBlock1 yıl önce
For more detailed discussion and setup configurations, check out our Reddit post here:

The Information1 yıl önce
Elon Musk claims to have finished a 100,000-strong H100 cluster in four months. How likely is that?

Anti-Federalist Caesar1 yıl önce
@nvidia cc @__tinygrad__

אלף1 yıl önce
@nvidia Awesome!

Novel Engineer1 yıl önce
Why not 3090 - twice the vram

TensorBlock1 yıl önce
Thanks for the feedback! While 3090s with NVLink would offer higher bandwidth, this experiment focuses on validating distributed inference with consumer GPUs - exploring possibilities for widespread LLM deployment. Excited to share more findings soon.

