
KyzoroX
@kyzoroXX • 1,228 subscribers
AI • Automation • Innovation
Shorts
Videos

RTX 3090 was never designed to be a server GPU. Yet someone just built a 4-GPU AI workstation around four of them. 4× RTX 3090 96GB VRAM 256GB ECC RAM AMD EPYC 7K62 2000W PSU 11 fans The interesting part isn't the performance. It's the memory. Four 24GB cards can be found for a fraction of what you'd pay for modern datacenter GPUs with similar total VRAM. The catch is obvious: 96GB doesn't behave like one 96GB GPU. You still have to deal with GPU-to-GPU communication, power and heat. But for local AI, this is exactly why the 3090 refuses to disappear. Would you build this today?
KyzoroX35,687 views • 2 days ago

Chinese factories are buying RTX 5090s in bulk, desoldering the chips, and rebuilding them as 128GB server cards. Roughly $4,000 each. Same money as a DGX Spark. Same 128GB. Six times the bandwidth. The reason this exists: export controls block Chinese labs from buying dedicated AI chips, so consumer gaming cards became the workaround. Rip the GPU and memory off the board, solder onto a custom server PCB, add a blower cooler so it fits a rack instead of a gaming case. Automated soldering lines, robotic arms. Started with the 5090, now spreading to the 5080, 5070 Ti, 5060 Ti. NVIDIA gets paid either way. They just lost control of what happens after the sale. Two things follow from this: Gamers are the collateral. Consumer cards are being bought in volume and stripped for AI, which is part of why you can't get one at a sane price. And the "$4,000 for 128GB" pitch that DGX and Strix Halo are built on now has a grey-market competitor with dramatically better bandwidth — assuming you can get one, trust the soldering, and live without support. The most interesting AI hardware of 2026 isn't being announced at keynotes. It's being reflowed in a workshop.
KyzoroX336,192 views • 1 month ago

"Your Mac is useless for AI without CUDA, buy the $4K DGX." Half right — and the half that's wrong costs you $4,000. CUDA matters for training and some frameworks. Real. But "running local models"? I benchmarked a Mac Mini against the DGX on the same 30B: generation was 56 vs 84 t/s. Not useless — usable, faster than you read, no CUDA required. llama.cpp and Ollama don't care what logo is on the chip. Here's the honest split the video skips: - Commercial fine-tuning, CUDA-only pipelines, big-context prefill → yes, the DGX earns its price - Running and chatting with local models → your Mac already does this fine The DGX isn't overpriced. It's mis-pitched. It's a prefill-and-training machine, not a "your Mac is trash" machine. Buy it for the workload that needs it — not out of CUDA FOMO. Full 3-way benchmark (DGX vs Strix Halo vs Mac Mini, prefill vs generation) pinned.
KyzoroX263,861 views • 1 month ago

Someone bought a Mac Mini for local AI. 64GB of unified memory seemed like the obvious choice. More memory. Bigger models. No VRAM limits. Then they put it against an RTX 3060. And things got weird. The Mac Mini M4 Pro pushed Qwen 2.5 at 38 tok/s. The old RTX 3060 with just 12GB VRAM hit 52 tok/s. That’s the interesting part about local AI. Apple and NVIDIA are solving different problems. Apple gives you a huge pool of memory. NVIDIA gives you dedicated compute and CUDA. One helps you fit bigger models. The other can run smaller models faster. Memory gets the model in. Compute gets the tokens out.
KyzoroX26,269 views • 23 days ago

$15,000 for a Mac sounds insane. Until you realize what Apple is actually selling. 256GB of unified memory. Not a huge GPU farm. Not an NVIDIA rack. Not 4 different machines. Just one box that can keep multiple large LLMs in memory and run them locally. That changes the equation for local AI. NVIDIA wins when you need raw GPU compute. Apple gets interesting when the bottleneck is memory capacity. Because once your model doesn't fit into 24GB, 32GB or 48GB of VRAM, buying a faster GPU doesn't solve the problem. You need more memory. And that's where machines like the Mac Studio start looking less crazy. The interesting question isn't: “Is a $15K Mac faster than an NVIDIA GPU?” It isn't. The real question is: “How much are you willing to pay to make huge local models actually fit?” That’s where unified memory gets interesting.
KyzoroX19,536 views • 25 days ago
No more content to load