
KyzoroX
@kyzoroXX • 1,297 subscribers
AI • Automation • Innovation
Shorts
Videos

We used to buy more VRAM. Now we’re finding ways around the VRAM limit. This little DDR5 SO-DIMM setup is paired with an old NVIDIA Tesla P40 in the video. The P40 is an old datacenter GPU with 24GB of VRAM. By today’s standards, it’s not fast. But combine that with a machine carrying 128GB of system memory, and suddenly you have a very different local AI box. The GPU handles what it’s good at. The system memory gives the model somewhere to go when VRAM runs out. You obviously pay for that with bandwidth and inference speed. But this is the part I find interesting: Local AI is making “how much memory do I have?” almost as important as “how powerful is my GPU?” We’ve spent years chasing faster GPUs. Maybe the next wave of local AI hardware will be about finding smarter ways to use all the memory we already have. More VRAM is still better. But it’s no longer the only way to run a bigger model.
KyzoroX35,421 görüntüleme • 1 gün önce

RTX 3090 was never designed to be a server GPU. Yet someone just built a 4-GPU AI workstation around four of them. 4× RTX 3090 96GB VRAM 256GB ECC RAM AMD EPYC 7K62 2000W PSU 11 fans The interesting part isn't the performance. It's the memory. Four 24GB cards can be found for a fraction of what you'd pay for modern datacenter GPUs with similar total VRAM. The catch is obvious: 96GB doesn't behave like one 96GB GPU. You still have to deal with GPU-to-GPU communication, power and heat. But for local AI, this is exactly why the 3090 refuses to disappear. Would you build this today?
KyzoroX36,030 görüntüleme • 5 gün önce

Chinese factories are buying RTX 5090s in bulk, desoldering the chips, and rebuilding them as 128GB server cards. Roughly $4,000 each. Same money as a DGX Spark. Same 128GB. Six times the bandwidth. The reason this exists: export controls block Chinese labs from buying dedicated AI chips, so consumer gaming cards became the workaround. Rip the GPU and memory off the board, solder onto a custom server PCB, add a blower cooler so it fits a rack instead of a gaming case. Automated soldering lines, robotic arms. Started with the 5090, now spreading to the 5080, 5070 Ti, 5060 Ti. NVIDIA gets paid either way. They just lost control of what happens after the sale. Two things follow from this: Gamers are the collateral. Consumer cards are being bought in volume and stripped for AI, which is part of why you can't get one at a sane price. And the "$4,000 for 128GB" pitch that DGX and Strix Halo are built on now has a grey-market competitor with dramatically better bandwidth — assuming you can get one, trust the soldering, and live without support. The most interesting AI hardware of 2026 isn't being announced at keynotes. It's being reflowed in a workshop.
KyzoroX336,192 görüntüleme • 1 ay önce

"Your Mac is useless for AI without CUDA, buy the $4K DGX." Half right — and the half that's wrong costs you $4,000. CUDA matters for training and some frameworks. Real. But "running local models"? I benchmarked a Mac Mini against the DGX on the same 30B: generation was 56 vs 84 t/s. Not useless — usable, faster than you read, no CUDA required. llama.cpp and Ollama don't care what logo is on the chip. Here's the honest split the video skips: - Commercial fine-tuning, CUDA-only pipelines, big-context prefill → yes, the DGX earns its price - Running and chatting with local models → your Mac already does this fine The DGX isn't overpriced. It's mis-pitched. It's a prefill-and-training machine, not a "your Mac is trash" machine. Buy it for the workload that needs it — not out of CUDA FOMO. Full 3-way benchmark (DGX vs Strix Halo vs Mac Mini, prefill vs generation) pinned.
KyzoroX263,861 görüntüleme • 1 ay önce

Someone built a workstation that needs 4.5 kilowatts. That's more than any standard wall circuit can deliver. Renderboxes Molecule Air: eight RTX 6000 Ada, a Threadripper PRO 7975WX, 768GB of DDR5 ECC 5800, and three Corsair HX1500i supplies stacked together. 🔴 US 15A circuit → 1,800W 🔴 US 20A circuit → 2,400W 🔴 EU 16A circuit → 3,680W 🟢 This machine → 4,500W The PSU rating isn't the draw, though. Eight cards at 300W plus the CPU lands near 3kW sustained. That's three space heaters running in the room, and the cooling to remove that heat costs more power on top. Note: 384GB of VRAM total, but it's eight separate 48GB pools talking over PCIe. Not one pool of 384. This is what "local AI" looks like at the top end. A datacenter rack that happens to be in a room. So before the GPU budget — what does your panel actually deliver?
KyzoroX21,491 görüntüleme • 9 gün önce

Someone bought a Mac Mini for local AI. 64GB of unified memory seemed like the obvious choice. More memory. Bigger models. No VRAM limits. Then they put it against an RTX 3060. And things got weird. The Mac Mini M4 Pro pushed Qwen 2.5 at 38 tok/s. The old RTX 3060 with just 12GB VRAM hit 52 tok/s. That’s the interesting part about local AI. Apple and NVIDIA are solving different problems. Apple gives you a huge pool of memory. NVIDIA gives you dedicated compute and CUDA. One helps you fit bigger models. The other can run smaller models faster. Memory gets the model in. Compute gets the tokens out.
KyzoroX26,269 görüntüleme • 26 gün önce

$15,000 for a Mac sounds insane. Until you realize what Apple is actually selling. 256GB of unified memory. Not a huge GPU farm. Not an NVIDIA rack. Not 4 different machines. Just one box that can keep multiple large LLMs in memory and run them locally. That changes the equation for local AI. NVIDIA wins when you need raw GPU compute. Apple gets interesting when the bottleneck is memory capacity. Because once your model doesn't fit into 24GB, 32GB or 48GB of VRAM, buying a faster GPU doesn't solve the problem. You need more memory. And that's where machines like the Mac Studio start looking less crazy. The interesting question isn't: “Is a $15K Mac faster than an NVIDIA GPU?” It isn't. The real question is: “How much are you willing to pay to make huge local models actually fit?” That’s where unified memory gets interesting.
KyzoroX19,536 görüntüleme • 28 gün önce
Daha fazla içerik yok.