Loading video...

Video Failed to Load

Go Home

136 TB in 1 desktop box, and the cloud subscription for that much space costs more per year than the box itself. Most creators run 2 separate setups. A slow archive drive for finished work, and a fast NVMe drive for the project open right now. Two boxes, two...

12,351 views • 19 days ago •via X (Twitter)

5 Comments

Cartel's profile picture
Cartel19 days ago

Damn, 136TB in one box with both fast NVMe and archive drives. Perfect for editors, no more juggling two separate setups.

0xShukshin's profile picture
0xShukshin19 days ago

worth noting: this is software RAID, and RAID 0 across those bays means zero redundancy. one drive fails, you lose the "fast" and "archive" data in the same event. one enclosure isn't the same as one backup

AiMind's profile picture
AiMind19 days ago

Fair point. You can set software RAID to 1 or 5 for redundancy, not just 0. Still keep a separate backup though.

Nexora's profile picture
Nexora18 days ago

Thanks! Now I know where I can save money.

wenWeights's profile picture
wenWeights18 days ago

Wow, 136TB in one box with fast NVMe and SATA archive. Perfect for editors, no more juggling two separate setups and cables. Cloud storage for that much space costs way more per year than the enclosure itself

Related Videos

APPLE SOLD THIS 39.9-POUND TOWER FOR $2,499 IN 2009 - ONE LATCH TURNED THE ENTIRE MACHINE INTO A WORKBENCH. this is the Early 2009 Mac Pro. the video shows the part spec sheets rarely capture. pull one latch and the aluminum side panel comes away. the memory, graphics card, drive bays and expansion slots are immediately visible. no hidden screws before you can inspect the machine. no loose SATA cables hanging from every drive. the base model shipped with: 2.66GHz quad-core Xeon 3GB of 1066MHz DDR3 ECC memory 640GB SATA hard drive GeForce GT 120 with 512MB of memory four PCIe 2.0 slots. the quad-core model supported up to 16GB of RAM. storage was handled by four cable-free carriers that connected directly inside the chassis. the case also left space for two optical drives. the rear panel carried hardware that now feels like an archive: dual Gigabit Ethernet optical audio USB 2.0 FireWire 800. the GT 120 offered Mini DisplayPort and dual-link DVI. install four of those cards and Apple rated the tower for up to eight 30-inch displays. the enclosure measured 20.1 inches tall and weighed almost 40 pounds before serious upgrades. this was not a compact desktop. it was a machine designed to be opened, understood and rebuilt. honest line: the Xeon, GT 120, SATA drive and USB 2.0 ports are ancient by modern standards. beautiful serviceability does not make old hardware fast. those parts aged. the enclosure did not. 17 years later, the most impressive component is still the case. bookmark & watch today ↓

Grimmer

224,481 views • 28 days ago

This Chinese developer linked two $2,999 NVIDIA DGX Sparks into one box and runs the full Qwen3-235B at home, after dropping his $1,999-a-month cloud bill to zero. He wired 2 small boxes into a single computer, split a giant 235-billion-parameter model in half between them, and serves it across his own network at about 10 tokens a second, with no internet, no cloud, right there on the desk. No data center, no thousand-dollar graphics cards, no monthly cloud bill. Just him, 2 gold boxes the size of a sandwich, one cable between them, and 1 power strip. And here is the whole payoff. He used to pay the cloud $1,999 a month for the same model, and the meter ticked on every request. Now he paid $5,998 once for 2 boxes, they covered their cost in 3 months, and after that he sends as many requests as he wants for free, only electricity. The two Sparks talk over one fast cable, each holds 128GB of memory, and together they carry the whole model, about 73GB loaded per box, with the chip inside pinned near the limit at 96%. Both boxes work as one and keep trading data over the cable, with no cloud in the loop and no single word leaking out. The ready model sits on one local address, and any app on his network calls it as easily as ChatGPT. And here is how he described, in plain words, what this pair of boxes does: "this is a pair of boxes that holds the huge Qwen3-235B model and serves it to one network. the model is split in half, and each box owns its half. parts: // Box 1 (holds the first half of the model and starts the answer fast, the first word appears in under a second) // Box 2 (holds the second half and writes out the rest, about 10 tokens a second) // Cable (connects the 2 boxes and moves data between them on every step, with no lag) // Address (one local address where any app sends its request, like to a cloud model) // Test (a script that runs big prompts through and measures speed and delays) // Monitor (checks temperature, power draw, and load on both boxes every 2 seconds). the model never goes to the cloud. he only steps in when a box runs hotter than 80 degrees or the cable between them starts dropping data." So the system knows exactly what it is, what it is for, and where its limits are. It knows it has to hold the whole huge model across 2 boxes on its own. It knows it has to answer every request locally, with no meter, no limits, and no internet. It knows the human is only needed when a box overheats or the link between them stalls. → The setup runs around the clock on 2 boxes, each pulling under 60 watts → However many requests he sends, the monthly bill is $0, only electricity → The first box starts the answer in under a second → The second writes text at about 10 tokens a second → One request at a time: 838 tokens in 85 seconds, first word in 0.8s → Two requests at once: 697 tokens in 108 seconds, first word in 0.7s → Both boxes sit at 96% load and warm up to 76-78 degrees And only when a chip in a box runs hotter than 80 degrees or the cable between the 2 Sparks drops data does the system call the owner. And when he himself is out on a run or in a coffee shop, he still reaches his own model at home from his phone: sends a big prompt to the local Qwen3-235B, gets the full answer back in under a minute and a half, with no token meter ticking and no limit to hit. Here is what the test shows on his screen during one of the night runs: "one request at a time: 838 tokens in 84.9 seconds, first word in 0.8s, then 0.1s per token." "two requests at once: 697 tokens in 107.6 seconds, first word in 0.7s, then 0.15s per token." "Box 1: chip at 96% load, 76 degrees, 56 watts, 73GB used in memory." "Box 2: chip at 96% load, 78 degrees, 56 watts, the Qwen3-235B model fully loaded." And while everyone around is paying for AI by the month and bumping into limits, his top-tier model just sits on the desk and works as much as he wants: his own little power plant instead of a forever meter. He has no server rack of his own and no cloud account behind it. Just 2 DGX Spark boxes on a desk, one model split in half between them, one local address, and a folder of prompts next to it. Out of everything I have seen this year, this is the cleanest way to stop paying for AI: $5,998 of hardware on the desk once, $0 a month to the cloud, unlimited forever, and between them 2 gold boxes, 1 cable, and the full Qwen3-235B answering at home with no internet.

Blaze

93,871 views • 3 months ago

NVIDIA quietly built two desktop boxes that delete a $25,000/year AI subscription bill You don't rewrite your stack, you don't rent another data center, you just plug both into the wall and switch one line of code One looks like a deck of cards, the other like a hardback novel, together they replace ChatGPT Plus, Claude Pro, Cursor Pro, the OpenAI API meter, and every cloud GPU you were renting for fine-tunes It's built on the same CUDA stack the data center runs, which means once you migrate one workflow the rest follow on the same code path The reason NVIDIA shipped this is simple The bigger you scale on cloud AI, the harder you get taxed, and a one-person operator paying $2,100/month is producing exactly $0 of asset value at the end of every month And their solution is to skip the rental meter entirely, push inference back onto your desk, and let you loop agents overnight without watching a number tick on someone else's invoice This is much cheaper, faster, and pays itself back in 6 weeks for anyone already running AI for work But there is still a question nobody has answered yet, what happens when the next frontier model drops and your local 70B falls 6 months behind mid-quarter Also, technically a stack of four of the big box runs a 1.6 trillion parameter model on a desk for under $12,000 Even a fraction of that compute is more than most people will ever need in a year Bookmark this, it's worth coming back to when you have time 👇

ZEUS⚡️

108,199 views • 3 months ago