正在加载视频...

视频加载失败

FOUR BOXES ON A DESK ARE NOW A FRONTIER AI CLUSTER This is four NVIDIA DGX Spark units wired together, sitting on a single desk and running Qwen3-235B end to end. Each box holds 128GB of unified memory. Stacked together, that is half a terabyte of memory pooled across...

19,473 次观看 • 3 个月前 •via X (Twitter)

35 条评论

leopardracer 的头像
leopardracer3 个月前

nice setup!

Asher 🪺 的头像
Asher 🪺3 个月前

the frontier keeps shrinking until it's just a guy, four boxes, and a power strip he's afraid to unplug

Paulera 的头像
Paulera3 个月前

What an idiot runs qwen3-235B on such setup lol

Ridark 的头像
Ridark3 个月前

Sat through a long flight watching this in two parts

Nahid 的头像
Nahid3 个月前

four boxes for a 235B model is wild pure future vibes

painn 的头像
painn3 个月前

i need these boxes too

Fokki 的头像
Fokki3 个月前

this is huge

Abo Saud 的头像
Abo Saud3 个月前

I'm excited to see same thing with competitive AMD device, especially considering its low price. 🔥🔥🔥

Jamie 的头像
Jamie3 个月前

Frontier class is debatable

Alpha Batcher 的头像
Alpha Batcher3 个月前

i guess i need these boxes too, for myself

Nick Venturi 的头像
Nick Venturi3 个月前

one coffee spill and its over

Craig|A.R. 的头像
Craig|A.R.3 个月前

At 10 t/s…

Bizzi 的头像
Bizzi3 个月前

Powerful stack

2ndClassCitizen 的头像
2ndClassCitizen3 个月前

Thats $16K sitting next to that "coffee cup on a desk"

Jan Swanepoel 的头像
Jan Swanepoel3 个月前

You forgot to mention how slow it runs that big model..

Yarchi 的头像
Yarchi3 个月前

this setup is so clean

Vital 👁️ 的头像
Vital 👁️3 个月前

How many t/s?

Antid 的头像
Antid3 个月前

bookmarked, sharding is the whole trick

shmidt 的头像
shmidt3 个月前

the fact this fits on a desk is actually insane

Deydara 的头像
Deydara3 个月前

These combinations are incredible

Inevitable East 的头像
Inevitable East3 个月前

What would happen if you put this machines working together into sewing a full lenght film together?

0xSlyth 的头像
0xSlyth3 个月前

this changes everything

vijn 的头像
vijn3 个月前

it's all NVIDIA?

Gipp 🦅 的头像
Gipp 🦅3 个月前

Wow, these 4 boxes are just crazy

Alexandru Bogdan 的头像
Alexandru Bogdan3 个月前

These are the new Lambo Photos?

0xbobaa 的头像
0xbobaa3 个月前

how much does a setup like this cost?

monokern 的头像
monokern3 个月前

this system look impressive

cristal💎 的头像
cristal💎3 个月前

half a terabyte locally goes hard

Atenov int. 的头像
Atenov int.3 个月前

But doesnt it need a ton of power to run

hrabiapolski 的头像
hrabiapolski3 个月前

I want this setup, its really good

Mykhailo Sorochuk 的头像
Mykhailo Sorochuk3 个月前

compact power move, impressive work

starmex 的头像
starmex3 个月前

thx for the share!

Belcebuu 的头像
Belcebuu3 个月前

Retarded waste of money

Nekt0 的头像
Nekt03 个月前

Very interesting

Adel Bucetta 的头像
Adel Bucetta3 个月前

the honest answer is that we're just starting to scratch the surface of what's possible when you combine massive memory with distributed computing and the latest in ai architecture

相关视频

This Chinese developer linked two $2,999 NVIDIA DGX Sparks into one box and runs the full Qwen3-235B at home, after dropping his $1,999-a-month cloud bill to zero. He wired 2 small boxes into a single computer, split a giant 235-billion-parameter model in half between them, and serves it across his own network at about 10 tokens a second, with no internet, no cloud, right there on the desk. No data center, no thousand-dollar graphics cards, no monthly cloud bill. Just him, 2 gold boxes the size of a sandwich, one cable between them, and 1 power strip. And here is the whole payoff. He used to pay the cloud $1,999 a month for the same model, and the meter ticked on every request. Now he paid $5,998 once for 2 boxes, they covered their cost in 3 months, and after that he sends as many requests as he wants for free, only electricity. The two Sparks talk over one fast cable, each holds 128GB of memory, and together they carry the whole model, about 73GB loaded per box, with the chip inside pinned near the limit at 96%. Both boxes work as one and keep trading data over the cable, with no cloud in the loop and no single word leaking out. The ready model sits on one local address, and any app on his network calls it as easily as ChatGPT. And here is how he described, in plain words, what this pair of boxes does: "this is a pair of boxes that holds the huge Qwen3-235B model and serves it to one network. the model is split in half, and each box owns its half. parts: // Box 1 (holds the first half of the model and starts the answer fast, the first word appears in under a second) // Box 2 (holds the second half and writes out the rest, about 10 tokens a second) // Cable (connects the 2 boxes and moves data between them on every step, with no lag) // Address (one local address where any app sends its request, like to a cloud model) // Test (a script that runs big prompts through and measures speed and delays) // Monitor (checks temperature, power draw, and load on both boxes every 2 seconds). the model never goes to the cloud. he only steps in when a box runs hotter than 80 degrees or the cable between them starts dropping data." So the system knows exactly what it is, what it is for, and where its limits are. It knows it has to hold the whole huge model across 2 boxes on its own. It knows it has to answer every request locally, with no meter, no limits, and no internet. It knows the human is only needed when a box overheats or the link between them stalls. → The setup runs around the clock on 2 boxes, each pulling under 60 watts → However many requests he sends, the monthly bill is $0, only electricity → The first box starts the answer in under a second → The second writes text at about 10 tokens a second → One request at a time: 838 tokens in 85 seconds, first word in 0.8s → Two requests at once: 697 tokens in 108 seconds, first word in 0.7s → Both boxes sit at 96% load and warm up to 76-78 degrees And only when a chip in a box runs hotter than 80 degrees or the cable between the 2 Sparks drops data does the system call the owner. And when he himself is out on a run or in a coffee shop, he still reaches his own model at home from his phone: sends a big prompt to the local Qwen3-235B, gets the full answer back in under a minute and a half, with no token meter ticking and no limit to hit. Here is what the test shows on his screen during one of the night runs: "one request at a time: 838 tokens in 84.9 seconds, first word in 0.8s, then 0.1s per token." "two requests at once: 697 tokens in 107.6 seconds, first word in 0.7s, then 0.15s per token." "Box 1: chip at 96% load, 76 degrees, 56 watts, 73GB used in memory." "Box 2: chip at 96% load, 78 degrees, 56 watts, the Qwen3-235B model fully loaded." And while everyone around is paying for AI by the month and bumping into limits, his top-tier model just sits on the desk and works as much as he wants: his own little power plant instead of a forever meter. He has no server rack of his own and no cloud account behind it. Just 2 DGX Spark boxes on a desk, one model split in half between them, one local address, and a folder of prompts next to it. Out of everything I have seen this year, this is the cleanest way to stop paying for AI: $5,998 of hardware on the desk once, $0 a month to the cloud, unlimited forever, and between them 2 gold boxes, 1 cable, and the full Qwen3-235B answering at home with no internet.

Blaze

93,871 次观看 • 4 个月前

FOUR DIFFERENT VENDORS ARE NOW SHIPPING GB10 MINI PCs WITH 128GB UNIFIED MEMORY, AND ONE MICROTIK CRS 804 SWITCH CAN CONNECT UP TO EIGHT OF THEM INTO A 1 TERABYTE LOCAL AI CLUSTER 00:00 he points at the MikroTik CRS 804, "you need some kind of switch that'll handle QSFP56 ports like these", the interconnect that makes the whole cluster possible the GB10 ecosystem is no longer just Nvidia. Dell Pro Max GB10, ASUS Ascent GX10, and MSI Edge Expert all ship the same Grace Blackwell Superchip with 128GB of coherent memory. same silicon, different cases, same 200 gigabit ports on the back the CRS 804 is what connects them at prosumer prices. four 400 gigabit QSFP56 ports on one 1U chassis, breakout cables that split each port into two 200 gigabit lanes. one switch drives eight GB10 units in parallel do the math. eight nodes at 128GB each equals 1024GB of pooled unified memory across the cluster. run vLLM, shard a frontier model across all eight, and inference happens locally on hardware that fits in half a rack the real limiter revealed in the stress test was never throttling. it was interconnect topology, exactly the layer this switch fixes at a fraction of enterprise switch pricing $400 a month for combined chatgpt pro and claude code max hits $4,800 a year per developer. a small team of five running through this cluster pays back inside eight months and never expires the article covers the buying ladder for a single desk. this post is proof of the cluster ladder that starts where the desk one ends save this before the GB10 lineup grows past four vendors and prosumer cluster switches move upmarket

NO1ennn

59,725 次观看 • 3 个月前

NVIDIA quietly built two desktop boxes that delete a $25,000/year AI subscription bill You don't rewrite your stack, you don't rent another data center, you just plug both into the wall and switch one line of code One looks like a deck of cards, the other like a hardback novel, together they replace ChatGPT Plus, Claude Pro, Cursor Pro, the OpenAI API meter, and every cloud GPU you were renting for fine-tunes It's built on the same CUDA stack the data center runs, which means once you migrate one workflow the rest follow on the same code path The reason NVIDIA shipped this is simple The bigger you scale on cloud AI, the harder you get taxed, and a one-person operator paying $2,100/month is producing exactly $0 of asset value at the end of every month And their solution is to skip the rental meter entirely, push inference back onto your desk, and let you loop agents overnight without watching a number tick on someone else's invoice This is much cheaper, faster, and pays itself back in 6 weeks for anyone already running AI for work But there is still a question nobody has answered yet, what happens when the next frontier model drops and your local 70B falls 6 months behind mid-quarter Also, technically a stack of four of the big box runs a 1.6 trillion parameter model on a desk for under $12,000 Even a fraction of that compute is more than most people will ever need in a year Bookmark this, it's worth coming back to when you have time 👇

ZEUS⚡️

118,817 次观看 • 4 个月前

The creator of High Bandwidth Memory (HBM) put a number on the AI build that should stop every infra investor cold. A cluster of a million GPUs runs at roughly 10-20% utilization (Save this). Kim Jung-ho spent thirty years building what feeds the GPU, and his claim is that the GPU is barely working. Here is what is actually happening. Every time a model generates output, the data has to be read out of memory, computed, and written back. The read and the write swallow almost the entire cycle. While that data moves, the GPU does nothing. It sits there, fully powered, fully paid for, waiting. By Kim's estimate the memory is doing only about 30 percent of the work it needs to do. The processor idles the rest. So a million installed GPUs run at 10 to 20 percent. You are not compute constrained. You are memory constrained, and the expensive part is standing around. Adding more GPUs does not fix this. It gives you more processors starving for the same data. Here is the part that decides the next decade. Memory can grow. When a cell cannot shrink any further, you stack it into a high-rise, layer on layer. A GPU cannot be stacked. It runs too hot and needs a cooler bolted to its back, so the one move that rescues memory is closed to the processor. The thing that can keep stacking compounds. The thing that cannot plateaus. The marginal dollar in an AI build now buys more by fixing the memory path than by bolting on another idle GPU. Which is why the companies that control memory bandwidth and supply are not suppliers to the AI trade. They are the AI trade.

Fireside Alpha

38,370 次观看 • 3 个月前