正在加载视频...

视频加载失败

THIS DEVELOPER CONNECTED 8 NVIDIA DGX SPARKS INTO ONE CLUSTER - AND RAN AN 800GB MODEL THAT MADE HIM 10X MORE PRODUCTIVE 21:47 he says it straight - "this is a terabyte of VRAM - we ran Quen 3.5, 800GB on disk, a model that doesn't even fit on...

186,592 次观看 • 3 个月前 •via X (Twitter)

36 条评论

Jared Schnelle 的头像
Jared Schnelle3 个月前

Cool video but pretty shitty of you to not even link to the dude's YouTube channel or even say his name. You're just stealing content by re-uploading his video.

Sudo su 的头像
Sudo su3 个月前

this is the spark doing what it was built for. people see one desk box, but it ships ConnectX networking specifically so you can pool them. that's why 8 hit 24 tok/s when one does 3, near linear over RDMA, not duct tape. i run one, clustering a second is the plan. own the compute, don't rent it.

Crowdsource The Truth 的头像
Crowdsource The Truth3 个月前

I think the average person, and even many sophisticated users would have a difficult time setting this up properly.

Jeremy Combs 的头像
Jeremy Combs3 个月前

You should tag and credit @digitalix if you’re going to post his content. @nikitabier is this grounds for demonetization?

GrainStats 🌾 的头像
GrainStats 🌾3 个月前

i see @digitalix video with no attribution to them 🤔

leopardracer 的头像
leopardracer3 个月前

worth watching

Noisy 的头像
Noisy3 个月前

thanks leopard you're a legend

Noisy 的头像
Noisy3 个月前

video from:

Roman 的头像
Roman3 个月前

That's about $38,000. That's based on the current Amazon price of a dgx spark. + The estimate on the cables. The cables were running 100 but they were going up last time I bought one..

Legion 的头像
Legion3 个月前

How do you win the Spark?

karla ❀ 的头像
karla ❀3 个月前

With claude is easy running the developer

Noisy 的头像
Noisy3 个月前

yeah facts, we live in simple times

Mark L Devine 的头像
Mark L Devine3 个月前

Maybe add another 50% to that so that it has a ceiling to do more than just stand up? It doesn't make sense. 24 t/s with no headroom. Just a PoC.

🇫🇷 Fearil DiabloBros 🇷🇺🌿🇵🇸 的头像
🇫🇷 Fearil DiabloBros 🇷🇺🌿🇵🇸3 个月前

8 GPU RTX PRO 6000 Blackwell 768go Vram

James Aita 的头像
James Aita3 个月前

Oh yeah no big deal just 8sparks and an insane switch. Everyone has those lying around the house.

Cryton 的头像
Cryton3 个月前

The crazy part isn't the model size. It's that local AI hardware is starting to compete with infrastructure that used to require a cloud budget.

ext87 的头像
ext873 个月前

Still nearly a 2 year ROI at 2k/mo for cloud opex vs local capex, and this assumes zero other costs like power or other cloud costs

Ventry 的头像
Ventry3 个月前

that's just incredible power

henny 的头像
henny3 个月前

Truly one of the dumbest posts I’ve ever seen. $40,000 for <20 tps?

dunik 的头像
dunik3 个月前

it’s crazy that you can now build your own AI system right at home.

qurool 的头像
qurool3 个月前

it worth my time for sure

TigerAIHQ 的头像
TigerAIHQ3 个月前

"This is actually insane… 8 DGX Sparks clustered together running an 800GB model at 24 tokens/s? A few months ago something like this would only be possible on big cloud clusters. Now regular people are just wiring up hardware at home. The fact that he even used Claude to set up the networking and SSH makes it even crazier 😂

rewind 的头像
rewind3 个月前

home clusters are getting serious

Noisy 的头像
Noisy3 个月前

yeah a cluster of 4 chips is a really powerful thing

The Foxx 的头像
The Foxx3 个月前

I barely have $10, let alone $10,000.

Ben White 的头像
Ben White3 个月前

You can just use cloud Qwen unless you need like 10 billion tokens.

Solvaix 的头像
Solvaix3 个月前

Not cheap yet, but the fact it works locally is impressive

Yeab - engg / core was in Palo Alto🌴 的头像
Yeab - engg / core was in Palo Alto🌴3 个月前

cool

Avid 的头像
Avid3 个月前

Nice

BeanBee 🐝 的头像
BeanBee 🐝3 个月前

A model that was previously cloud-only is becoming locally accessible.

starmex 的头像
starmex3 个月前

bookmarked, terabyte of vram in one cluster is wild

Billionaire Protocol 的头像
Billionaire Protocol3 个月前

24 tps = dogshit

Ramirez Eltoro 的头像
Ramirez Eltoro3 个月前

Amazing

Amarule 的头像
Amarule3 个月前

The Future That Has Already Arrived🕒

Yarchi 的头像
Yarchi3 个月前

this is the last straw, im gonna buy it rn

PLN.ALGO 的头像
PLN.ALGO3 个月前

HE SPEND SUCH MONEY ON that, and i can live it for next 10 year for that cash..

相关视频

FOUR DIFFERENT VENDORS ARE NOW SHIPPING GB10 MINI PCs WITH 128GB UNIFIED MEMORY, AND ONE MICROTIK CRS 804 SWITCH CAN CONNECT UP TO EIGHT OF THEM INTO A 1 TERABYTE LOCAL AI CLUSTER 00:00 he points at the MikroTik CRS 804, "you need some kind of switch that'll handle QSFP56 ports like these", the interconnect that makes the whole cluster possible the GB10 ecosystem is no longer just Nvidia. Dell Pro Max GB10, ASUS Ascent GX10, and MSI Edge Expert all ship the same Grace Blackwell Superchip with 128GB of coherent memory. same silicon, different cases, same 200 gigabit ports on the back the CRS 804 is what connects them at prosumer prices. four 400 gigabit QSFP56 ports on one 1U chassis, breakout cables that split each port into two 200 gigabit lanes. one switch drives eight GB10 units in parallel do the math. eight nodes at 128GB each equals 1024GB of pooled unified memory across the cluster. run vLLM, shard a frontier model across all eight, and inference happens locally on hardware that fits in half a rack the real limiter revealed in the stress test was never throttling. it was interconnect topology, exactly the layer this switch fixes at a fraction of enterprise switch pricing $400 a month for combined chatgpt pro and claude code max hits $4,800 a year per developer. a small team of five running through this cluster pays back inside eight months and never expires the article covers the buying ladder for a single desk. this post is proof of the cluster ladder that starts where the desk one ends save this before the GB10 lineup grows past four vendors and prosumer cluster switches move upmarket

NO1ennn

59,725 次观看 • 3 个月前

This Chinese developer linked two $2,999 NVIDIA DGX Sparks into one box and runs the full Qwen3-235B at home, after dropping his $1,999-a-month cloud bill to zero. He wired 2 small boxes into a single computer, split a giant 235-billion-parameter model in half between them, and serves it across his own network at about 10 tokens a second, with no internet, no cloud, right there on the desk. No data center, no thousand-dollar graphics cards, no monthly cloud bill. Just him, 2 gold boxes the size of a sandwich, one cable between them, and 1 power strip. And here is the whole payoff. He used to pay the cloud $1,999 a month for the same model, and the meter ticked on every request. Now he paid $5,998 once for 2 boxes, they covered their cost in 3 months, and after that he sends as many requests as he wants for free, only electricity. The two Sparks talk over one fast cable, each holds 128GB of memory, and together they carry the whole model, about 73GB loaded per box, with the chip inside pinned near the limit at 96%. Both boxes work as one and keep trading data over the cable, with no cloud in the loop and no single word leaking out. The ready model sits on one local address, and any app on his network calls it as easily as ChatGPT. And here is how he described, in plain words, what this pair of boxes does: "this is a pair of boxes that holds the huge Qwen3-235B model and serves it to one network. the model is split in half, and each box owns its half. parts: // Box 1 (holds the first half of the model and starts the answer fast, the first word appears in under a second) // Box 2 (holds the second half and writes out the rest, about 10 tokens a second) // Cable (connects the 2 boxes and moves data between them on every step, with no lag) // Address (one local address where any app sends its request, like to a cloud model) // Test (a script that runs big prompts through and measures speed and delays) // Monitor (checks temperature, power draw, and load on both boxes every 2 seconds). the model never goes to the cloud. he only steps in when a box runs hotter than 80 degrees or the cable between them starts dropping data." So the system knows exactly what it is, what it is for, and where its limits are. It knows it has to hold the whole huge model across 2 boxes on its own. It knows it has to answer every request locally, with no meter, no limits, and no internet. It knows the human is only needed when a box overheats or the link between them stalls. → The setup runs around the clock on 2 boxes, each pulling under 60 watts → However many requests he sends, the monthly bill is $0, only electricity → The first box starts the answer in under a second → The second writes text at about 10 tokens a second → One request at a time: 838 tokens in 85 seconds, first word in 0.8s → Two requests at once: 697 tokens in 108 seconds, first word in 0.7s → Both boxes sit at 96% load and warm up to 76-78 degrees And only when a chip in a box runs hotter than 80 degrees or the cable between the 2 Sparks drops data does the system call the owner. And when he himself is out on a run or in a coffee shop, he still reaches his own model at home from his phone: sends a big prompt to the local Qwen3-235B, gets the full answer back in under a minute and a half, with no token meter ticking and no limit to hit. Here is what the test shows on his screen during one of the night runs: "one request at a time: 838 tokens in 84.9 seconds, first word in 0.8s, then 0.1s per token." "two requests at once: 697 tokens in 107.6 seconds, first word in 0.7s, then 0.15s per token." "Box 1: chip at 96% load, 76 degrees, 56 watts, 73GB used in memory." "Box 2: chip at 96% load, 78 degrees, 56 watts, the Qwen3-235B model fully loaded." And while everyone around is paying for AI by the month and bumping into limits, his top-tier model just sits on the desk and works as much as he wants: his own little power plant instead of a forever meter. He has no server rack of his own and no cloud account behind it. Just 2 DGX Spark boxes on a desk, one model split in half between them, one local address, and a folder of prompts next to it. Out of everything I have seen this year, this is the cleanest way to stop paying for AI: $5,998 of hardware on the desk once, $0 a month to the cloud, unlimited forever, and between them 2 gold boxes, 1 cable, and the full Qwen3-235B answering at home with no internet.

Blaze

93,871 次观看 • 4 个月前

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

344,670 次观看 • 1 个月前

single most useful thing for my local ai setup that made life easy is my 2x DGX Spark becoming one serving box for every device i own here is how i do it: > 1. the two sparks run one vLLM server together with one endpoint, i use dgx sparks, you can use any nodes that can run an llm > 2. install tailscale to have all your nodes and machines on the same tailnet, the endpoint lives there so every device reaches it from anywhere and no port is open to the internet > 3. that one tailnet endpoint with the exposed port is your base url for everything, it speaks OpenAI chat, OpenAI Responses and Anthropic Messages, so any chat app, coding agent or phone bot you point at it talks to the same model > 4. the server also answers to the name "local", so clients can ask for "local" instead of a model name and when i swap the model on the servers nothing on the devices needs a config change. right now it's serving GLM 5.3-Flash, next week it can be something else this video below is the whole thing running end to end, orange dots are requests going in, white and green are tokens coming back it takes however many requests you throw at it, your config decides how many run at once and the rest wait their turn, when several run together each request's tok/s drops a bit while the box moves more tokens in total. mine is set to one at a time right now, so when every device fires together they queue up this setup really improved my quality of life with local ai, if you get stuck anywhere setting it up leave a comment and i'll help you

Sudo su

32,379 次观看 • 9 天前