正在加载视频...
视频加载失败
Here's a video of our Dell Technologies 7975 Precision workstation running local AI for coding agents. It features dual RTX6000 Blackwell GPU w/ 96GB each, so I'm running Qwen 3B on it. I had ChatGPT's desktop app set it all up and configure the integration into VS Code as... show more
25,558 次观看 • 8 天前 •via X (Twitter)
25 条评论

@Dell when you say "Qwen 3B" are you referring to their ancient Qwen2.5-3B model older than Jesus? o__O

@Dell Sorry, it's Qwen3.8 Flash Extra High. Not sure the size/quant.

@Dell I wish I was young enough to know what any of this meant, but I’m excited!

@Dell Pretty nice Dave. Best I have at the moment is a HP Z8, dual 8160's, 256GB with dual Nvida Quadro RTX4000's /w 8GB each. I'll get there one of these days. Cheers!

@Dell are you running llama.cpp? there's a fork that makes better use of multiple GPUs by splitting the graph differently (split mode "graph", more efficient than layer or tensor split)

@Dell 192GB of VRAM and a 3B model? that box wants a 30B coder

@Dell I'd copy the review-only part. A small local model earns its place on the work where slower doesn't matter and the code never leaves the box. Review is that job, since you're checking code that already exists.

@Dell Local tools are a different conversation from untethered systems — useful, inspectable, and with a clearer off switch. The governance question shifts as soon as the agent can act outside the screen.

@Dell 2x RTX6000's are like $50,000 here (NZD). Still makes no sense to do local AI until the models get better and hardware comes down in price.

@Dell The local-first boundary is compelling: keeping the agent in review mode makes the privacy and failure costs easier to contain, while the VS Code integration still removes a lot of setup friction.

@Dell Looks like whatever model your using is offloading to CPU as its only using 150watts max on each GPU.

@Dell Ohnoes. REGULATION! After Altman is done talking to the UN you'll be categorized as a terrorist my friend.

@Dell A local review agent is like a second pair of eyes that never needs to see the source code leave the room. Slower is a fair trade when the task is checking, not creating.

@Dell local models make way more sense for unlimited review runs. no token bill makes "very high" a lot easier to justify

@Dell keep tweaking the settings, you should be able to crank qwen3.8 flash next to 250 t/s+ on a rig like that.

@Dell Geht so. Ganz nett. GLM 5.3 Flash lacht.

@Dell Cost of the system?

@Dell Code reviews must be 1/3 of my token usage. Expensive.

@Dell getting chatgpt to wire the local model into vs code is the smart move. 96GB per card leaves room to run much bigger models later.

@Dell Epic! Hows that setup with power draw?? I’m on an rtx6000 ampere card with 3.8:27b. Very new to self-hosting models.

@Dell Like having your own little code monkey!

@Dell nope, I can't

@Dell hardware hardware blah

@Dell review-only on a 3B is the interesting bit. does it catch the same class of bugs Codex would, or only the ones that already look like lint?

@Dell Qwen 3B??? Why not qwen3.8:27b, or something bigger?
