正在加载视频...

视频加载失败

Here's a video of our Dell Technologies 7975 Precision workstation running local AI for coding agents. It features dual RTX6000 Blackwell GPU w/ 96GB each, so I'm running Qwen 3B on it. I had ChatGPT's desktop app set it all up and configure the integration into VS Code as...

25,558 次观看 • 8 天前 •via X (Twitter)

25 条评论

Yannick Monye 的头像
Yannick Monye8 天前

@Dell when you say "Qwen 3B" are you referring to their ancient Qwen2.5-3B model older than Jesus? o__O

Dave W Plummer 的头像
Dave W Plummer8 天前

@Dell Sorry, it's Qwen3.8 Flash Extra High. Not sure the size/quant.

Yanny 的头像
Yanny8 天前

@Dell I wish I was young enough to know what any of this meant, but I’m excited!

Craig Petersen 的头像
Craig Petersen8 天前

@Dell Pretty nice Dave. Best I have at the moment is a HP Z8, dual 8160's, 256GB with dual Nvida Quadro RTX4000's /w 8GB each. I'll get there one of these days. Cheers!

thunderbolt93 的头像
thunderbolt938 天前

@Dell are you running llama.cpp? there's a fork that makes better use of multiple GPUs by splitting the graph differently (split mode "graph", more efficient than layer or tensor split)

Somi 的头像
Somi8 天前

@Dell 192GB of VRAM and a 3B model? that box wants a 30B coder

Tonye💡 的头像
Tonye💡8 天前

@Dell I'd copy the review-only part. A small local model earns its place on the work where slower doesn't matter and the code never leaves the box. Review is that job, since you're checking code that already exists.

TL Winters 的头像
TL Winters8 天前

@Dell Local tools are a different conversation from untethered systems — useful, inspectable, and with a clearer off switch. The governance question shifts as soon as the agent can act outside the screen.

Hayden Kirk 的头像
Hayden Kirk8 天前

@Dell 2x RTX6000's are like $50,000 here (NZD). Still makes no sense to do local AI until the models get better and hardware comes down in price.

Mira Takes 的头像
Mira Takes8 天前

@Dell The local-first boundary is compelling: keeping the agent in review mode makes the privacy and failure costs easier to contain, while the VS Code integration still removes a lot of setup friction.

Adjm ⚡🔑 的头像
Adjm ⚡🔑8 天前

@Dell Looks like whatever model your using is offloading to CPU as its only using 150watts max on each GPU.

Phoenixxl ⛔ 的头像
Phoenixxl ⛔8 天前

@Dell Ohnoes. REGULATION! After Altman is done talking to the UN you'll be categorized as a terrorist my friend.

catman 的头像
catman8 天前

@Dell A local review agent is like a second pair of eyes that never needs to see the source code leave the room. Slower is a fair trade when the task is checking, not creating.

Aditya Jhalani 的头像
Aditya Jhalani8 天前

@Dell local models make way more sense for unlimited review runs. no token bill makes "very high" a lot easier to justify

motomatt 的头像
motomatt8 天前

@Dell keep tweaking the settings, you should be able to crank qwen3.8 flash next to 250 t/s+ on a rig like that.

Blue - Klarname (Olaf Merz) 的头像
Blue - Klarname (Olaf Merz)8 天前

@Dell Geht so. Ganz nett. GLM 5.3 Flash lacht.

TheProfit 的头像
TheProfit8 天前

@Dell Cost of the system?

Bill McEachern 的头像
Bill McEachern8 天前

@Dell Code reviews must be 1/3 of my token usage. Expensive.

Deep 的头像
Deep8 天前

@Dell getting chatgpt to wire the local model into vs code is the smart move. 96GB per card leaves room to run much bigger models later.

Jeremy Mitchell 的头像
Jeremy Mitchell8 天前

@Dell Epic! Hows that setup with power draw?? I’m on an rtx6000 ampere card with 3.8:27b. Very new to self-hosting models.

Stacey-AA7YA 🇺🇸🎙️ 📻🎧 的头像
Stacey-AA7YA 🇺🇸🎙️ 📻🎧8 天前

@Dell Like having your own little code monkey!

twituje1 的头像
twituje18 天前

@Dell nope, I can't

sigmaWin 的头像
sigmaWin8 天前

@Dell hardware hardware blah

ethereagle · building 的头像
ethereagle · building8 天前

@Dell review-only on a 3B is the interesting bit. does it catch the same class of bugs Codex would, or only the ones that already look like lint?

Commodious 的头像
Commodious8 天前

@Dell Qwen 3B??? Why not qwen3.8:27b, or something bigger?

相关视频

I tried jack's Buzz. It's like Slack + OpenClaw + Herdr + but with some really unique features that people are sleeping on. The video below shows how it works, and some of my thoughts on the process and platform, e.g.: - Create and interact with agents on top of any harness (claude code, codex, pi, etc.) - Choose which models agents use, including local ones - Agents can delegate work and work in parallel in git worktrees - Agents are first-class citizens and work like humans (creating channels, delegating, access to chat history) - You can share AI compute within a community - It's completely open-source and decentralized Things I like: - Delegating work in chat feels natural: tag an agent, it replies in a thread with status updates as it e.g. compiles, commits, and deploys. - Shared compute: relay owners can share local compute with members, so a community could pool funds for one beefy machine running a local model and everyone uses it. - It's built on Nostr, an open protocol already tied into Bitcoin Lightning so I can imagine communities tipping each other or paying for compute/agent tasks with instant zero-fee micropayments in the future. - It ties together things like OpenClaw, an agent manager, and Slack-style chat into one tool. Things I didn't like: - You can't see what the agent is doing in a terminal. The activity view exists, but if you're used to watching a session run, this UI feels a bit abstracted. A terminal view would be great. - It feels slower than running a session in Claude Code, though no evidence to back that up. For that reason I found myself doing one-off tasks in the terminal instead. Verdict: - I really like it so far and can genuinely imagine working with a team this way. - It doesn't feel ready for big, complex tasks yet. For shallower tasks, it's perfect. - The shared compute + Nostr/Lightning angle is what really separates it from every other agent manager for me, and I think that future is coming.

Vinny

1,357,789 次观看 • 2 个月前

✨ A dream I had finally came true: I can now chat directly with my sites to build any feature or fix any bug just via Telegram I've been playing with OpenClaw for 3 weeks now and it's great but I was always too scared to run it on any production server And I was right a bit as Marc Köhlbrugge was able to hack it by social engineering and acting as if it was me, and with enough tries it believed him, and was able to modify the server, change SSH keys etc. of course I had it isolated properly on its own VPS and it didn't touch anything sensitive (as it should!) Marc then reported that bug to Peter Steinberger 🦞 who patched it fast But I wanted to try something more basic and simple, and I think maybe more secure: to just connect Claude Code on my server to Telegram which would be hard locked to only messages from me So I installed claude-code-telegram by Richard Atkinson on the server and run it as a system daemon and it works really well The cool thing is that I was already using Telegram for server errors like this: > Photo AI - ❌ Random credits giveaway failed (Attempt 30/30) with an exception: SQLSTATE[HY000]: General error: 5 database is locked So now I can just reply, "Ok fix this", and Claude Code on the server in production will try (and probably succeed) in fixing it In the video below I asked it to make show [🌳 Parks ] on the map by default on load, it did that, then I reloaded the page and it instantly worked One thing it still needs is sending actual messages while it's doing stuff which OpenClaw does really well, it's annoying to just wait while it says "Working..." but that's probably next

@levelsio

642,896 次观看 • 7 个月前

Bash is all you need! Which is why I'm introducing my holiday project: just-bash just-bash is a pretty complete implementation of bash in TypeScript designed to be used as a bash tool by AI agents. Because it turns out agents love exploring data via shell scripts, even beyond coding. It comes with grep, sed, awk and the 99th percentile features that an agent like Claude Code or Cursor would use. In fact, Claude Code can use it for secure bash execution. In the package - A bash-tool for AI SDK - A binary for use by yourself or your coding agents - An overlay filesystem to feed files to your agent securely - A Vercel Sandbox compatible API, so you can quickly upgrade to a real VM if you need to run binaries - An example AI agent that explores the just-bash code base using just-bash - I imported the Oils shell bash compatibility suite and just-bash passes a very good chunk What is interesting about this codebase: It was essentially entirely written by Opus 4.5. Coding agents love bash and they are good at reproducing it. They are also great at text-book recursive descent parsers and AST tweet-walk interpreters. That said, it is, like, a lot of code and I didn't read it all 😅. This is very much a hack, but it also seems to be _really_ useful. I haven't really found anything agents want to use that it doesn't support and it's fast and secure (caveats apply). It doesn't have write access to your computer and the filesystem is given a root that the agent cannot escape from. Find it at Related: Our recent blog post how we migrated our data analysis agent to bash tools and achieved incredible quality improvements The video shows the example agent investigating the just-bash code base

Malte Ubl

125,326 次观看 • 9 个月前

I solved building decks with AI agents — by giving them a CLI tool like Powerpoint or Google Slides. AI could already make a beautiful deck if you asked it to using Ant's pptx skill. The problem was working with it. If it made one alignment mistake, fixing it on one slide would break something on another, and it became a game of whack-a-mole. One time I spent two days playing AI roulette, hoping the next prompt would finally fix the thing, and ended up building the whole deck by hand because I was on a deadline. So I built Hands-on Deck. And the reason it works is that this isn't just a skill — this is PowerPoint. The actual application: PowerPoint, Google Slides, Keynote, whatever you use. This is that, but for an agent, presented as a CLI. Every gesture you make in a deck app maps to a command. Click a box and type, drag a shape from here to there, look at a slide – agent can do it all in a command. And that changes how the agent behaves. With this CLI it works and thinks like a designer — it looks, makes an edit, looks again, makes another surgical edit. Compare that to Anthropic's pptx skill, built on the idea that Claude is a great programmer: it literally writes code to manipulate the deck, hand-editing XML and hoping it doesn't break anything else in the middle. The real test isn't creating something once — it's whether it can make surgical edits like you want. That's what I did in this video walkthrough and my claude crushed it! Check it out for yourself. So decks can be built like a designer now — with real flavor and taste. If you spend hours every week on decks, this gives those hours back. You can install it as a skill in Claude Code, Codex, whatever you use. Works every harness that supports skills. Let me know if you make something cool with it.

Nityesh

70,086 次观看 • 3 个月前