Video wird geladen...
Video konnte nicht geladen werden
I rebuilt much of my OpenClaw stack to run on local models. Getting this right is harder than it looks. I partnered with @NVIDIA_AI_PC to show you exactly how my hybrid local/hosted architecture works:
167,611 Aufrufe • vor 5 Monaten •via X (Twitter)
41 Kommentare

awesome 🔥

@NVIDIA_AI_PC Nice video, Matthew!

@NVIDIA_AI_PC thanks!

@NVIDIA_AI_PC This is the way.

@NVIDIA_AI_PC THat's so cool! You should invite @LyalinDotCom he did something similar (I copied him)

@NVIDIA_AI_PC local models can be a pain to get right. nice to see someone sharing the process.

@NVIDIA_AI_PC That 'harder than it looks' resonates. Balancing local and hosted models is a deep rabbit hole. Curious to dive into the architecture.

@NVIDIA_AI_PC The toughest part for me has just been switching between models. Things have been forgotten. But I think moving to using Obsidian will help.

@NVIDIA_AI_PC you need to 1) have a single place for memory and 2) have multiple prompt versions if you're swapping models for a single use case

@NVIDIA_AI_PC ✍️✍️ Always appreciate these drops

@NVIDIA_AI_PC thanks :)

@NVIDIA_AI_PC thank you! i think ollama allows this as well by running "ollama launch openclaw"

@NVIDIA_AI_PC Future is in local, small models. It's a myth you need expensive hardware. Useful models like Qwen 9b runs on medicore laptops or macs mini. :) For better results go with Qwen 3.5 27b or Gemma 4 31b.

@NVIDIA_AI_PC I don't get the local model move at this point other than for privacy reasons. Froniter models are just so much better.

@NVIDIA_AI_PC because you dont need frontier models for most things. and you save on cost, privacy, and sometimes latency.

If you set privacy aside, the other 99,9% of use cases, local llms are neither cheaper nor better. I'd love to be convinced otherwise, because I just bought a MacBook Pro M5 Max 128GB. But to me, a poweruser of OpenClaw for over two months, the whole local llm thing is yet to be proven ...

@NVIDIA_AI_PC My OpenClaw has been total 💩since we lost Anthropic. I’ve tried all of the open source models and all of them are worthless running OpenClaw.

@NVIDIA_AI_PC great thing appreciate

@NVIDIA_AI_PC Ty

@NVIDIA_AI_PC The hybrid local-hosted split is the only way to keep latency down. Smart move.

Literally watching this on YouTube RN and saw the post drop 😆 The Open Source community is so great. Every time a useful idea pops ups, thousands of people instantly jump on the idea and start testing it. Within days there are hundreds of resources and dozens of informative videos. Thankyou for staying on top of things and sharing your ideas and insights!

@NVIDIA_AI_PC local models are like meal prep, sounds great until you're debugging CUDA drivers at 2am. hybrid is the move: keep the easy stuff local, throw the chaos at the cloud

@NVIDIA_AI_PC Looking forward to apply all your wisdom you’ve been dropping , thanks @MatthewBerman

@NVIDIA_AI_PC splitting tasks between RTX for speed and cloud for context is the only way.

@garrytan @NVIDIA_AI_PC Me too 100% local for me

@NVIDIA_AI_PC Great video! I’ve been looking for info on how to run hybrid Openclaw workflows. ❤️ local LLMs

@NVIDIA_AI_PC Thanks for sharing. I have a bunch of AI video production (100% autonomous end to end) and I was thinking about trying to move it local. I did not realize how simple it is. Thanks again.

@NVIDIA_AI_PC @memexgarden summary

@hnshah @NVIDIA_AI_PC Try

@NVIDIA_AI_PC Fun topic rn great vid Matt

@NVIDIA_AI_PC The hybrid approach is underrated. Local for privacy/speed, hosted for heavy lifting. Best of both worlds.

@NVIDIA_AI_PC Would love to see how this applies to mobile applications...could really enhance user experience!

@NVIDIA_AI_PC Metasync remembers end to end ^ I don’t have those issues

@NVIDIA_AI_PC Hey there. Like OpenClaw but specializes in local config. Your feedback would be amazing.

Great video, Matt. For the budget minded, I would suggest a potentially different path to local models. For example, you still start with cloud frontier for building/experimentation as you suggest but then introduce a new middle step of using smaller cloud models for less complex tasks (e.g. Gemini Flash / Flash Lite; Claude Haiku, etc.) And then only move to local as a 3rd step. Especially for people who have weaker systems (e.g., Mac Minis with less RAM), this is a nice transition ahead of shelling out $3-$4k (or a lot more) for a workstation that can run capable local models.

@NVIDIA_AI_PC @MatthewBerman Please request a free DGX Spark from @NVIDIA_AI_PC - Love from Kenya

@NVIDIA_AI_PC finally someone admitting local models are the bottleneck instead of the solution just swap your hybrid setup for a single ai workflow and save the headache

@NVIDIA_AI_PC Local LLMs are the future.

@NVIDIA_AI_PC Hybrid is the move. Full local sounds cool until you're waiting 3 minutes for a response!

@NVIDIA_AI_PC What challenges did you face with local models, and how did you overcome them?

@NVIDIA_AI_PC top video about Nvidia, Matthew
