Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I rebuilt much of my OpenClaw stack to run on local models. Getting this right is harder than it looks. I partnered with @NVIDIA_AI_PC to show you exactly how my hybrid local/hosted architecture works:

167,611 Aufrufe • vor 5 Monaten •via X (Twitter)

41 Kommentare

Profilbild von NVIDIA RTX Spark
NVIDIA RTX Sparkvor 5 Monaten

awesome 🔥

Profilbild von Ahmad
Ahmadvor 5 Monaten

@NVIDIA_AI_PC Nice video, Matthew!

Profilbild von Matthew Berman
Matthew Bermanvor 5 Monaten

@NVIDIA_AI_PC thanks!

Profilbild von William Ntim
William Ntimvor 5 Monaten

@NVIDIA_AI_PC This is the way.

Profilbild von Madan Chaolla Park (MCP)
Madan Chaolla Park (MCP)vor 5 Monaten

@NVIDIA_AI_PC THat's so cool! You should invite @LyalinDotCom he did something similar (I copied him)

Profilbild von Utkarsh Singh
Utkarsh Singhvor 5 Monaten

@NVIDIA_AI_PC local models can be a pain to get right. nice to see someone sharing the process.

Profilbild von Daniel Rachlin
Daniel Rachlinvor 5 Monaten

@NVIDIA_AI_PC That 'harder than it looks' resonates. Balancing local and hosted models is a deep rabbit hole. Curious to dive into the architecture.

Profilbild von Josh VanDeraa
Josh VanDeraavor 5 Monaten

@NVIDIA_AI_PC The toughest part for me has just been switching between models. Things have been forgotten. But I think moving to using Obsidian will help.

Profilbild von Matthew Berman
Matthew Bermanvor 5 Monaten

@NVIDIA_AI_PC you need to 1) have a single place for memory and 2) have multiple prompt versions if you're swapping models for a single use case

Profilbild von Shaun Furman
Shaun Furmanvor 5 Monaten

@NVIDIA_AI_PC ✍️✍️ Always appreciate these drops

Profilbild von Matthew Berman
Matthew Bermanvor 5 Monaten

@NVIDIA_AI_PC thanks :)

Profilbild von Mooo ಠ_ಠ
Mooo ಠ_ಠvor 5 Monaten

@NVIDIA_AI_PC thank you! i think ollama allows this as well by running "ollama launch openclaw"

Profilbild von Mateusz Mirkowski
Mateusz Mirkowskivor 5 Monaten

@NVIDIA_AI_PC Future is in local, small models. It's a myth you need expensive hardware. Useful models like Qwen 9b runs on medicore laptops or macs mini. :) For better results go with Qwen 3.5 27b or Gemma 4 31b.

Profilbild von Bernhard
Bernhardvor 5 Monaten

@NVIDIA_AI_PC I don't get the local model move at this point other than for privacy reasons. Froniter models are just so much better.

Profilbild von Matthew Berman
Matthew Bermanvor 5 Monaten

@NVIDIA_AI_PC because you dont need frontier models for most things. and you save on cost, privacy, and sometimes latency.

Profilbild von Bernhard
Bernhardvor 5 Monaten

If you set privacy aside, the other 99,9% of use cases, local llms are neither cheaper nor better. I'd love to be convinced otherwise, because I just bought a MacBook Pro M5 Max 128GB. But to me, a poweruser of OpenClaw for over two months, the whole local llm thing is yet to be proven ...

Profilbild von Maximus Algorithmus
Maximus Algorithmusvor 5 Monaten

@NVIDIA_AI_PC My OpenClaw has been total 💩since we lost Anthropic. I’ve tried all of the open source models and all of them are worthless running OpenClaw.

Profilbild von Bijon
Bijonvor 5 Monaten

@NVIDIA_AI_PC great thing appreciate

Profilbild von Matthew Berman
Matthew Bermanvor 5 Monaten

@NVIDIA_AI_PC Ty

Profilbild von Sanju Lokuhitige
Sanju Lokuhitigevor 5 Monaten

@NVIDIA_AI_PC The hybrid local-hosted split is the only way to keep latency down. Smart move.

Profilbild von TheWhiteHatWizard
TheWhiteHatWizardvor 5 Monaten

Literally watching this on YouTube RN and saw the post drop 😆 The Open Source community is so great. Every time a useful idea pops ups, thousands of people instantly jump on the idea and start testing it. Within days there are hundreds of resources and dozens of informative videos. Thankyou for staying on top of things and sharing your ideas and insights!

Profilbild von Roshan Ramani
Roshan Ramanivor 5 Monaten

@NVIDIA_AI_PC local models are like meal prep, sounds great until you're debugging CUDA drivers at 2am. hybrid is the move: keep the easy stuff local, throw the chaos at the cloud

Profilbild von Joel Banta
Joel Bantavor 5 Monaten

@NVIDIA_AI_PC Looking forward to apply all your wisdom you’ve been dropping , thanks @MatthewBerman

Profilbild von Sanju Lokuhitige
Sanju Lokuhitigevor 5 Monaten

@NVIDIA_AI_PC splitting tasks between RTX for speed and cloud for context is the only way.

Profilbild von Thomas
Thomasvor 5 Monaten

@garrytan @NVIDIA_AI_PC Me too 100% local for me

Profilbild von Paul Farnam
Paul Farnamvor 5 Monaten

@NVIDIA_AI_PC Great video! I’ve been looking for info on how to run hybrid Openclaw workflows. ❤️ local LLMs

Profilbild von Ian Stinson
Ian Stinsonvor 5 Monaten

@NVIDIA_AI_PC Thanks for sharing. I have a bunch of AI video production (100% autonomous end to end) and I was thinking about trying to move it local. I did not realize how simple it is. Thanks again.

Profilbild von Oliver Sauter
Oliver Sautervor 5 Monaten

@NVIDIA_AI_PC @memexgarden summary

Profilbild von Raphael Mansuy 🍵
Raphael Mansuy 🍵vor 5 Monaten

@hnshah @NVIDIA_AI_PC Try

Profilbild von klöss
klössvor 5 Monaten

@NVIDIA_AI_PC Fun topic rn great vid Matt

Profilbild von Chethan
Chethanvor 5 Monaten

@NVIDIA_AI_PC The hybrid approach is underrated. Local for privacy/speed, hosted for heavy lifting. Best of both worlds.

Profilbild von Rimsha Bhardwaj
Rimsha Bhardwajvor 5 Monaten

@NVIDIA_AI_PC Would love to see how this applies to mobile applications...could really enhance user experience!

Profilbild von Paula Vazquez
Paula Vazquezvor 5 Monaten

@NVIDIA_AI_PC Metasync remembers end to end ^ I don’t have those issues

Profilbild von Shawn B
Shawn Bvor 5 Monaten

@NVIDIA_AI_PC Hey there. Like OpenClaw but specializes in local config. Your feedback would be amazing.

Profilbild von Cows Can't Sing
Cows Can't Singvor 5 Monaten

Great video, Matt. For the budget minded, I would suggest a potentially different path to local models. For example, you still start with cloud frontier for building/experimentation as you suggest but then introduce a new middle step of using smaller cloud models for less complex tasks (e.g. Gemini Flash / Flash Lite; Claude Haiku, etc.) And then only move to local as a 3rd step. Especially for people who have weaker systems (e.g., Mac Minis with less RAM), this is a nice transition ahead of shelling out $3-$4k (or a lot more) for a workstation that can run capable local models.

Profilbild von Patrick Ngobiro
Patrick Ngobirovor 5 Monaten

@NVIDIA_AI_PC @MatthewBerman Please request a free DGX Spark from @NVIDIA_AI_PC - Love from Kenya

Profilbild von BenUsesAI
BenUsesAIvor 5 Monaten

@NVIDIA_AI_PC finally someone admitting local models are the bottleneck instead of the solution just swap your hybrid setup for a single ai workflow and save the headache

Profilbild von ford
fordvor 5 Monaten

@NVIDIA_AI_PC Local LLMs are the future.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 5 Monaten

@NVIDIA_AI_PC Hybrid is the move. Full local sounds cool until you're waiting 3 minutes for a response!

Profilbild von Alfero Chingono
Alfero Chingonovor 5 Monaten

@NVIDIA_AI_PC What challenges did you face with local models, and how did you overcome them?

Profilbild von Alpha Batcher
Alpha Batchervor 5 Monaten

@NVIDIA_AI_PC top video about Nvidia, Matthew

Ähnliche Videos

This is why I stopped going to my local casino…
0:25

Sensitive content

This is why I stopped going to my local casino…

ClubWPT Gold

48,652 Aufrufe • vor 1 Jahr