Loading video...
Video Failed to Load
IT WORKED!! Here's how you run Claude Code + Local LLM on your own machine in under 90 seconds
466,783 views • 5 months ago •via X (Twitter)
82 Comments

Just because you ran it doesn’t mean it worked. Claude code uses a lot of context tokens, you will need to find the perfect models with enough context windows

I showed the possibility. Now it's your turn to make your own Claude code.

@MoTheAgent Bold of you to assume the people that pass judgement are actually building anything let alone showing an ounce of appreciation. I will try this out and see what I can build on top , thank you.

@MoTheAgent Let’s go!!! Show me the sauce haha 😆

The harness is the easy part though. Where local models still fall apart is the ambiguous multi-step reasoning: 'refactor this module but don't break the tests I didn't write yet.' Would love to see someone benchmark real-world multi-file tasks, not just 90-second demos.

Haha maybe I should do it too

@boyuan_chen You think you can show me how it’s done ?

@boyuan_chen About how to set out one or

90 seconds to run it, 3 hours to debug the python path variables

Alternative with llama.cpp

Congrats! Which local LLM did you pair with Claude Code?

I have Qwen3.5 0.8B setup. However it's because I have the Mac mini with 16GB only.

@LCHHAB_9 What model would you recommend if someone had more resources?

@LCHHAB_9 Minimax 2.7? Time will change any new model could come out

Slow as fuck. Unless your setup at home is super duper powerful. I mean server level memory and gpu. Else you’re wasting your time.

Nah I have Mac mini with 0.8 qwen 3.5 has 50 tokens / sec Pretty fast

16 GB by the way

Awesome!!

Let's go! Time to set up

I am starting in an hour so definitely setting it up!

Can we put azure open ai api and use it?

For now I don't think Leon has that setup yet.

Is this better? I tried to use this initially but couldn’t run it

Great one, but the core essence is the model than the md instructions it routes to

😂😂 I agree. Now we need another spy. All they need to do is putting their model in other source maps.

Claude Code is open source now so you don't need to use some shady repos

🤣🤣🤣 I tried with the src code it's very buggy. Asked Gemini, Codex and Minimax they all said not able to reverse engineer it...

I think it’s the client side code

So this is what got leaked, Claude Code client and then they open-sourced it.

Nah it’s been open sourced that’s why we can download with npm package

Just type ollama launch claude , it will do similar work as this man suggested, why everyone going crazy about leaks it's just shell not llm.

Oh, it was the source code itself. And someone wrapped it into a shell. So you can edit the code right away and build your own Claude code.

What's the purpose of building own claude code without LLM ?? And if using open source self trained model or distilled model, then only it makes sense.

Giving people an option. You never know what others can build with this.

Kudos to those individuals in advance 👏.

💪💪

Any guide to do it with openclaw? I don’t what my openclaw to cost me few thousands for simply prompts

I built one before You just need to change the config to the current model. If you want to see like a local LLM for Ollama I might come out a new video. This time can be a screen recording haha 😆

It doesn't work anymore cuz the repo is gone. But now check out this to implement with Claw Code

Nothing knew here

Haha, the local models are not new. But the Claude code is. We can now build our own custom Claude code.

cool hack, but 90 seconds setup is irrelevant if your machine overheats constantly

😂😂 that’s out of the concept. Haha

The 90-second setup is what matters. I've tried so many local LLM tutorials that turn into 3-hour rabbit holes with dependency hell. Actually timing yourself forces you to write instructions that work for real people, not just the person who already has everything configured.

💪💪 glad you think it can help.. Time to cook some sauce

thank you bing chilling

💪💪

ask him to make an app to record your screen

😂😂 yo i will do screen recording next time haha

We could already run this long before it was leaked

We can run, but we can’t build. Now we can build above on the Claude code and make your own Claude code

Or you could just use the original and change the env var ANTHROPIC_BASE_URL to point to your local setup 🤷♂️

Yea for local model! But this time we can edit the Claude Code and make your own Claude Code 😂

we need its brain not its shell

😂😂 who doesn’t.. we just need another spy to jump in there

@ziwenxu_ that’s dope! I’ve been wanting to try this. 90 seconds is wild, can’t wait to give it a shot!

Let me know if it works !!

@ziwenxu_ will do! have you tried it out yet?

thank you bing chilling

Time to cook your own Claude code haha 😆

The idea is not new , it doesn't offer value ,it's just another approach for a repeated work. I wished the claude models to be free to use after the leak , but it turns out this is not possible

🤣🤣 Yo we need to send some spy in anthropic.

Hahahaha 😄 007

Nothing unique just put it on the source map again 🤣🤣

Je suis en train de configurer le mien sous win 7 > llama-b8555-bin-win-cpu-x64. Et quand j'entend parler des IA ou grok qui se prends pour une iA j'en pisse de rire ..

Llama is bad 🤣🤣

@grok is this Claude code or claw?? What’s he saying exactly?? Based on how powerful Claude code is can it even be run in a single device??

BUY A PHONE STAND CHEAPO!

will do!!

Hell yeah! Thanks for posting!

谢谢你

💪💪💪 no problem

@grok peux tu tester et me dire si ça marche

You can already use your own llm using environmental variables on the closed source version.

Yup, but this time I set up with the open source version.

Drop full bundle so we can plug in and play!

😂 full bundle in the video

You can even try Echo AI also, if you want another serious coding assistant to compare. Directly download VS Code extension or npm install -g echoai

still works?

Yea it’s only yesterday 😂

okay am a give it a try, thanks for the info

💪💪
