Loading video...
Video Failed to Load
single RTX 3090. 24 GB VRAM. Qwen3.5-35B-A3B. 4-bit quant, 113 tokens per second at full 262K context harnessing Claude Code locally with no API, no subscription, no proxy. told it what it is. 30 Mamba2 layers, 10 attention, 256 experts, 8 active per token. said "build something that shows... show more
110,464 views • 7 months ago •via X (Twitter)
45 Comments

this model won't stop. hit a Three.js bug and wrong API call. found it, fixed it, kept going.

RTX 5070, 12GB VRAM + DDR5 partial offload. Last night: refactored 17,800 lines → 1,127. Production code. Enterprise SaaS. From a camper. this is acceleration. 🫡

It's an amazing model, i run it on cheap hardware with 131k context and still get 20tok/s. Image modality is also very good! It does think a lot though, but in this case it might be good.

what GPU are you running it on?

hehehe please don't laugh! A tesla p4, 8gb vram. Pascal generation.

@sudoingX what settings are you using for that? I also have an 8gb vram card and I thought I'd have to sit this one out

@sudoingX I compiled llama.cpp and use a pretty plain command to launch llama-server. Llama is good for figuring out how many layers on gpu/cpu I’ll post more details later

@Laythe_li_suwi @sudoingX please release specs I would love more info. That's amazing

@Laythe_li_suwi @sudoingX I don't know what is standard practice for specs, i can tell you i have a i7-8700 with 64gb of ddr4, a tesla p4. I run a ubuntu VM, compiled llama.cpp, and run with those args (unsloth reccomendations). Had it guess the movie from a large 1080p still and it got it

@Laythe_li_suwi @sudoingX take note i am running this at 131k context, which is mind boggling considering my hardware... I am trying coding with opencode and, it is a bit slow compared to APIs but it is definetely useable!!!

113 tok/s at full 262K on a single 3090 is wild. how do the mamba2 layers hold up on code though? i'd expect them to struggle with long-range dependencies vs pure attention. 256 experts / 8 active is a nice ratio for keeping latency flat

Qwen/Llama.cpp are making LLMs entry level. Props. Now time to stock up on gpus and ram.

Can you please show how you configure Claude Code

writing up the updated setup soon. llama.cpp merged native Anthropic API support so the stack is even simpler now. no proxy needed. stay tuned.

This is wild throughput for full 262k context on a 3090. If you have logs, would love to see latency split (prefill vs decode) across context lengths. That’s usually where “flat line” claims break — super interesting that yours didn’t.

Same.

@BobSummerwill Qwen-3 with OpenCode is working pretty good for me on an old Nvidia Tesla P40 with 24 gb vram

@BobSummerwill P40 is a 2016 card. the fact this model runs on 8 year old hardware says everything about where MoE architecture is heading. what tok/s are you getting?

@patomation @BobSummerwill I have a spare Tesla K80 I keep wanting to rig up somewhere. My main setup is x2 3090s so I am loving your research here. Have you optimized the 27B? I like MOEs too, but the 27B looks like the truth.

@sudoingX @BobSummerwill I have used 32b models that are quantized to 4 I think a 27b model quantized will be ok

Still testing Qwen3-Coder-Next on a DGX Spark with vllm. It's insane. "Overloaded" it actually with MCP-tools as I have dozens in MCPHub and it just doesn't care. Always picks the right one. Fast, reliably. This is ChatGPT-Quality from 6 months ago for free.

You inspired me. I am running the 27b version and getting horny!

4060 ti 16 gb can run 41 tokens/sec with the settings below

Building complex models locally can be a game changer, I've seen significant speedups with Claude Code on my own projects. Saying "build something" and seeing what happens is often the most exciting part. Usually leads to some surprising discoveries.

Great coverage. Thanks! I've used Aider with Qwen 3 coder and also claude code with its $20 sub as well as configured for local LLM. Do you think Claude Code is better than other agents for local LLM or it's the LLM that matters?

What quant are you using? I tried unsloth's UD-Q4_K_XL with llama.cpp master, but it gets stuck in infinite thinking loop on a simple "hi" (and it seems I'm not alone). Should be an amazing model - in theory.

Cool! Can you show the code?

Are you using the mxfp4 model? Got the abliterad mxfp4 version, with q4 kv cache, seems to run on 24GB VRAM

@grok suggest hardware for this, preferrably under $1500

@Alibaba_Qwen needs to see this. Great work
Ok I have to ask, how the hell are you even loading this with 262k context and getting that kind of performance? It should require like 80 gigs of VRAM to load it all. I have a 3090 also and cap out at like 24k tokens before it bleeds into disk or regular memory.

Ohh my god 😍😍😍

How does someone install this on a PC locally so Open Claw can use it?

This is a really nice example

This is absolutely essential for incredible

its now likely that openclaw/tinyclaw local-cloud multi agent autonomous coding bots could cause financial crisis in the cloud giants businesse plans. Its now looking very iffy that the projected cloud AI market the tech titans are banking on exists with the emerging power of local models. Correct?

So sick. I’m jealous. I need more power.

I can't do this on my 4090, sad - 32K — 130 tok/s, 4.7 GB free. Full speed. - 65K — 127 tok/s, 548 MB free. Max practical on your 4090. - 262K — 4 tok/s, 417 MB free. GPU memory-starved, unusable.

I'm getting nowhere that speed on a 3090 lmstudio, more like 20tk/s.

@grok What's the best software setup to load this model an use it in VSCode or Claude Code?

Same

Fuck maybe I’ll keep my gear then

How's it compare to Qwen Coder Next for speed/quality/reasoning?

very nice!

Do you have setup configuration please?
