Loading video...
Video Failed to Load
MLX + OpenCode + Qwen3.5-122B-A10B-4bit on M3 Ultra created a great snake game! Work zero-shot. Video clearly in super fast mode during generation. I generated the prompt using Grok 4.20, it's in the article.
75,087 views • 7 months ago •via X (Twitter)
15 Comments

I've used this settings specific for coding: Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

This command to start mlx_lm server: mlx_lm.server --model mlx-community/Qwen3.5-122B-A10B-4bit --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0.0

I know it is just a benchmark but don't you think these models have this type of application memorized by now? I use local models every day and while they can pop out snake games like nothing, they struggle with building real production software without supervision of a cloud model. If you try to build something complex or intricate, you can see there is still a ways to go. Likely 1-3 more iterations. Wondering your thoughts and what is the most complex software you've built with a local model?

I bet there is something wrong with the 10B, it’s too slow.

Wild to see Grok feeding MLX and Qwen like this—stack synergy is getting real.

Apple Silicon shows just 14% utilization running Qwen3.5-122B - does this hold at larger batch sizes? Efficient demo.

Can’t wait Mac mini m5

Tried MLX 4b in lmstudio and it was unusable on m4 pro. Very slow. 40!seconds waiting for reply.

So that's how you get a snake to move at light speed. Meanwhile my code still thinks "zero-shot" means missing the target completely. 🐍

ollama needs to bring in that 27B model fast ! there are conmunity ones though

Great zero shot result on M3 Ultra, especially with 4 bit weights. Reporting token throughput, compile warmup time, and memory pressure would make the setup much easier to compare against server GPUs.

@cherry_cc12 a snake game? seriously? 😐

Zero-shot game generation with a 122B model on local hardware is wild. The M3 Ultra + MLX combination is quietly becoming the best local AI dev setup available.

been running qwen MoE models on M4 max and they're surprisingly usable for agent loops. what's the token/s like on M3 ultra with the 4bit quant? 122B with only 10B active feels like the right size for local inference

neat can you test the 33B3A also
