Loading video...

Video Failed to Load

Go Home

MLX + OpenCode + Qwen3.5-122B-A10B-4bit on M3 Ultra created a great snake game! Work zero-shot. Video clearly in super fast mode during generation. I generated the prompt using Grok 4.20, it's in the article.

75,087 views • 7 months ago •via X (Twitter)

15 Comments

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ7 months ago

I've used this settings specific for coding: Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ7 months ago

This command to start mlx_lm server: mlx_lm.server --model mlx-community/Qwen3.5-122B-A10B-4bit --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0.0

Roy Jossfolk Jr.'s profile picture
Roy Jossfolk Jr.7 months ago

I know it is just a benchmark but don't you think these models have this type of application memorized by now? I use local models every day and while they can pop out snake games like nothing, they struggle with building real production software without supervision of a cloud model. If you try to build something complex or intricate, you can see there is still a ways to go. Likely 1-3 more iterations. Wondering your thoughts and what is the most complex software you've built with a local model?

Ivan Fioravanti ᯅ's profile picture
Ivan Fioravanti ᯅ7 months ago

I bet there is something wrong with the 10B, it’s too slow.

Martin Szerment | Practical AI's profile picture
Martin Szerment | Practical AI7 months ago

Wild to see Grok feeding MLX and Qwen like this—stack synergy is getting real.

Sven Nachtzeit's profile picture
Sven Nachtzeit7 months ago

Apple Silicon shows just 14% utilization running Qwen3.5-122B - does this hold at larger batch sizes? Efficient demo.

La Voix du Chat Artiste's profile picture
La Voix du Chat Artiste7 months ago

Can’t wait Mac mini m5

RM's profile picture
RM7 months ago

Tried MLX 4b in lmstudio and it was unusable on m4 pro. Very slow. 40!seconds waiting for reply.

M.en's profile picture
M.en7 months ago

So that's how you get a snake to move at light speed. Meanwhile my code still thinks "zero-shot" means missing the target completely. 🐍

kaiyes's profile picture
kaiyes7 months ago

ollama needs to bring in that 27B model fast ! there are conmunity ones though

Donny Li's profile picture
Donny Li7 months ago

Great zero shot result on M3 Ultra, especially with 4 bit weights. Reporting token throughput, compile warmup time, and memory pressure would make the setup much easier to compare against server GPUs.

Sean's profile picture
Sean7 months ago

@cherry_cc12 a snake game? seriously? 😐

Denis Wachter's profile picture
Denis Wachter7 months ago

Zero-shot game generation with a 122B model on local hardware is wild. The M3 Ultra + MLX combination is quietly becoming the best local AI dev setup available.

Boyuan (Nemo) Chen's profile picture
Boyuan (Nemo) Chen7 months ago

been running qwen MoE models on M4 max and they're surprisingly usable for agent loops. what's the token/s like on M3 ultra with the 4bit quant? 122B with only 10B active feels like the right size for local inference

Alan's profile picture
Alan7 months ago

neat can you test the 33B3A also

Related Videos