正在加载视频...

视频加载失败

MLX + OpenCode + Qwen3.5-122B-A10B-4bit on M3 Ultra created a great snake game! Work zero-shot. Video clearly in super fast mode during generation. I generated the prompt using Grok 4.20, it's in the article.

75,087 次观看 • 7 个月前 •via X (Twitter)

15 条评论

Ivan Fioravanti ᯅ 的头像
Ivan Fioravanti ᯅ7 个月前

I've used this settings specific for coding: Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Ivan Fioravanti ᯅ 的头像
Ivan Fioravanti ᯅ7 个月前

This command to start mlx_lm server: mlx_lm.server --model mlx-community/Qwen3.5-122B-A10B-4bit --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0.0

Roy Jossfolk Jr. 的头像
Roy Jossfolk Jr.7 个月前

I know it is just a benchmark but don't you think these models have this type of application memorized by now? I use local models every day and while they can pop out snake games like nothing, they struggle with building real production software without supervision of a cloud model. If you try to build something complex or intricate, you can see there is still a ways to go. Likely 1-3 more iterations. Wondering your thoughts and what is the most complex software you've built with a local model?

Ivan Fioravanti ᯅ 的头像
Ivan Fioravanti ᯅ7 个月前

I bet there is something wrong with the 10B, it’s too slow.

Martin Szerment | Practical AI 的头像
Martin Szerment | Practical AI7 个月前

Wild to see Grok feeding MLX and Qwen like this—stack synergy is getting real.

Sven Nachtzeit 的头像
Sven Nachtzeit7 个月前

Apple Silicon shows just 14% utilization running Qwen3.5-122B - does this hold at larger batch sizes? Efficient demo.

La Voix du Chat Artiste 的头像
La Voix du Chat Artiste7 个月前

Can’t wait Mac mini m5

RM 的头像
RM7 个月前

Tried MLX 4b in lmstudio and it was unusable on m4 pro. Very slow. 40!seconds waiting for reply.

M.en 的头像
M.en7 个月前

So that's how you get a snake to move at light speed. Meanwhile my code still thinks "zero-shot" means missing the target completely. 🐍

kaiyes 的头像
kaiyes7 个月前

ollama needs to bring in that 27B model fast ! there are conmunity ones though

Donny Li 的头像
Donny Li7 个月前

Great zero shot result on M3 Ultra, especially with 4 bit weights. Reporting token throughput, compile warmup time, and memory pressure would make the setup much easier to compare against server GPUs.

Sean 的头像
Sean7 个月前

@cherry_cc12 a snake game? seriously? 😐

Denis Wachter 的头像
Denis Wachter7 个月前

Zero-shot game generation with a 122B model on local hardware is wild. The M3 Ultra + MLX combination is quietly becoming the best local AI dev setup available.

Boyuan (Nemo) Chen 的头像
Boyuan (Nemo) Chen7 个月前

been running qwen MoE models on M4 max and they're surprisingly usable for agent loops. what's the token/s like on M3 ultra with the 4bit quant? 122B with only 10B active feels like the right size for local inference

Alan 的头像
Alan7 个月前

neat can you test the 33B3A also

相关视频