Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

MLX + OpenCode + Qwen3.5-122B-A10B-4bit on M3 Ultra created a great snake game! Work zero-shot. Video clearly in super fast mode during generation. I generated the prompt using Grok 4.20, it's in the article.

75,087 görüntüleme • 7 ay önce •via X (Twitter)

15 Yorum

Ivan Fioravanti ᯅ profil fotoğrafı
Ivan Fioravanti ᯅ7 ay önce

I've used this settings specific for coding: Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Ivan Fioravanti ᯅ profil fotoğrafı
Ivan Fioravanti ᯅ7 ay önce

This command to start mlx_lm server: mlx_lm.server --model mlx-community/Qwen3.5-122B-A10B-4bit --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0.0

Roy Jossfolk Jr. profil fotoğrafı
Roy Jossfolk Jr.7 ay önce

I know it is just a benchmark but don't you think these models have this type of application memorized by now? I use local models every day and while they can pop out snake games like nothing, they struggle with building real production software without supervision of a cloud model. If you try to build something complex or intricate, you can see there is still a ways to go. Likely 1-3 more iterations. Wondering your thoughts and what is the most complex software you've built with a local model?

Ivan Fioravanti ᯅ profil fotoğrafı
Ivan Fioravanti ᯅ7 ay önce

I bet there is something wrong with the 10B, it’s too slow.

Martin Szerment | Practical AI profil fotoğrafı
Martin Szerment | Practical AI7 ay önce

Wild to see Grok feeding MLX and Qwen like this—stack synergy is getting real.

Sven Nachtzeit profil fotoğrafı
Sven Nachtzeit7 ay önce

Apple Silicon shows just 14% utilization running Qwen3.5-122B - does this hold at larger batch sizes? Efficient demo.

La Voix du Chat Artiste profil fotoğrafı
La Voix du Chat Artiste7 ay önce

Can’t wait Mac mini m5

RM profil fotoğrafı
RM7 ay önce

Tried MLX 4b in lmstudio and it was unusable on m4 pro. Very slow. 40!seconds waiting for reply.

M.en profil fotoğrafı
M.en7 ay önce

So that's how you get a snake to move at light speed. Meanwhile my code still thinks "zero-shot" means missing the target completely. 🐍

kaiyes profil fotoğrafı
kaiyes7 ay önce

ollama needs to bring in that 27B model fast ! there are conmunity ones though

Donny Li profil fotoğrafı
Donny Li7 ay önce

Great zero shot result on M3 Ultra, especially with 4 bit weights. Reporting token throughput, compile warmup time, and memory pressure would make the setup much easier to compare against server GPUs.

Sean profil fotoğrafı
Sean7 ay önce

@cherry_cc12 a snake game? seriously? 😐

Denis Wachter profil fotoğrafı
Denis Wachter7 ay önce

Zero-shot game generation with a 122B model on local hardware is wild. The M3 Ultra + MLX combination is quietly becoming the best local AI dev setup available.

Boyuan (Nemo) Chen profil fotoğrafı
Boyuan (Nemo) Chen7 ay önce

been running qwen MoE models on M4 max and they're surprisingly usable for agent loops. what's the token/s like on M3 ultra with the 4bit quant? 122B with only 10B active feels like the right size for local inference

Alan profil fotoğrafı
Alan7 ay önce

neat can you test the 33B3A also

Benzer Videolar