Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

MLX + OpenCode + Qwen3.5-122B-A10B-4bit on M3 Ultra created a great snake game! Work zero-shot. Video clearly in super fast mode during generation. I generated the prompt using Grok 4.20, it's in the article.

75,087 Aufrufe • vor 7 Monaten •via X (Twitter)

15 Kommentare

Profilbild von Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅvor 7 Monaten

I've used this settings specific for coding: Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0

Profilbild von Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅvor 7 Monaten

This command to start mlx_lm server: mlx_lm.server --model mlx-community/Qwen3.5-122B-A10B-4bit --temp 0.7 --top-p 0.95 --top-k 20 --min-p 0.0

Profilbild von Roy Jossfolk Jr.
Roy Jossfolk Jr.vor 7 Monaten

I know it is just a benchmark but don't you think these models have this type of application memorized by now? I use local models every day and while they can pop out snake games like nothing, they struggle with building real production software without supervision of a cloud model. If you try to build something complex or intricate, you can see there is still a ways to go. Likely 1-3 more iterations. Wondering your thoughts and what is the most complex software you've built with a local model?

Profilbild von Ivan Fioravanti ᯅ
Ivan Fioravanti ᯅvor 7 Monaten

I bet there is something wrong with the 10B, it’s too slow.

Profilbild von Martin Szerment | Practical AI
Martin Szerment | Practical AIvor 7 Monaten

Wild to see Grok feeding MLX and Qwen like this—stack synergy is getting real.

Profilbild von Sven Nachtzeit
Sven Nachtzeitvor 7 Monaten

Apple Silicon shows just 14% utilization running Qwen3.5-122B - does this hold at larger batch sizes? Efficient demo.

Profilbild von La Voix du Chat Artiste
La Voix du Chat Artistevor 7 Monaten

Can’t wait Mac mini m5

Profilbild von RM
RMvor 7 Monaten

Tried MLX 4b in lmstudio and it was unusable on m4 pro. Very slow. 40!seconds waiting for reply.

Profilbild von M.en
M.envor 7 Monaten

So that's how you get a snake to move at light speed. Meanwhile my code still thinks "zero-shot" means missing the target completely. 🐍

Profilbild von kaiyes
kaiyesvor 7 Monaten

ollama needs to bring in that 27B model fast ! there are conmunity ones though

Profilbild von Donny Li
Donny Livor 7 Monaten

Great zero shot result on M3 Ultra, especially with 4 bit weights. Reporting token throughput, compile warmup time, and memory pressure would make the setup much easier to compare against server GPUs.

Profilbild von Sean
Seanvor 7 Monaten

@cherry_cc12 a snake game? seriously? 😐

Profilbild von Denis Wachter
Denis Wachtervor 7 Monaten

Zero-shot game generation with a 122B model on local hardware is wild. The M3 Ultra + MLX combination is quietly becoming the best local AI dev setup available.

Profilbild von Boyuan (Nemo) Chen
Boyuan (Nemo) Chenvor 7 Monaten

been running qwen MoE models on M4 max and they're surprisingly usable for agent loops. what's the token/s like on M3 ultra with the 4bit quant? 122B with only 10B active feels like the right size for local inference

Profilbild von Alan
Alanvor 7 Monaten

neat can you test the 33B3A also

Ähnliche Videos