Loading video...

Video Failed to Load

Go Home

Wait Xiaomi has also released the best 9B model in the world?! This model is based on Qwen-9B... And you can easily run it even on 8GB of memory! - MIT license - Multimodal (image + text) - Solid for coding/agentic tasks Perfect to use locally and offline... even...

82,735 views • 1 day ago •via X (Twitter)

24 Comments

Paul Couvert's profile picture
Paul Couvert1 day ago

Base model weights: Quantized GGUF:

Deepakraj Solanki's profile picture
Deepakraj Solanki1 day ago

Stop making wrong claims and that 8GB RAM thing is just so stupid about you

Paul Couvert's profile picture
Paul Couvert1 day ago

Fits using a Q4 version, especially if you're offloading the context in the ram

Nikolaus Kern's profile picture
Nikolaus Kern1 day ago

Amazing. Now let's see your tests on 8gb VRAM 🙃

Antifield's profile picture
Antifield1 day ago

我认为并不是,Ornith-1.5-9B在编码任务上明显更强

Mac's profile picture
Mac1 day ago

@grok will this work on my pc? I have rtx 3060 and 32 GB RAM

Jimmy Otis #TruthMatters #EndCronyism's profile picture
Jimmy Otis #TruthMatters #EndCronyism1 day ago

max-distilled fable.... nice.

Paul Couvert's profile picture
Paul Couvert1 day ago

Everyone is distilling from everyone... and even closed labs are integrating open research.

Jimmy Otis #TruthMatters #EndCronyism's profile picture
Jimmy Otis #TruthMatters #EndCronyism1 day ago

yes. IMO anthropic has been distilled most. Meta distills least from other models. A bit like inbreeding. We need more diversity!

Mark Nisin's profile picture
Mark Nisin1 day ago

This one is very good and high quality!

AI News Daily's profile picture
AI News Daily1 day ago

The 8GB angle is what makes this genuinely useful, not just impressive on a benchmark. I am curious how it holds up on longer agentic tasks once context and tool calls start competing for memory.

Fajar M Reza's profile picture
Fajar M Reza1 day ago

Eight-gigabyte multimodal inference makes privacy concrete, but leaves little context headroom.

Simon Bullows's profile picture
Simon Bullows1 day ago

@grok how would this compare to spark 2.5 4b in terms of performance and requirements? If I have 12bg if VRAM, what tokens a second could I expect?

Alex Radulescu 🇸🇬's profile picture
Alex Radulescu 🇸🇬1 day ago

Have you tested it?

Paul Couvert's profile picture
Paul Couvert1 day ago

On it. So far so good.

Crio Songo's profile picture
Crio Songo1 day ago

That's really impressive. Open-source, low-resource requirement, multi-modal, great for local deployment.

NotaDEV's profile picture
NotaDEV1 day ago

can it actually hold a long agentic loop or does it fall apart after step three

Quinn’s Neural Pathways's profile picture
Quinn’s Neural Pathways1 day ago

GPUs go burrr is optional now. A 9B MIT-licensed agentic model that runs offline on 8GB of unified memory is the actual news here.

Abhiram.N.S's profile picture
Abhiram.N.S1 day ago

I think bonsai 2 will still be on top due to the 27B representational space .

ఏది అయితే ఏంటి's profile picture
ఏది అయితే ఏంటి1 day ago

Will it work on non Mac

Snapolino's profile picture
Snapolino1 day ago

unfortunately the RL version is not yet released :-) only the SFT versino

Brjan | AI Builder's profile picture
Brjan | AI Builder1 day ago

how does it handle complex coding tasks compared to other models?

Enoch Bowden's profile picture
Enoch Bowden1 day ago

Wait... Fuck off

Laythe's profile picture
Laythe1 day ago

is it actually any good though

Related Videos

Alibaba just released a coding model that hits 82 percent on SWE-Bench Verified. That is the highest score ever published for an open-source model. The weights are free. The license is Apache 2.0. You can run it today. The model is Qwen 4 Coder 32B. Here is what 82 percent on SWE-Bench Verified actually means. SWE-Bench Verified tests whether an AI can autonomously resolve real bugs pulled from real production GitHub repositories. Not synthetic exercises. Real open-source projects that real teams depend on. A model gets a bug report, reads the code, writes a fix, and either passes the test suite or it does not. At 82 percent, Qwen 4 Coder 32B resolves 82 out of every 100 real production bugs it is given. Without a human guiding it. On code it has never seen before. For comparison: Qwen 4 Coder 32B: 82 percent SWE-Bench Verified. Open source. Apache 2.0. Claude Fable 5: 80.3 percent SWE-Bench Pro. $10 input / $50 output per million tokens. Currently suspended. GPT-5.6 Sol: Competitive on Terminal-Bench. $5 input / $30 output per million tokens. An open-weight model that you can download and run for free just beat both of them on the benchmark designed to measure real software engineering capability. Here is the architecture. Qwen 4 Coder 32B is a 32 billion parameter dense model. Not a Mixture-of-Experts. Every parameter is active on every request. This matters for inference: a dense 32B model runs on 22 gigabytes of VRAM, which fits on a single high-end consumer GPU or a MacBook Pro with 64GB of unified memory. The smaller variant, Qwen 4 Coder 4B, runs at approximately 135 tokens per second on an M5 Max and fits inside 8 gigabytes of RAM. For a model with usable coding capability, that is a new bar for what fits in a single laptop. The training methodology continued Alibaba's approach of reinforcement learning on verifiable coding tasks. The model gets rewarded when its code passes tests. It gets penalized when it fails. Over millions of training steps, the model learns to write code that actually runs rather than code that looks plausible. License: Apache 2.0. Full commercial use. No attribution requirement. No revenue threshold. No monthly active user ceiling. Weights: Hugging Face, available today. Runs on: vLLM, Ollama, SGLang, and any standard GGUF-compatible inference engine. Qwen 4 32B also runs at approximately 135 tokens per second on an M5 Max chip, setting a new bar for what a sub-8GB model can do on Apple Silicon. The open-source coding model just beat the best closed-source model in the world on the benchmark designed to test whether AI can actually do software engineering. The weights are free. The subscription is optional. Source: Autom8Labs AI Insight July 2026, State of Open Source LLMs June 2026, Kunal Ganglani blog June 2026.

Harman

41,278 views • 2 months ago