Loading video...

Video Failed to Load

Go Home

We also tested Ternary Bonsai 2 27B on the 2026 IMO problems against full-precision Qwen3.8 27B (54GB) and Gemma 4 12B QAT (~7GB). No internet. No tools. 131K-token reasoning budget. Bonsai scored in the upper end of the human bronze-medal range and retained 95% of Qwen’s IMO score, while...

34,109 views • 9 days ago •via X (Twitter)

7 Comments

PrismML's profile picture
PrismML9 days ago

First up: building a 3D skateboard game from scratch in a single HTML file. The first attempt isn't perfect. Instead of restarting a better generation, we keep the same conversation going, point out the issues, and give the model feedback. It progressively fixes the gameplay, visuals, and controls over several iterations. That iterative loop being able to take feedback, debug its own work, and improve an existing artifact is much closer to how we think language models are most useful in practice. Code:

PrismML's profile picture
PrismML9 days ago

We have been very excited to see how the community has been utilizing our Bonsai 2 model, well beyond use cases that we had originally envisioned it for. We’ve spent the last few days putting Bonsai 2 27B through various examples inspired by what the community has shown us. These examples showcase the strengths, and some of the shortcomings–particularly on longer multi-turn agentic workflows, and will be extremely helpful as we continue to improve. We’ll be sharing a few of those demos, along with some practical guidance on the settings and prompting that get the best results from the model. Demo repo:

PrismML's profile picture
PrismML9 days ago

Starting from an empty workspace, Bonsai 2 27B built a browser-based desktop in a single HTML file, then iteratively fixed issues across multiple rounds of feedback, from broken windows and UI behavior to complete functionality and final styling. Code:

PrismML's profile picture
PrismML9 days ago

We’ll keep sharing both what the model does well and where we’re working to make it better. More updates soon.

BirchEater67's profile picture
BirchEater679 days ago

What is going on in Gemma's CoT lol

luresh's profile picture
luresh9 days ago

When availability on lmstudio?

Craig Van's profile picture
Craig Van9 days ago

Are you gonna do one of the MoE versions anytime?

Related Videos

PrismML Releases Bonsai 27B: 1-bit and Ternary Builds of Qwen3.6-27B Hitting 89.5% of FP16 at 3.9GB. No new pretrain. No higher-precision escape hatches. No multi-GPU rig. Here's how it works. 👇 1: Codes, not floats Every weight becomes a code, with one shared FP16 scale per group of 128. Ternary is {−1, 0, +1}, binary is {−1, +1}. Sharing the scale across 128 weights keeps its cost at 16/128 = 0.125 bits. → Ternary: log2(3) + 16/128 ≈ 1.71 bits/weight → 5.9GB → Binary: 1 + 16/128 = 1.125 bits/weight → 3.9GB 2: Post-training, not from scratch No BitNet-style low-bit pretrain. It starts from off-the-shelf Qwen3.6-27B, architecture unchanged. The representation runs end to end across embeddings, attention projections, MLP projections, and the LM head. → 9.4× (ternary) and 14.2× (binary) vs the 54GB FP16 baseline 3: Labels are not bit-widths Conventional low-bit builds are mixed-precision by construction. The advertised name describes the most-compressed tensors, not the model. → Q4_K_XL, labeled "4-bit," is really 5.2 bits/weight at 17.6GB → IQ2_XXS, labeled "2-bit," is really 2.8 bits/weight at 9.4GB 4: Fitting a phone is two budgets iOS caps a single app near half of RAM, so a 12GB iPhone exposes ~6GB. The KV cache grows on top. Hybrid attention at ~75% linear means only 16 of 64 layers cache. → 4-bit KV: 4.3GB at 262K context, down from 17.2GB → 11.0 tok/s on iPhone 17 Pro Max 5: The numbers (15 benchmarks, thinking mode) → Ternary: 80.49 avg at 5.9GB — 94.6% of FP16 → 1-bit: 76.11 avg at 3.9GB — 89.5% of FP16 → IQ2_XXS falls to 57.5 on AIME26 while still scoring 88.93 on MMLU-Redux The key takeaway: 27B-class reasoning without the 54GB checkpoint — group-wise ternary and binary codes, an end-to-end low-bit language stack, 4-bit KV, on one phone. Full analysis: Repo: Model weight: Technical details: PrismML

Marktechpost AI

31,860 views • 2 months ago

bonsai 2 27b on an rtx 3060 12gb, the full receipt sheet. save this one, the 12gb row of the small gpu guide is built from it. speed by depth, then what context costs, live server, thinking on, real sessions > 7k deep: 24.4 tok/s > 12k deep: 21.9 tok/s > 35k deep: 17.8 tok/s > 77k deep: 13.0 tok/s > 64k window: 7.3gb resident > 128k window: 8.8gb resident > 192k window: 10.2gb resident > 262k window: 11.7gb resident, 0.6gb to spare, the whole native window on a 12gb card > every 64k of context costs 1.47gb, so 327k would not fit > a 41,312 token build session from 35k to 77k of context averaged 15.0 tok/s across 46 minutes > prefill 295 tok/s at 2k of context, 243 tok/s at 35k, first token in 0.6 seconds, fresh decode 26.1 tok/s > the card pinned 149 of 150 w the entire time, 78c, fan at 80%, power bound, not heat bound > 0.158 tok/s per watt at the fresh end the setup > model: ternary bonsai 2 27b, PTQ1_0, 1.75 bits per weight, 5.95gb on disk, base qwen 3.8 27b, apache 2.0 > runtime: prismml llama.cpp fork, prebuilt cuda 12.4 binary, no compile > serve: full 262k native window resident, q4 kv cache, flash attention, one slot, 11.7 of 12gb in use for anyone who followed bonsai 1 in july, that was the 3.9gb 1bit file at 42 tok/s on a 3060 ti, a faster card and a smaller file, so the same card comparison is not on the table yet, it comes with the 8gb test. what changed is the base, qwen 3.8 instead of 3.6, and the retention, 98.2% on their suite instead of 95%, and the whole 262k window fitting on 12gb.

Sudo su

31,174 views • 16 days ago

a new 8GB VRAM GPU dense Local LLM leader was born yesterday runs on: RTX 4060 / RTX 3070 / RTX 2080. any 8GB card Qwen 3.5 9B (dense) was the go to for 6-8GB VRAM builds. Gemma 4 12B QAT (dense) just changed that. same llama.cpp + cuda 13.2. i7 12700H. 16GB RAM. same -ngl 99 flags. same 48k context. unsloth gemma-4-12b-it-Q4_K_M.gguf → 15 tok/sec @ 48k ctx unsloth gemma-4-12B-it-qat-UD-Q4_K_XL.gguf → 32 tok/sec @ 48k ctx → 26 tok/sec @ 64k ctx 64k context is a big deal. Hermes 3 agent requires 64k minimum to run. you're now getting full hermes compatible context on a budget consumer GPU at 26 tok/sec locally. 2.1x faster on identical hardware. and here's the part that breaks your brain: the QAT-UD-Q4_K_XL is actually SMALLER than the Q4_K_M "XL" why? QAT = Quantization Aware Training Google didn't train the model first and compress it later they trained it to be quantized from day one the weights already know how to survive low precision that's why you get more quality per byte llamacpp flags: -m gemma-4-12B-it-qat-UD-Q4_K_XL.gguf -cnv -ngl 99 -c 48000 -v fits in 8GB VRAM clean. no API. no cloud. no subscription. and this isn't even the MTP variant yet Gemma-4-E2B QAT runs on 3GB RAM, E4B on 5GB, 12B on 7GB, 26-A4B on 15GB and 31B on 18GB. I have benchmarked the 26b and 31b qat as well on a single RTX 4090, checkout the comments for details. If you have a 6GB or 8GB VRAM GPU, post your numbers. more benchmarks and configs coming soon

Alok

264,062 views • 4 months ago