Video yükleniyor...
Video Yüklenemedi
We also tested Ternary Bonsai 2 27B on the 2026 IMO problems against full-precision Qwen3.8 27B (54GB) and Gemma 4 12B QAT (~7GB). No internet. No tools. 131K-token reasoning budget. Bonsai scored in the upper end of the human bronze-medal range and retained 95% of Qwen’s IMO score, while... show more
34,109 görüntüleme • 9 gün önce •via X (Twitter)
7 Yorum

First up: building a 3D skateboard game from scratch in a single HTML file. The first attempt isn't perfect. Instead of restarting a better generation, we keep the same conversation going, point out the issues, and give the model feedback. It progressively fixes the gameplay, visuals, and controls over several iterations. That iterative loop being able to take feedback, debug its own work, and improve an existing artifact is much closer to how we think language models are most useful in practice. Code:

We have been very excited to see how the community has been utilizing our Bonsai 2 model, well beyond use cases that we had originally envisioned it for. We’ve spent the last few days putting Bonsai 2 27B through various examples inspired by what the community has shown us. These examples showcase the strengths, and some of the shortcomings–particularly on longer multi-turn agentic workflows, and will be extremely helpful as we continue to improve. We’ll be sharing a few of those demos, along with some practical guidance on the settings and prompting that get the best results from the model. Demo repo:

Starting from an empty workspace, Bonsai 2 27B built a browser-based desktop in a single HTML file, then iteratively fixed issues across multiple rounds of feedback, from broken windows and UI behavior to complete functionality and final styling. Code:

We’ll keep sharing both what the model does well and where we’re working to make it better. More updates soon.

What is going on in Gemma's CoT lol

When availability on lmstudio?

Are you gonna do one of the MoE versions anytime?

