
PrismML
@PrismML • 28,279 subscribers
Centering AI research on efficiency. https://t.co/88MQHGCeFD
Videos

We also tested Ternary Bonsai 2 27B on the 2026 IMO problems against full-precision Qwen3.8 27B (54GB) and Gemma 4 12B QAT (~7GB). No internet. No tools. 131K-token reasoning budget. Bonsai scored in the upper end of the human bronze-medal range and retained 95% of Qwen’s IMO score, while completing the problems in 70% of the time. Code:
PrismML34,109 次观看 • 9 天前

Here is Ternary Bonsai 27B running an end-to-end agentic workflow locally with Hermes on an NVIDIA GeForce RTX 5090 GPU. The model reasons, calls tools, reads outputs, modifies files, and surfaces insights - all on consumer hardware, while all private files, intermediate states, and iterations remain local.
PrismML154,473 次观看 • 2 个月前

The phone threshold is even harder than the storage number suggests. A phone exposes only part of its memory to an application, and the model must share that budget with its KV cache, activations, runtime, and the rest of the product. At 3.9 GB, 1-bit Bonsai 27B clears that threshold with room to work, making 27B-class local AI possible on a phone for the first time. (Demo Mode: cached & prefilled image context)
PrismML33,665 次观看 • 2 个月前
没有更多内容可加载