Back
PrismML releases Bonsai 2 27B, a 27B model compressed 9x with near-full scores
SiTech AI Team3 წთ. საკითხავი

PrismML releases Bonsai 2 27B, a 27B model compressed 9x with near-full scores

PrismML has released Ternary Bonsai 2 27B, a 5.9GB model built on Qwen3.8 27B that is more than nine times smaller than the full-precision version while retaining 98.2% of its aggregate benchmark scores.

PrismML, a company formed by a team of Caltech researchers, released Ternary Bonsai 2 27B on September 17: a multimodal model built on Qwen3.8 27B that needs far less memory than the full-precision version it derives from. The company describes it as the biggest step yet in its line of models designed to run on local devices, following the first Bonsai 27B release two months earlier.

The model uses ternary weights — values of minus one, zero or one — with FP16 group-wise scaling, giving 1.76 effective bits per weight. The total footprint is 5.9GB, the context window reaches 262K tokens, and the release handles text and image input. It ships under the Apache 2.0 licence, with weights available now.

Benchmark results

Across a suite spanning reasoning, maths, coding, instruction following, vision and agentic tool use, Ternary Bonsai 2 27B scores 83.9 and retains 98.2% of the aggregate performance of Qwen3.8 27B. For comparison, the full-precision Qwen3.8 27B scores 85.4 and the earlier Qwen3.6 27B scores 83.6.

Category by category, instruction following is where the compressed model comes out ahead: 82.66 against 81.25 for the full-precision model and 74.53 for Qwen3.6 27B. Maths lands at 96.57 (97.06 at full precision), coding at 81.58 (82.17), agentic and tool-calling tasks at 77.57 (79.74), knowledge and reasoning at 83.95 (86.66) and vision at 78.59 (81.64).

Why the compression matters

PrismML says its earlier Bonsai 27B kept 95% of the full-precision model's performance; the new version pushes that retention above 98%, which the company calls practically lossless. Coding agents, tool-use systems, multimodal workflows and long-horizon tasks are especially sensitive to model degradation because small errors compound over many steps, and PrismML argues those are exactly the areas where Bonsai 2 27B holds most of its capability at a fraction of the memory.

On speed, the model reaches up to 143 tokens per second on an NVIDIA GeForce RTX 5090 and 46.8 tokens per second on an M5 Max. On an RTX 4090 it draws 0.714 mWh per token, which PrismML says makes it 40% more energy-efficient than an 8B model running at full precision. It runs on NVIDIA GPUs through CUDA and on Apple devices — Mac, iPhone and iPad — through MLX, using custom low-bit kernels.

Background

The previous generation, Bonsai 27B, arrived in July 2026: a ternary build at 5.9GB for laptops and a 1-bit build at 3.9GB for an iPhone 17 Pro, both with a 262K-token context window. PrismML says it was founded with support from Khosla Ventures, Cerberus and Google, with continuing support from Samsung, and that it has spent years working on compressing neural networks without sacrificing reasoning ability.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.