Illustration for: PrismML's Ternary-Weight LLM Fits A 27B Model In 5.9GB

PrismML's Ternary-Weight LLM Fits A 27B Model In 5.9GB

PrismML, backed by Khosla Ventures and Caltech on a $22.25 million seed, compressed a 27-billion-parameter model to 5.9GB using ternary weights while retaining 98% of the original's benchmark performance.

By the Numbers

$22.25M
Seed round
5.9GB (27B params)
Compressed size
98% of Qwen3.8
Benchmark retention
11M+ (base model)
Downloads
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

PrismML's Bonsai 2 27B compresses Qwen3.8 27B down to 5.9GB using 'ternary' weights -- simplifying standard 16-bit weights to just three values (+1, -1, 0) -- while matching 98% of the original's aggregate benchmark scores.

2

That compression ratio is large enough to run a genuinely capable reasoning model locally on PCs and smartphones rather than requiring a cloud API call, a meaningful shift for cost, latency and data-privacy-sensitive use cases.

3

11 million-plus downloads of the base model, plus 2.6 million more for smaller variants, is a real distribution signal for a startup this young -- most model-compression research doesn't reach this kind of adoption outside academic benchmarking.

4

Khosla Ventures and Caltech backing a hardcore model-compression bet, rather than an application-layer AI startup, is a reminder that infrastructure-level differentiation is still getting funded even as most 2026 venture dollars chase agents and applications.

TC

The VC Read · Trace's Take

Trace Cohen

11 million downloads is real distribution, not a benchmark-chasing paper -- the diligence item is whether ternary compression holds up past 27B parameters, where most of the actual frontier-capability gap lives. If it degrades at scale, this is a durable on-device niche, not a challenger to frontier labs.

Analysis

PrismML raised a $22.25 million seed round backed by Khosla Ventures, Cerberus Capital and Caltech, and released Bonsai 2 27B, its latest compressed reasoning model, Techmeme and the Wall Street Journal reported, with additional detail from TechCrunch.

How The Compression Actually Works

PrismML's technical approach, called "ternary" weighting, simplifies the standard 16-bit weights used in most large language models down to just three possible values: +1, -1 or 0. That dramatic simplification shrinks model size substantially while -- according to PrismML's own benchmarks -- preserving most of the underlying model's capability. Bonsai 2 compresses Qwen3.8 27B, a 27-billion-parameter open model, down to 5.9GB, and matches 98% of Qwen's aggregate benchmark scores in that compressed form.

Bonsai 2 compresses Qwen3.8 27B, a 27-billion-parameter open model, down to 5.9GB, and matches 98% of Qwen's aggregate benchmark scores in that compressed form.

Why On-Device Matters

Running a model locally rather than through a cloud API changes the economics and privacy profile of AI applications meaningfully: no per-token API cost, no round-trip latency to a remote server, and no user data leaving the device by default. PrismML's pitch -- that capable reasoning models don't have to be large -- directly targets use cases where those three factors matter more than having access to the absolute largest frontier model. This is a different competitive lane than Google DeepMind's newly launched AGI Institute or the trillion-dollar valuations attached to OpenAI and Anthropic -- PrismML isn't competing on raw frontier capability, it's competing on deployability.

Adoption Numbers Worth Taking Seriously

The original Bonsai model has been downloaded more than 11 million times, with smaller variants adding another 2.6 million downloads -- a genuinely large adoption number for a young, infrastructure-focused startup, and a signal that developers are actively choosing PrismML's compressed models over full-size alternatives for real deployments rather than just academic curiosity. That kind of organic pull is harder to manufacture than benchmark scores alone.

Does Ternary Survive Frontier Scale?

The open question is whether ternary-weight compression holds up as models scale further -- Bonsai 2's 98% benchmark retention is measured against a 27-billion-parameter base model, and it's not yet demonstrated whether the same compression ratio and performance retention holds at 70 billion, 100 billion or larger parameter counts, where most of the frontier-capability gap with proprietary labs actually lives. If ternary compression degrades faster at larger scales, PrismML's technique may end up a durable niche for smaller, on-device models rather than a genuine alternative to frontier-scale reasoning.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by TechCrunch · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.