Tech
EN AZ
PrismML hopes its tiny LLM will change how we all use AI

PrismML hopes its tiny LLM will change how we all use AI

techcrunch.com 18.09.2026 00:34 3 views
If AI lab PrismML isn't on your radar yet, it should be.

If AI lab PrismML isn’t on your radar yet, it should be — not because it’s raised gobs of money (it hasn’t yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it’s developing. PrismML is betting that capable, high-performing, reasoning large language models don’t, in fact, have to be large. It is making reasoning models so small they can fit on PCs and smartphones. (It’s even rumored to be in talks with Apple, though CEO Babak Hassibi declined to comment on that to TechCrunch.) On Thursday, PrismML released Bonsai 2 27B, its latest in a family of models, which compresses Qwen3.8 27B, a widely used open-source model from Alibaba, down to 5.9 GB.

That’s small enough to fit on a PC and, possibly, a high-end smartphone. It’s a 9x to 10x reduction in memory versus the original. PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies.

The startup also counts Ion Stoica as an advisor. Stoica is a co-founder of Databricks (and other companies) and the director of Berkeley’s famed Sky Computing Lab, which has birthed many technologies and startups, from Letta to SGLang. PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech.

This startup is certainly not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spain’s Donostia International Physics Center, is another. (And Multiverse Computing has raised gobs of cash.) But Hassibi says that PrismML’s compression tech is unique because its LLMs have lost virtually no performance compared with the originals. Bonsai 2 matches 98% of Qwen’s aggregate benchmark scores.

That’s up from the first Bonsai, released a couple of months ago in March, that matched 95%. That original model has already been downloaded over 11 million times, and PrismML’s even smaller models have been downloaded another 2.6 million times, the company says. So this shows that PrismML’s compression results have improved from one release to the next.

Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always have some impact, Hassibi says. Still, perfect benchmark parity is fairly academic anyway.

Extract — continue reading at the source.

Read full story