PrismML squeezes a 27B Qwen reasoning model into 5.9 GB with its Bonsai 2 release
The Caltech-founded startup, backed by a $22.25 million seed round and advised by Ion Stoica, wants reasoning models to run on PCs and high-end phones.

PrismML on Thursday released Bonsai 2 27B, a compressed version of Alibaba's widely used open source Qwen3.8 27B model that fits in 5.9 GB of memory. TechCrunch reports that is a 9x to 10x reduction compared with the original, small enough for a PC and possibly a high-end smartphone. The startup's bet is that capable reasoning models do not have to be large, and it is building a family of models small enough to run locally on consumer hardware. PrismML was founded by a group of Caltech researchers and is led by CEO Babak Hassibi, a Caltech professor and compression expert. Ion Stoica, a co-founder of Databricks and director of Berkeley's Sky Computing Lab, is an adviser. The company has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital and Caltech. It is rumored to be in talks with Apple, though Hassibi declined to comment. PrismML is not alone in model compression; Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, is among the competitors. If the compressed models hold their reasoning quality, developers could ship on-device assistants without cloud inference costs, and users would gain privacy and offline access.