AI & Technology

PrismML: Small AI Models That Run on Your Phone Without the Cloud

DROPIDEA By Admin
September 18, 2026 12 views
DROPIDEA | دروب ايديا - PrismML: Small AI Models That Run on Your Phone Without the Cloud

For many, the development of large language models (LLMs) has long been tied to the necessity of massive infrastructure and expensive cloud servers. But a startup called "PrismML" is trying to turn this equation on its head, by proving that models capable of logical reasoning don't have to be large in size to deliver high performance.

Although the company has not yet raised huge funding (settling for an initial round of $22.25 million), what makes it worth following are the technical minds behind it and the technology that could bring about a real transformation in how we use artificial intelligence.

AI in Your Pocket

PrismML is betting on a simple yet ambitious idea: shrinking language models to a degree that allows them to run directly on personal computers, and even on advanced smartphones. On Thursday, the company launched its latest model, "Bonsai 2 27B," a compressed version of Alibaba's open-source model "Qwen3.8 27B."

The remarkable achievement here lies in the numbers; the model was compressed to a size of just 5.9 gigabytes, representing a reduction of 9 to 10 times compared to the memory required for the original model. This size makes it suitable for running on personal computers, and possibly on high-end smartphones.

Who Is Behind the Company?

PrismML was founded by a group of researchers from the California Institute of Technology (Caltech), and is led by Babak Hassibi, a professor at the same institute and an expert in data compression technologies. The company also counts among its advisors Ion Stoica, co-founder of "Databricks" and director of the renowned "Sky Computing Lab" at UC Berkeley.

The company has the backing of prominent investors, including:

  • Khosla Ventures
  • Cerberus Capital
  • California Institute of Technology (Caltech)

Performance That Nearly Matches the Original

There's no doubt that PrismML is not the only company working in the field of language model compression, as there is strong competition from firms like "Multiverse Computing." But Hassibi asserts that what distinguishes his company's technology is maintaining performance levels with virtually no notable decline.

According to the company, the "Bonsai 2" model achieves 98% of the total benchmark performance scores of the original model, compared to 95% achieved by the first version released two months ago. The numbers point to notable commercial success, with the first model downloaded more than 11 million times, while downloads of the smaller models exceeded 2.6 million.

Hassibi believes that reaching a full 100% match remains more theoretical than practical, since uncompressed models themselves are not entirely accurate, and standard benchmarks don't necessarily reflect performance on real-world tasks. Therefore, a 2% difference is not expected to have a tangible impact on actual use.

How Is This Technology Achieved?

The company's approach relies on shrinking the "weights" that make up the model, which are the pieces of information the model learns and stores during training. Normally, each weight requires 16 bits to store, but PrismML's technique, known as "ternary weights," reduces this to just three values: (+1), (−1), or (0). As a result of storing much smaller values for each weight, the model takes up dramatically less space.

The Next Ambition: Larger Models

The company's sights are now set on applying compression technology to much larger models. Hassibi says the upcoming models will be in the range of hundreds of billions of parameters, expecting that preserving a model's "intelligence" will be easier as its size increases, since there is more room for compression without losing capabilities.

Advisor Ion Stoica expresses his enthusiasm for this technology, considering that it will make running advanced models possible directly on users' devices. He sums up his vision by saying that intelligence will become "at your fingertips," and free because it runs on a device you've already purchased, in addition to being more private since your data will not be sent to the cloud.

✦ بقلم فريق دروب أيديا

DROPIDEA

We hope this article has added real value to you. At DROPIDEA, we always strive to deliver high-quality content that helps you grow and evolve in the digital space. Follow us for more useful articles and guides.

Tags

#الذكاء الاصطناعي #نماذج اللغة #PrismML #ضغط النماذج

Share Article