Skip to main content
Back to AI NewsNews

Apple in Talks With AI Compression Startup PrismML for iPhone

Apple is in talks with PrismML, whose compressed AI models use up to 15 times less memory on iPhone hardware.

cueball EditorialTuesday, 14 July 2026 4 min read

What Happened

Apple is in active talks with PrismML, a startup specializing in AI model compression, according to a report published Monday by CNBC. PrismML has developed a compressed version of Alibaba's Qwen large language model that uses up to 15 times less memory than the original, a capability that could allow more powerful AI features to run directly on iPhone hardware without relying on cloud servers.

Background

Apple has been expanding its on-device AI strategy under its Apple Intelligence initiative, which the company introduced in 2024. The initiative aims to run AI workloads locally on Apple silicon chips, reducing latency and preserving user privacy by keeping data off remote servers. Apple Intelligence has faced criticism for rolling out features more slowly than competitors, and the company has been under pressure to accelerate its AI capabilities on consumer hardware.

PrismML is a startup focused on model compression, a technique that reduces the size and computational requirements of AI models while preserving their functional accuracy. The company's compressed Qwen model is built on Alibaba's open-weight Qwen series, which has become a widely used base model in third-party AI development. PrismML's process targets memory footprint specifically, addressing one of the central constraints of running large language models on mobile devices, where available RAM is measured in gigabytes rather than the terabytes available in data center environments.

What the Technology Does

On-device AI deployment requires models to fit within the memory constraints of consumer hardware. A standard large language model can require tens of gigabytes of memory to operate, making direct deployment on a smartphone impractical without compression. PrismML's reported 15-times memory reduction, if validated in production, would bring models that currently require data center infrastructure closer to the operational range of Apple's current iPhone lineup.

Model compression techniques generally involve methods such as quantization, pruning, and knowledge distillation, each of which reduces model size with varying trade-offs in performance. CNBC's report did not specify which compression methods PrismML employs, nor did it detail how the company's approach affects model accuracy or benchmark performance relative to the original Qwen model.

The Broader Context

Apple's interest in PrismML reflects a pattern of activity across the technology industry, where major hardware and software companies are seeking to move AI inference workloads from the cloud onto edge devices. On-device processing reduces costs associated with server infrastructure, lowers response latency for users, and eliminates the need to transmit sensitive data over networks, a consideration Apple has emphasized publicly in its AI communications.

Alibaba's Qwen models have seen broad adoption as a foundation for third-party AI applications since the series became available under open licensing. Using Qwen as a base, rather than building a proprietary model from scratch, allows a compression-focused startup like PrismML to demonstrate its technology against a well-documented, widely benchmarked system.

Apple regularly engages in acquisition talks and partnership discussions with smaller technology companies, and not all such conversations result in completed deals. The company has not publicly confirmed the discussions reported by CNBC, and no terms, timeline, or deal structure were disclosed in the report.

What It Means in Practice

If an agreement is reached, PrismML's compression technology could be applied to AI models Apple deploys through Apple Intelligence, potentially enabling capabilities on current or future iPhone models that would otherwise require a server connection. The practical scope of any integration would depend on how PrismML's compression performs across the range of tasks Apple Intelligence is designed to handle, including text summarization, writing assistance, and image generation.

Apple is scheduled to present further details on Apple Intelligence features at upcoming developer and product events, where the company is expected to outline the next phase of its on-device AI rollout.

Get our editors' take on what it all means. Read the Editor's Blog →