In brief
Since 2024, high-end smartphones have embedded chips capable of running a small language model locally, without going through the cloud. The Galaxy S25 translates calls in real time on the device. Apple Intelligence summarizes messages without sending anything to a server. This is embedded inference — or edge inference. It is not a promise. It is already in production, on hundreds of millions of devices.
But the constraints are real. Limited memory, battery life, heat, and reduced compute power compared to a datacenter: every trade-off has a cost somewhere. This article details what concretely happens on the inside.
The AI chip in your phone: what is it?
Every recent smartphone contains several types of processors. The CPU handles general tasks. The GPU manages graphics. And for the past few years, a third component has appeared: the NPU (Neural Processing Unit), designed specifically for the matrix computations required by neural network inference.
NPUs did not emerge with LLMs. Apple has integrated them in iPhones since the A11 Bionic in 2017. Google puts one in its Pixel devices via the Tensor chip. Qualcomm has its Hexagon in Snapdragon SoCs. MediaTek follows suit on mid-range devices.
What changed in 2024–2025 is their power. Qualcomm’s Hexagon NPU reports announced gains of ×3 on AI inference compared to the previous generation. The Apple Neural Engine on M-series chips dominates performance-per-watt benchmarks. These chips can now run, locally, models that would have required a server three years ago.
But “run” does not mean “run anything.”
Memory: the structural constraint
An LLM is essentially a collection of numbers — billions of parameters. GPT-4 or Llama 3 70B weigh tens to hundreds of gigabytes at full precision (FP32). A high-end smartphone has 8 to 12 GB of RAM shared between the OS, applications, and the model.
To fit within this envelope, on-device models must be quantized: instead of storing each parameter on 32 or 16 bits, it is compressed to 8 bits (INT8) or 4 bits (INT4). INT4 quantization reduces memory footprint by 75% compared to FP16. A 7-billion-parameter model that weighed 14 GB drops to under 4 GB.
Quality loss is generally limited for INT4 and INT8. It becomes significant below that: INT2 and INT3 can fit 20-billion-parameter models within 8 GB, but the degradation is documented — and an additional risk is identified in the literature: aggressive compression can impair moderation filters and increase the probability of problematic outputs.
No standardized benchmark yet exists to measure this degradation from the perspective of the real user, rather than academic benchmarks. This is a blind spot in the field.
In short: quantization is compressing a model the way you compress a photo. INT8 and INT4 are reasonable compressions — minimal loss, 50-75% memory gain. INT2 and below is the photo run through ten generations of copies: quality drops, and certain details — including moderation filters — can be the first sacrificed.
Compact models: which LLMs for the edge?
Several models are designed or optimized for on-device deployment:
Meta’s Llama 3.2 comes in 1B and 3B parameter versions. The 1B model runs on low-end devices. The 3B outperforms heavier models (Gemma 2 2.6B, Phi 3.5-mini) on instruction-following and tool use — making it the practical reference for edge deployments in 2025.
Microsoft’s Phi-4 mini (3.8 billion parameters) reaches performance comparable to 7–9 billion parameter models on reasoning tasks, with coverage of over 20 languages. Its advantage: versatility + lightness.
Google’s Gemma 3n is architecturally designed for mobile. It integrates vision and audio encoders optimized for real time, with E2B (2 billion effective parameters) and E4B (4 billion) variants.
HuggingFace’s SmolLM2 goes down to 135 million parameters — the ultra-lightweight segment, suited for IoT devices and microcontrollers.
A trend is asserting itself in 2026: natively multimodal models for the edge, capable of processing text, image, and audio within a single compact model, without needing to orchestrate three separate components.
| Model | Parameters | INT4 memory | Target device | Main strength |
|---|---|---|---|---|
| SmolLM2 (HuggingFace) | 135M | ~ 100 MB | Microcontrollers, IoT | Ultra-light, basic versatility. |
| Llama 3.2 1B (Meta) | 1 billion | ~ 700 MB | Low/mid-range smartphone | Simple text, decent instruction-following. |
| Llama 3.2 3B (Meta) | 3 billion | ~ 2 GB | High-end smartphone | The practical 2025 reference, good size/quality ratio. |
| Phi-4 mini (Microsoft) | 3.8 billion | ~ 2.5 GB | Smartphone/laptop | Reasoning, multilingual across 20+ languages. |
| Gemma 3n (Google) | 2 to 4 billion effective | ~ 1.5-3 GB | Mobile, multimodal | Native vision + audio in real time. |
Frameworks: a fragmented ecosystem
Running a model on an NPU is not improvised. It requires a runtime — a software layer that translates model operations into instructions optimized for the target chip. Several frameworks coexist, with no native interoperability between them.
ExecuTorch (Meta, October 2025) is the official PyTorch runtime for the edge. It supports more than 12 hardware backends — Apple, Qualcomm, Arm, MediaTek, Vulkan — with a baseline footprint of 50 kilobytes. More than 80% of popular edge LLMs on HuggingFace are supported.
MLX (Apple) delivers the best raw performance on Apple Silicon: approximately 230 tokens per second on M2 Ultra in a comparative study (arXiv 2511.05502). Limitation: exclusive to macOS/iOS.
MLC-LLM achieves approximately 190 tokens per second on M2 Ultra with a lower time-to-first-token (TTFT) than MLX, at the cost of higher memory consumption.
llama.cpp is the reference engine of the open ecosystem. Written in C++, lightweight, CPU-optimized, with minimal memory footprint. Less performant for mobile GPU prefill, but ported to nearly everything.
Core ML (Apple), ONNX Runtime (Microsoft), LiteRT (Google, formerly TensorFlow Lite) round out the picture. Five major runtimes, little interoperability. A model optimized for Core ML may require significant re-optimization to run on ONNX Runtime. This is a real barrier for developers who want to deploy across multiple platforms.
What happens during inference
When the model generates text, it produces one token at a time. Each token requires reading all the model parameters and computing attention vectors over the context already produced. This context — the KV-cache (Key-Value cache) — grows with each token generated and consumes memory as the conversation lengthens.
This is the second bottleneck after model parameters. For a context of several thousand tokens, the KV-cache can exceed available memory. Specialized techniques exist: cache quantization (KVTuner achieves 3.25 bits in mixed-precision quasi-lossless on Llama-3.1-8B), adaptive per-layer allocation (DynamicKV: 85% of full-KV performance with 1.7% of original KV memory). These techniques allow extending the usable context length on memory-constrained devices.
Speculative decoding is another approach: a lightweight model proposes several tokens ahead, a main model verifies and accepts or rejects. Gains can reach ×2.8 on some configurations. But on smartphones without a dedicated GPU, the trade-off is not obvious: speculative decoding increases the total number of operations, which can worsen battery consumption and heat. The problem remains open for purely on-device deployments.
The hybrid model: between device and cloud
Not all queries require the same computing power. Summarizing a short text message: a 1B model is sufficient. Answering a complex legal question: the local model will fall short.
Adaptive routing — dynamically deciding whether a request stays on-device or goes to the cloud — is an active research direction (arXiv 2507.16731). Apple already implements this partially with Private Cloud Compute: lightweight tasks stay on the device, heavy tasks go to Apple servers with specific privacy guarantees.
But the routing criteria — at what complexity threshold to switch? what information can be sent to the cloud without compromising privacy? — are not the subject of any standard. Each actor implements its own logic.
What on-device inference does not solve
Two structural limitations remain.
Data freshness. An on-device model has a knowledge cutoff. It cannot update itself. On a device that is mostly offline, it does not know what happened after its last update. On-device fine-tuning via LoRA exists experimentally, but its deployment at scale — battery impact, privacy of adaptation data — remains unresolved in the literature.
Cross-device comparison. No standardized cross-hardware benchmark exists for LLMs on mobile. Figures published in tokens per second, TTFT, or memory consumption vary depending on measurement conditions. Comparing a Snapdragon 8 Elite to an Apple A18 Pro on an LLM inference task is non-trivial.
Key takeaways
- Recent smartphones have chips (NPUs) designed for neural network inference. This is not an option — it is integrated into virtually all high-end devices released since 2024.
- On-device inference relies on compact models (1B to 4B parameters) and compression techniques (INT4/INT8 quantization) to fit within device memory constraints.
- The KV-cache is the second bottleneck: its management determines the usable context length.
- The framework ecosystem is fragmented (ExecuTorch, MLX, llama.cpp, Core ML, LiteRT) with no native interoperability.
- The real limits are model freshness (no continuous updates), quality degradation under aggressive compression, and the absence of comparable cross-hardware benchmarks.
- The hybrid edge-cloud model is in development but without an established routing standard.