AI Frontiers, part 46: Small models on device — NPUs, memory budgets, and privacy
Part 46from the AI Frontiers series · 65 parts in all
Part 9 made the case that a well-trained 3B model is a useful product tier. This entry is about what happens when that tier has to run on hardware you do not own, cannot upgrade, and must not overheat — a phone, a laptop, a kiosk, a set-top box, a vehicle head unit. It is a different problem than serving small models from a container, and the difference is mostly not about the model. It is about memory bandwidth, thermal envelopes, and the fact that "on-device" is simultaneously a performance decision and a legal one.
The pitch for local inference has been stable for years: no network round-trip, no per-token cost, no data leaving the device, works offline. What changed in 2024 is that the first two stopped being aspirational. Quantization and distillation put genuinely useful models inside a phone's memory budget, and the silicon caught up enough to run them at conversational speed — with caveats that the marketing materials rarely include.
Decoding is a memory-bandwidth problem
Start with the arithmetic intensity of autoregressive decoding, because it explains almost every on-device performance observation. Generating one token requires reading the entire model once — every weight matrix, every time. The number of arithmetic operations per byte of weight read is roughly one. That is the definition of a memory-bound workload: the compute units sit idle waiting for weights to arrive, and throughput is governed by bytes per second over whatever memory the weights live in.
The consequence is a back-of-envelope formula you can use before writing any code. A 4-bit 3B model is about 1.7 GB of weights. Reading that from LPDDR5 at roughly 60–100 GB/s (the order of magnitude for a recent phone SoC) gives an upper bound near 35–60 tokens per second, and that bound counts nothing else — not the KV cache, not the OS, not the browser, not a second app. Real throughput lands well below it. Now repeat the calculation for a 7B model at 4 bits: 3.5 GB of weights, and a memory bus that is doing everything else the phone wants. The reason a phone comfortably runs a 1–3B model at conversational speed and struggles with a 7B is not that the 7B is 2.3× "harder" — it is that weight reading dominates and the ratio of weights to bandwidth did not improve.
This is also why quantization is not an optimization on-device but a precondition. Four-bit weight-only quantization cuts the dominant cost by roughly four times with surprisingly small quality loss, which is what turned the class from a demo into a product. GPTQ, AWQ, and their descendants all target the same bottleneck (Frantar et al.; Lin et al.), and the difference between them matters mostly for accuracy at low bit widths. The published surveys are worth reading before choosing a quantizer for a shipped build, because the trade-offs are systematic rather than random (Gholami et al.).
What the NPU is actually for
The neural accelerator has a longer history than most people assume — the DianNao work out of Tsinghua in 2014 established the architectural pattern of a small dedicated matrix engine with an on-chip buffer, and today's NPUs are direct descendants (Chen et al.). The modern question is not whether NPUs can multiply; it is whether they help for this workload.
Because generation is memory-bound, an accelerator with enormous TOPS and a modest share of memory bandwidth adds limited throughput for decoding. Where NPUs earn their place is prefill — processing the prompt, which is compute-bound — and the vision or audio front end of a multimodal pipeline. Those are dense, parallel, and bandwidth-light. So the pattern that emerged is a split: NPU or GPU for prefill and encoder work, CPU or unified memory for the decode loop, and a scheduler handing off between them. Apple's unified memory architecture is the cleanest expression of this, because weights live in one address space that both the GPU and the CPU can read, eliminating the copy that dominates discrete-memory designs (Apple, "Introducing Apple's On-Device and Server Foundation Models").
The practical exposure is the runtime stack, not the silicon. Core ML, Google's AI Edge, Qualcomm's AI Hub, ONNX Runtime, and PyTorch's ExecuTorch each have their own notion of which operators accelerate, and the differences are large enough that the same model runs at wildly different speeds across them. The reliable workflow is boring: build once, benchmark on the actual target devices, and accept that the fastest path is usually the one whose operator coverage matches your model's architecture. Published numbers from a different phone model tell you nothing.
The memory budget, honestly
Phone apps do not get the whole RAM budget, and a demo that runs fine in isolation has a habit of being killed when the camera, a messaging app, and a game are resident. A realistic plan for a foreground app on a recent flagship is roughly 1–2 GB for everything the app needs, plus a short-lived peak during load. That leaves room for a small quantized model and a modest KV cache, and not much else.
Three budget lines are usually underweighted. Context costs memory even on-device: the KV cache grows linearly with tokens held, and a 4K-token conversation can consume a meaningful fraction of a 1.5 GB allotment depending on architecture. Load time is a real budget item, since memory-mapping several gigabytes from storage takes seconds and users notice. And thermals are the constraint nobody models: sustained token throughput on a phone is lower than burst throughput, sometimes by half, because the SoC throttles. Demo on the desk, suffer on the train, hear about it in reviews.
The design answer is to make the local model's job narrow. Classification, extraction, summarization of short text, intent parsing, suggestion generation, first-draft retrieval ranking — tasks with bounded inputs where a 1–3B model is sufficient and where a wrong answer is cheap to correct. Anything open-ended goes to the network, and a router decides. This is the two-tier architecture from part 9, restated for the constraints of a device: the local model is not a smaller version of the cloud model, it is a different component with a different specification.
Privacy may be the stronger argument
Every deployment conversation I have had about on-device inference starts with latency and budget and ends with data governance. That ordering is backwards. If the input is health information, a financial record, a child's photograph, or employee communication, the question "may this leave the device" has a legal answer that arrives before any engineering question, and local inference is one of the few architectures that answers it cleanly.
Regulatory regimes built on data minimization and purpose limitation — the EU's General Data Protection Regulation is the canonical example (European Parliament and Council) — treat transfer of personal data to a third-party processor as a distinct event requiring a legal basis, a contract, and often an assessment. Inference that never leaves the device does not create that event. For a European deployment this is not a nice-to-have; it can be the difference between shipping and not shipping. The same logic drives the sovereign-cloud conversations covered in part 39 from the other direction: where the computation happens is increasingly part of the product specification.
The federated-learning and differential-privacy literature is the research complement to this — training on signals that never leave the device, with formal privacy guarantees (McMahan et al.; Abadi et al.) — but its near-term relevance to product teams is limited. The 2004-era insight that matters today is simpler: not sending the data is the strongest privacy control available, and it costs nothing to audit.
Three shipping shapes
What does not exist is a single deployment pattern. The three I see in practice differ mostly in who controls the runtime. A shared local runtime — an app shipping llama.cpp or MLX in-process, or calling an OS-level model API — is the simplest and the one most teams should start with, because it puts the model version under your release process. A platform integration — a system model exposed through the OS, as Apple and Google both began offering in 2024 — trades control for zero maintenance and a much better battery profile, at the cost of being subject to someone else's capability roadmap and policy. An in-browser deployment using WebGPU is the newest shape and the most attractive for zero-install web apps, with the caveat that you are shipping model bytes over the network to every visitor unless you are careful with caching.
Pick by asking where the model is updated from and who is accountable when it misbehaves. Those two questions determine which shape is viable long before throughput does.
What I would put in the plan
Choose the model by memory, not by benchmark: measure bytes and tokens per second on the target device class, and treat the published leaderboard as a starting shortlist. Quantize deliberately, and re-run the quality suite after quantization rather than assuming four bits is free. Bound the task so the model's failure mode is survivable. Budget for load time and thermal throttling as product requirements, not engineering footnotes. And write the data-flow diagram before the architecture diagram, because on-device inference is often a compliance decision in disguise. Done well, it is the least glamorous and most durable tier in the stack — the model that is just there, on the device, doing a small job quickly and never phone home. That is a good place to be.
Measuring on the device, not on the spec sheet
The measurement discipline for on-device inference is different from server benchmarking in one important way: the numbers that matter are the ones that persist. A harness I would build before optimizing anything has six measurements, all taken on the slowest device you intend to support. Cold start to first token, which includes loading weights from storage and is often the worst part of the user experience. Time to first token on a warm run, which bounds perceived responsiveness. Steady-state throughput, measured over several minutes rather than several seconds, because that is where throttling appears. Peak resident memory, since the operating system kills applications that exceed their allotment rather than degrading them. Energy per thousand tokens, which the platform profilers report in milliwatt-hours and which determines whether the feature is usable on a forty-percent battery. And quality on your task, because a model that generates tokens quickly and extracts the wrong field has optimized the wrong variable.
The energy line deserves more attention than it usually gets. Inference at scale has a real power cost, and the literature quantifying it — Watts per query, kilowatt-hours per million queries, the difference between a large model and a distilled successor (Luccioni et al.; Patterson et al.) — is now developed enough to inform design decisions rather than just policy arguments. On a device the same arithmetic shows up as battery drain, and it is the constraint users notice first and forgive least.
The portability tax
The last cost of on-device work is the one nobody budgets: the same model has to be built for each runtime you support, and the runtimes do not agree on anything. Weight formats differ between GGUF, MLX, ONNX, and the platform-native bundles; quantization schemes differ in which layers are quantized and how the scales are stored; operator coverage differs enough that a fused attention implementation available on one runtime does not exist on another. A model that is a single artifact in a server deployment becomes five artifacts in a client application, each of which must be validated separately.
Two practices reduce the pain. Keep one source of truth for the model and its evaluation — a single upstream checkpoint plus the harness that scores it — and treat each platform export as a build output rather than an independent project. And pin the toolchain version in your build, because NPU compilers and runtime libraries change behavior between OS releases in ways that alter both performance and output. The teams that ship on-device inference successfully treat it as a release-engineering problem with a long validation tail, not as a model-integration problem, and they start that work earlier than feels necessary.
Works Cited
Abadi, Martin, et al. "Deep Learning with Differential Privacy." arXiv, 2016, arxiv.org/abs/1607.00133. Accessed 27 Nov. 2024.
Apple. "Introducing Apple's On-Device and Server Foundation Models." Apple Machine Learning Research, 10 June 2024, machinelearning.apple.com/research/introducing-apple-foundation-models. Accessed 27 Nov. 2024.
Banbury, Colby, et al. "MLPerf Tiny Benchmark." arXiv, 2021, arxiv.org/abs/2106.07597. Accessed 27 Nov. 2024.
Chen, Tianshi, et al. "DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning." Proceedings of the 19th International Conference on Architectural Support for Programming Languages and Operating Systems, 2014. Accessed 27 Nov. 2024.
European Parliament and Council. "Regulation (EU) 2016/679 on the Protection of Natural Persons with Regard to the Processing of Personal Data." Official Journal of the European Union, 2016, eur-lex.europa.eu/eli/reg/2016/679/oj. Accessed 27 Nov. 2024.
Frantar, Elias, et al. "GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers." arXiv, 2022, arxiv.org/abs/2210.17323. Accessed 27 Nov. 2024.
Gholami, Amir, et al. "A Survey of Quantization Methods for Efficient Neural Network Inference." arXiv, 2021, arxiv.org/abs/2103.13630. Accessed 27 Nov. 2024.
Lin, Ji, et al. "AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration." arXiv, 2023, arxiv.org/abs/2306.00978. Accessed 27 Nov. 2024.
Luccioni, Alexandra Sasha, Yacine Jernite, and Emma Strubell. "Power Hungry Processing: Watts Driving the Cost of AI Deployment?" arXiv, 2023, arxiv.org/abs/2311.16863. Accessed 27 Nov. 2024.
McMahan, H. Brendan, et al. "Communication-Efficient Learning of Deep Networks from Decentralized Data." arXiv, 2016, arxiv.org/abs/1602.05629. Accessed 27 Nov. 2024.
Patterson, David, et al. "The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink." arXiv, 2022, arxiv.org/abs/2204.05149. Accessed 27 Nov. 2024.
PyTorch. "ExecuTorch: On-Device Inference." PyTorch, 2024, pytorch.org/executorch-overview. Accessed 27 Nov. 2024.
Wang, Chong, et al. "MobileLLM: Optimizing Sub-Billion Parameter Language Models for On-Device Use Cases." arXiv, 2024, arxiv.org/abs/2402.14905. Accessed 27 Nov. 2024.
Zhu, Ligeng, et al. "TinyChat: Efficient and Lightweight Chatbot with LLM." arXiv, 2023, arxiv.org/abs/2312.04985. Accessed 27 Nov. 2024.