AI Frontiers, part 9: Small language models — Phi, Gemma, and the desktop class
Part 9from the AI Frontiers series · 65 parts in all
The scaling-law story of part 2 had a quiet corollary that nobody priced in: if capability follows compute and data quality rather than raw parameter count, then small models trained exceptionally well should punch far above their size. In February 2024, Microsoft's Phi-2 (2.7B parameters) was matching models ten times its size on reasoning benchmarks. By April, Microsoft's Phi-3-mini (3.8B) scored 69% on MMLU — GPT-3.5 territory — while running on a phone (Abdin et al.). Google answered with Gemma (February) and Gemma 2 on the way, in 2B–27B sizes, explicitly built from Gemini's recipe (Gemma Team). This entry is about the small-model thesis: where these capabilities come from, what the models are actually for, and why I think every production LLM system will end up hybrid — small models doing most of the work, frontier models doing the hard percent.
The data-quality thesis: textbooks are all you need
The Phi line began, as a joke that stopped being one, with "Textbooks Are All You Need" — the August 2023 phi-1 paper, a 1.3B model trained mostly on synthetic "textbook-quality" Python exercises and explanations that scored state-of-the-art-for-size on HumanEval (Gunasekar et al.). The thesis: web-scale corpora spend most of their tokens on low-information text — boilerplate, duplication, noise — while a small model's compute budget deserves denser material. phi-1.5 scaled the recipe (Li et al.); Phi-2 (December 2023) pushed to 2.7B with a scaled mixture of synthetic and curated web data; Phi-3 (April 2024) brought the family to 3.8B/7B/14B with a 128K-context long-context variant, and the tech report is explicit that data filtering rigor — quality over quantity, heavily filtered web plus synthetic pedagogy — did the work (Abdin et al.).
The honest caveats came with the papers. phi-1's benchmark gains partly reflected contamination risk and narrow evaluation (the authors themselves flagged both); "compact models trained on curated data memorize benchmark-adjacent text more readily," and common-sense breadth lagged raw benchmark performance. Phi-3's report is admirably blunt about the trade-off: its 3.8B mini model beats much larger models on curated reasoning suites while losing to those same models on trivia and long-tail knowledge — exactly what the training recipe predicts. The skill is real; it is reasoning-shaped skill, not encyclopedic coverage. When you pick a small model for production, you are buying that trade.
Gemma: the same thesis, industrialized
Google's Gemma (February 2024), in 2B and 7B, was pitched as "Gemini's little siblings" — same tokenizer, same architecture family (multi-query attention, RoPE, GeGLU activations), distilled from the much larger Gemini line, with a technical report that leans on the word distillation: the small models initialize from and learn against the big model's output distributions, not just a curated corpus (Gemma Team). That is the second engine of the small-model revolution and it matters as much as data quality: a frontier model is a teacher, and its distilled students inherit much of the competence at a fraction of the serving cost. Add the ecosystem factor — Gemma shipped with quantized GGUF variants, Hugging Face integration, and Vertex hosting on day one — and the small-model class became a deployable product tier rather than a research curiosity.
The timing helps explain the rush. Llama 3 (April 2024, 8B and 70B, dense, 15T tokens) reset expectations for the 8B class within weeks of Phi-3; Mistral's April releases pushed the same direction. The pattern across all of them: frontier techniques — better data pipelines, distillation, the architectural refinements of part 8 — cascade down to the small classes within quarters. A 2020 rule of thumb ("7B models are toys") was, by mid-2024, simply wrong.
Distillation: the second engine
Data quality is only half the small-model story; the other half is distillation, and it deserves precision because the word gets loose usage. Classical knowledge distillation (Hinton et al.) trains a student on a teacher's output distributions — the soft probabilities over the vocabulary carry richer signal than hard labels, including the teacher's second-choice knowledge. The 2024 practice is broader: generate training data with a frontier model and train the small model on it (synthetic instruction data — the Alpaca pattern of part 3, industrialized), train on the teacher's logits where available, or initialize from the teacher's weights and compress. What unifies the variants is the transfer of behavior from a model too expensive to run everywhere to models cheap enough to run anywhere. The frontier model becomes a data factory for its own successors — which is why every frontier lab now ships a small-model family alongside the flagship: the marginal cost of distilling is trivial next to pre-training, and the small models capture the workload the flagship is overkill for.
What the desktop class is for, part two: on-device economics
The on-device tier deserves numbers, because intuition lags reality by about a year. A 4-bit 3.8B model is roughly 2.3GB of weights — it fits in a phone's memory budget beside camera stacks and games, runs at usable tokens-per-second on last year's flagship NPUs, and answers within milliseconds because there is no round-trip. Google's Gemini Nano is already announced for the Pixel 8 Pro and Chrome's roadmaps, Apple's on-device work is the designated guess of every analyst watching Cupertino, and the OEM cohort is following — because the reasoning is arithmetic, not fashion: a billion small calls between them cost fractions of a cent; the same volume through an API costs millions of dollars plus a privacy review per feature. The constraint is the one this entry already named — capability floor, not capability ceiling — which is why the on-device pattern is narrow by design: intent detection, extraction, summarization, suggestions, first-pass drafts, with the frontier model a network call away for the hard cases and the router (part 8's theme again) deciding locally.
There is a second-order effect worth flagging for anyone building developer tooling: the desktop class made local evaluation respectable. When the model under test runs on your own hardware, you can afford to run your eval suite continuously — every commit, every prompt change — instead of rationing API spend. Several teams found that moving evaluation to small local models for the routine cases improved their larger deployments too, because evaluation stopped being the bottleneck it had silently been. The tooling ecosystem around small models — llama.cpp servers, GGUF registries, OpenAI-compatible local endpoints — quietly became the standard harness for AI engineering practice, not just a hobbyist curiosity.
The 8B class, one quarter later: Llama 3 and the moving floor
Any snapshot of the small-model tier needs its volatility stated, because the tier moves faster than any other. Llama 3's 8B — dense, 15 trillion tokens, grouped-query attention, released April 18, one month before this entry's date — illustrates the motion precisely: its benchmark profile resembles the 70B class of eighteen months earlier, and its MMLU lands within a few points of Phi-3-mini's despite the Phi line's radically different data recipe. Two very different training approaches (frontier-scale filtered web versus textbook-quality curation) converging on similar small-model capability is itself the finding: by 2024, the 8B class was systems — architecture settled, data machinery understood, training compute affordable — rather than research. When a problem becomes a system, capability stops being a mystery and becomes a product decision, which is why the 8B tier now ships from every lab with a schedule rather than a paper.
The moving floor has a production consequence worth internalizing: any capability you certify on a small model today will be available in a cheaper, faster, quantized form within two quarters. Contracts, evaluation suites, and prompt caches built around a specific small model should therefore be built around its interface — same tokens-in/tokens-out contract, swappable behind a router — rather than its identity. Teams that hard-wired model names into their pipelines in 2023 spent 2024 re-plumbing; teams that routed by capability tier swapped the underlying model by editing a config.
The economic consequence follows from part 8's reasoning. MoE pushes frontier serving toward batch efficiency and memory abundance; distillation and the Phi-style data pipeline push the workload toward the small end, where inference is cheap, local, and private. The equilibrium that was forming by mid-2024 — frontier models as teachers and hard-case arbiters, small models as the workforce — is the stable two-tier market this series will track through the rest of its run.
What the desktop class is for
Small models are not cheaper ChatGPTs. Their economics and their privacy story open different doors, and by May 2024 four uses had shaken out in production practice.
1. High-volume classification and extraction. Routing support tickets, extracting structured fields from documents, tagging content. These are billion-call-a-year workloads where GPT-4-class intelligence is waste; a 4–8B model at 10–30× lower cost per call, hosted on your own hardware, changes the unit economics from "pilot" to "ship."
2. Privacy-constrained inference. Healthcare notes, legal documents, financial records: data that cannot leave the building, or that procurement will not let near a third-party API. A 7B model on a $500 GPU is the compliance answer, and the quality gap on constrained extraction tasks is small enough to clear the bar.
3. On-device. The phone/laptop tier: llama.cpp-class runtimes (part 3) plus 4-bit quantization put 3–8B models on commodity hardware — Phi-3-mini runs on an iPhone at usable speeds. Latency stops being a network round-trip; offline stops being a bug.
4. The pattern part 3 built. LoRA fine-tuning shines brightest on small bases: a task-tuned 7B beats a prompted 70B on narrow work at a fraction of the cost. The adapter-swap serving pattern — one base, many task LoRAs — is the small-model production idiom, and it composes with the RAG stack of parts 5–6 rather than competing with it.
Routing: the architecture that makes hybrids real
The two-tier market needs a switch, and by May 2024 the router had become its own production component with a small design space. Cascade routing: try the small model, verify the output cheaply (self-consistency, a verifier model, or a grader prompt), and escalate to the frontier only on failure — cheapest when most calls are easy, with the verification itself as the overhead to tune. Classifier routing: a tiny model (or mere heuristics — length, category, user tier) picks the tier up front; deterministic, cheap, and only as good as its training. Cost-cap routing: run the small model with a token/time budget and treat budget exhaustion as the escalation signal. Each scheme embeds the same trade — verification cost versus escalation cost — and the right choice depends on how skewed your workload is: the more routine your traffic, the more aggressive the cascade. The quiet enabler is that routing decisions need confidence, not intelligence — part 8's MoE routers are the same primitive inside the weights — and small models are precisely the cheapest source of calibrated-enough confidence.
What they still cannot do, stated plainly: long-tail knowledge, genuinely open-ended reasoning, and anything where being confidently wrong is expensive. The frontier remains a different class — part 11 of this series will get to why (inference-time compute changes the frontier's economics again). The 2024 architecture that works: small models for the 95% of calls that are routine, frontier models for the 5% that are hard, a router deciding which is which — and the router itself, fittingly, is a small-model workload.
A closing note on the series so far
Nine parts in, the arc of 2023 deserves one paragraph of synthesis, because this entry completes the year-and-a-half from the transformer's consequences to the market's new shape. The architecture settled (part 1); capability arguments were had and partially resolved (part 2); the means of production escaped the labs (part 3); the weights went legitimately commercial (part 4); knowledge moved into systems (part 5); infrastructure crystallized around vectors (part 6); context became a purchaseable resource with engineering disciplines attached (part 7); sparsity restructured the frontier's economics (part 8); and now capability itself has been miniaturized (this part). Every step traded something central for something distributed: parameters for systems, clusters for cards, APIs for weights, big models for routed hybrids. The next nine parts follow that distribution into its consequences — multimodal models, inference-time reasoning, the agent protocols, and the quantization and cache engineering that make all of it payable. The theme, stated once and kept: the frontier is where the ideas are born; the industry is where they get cheap.
Works Cited
Abdin, Marah, et al. "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone." arXiv.org, 2024, arxiv.org/abs/2404.14219. Accessed 16 May 2024.
Gemma Team. "Gemma: Open Models Based on Gemini Research and Technology." arXiv.org, 2024, arxiv.org/abs/2403.08295. Accessed 16 May 2024.
Hinton, Geoffrey, Oriol Vinyals, and Jeff Dean. "Distilling the Knowledge in a Neural Network." arXiv.org, 2015, arxiv.org/abs/1503.02531. Accessed 16 May 2024.
Gunasekar, Suriya, et al. "Textbooks Are All You Need." arXiv.org, 2023, arxiv.org/abs/2306.11644. Accessed 16 May 2024.
Li, Yuanzhi, et al. "Textbooks Are All You Need II: phi-1.5 Technical Report." arXiv.org, 2023, arxiv.org/abs/2309.05463. Accessed 16 May 2024.
Meta AI. "Introducing Meta Llama 3: The Most Capable Openly Available LLM to Date." Meta AI, 18 Apr. 2024, ai.meta.com/blog/meta-lama-3/. Accessed 16 May 2024.