Trust is earned, not given

A different perspective

2025-04-03 · Projects

AI Frontiers, part 50: Inference cost engineering — caching, routing, and unit economics

Part 50from the AI Frontiers series · 65 parts in all

There is a moment in the life of every AI feature when someone looks at the monthly bill and asks why the invoice looks like a second cloud migration. The answer is almost never that the model is expensive. It is that the architecture pays for tokens it did not need to send, in a configuration that costs more per token than necessary, on requests that a tenth of the model could have handled. Inference cost is not a line item you optimize at the end; it is a property of the system's design, and it is decided by a handful of choices made early.

This entry is a cost model, then a tour of the four levers that move it, then the arithmetic that turns token counts into a business decision. It follows directly from part 32, which was about how serving systems work, and from the efficiency techniques scattered through the series — quantization (part 17), speculative decoding (part 18), and the KV cache (part 19). Those were about making inference faster. This one is about making it cheaper, which is a related but distinct problem — and occasionally a conflicting one.

A cost model you can actually reason with

Start with the two-dimensional picture that every provider's pricing page reflects: input tokens and output tokens, with output priced several times higher. That asymmetry is not marketing. Prompt processing — prefill — happens in parallel across all prompt tokens, so it is compute-efficient and easily batched. Generation is sequential and memory-bandwidth-bound: every output token requires a full pass through the weights and re-reading the growing KV cache. The per-token cost of output is structurally higher, which is why the pricing gap exists and why it is stable across providers.

Three consequences follow immediately. First, prompt length is a real cost, and long system prompts plus retrieved context are the largest single line item in most systems. Second, output length matters more than it looks, because a model that rambles through four paragraphs to deliver two sentences is paying a premium for the rambling. Third, and least intuitive, retries are the hidden multiplier: a pipeline that retries on parse failure, or a research agent that loops, multiplies both dimensions invisibly. Cost-per-call monitoring will miss this entirely; cost-per-task makes it obvious.

So the first instrumentation decision is the one that matters: measure cost per completed task, including retries, and break it into input tokens, output tokens, and the number of calls. That single breakdown converts every subsequent optimization into a measurable claim.

Lever one: stop paying for the prefix twice

The largest, cheapest win in most systems is prompt caching. Because a transformer's prefill is prefix-determined, a shared prefix — the system prompt, a fixed instruction block, a stable document preamble — can have its attention state reused rather than recomputed. Providers exposed this as a product through 2024: cached input tokens are priced at a fraction of uncached input, with a short time-to-live and a minimum prefix length (Anthropic; OpenAI, "Prompt Caching"). Systems that put static instructions first and volatile content last can see input costs fall sharply with no quality change whatsoever, and the only work is ordering the prompt correctly.

Inside your own serving stack, the same idea has a deeper implementation. Radix-tree-based KV reuse in SGLang lets concurrent requests sharing a prefix reuse the computed attention state (Zheng et al.), and Prompt Cache generalized modular attention reuse across arbitrary prompt segments (Gim et al.), with CacheBlend addressing the harder case where only part of a prefix matches (Yao et al.). The design principle is identical at every level: put the stable parts first, the volatile parts last, and let the cache do the work. If your retrieval layer injects documents above the instructions, you have made caching impossible and paid for it every call.

Lever two: make each token cheaper

Token-level efficiency is the serving-stack lever, and 2024's toolkit is mature: continuous batching and paged memory (Kwon et al.), quantization to four bits where the quality cost is acceptable, speculative decoding to convert idle compute into throughput, and chunked-prefill scheduling so that long prompts do not block short generations (Agrawal et al.). On the operating side, batch APIs — offered at a substantial discount by the major providers for workloads with loose latency requirements (OpenAI, "Batch API") — are the simplest available saving and are chronically underused. Any workload that is not interactive should be a batch workload: evaluations, backfills, offline enrichment, nightly summarization.

The delivery mechanism for all of this, if you are self-hosting, is the inference server itself: vLLM, TensorRT-LLM, SGLang, TGI. The difference between a naive PyTorch loop and a proper serving engine is routinely an order of magnitude in throughput per dollar, which is covered in detail in part 32 and needs no restating. What does deserve restating is the counterweight: cheaper per token is not cheaper per task if quality falls enough to raise retries or escalations. Quantization is a trade, and the trade has to be measured.

Lever three: route to the cheapest model that works

Most systems spend most of their budget on calls that did not need their best model. The first formal treatment of this idea was FrugalGPT, which framed it explicitly as a cascade: try a cheap model, score the answer, escalate to a more expensive one only when the cheap answer is insufficient, and choose the sequence that minimizes cost subject to a quality constraint (Chen et al., "FrugalGPT"). The reported savings were substantial, and the structure — cheap first, verify, escalate — is the same one that appears in production routers today.

Routing has three implementable forms, covered in more detail in part 9's companion discussion and worth repeating as cost math: static routing by request type, which requires no model and no training but captures most of the win when the workload has obvious strata; cascade routing, where a cheap model answers and a verifier decides whether to escalate — cheap when most calls are easy, with the verifier as the overhead; and learned routing, where a small classifier or a compact model predicts which tier will suffice. Also cheap and effective: degraded-mode routing by user tier, so that a free-tier user's request goes to the small model while a paying user's goes to the frontier. Whatever the shape, the router needs a target metric and an evaluation, because a router without a scoreboard is just a bug distribution engine.

Lever four: spend fewer tokens on the same task

The lever with no infrastructure cost is prompt design. Four habits consistently cut token spend by a large fraction with no quality loss: cap the output contract (instruct and enforce a length; structured output does this naturally — see part 45); retrieve less and rank better (five well-chosen chunks beat twenty mediocre ones, and a reranker that cuts the context in half usually improves quality and cost); cache deterministic sub-answers in your own application layer, because a surprising share of traffic is near-duplicate queries; and replace a long reasoning prompt with a distilled small model where the task is narrow, which is exactly the cost argument behind part 48. None of these require a new component. All of them require that someone is looking at the per-task cost breakdown on a schedule.

The arithmetic that matters

Cost engineering becomes a business question the moment you divide by revenue. If a task costs a few cents and the customer pays a few dollars per seat, the feature has a healthy margin and the interesting question is quality. If the task costs a dime and the customer's willingness to pay works out to three cents per interaction, you have a feature that scales into a loss, and no amount of prompt tweaking fixes it — the architecture is wrong. The teams that got this right in 2024 and 2025 shared one habit: they computed cost per task before launch, listed the levers in order of expected impact, and set a budget with an alert. Everyone else discovered their unit economics from a finance review.

There is one cost line that is routinely omitted from these calculations and should not be: the evaluation pipeline itself. If your quality bar is enforced by running a few hundred test cases against two or three model configurations on every change, those runs are inference, they are not free, and at scale they can approach the cost of serving real traffic. That is not an argument against measuring — it is an argument for measuring efficiently, which in practice means caching fixtures, running the suite against small models for routine checks, reserving full frontier-model evaluation for release candidates, and treating the evaluation budget as a line item with an owner. The teams with the best cost discipline in 2025 were the ones that had applied it to their own test harnesses as well.

One closing caution about the incentives. Because price-per-token has fallen dramatically and continues to fall, it is tempting to conclude that efficiency does not matter. That conclusion is wrong for a structural reason: capability per token has been rising even faster than price has fallen, so systems keep consuming more inference as models get better — longer reasoning, more retrieved context, more agent steps. Efficiency work does not get obsoleted by cheaper tokens; it gets re-spent on capability. The systems with good cost engineering are not the ones that spend less; they are the ones that can afford to think harder on the calls that deserve it.

Two worked examples

The levers are easier to trust with arithmetic attached, so here are two workloads with illustrative numbers. The per-million-token prices are placeholders — the ratios matter, the dollar figures do not — and both examples assume the levers are applied in the order above.

A support triage pipeline. Two hundred thousand tickets a month. Each ticket is classified, routed to a queue, and answered with a draft. The naive version sends everything to the frontier model with a 1,200-token system prompt, eight thousand tokens of retrieved policy and history, and produces a 250-token draft. That is roughly two billion input tokens a month for the drafts alone, plus the same prompt again for the separate classification call, which is where the classic mistake lives: two calls, one of which needed nothing but a label.

The optimized version changes four things. The classification call moves to a small local model, which is a rounding error in cost and, being local, also removes a network round trip from the critical path. The system prompt is placed first and left byte-identical between calls so that prompt caching applies, which cuts the effective input price sharply on the repeated prefix. Retrieval stops fetching eight chunks and starts fetching three after a reranker, halving the context at a small quality cost that the evaluation set can adjudicate. And the frontier model is invoked only for the slice that the small model flagged as hard — routinely fifteen to twenty percent of traffic. The result is a pipeline that costs a fraction of the original while producing better drafts, because the expensive model is now doing only the work it is actually better at.

A document extraction backfill. Fifty thousand documents a month, each needing structured fields pulled out per the schema in part 45. This one is not interactive, which makes it the easiest win on the list: move it to the batch API for the provider's bulk discount, and the bill halves with no engineering. Then fix the prompt-ordering mistake that most extraction systems have, where the instruction block is repeated after each document instead of once at the top, which defeats prefix caching. Then reduce output size by requiring short field values and no explanation — the schema already enforces shape, and an output contract enforces length. Then measure per-document cost against per-document accuracy across a length distribution, because the long documents are where the cost concentrates and, frequently, where the accuracy collapses.

In both cases the largest single saving was not a model choice. It was noticing that the system was paying to say the same thing on every call.

Works Cited

Agrawal, Amey, et al. "Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve." arXiv, 2024, arxiv.org/abs/2403.02310. Accessed 3 Apr. 2025.

Anthropic. "Prompt Caching." Anthropic Documentation, 2024, docs.anthropic.com/en/docs/build-with-claude/prompt-caching. Accessed 3 Apr. 2025.

Chen, Lingjiao, Matei Zaharia, and James Zou. "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance." arXiv, 2023, arxiv.org/abs/2305.05176. Accessed 3 Apr. 2025.

Gim, In, et al. "Prompt Cache: Modular Attention Reuse for Low-Latency Inference." arXiv, 2023, arxiv.org/abs/2311.04934. Accessed 3 Apr. 2025.

Kwon, Woosuk, et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." arXiv, 2023, arxiv.org/abs/2309.06180. Accessed 3 Apr. 2025.

OpenAI. "Batch API." OpenAI Platform Documentation, 2024, platform.openai.com/docs/guides/batch. Accessed 3 Apr. 2025.

OpenAI. "Prompt Caching in the API." OpenAI, 1 Oct. 2024, openai.com/index/api-prompt-caching/. Accessed 3 Apr. 2025.

Pope, Reiner, et al. "Efficiently Scaling Transformer Inference." arXiv, 2022, arxiv.org/abs/2211.05102. Accessed 3 Apr. 2025.

Sheng, Ying, et al. "FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU." arXiv, 2023, arxiv.org/abs/2303.06865. Accessed 3 Apr. 2025.

Yao, Jiayi, et al. "CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion." arXiv, 2024, arxiv.org/abs/2405.16444. Accessed 3 Apr. 2025.

Zheng, Lianmin, et al. "SGLang: Efficient Execution of Structured Language Model Programs." arXiv, 2023, arxiv.org/abs/2312.07104. Accessed 3 Apr. 2025.