Trust is earned, not given

A different perspective

2025-11-20 · Projects

AI Frontiers, part 18: Speculative decoding — faster inference for free

Part 18from the AI Frontiers series · 65 parts in all

Almost every performance technique in machine learning involves a trade. Quantization buys memory and speed with quality (part 17). Distillation buys speed with capability. Truncating context buys latency with information. Speculative decoding is the rare exception: deployed correctly, it makes generation two to three times faster while producing exactly the same distribution over outputs as the unmodified model. Not similar. Not within a quality threshold. Identical to the sampling procedure, in the probabilistic sense — and in the greedy case, bit-for-bit identical.

That property is why it went from a 2022 research paper to a default feature in every serious serving stack within about two years, and it is worth understanding because it is a rare example of a clean algorithmic win in a field that mostly advances by making tradeoffs.

The trick, stated properly

Generation is sequential and that is the whole problem. A transformer producing token n+1 needs token n, so the accelerator sits mostly idle while it reads hundreds of gigabytes of weights to produce a single token. Decoding is memory-bandwidth bound, which means the multiplication capacity of the chip is going to waste. Speculative decoding exploits that waste.

The scheme: run a small, cheap draft model to generate k candidate tokens ahead. Then run the large target model once over the whole sequence — draft tokens included — which is a single forward pass costing roughly the same as producing one token at that sequence length, because the work was bandwidth bound anyway. The target model's outputs at each position are the probability distributions it would have produced had it been decoding those positions itself. Now compare.

For greedy decoding, the check is trivial: accept the draft tokens up to the first position where the target model's argmax differs, and continue from there. Every accepted token is exactly what the target model would have produced, and the rejected suffix cost you almost nothing. For sampling, the correct rule is more subtle, and this is the part of the deep-learning-systems literature I find genuinely elegant. Sample a candidate from the draft distribution q, accept it with probability min(1, p/q) where p is the target distribution, and if you reject, sample from the normalized residual distribution max(0, p − q) (Leviathan et al.; Chen et al., "Accelerating Large Language Model Decoding with Speculative Sampling"). The residual correction is what preserves exactness: the tokens you reject are exactly the ones the draft overproduced relative to the target, so the rejection step compensates. The result is a sample from the target distribution, not an approximation of one.

If that rule were hard to get right, the technique would be a research curiosity. In practice the last several years have produced a small industry of correct implementations, and the correctness is well enough established that "does spec decoding change my outputs?" now has a documented answer in the vLLM and TensorRT-LLM documentation rather than open debate.

Four ways to build a drafter

Everything interesting about the technique is in how you get candidate tokens cheaply and accurately.

A separate small model. The original approach. Pair a 7B model as drafter with a 70B target, both from the same family so their distributions are aligned. Simple, effective, and it doubles the memory footprint, which is the reason it did not immediately become universal.

Extra prediction heads. Medusa (Cai et al.) attaches several decoding heads to the target model itself, each trained to predict a token at a different offset, so one forward pass yields candidates for the next several positions. Hydra and related work refined the training. No second model, at the cost of training and a small amount of extra compute per step.

Feature-level autoregression. EAGLE (Li et al.) instead predicts the target model's hidden features at the next position with a small autoregressive head, then reuses the target's own output layer to convert features into tokens. The insight is that the output distribution is far less predictable than the feature that produces it, so drafting in feature space gets much higher acceptance rates for the same drafter size. EAGLE-2 added dynamic draft-tree construction, which adapts the number of candidates per step.

No drafter at all. Prompt lookup and n-gram decoding (Saxena) choose draft tokens by finding the current suffix somewhere in the prompt or context and copying what followed last time. It sounds crude and it is dramatically effective on exactly the workloads people run most: summarization, question answering with quoted context, code editing and refactoring, and any task where the output is largely copied from the input. No training, no extra parameters, no alignment problem. For a document-summarization endpoint I would try this before anything else.

The lineage starts earlier than most people assume: blockwise parallel decoding (Stern et al.) proposed predicting multiple future tokens in parallel in 2018, before large language models made the memory-bandwidth argument decisive.

Why it stops working, and when

The speedup is the number of accepted tokens per verification step divided by the cost of that step, and both terms are workload-dependent.

Acceptance rate is the dominant factor, and it is a measure of how well the drafter approximates the target. Same-family drafters accept more. Specialized drafters trained on the target's outputs accept more than general ones. And temperature matters enormously: at temperature zero, the draft only has to get the argmax right; at high temperature the target's distribution is flatter and the acceptance check rejects more often. A configuration that gives 3x at greedy decoding can give nearly nothing at temperature 1.0 with top-p sampling, which is why so many teams benchmark the wrong thing and conclude the technique does not work.

Batch size is the trap. Speculative decoding exploits an idle compute unit. Serve a single stream, or a small batch, and the chip is bandwidth bound with plenty of spare arithmetic capacity, so speculation is nearly free. Serve a batch of a hundred sequences and the accelerator becomes compute bound, at which point generating extra candidate tokens is genuinely extra work. Under high load, the correct configuration is often to disable speculation entirely and let continuous batching do its job. This is counterintuitive enough that it is worth stating as a rule: speculative decoding is a latency optimization, not a throughput optimization, and it belongs in the serving config of the latency-sensitive endpoints and not the batch endpoints.

Output predictability determines how much headroom exists. Code completions that continue existing patterns, structured output like JSON keys, and reasoning traces that follow learned formulas are all highly predictable, and predictably long. It is not a coincidence that reasoning models — whose verbosity part 16 turned into a product feature — are one of the best use cases for speculation. The very property that makes them expensive makes them compressible in this particular dimension.

How to measure it without fooling yourself

Most teams that conclude speculative decoding does not work have benchmarked it wrong, and the mistakes are consistent enough to enumerate.

Fix the sampling settings. An acceptance rate measured at temperature 1.0 with aggressive top-p is not comparable to one measured greedily, and the second is what production code-completion and structured-output endpoints actually run. Report the settings next to the speedup or the number means nothing.

Separate time to first token from inter-token latency. Speculative decoding does not help prefill at all; it improves the decode loop. Averages that blend the two hide both the win and the workload change. Total latency for a long generation improves a lot; time to first token does not move.

Measure at the concurrency you actually serve. This is the mistake that matters most. A benchmark on a single stream will show a large speedup and a production deployment behind a load balancer at batch size 64 will show none, because the compute that speculation was supposed to reclaim is already in use. Either characterize the speedup as a function of concurrency or accept that the number only applies to the latency-sensitive tier.

Instrument the acceptance rate directly. It is the single diagnostic that tells you whether a poor result is a bad drafter, an over-aggressive sampling configuration, or an unsuitable workload. Most serving stacks expose it; if yours does not, add the counter. A healthy rate for a same-family drafter at greedy decoding is well above half, and if you are seeing much less, the drafter is mismatched rather than the technique being useless.

Check the cost side. Running a draft model consumes memory that could hold more of the target model's KV cache, or a larger batch. On constrained hardware the trade can go either way, and the only way to settle it is to measure throughput per dollar under the real traffic mix rather than latency in isolation.

The rest of the free-lunch list

Speculative decoding is the most elegant member of a family of optimizations that buy performance without touching the model's outputs, and it is worth naming the others because they are usually what you should try first.

Prefix caching stores the KV cache of a shared prefix — a system prompt, a long context document, a conversation history — so it is computed once and reused. For any application with a fixed instruction preamble this is often the largest single win available, and it is free.

Continuous batching admits new requests into a running batch between decode steps instead of waiting for the whole batch to finish. It can multiply throughput on mixed workloads by a large factor, and it is now the default in modern servers rather than a feature (Kwon et al.).

Chunked prefill splits a long prompt into pieces and interleaves them with ongoing decoding work so that one huge prompt does not stall every other request. This matters enormously for retrieval-augmented applications, where arrival of a 30,000-token prompt is routine rather than exceptional.

Graph capture and kernel fusion eliminate per-step CPU overhead. The gains look small per token and compound over thousands of tokens, and they are pure engineering with no quality component whatsoever.

What the items on this list have in common is that none of them improve the model and all of them improve the product. That is a strange asymmetry in a field obsessed with model releases, and I suspect it persists because a serving optimization is invisible in a demo and enormous in a profit-and-loss statement. Two years of working on inference has convinced me of a simple heuristic: before spending money on a bigger model, spend a week finding out where the current one is wasting the hardware.

The serving stack it lives in

Paged attention and continuous batching (Kwon et al.) are what made it economically viable to put speculative decoding in a shared server, because those techniques are what allow the same GPU to serve a fast single-stream request and a batch of background jobs at once — speculation on for the former, off for the latter. Prefix caching (and its more general form in radix-tree attention, as in SGLang) matters for the same reason: the acceptance probability of prompt-lookup decoding depends on how much of the relevant text is already in context and cached.

The pattern across parts 17 and 18 is that the inference stack is now a stack of independently optional optimizations that compose: quantization reduces bytes per weight, paging and caching reduce wasted memory, batching improves utilization, and speculation spends idle compute to reduce latency. None of them is glamorous. Each of them, applied in isolation, is worth a modest factor. Applied together they are the difference between an inference bill your product can survive and one it cannot, and the fact that they are all invisible to the model's outputs is precisely why they get so little attention relative to whatever model was released this month.

That is the real lesson here. The most reliable way to make a system faster is usually not to make the model better — it is to stop wasting the hardware you already paid for. The model is the headline, but serving is where the product's unit economics are decided, and speculative decoding is the cleanest example in the field of a pure win: same answer, less time, no asterisk.

Works Cited

Cai, Tianle, et al. "Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads." arXiv, 2024, arxiv.org/abs/2401.10774. Accessed 20 Nov. 2025.

Chen, Charlie, et al. "Accelerating Large Language Model Decoding with Speculative Sampling." arXiv, 2023, arxiv.org/abs/2302.01318. Accessed 20 Nov. 2025.

Kwon, Woosuk, et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." arXiv, 2023, arxiv.org/abs/2309.06180. Accessed 20 Nov. 2025.

Leviathan, Yaniv, et al. "Fast Inference from Transformers via Speculative Decoding." Proceedings of the 40th International Conference on Machine Learning, 2023, arxiv.org/abs/2211.17192. Accessed 20 Nov. 2025.

Li, Yuhui, et al. "EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty." arXiv, 2024, arxiv.org/abs/2401.15077. Accessed 20 Nov. 2025.

---. "EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees." arXiv, 2024, arxiv.org/abs/2406.16858. Accessed 20 Nov. 2025.

Miao, Xupeng, et al. "SpecInfer: Accelerating Generative Large Language Model Serving with Tree-Based Speculative Inference and Verification." arXiv, 2023, arxiv.org/abs/2305.09781. Accessed 20 Nov. 2025.

Saxena, Apoorv. "Prompt Lookup Decoding." GitHub, github.com/apoorvumang/prompt-lookup-decoding. Accessed 20 Nov. 2025.

Stern, Mitchell, et al. "Blockwise Parallel Decoding for Deep Autoregressive Models." arXiv, 2018, arxiv.org/abs/1811.03115. Accessed 20 Nov. 2025.

Xia, Heming, et al. "Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding." arXiv, 2024, arxiv.org/abs/2401.07851. Accessed 20 Nov. 2025.

Zhang, Jun, et al. "Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding." arXiv, 2023, arxiv.org/abs/2309.08168. Accessed 20 Nov. 2025.

Zheng, Lianmin, et al. "SGLang: Efficient Execution of Structured Language Model Programs." arXiv, 2023, arxiv.org/abs/2312.07104. Accessed 20 Nov. 2025.