AI Frontiers, part 53: Reasoning distillation — putting thinking into small models
Part 53from the AI Frontiers series · 65 parts in all
When DeepSeek released R1 in January 2025 with open weights and, more importantly, open distillation data, the most consequential part of the release was not the 671-billion-parameter model anyone could read about. It was the small dense models trained on its reasoning traces. The R1-Distill series showed that a student model fine-tuned on a strong reasoner's chain-of-thought could inherit a surprising amount of that model's problem-solving behavior, scoring far above what the same architecture achieved with reinforcement learning on its own (DeepSeek-AI). Within weeks, "distill the reasoning" had gone from a research curiosity to a standard step in quantization and deployment pipelines.
This entry is about that technique: why it works, what actually transfers from teacher to student, the two papers that defined the practical recipe, and the failure modes that make a distilled reasoner worse than a plain small model for the tasks you were trying to serve. It sits between part 48, which is about deciding to fine-tune at all, and part 52, which is about the input side of the same budget.
Why it works at all
Instruction distillation — training a small model on a large model's generated answers — has been standard practice since Alpaca. Reasoning distillation adds one element that changes the character of the transfer: the student is trained on the process, not just the conclusion. That distinction was established well before the reasoning-model era. Guha and colleagues showed that adding rationales generated by a much larger model to the training data improved a small model's reasoning accuracy substantially, and that the benefit came from the rationales rather than from extra examples. Ho and colleagues extended it to NLP tasks with a teacher producing step-by-step demonstrations, and the Distilling Step-by-Step work showed that rationale-augmented distillation could beat a fine-tuned teacher with a fraction of the training data (Guha et al.; Ho et al.; Hsieh et al.). The self-taught reasoner line supplied the third piece: a model can generate its own rationales and learn from the ones that lead to correct answers, bootstrapping reasoning without any external traces (Zelikman et al.).
What R1 added was a teacher whose traces were themselves the product of large-scale reinforcement learning on verifiable problems, and therefore contained something earlier synthetic rationales did not: evidence of search. Backtracking, self-checking, trying a different decomposition after a dead end. A trace like that teaches a student a procedure, where a fluent explanation teaches it a tone.
What actually transfers
Honesty about this is the difference between a distillation project that pays off and one that produces a confident rambler. From the evidence and from what I have seen in practice, four things transfer and two do not.
Transfers: output format and structure. The student learns to lay out a derivation in steps, to restate the goal, to compute intermediates. This is worth more than it sounds, because on many tasks the ability to externalize intermediate values is the whole game — it turns an error-prone single-shot computation into a sequence of easier ones.
Transfers: domain-specific procedures. If the teacher's traces show a consistent decision procedure for your problem class — the order to check preconditions, which edge cases to consider — the student picks that up remarkably well, because it is a pattern rather than a fact.
Transfers: refusal and hedging toward verification. A student trained on traces containing self-checks often reproduces the checking behavior even when it cannot do the check, which changes the shape of the errors it makes.
Transfers: length and effort calibration, partially. The student learns roughly how much deliberation a problem deserves, which is a real capability and also the source of the worst failure mode below.
Does not transfer: latent knowledge. Distillation cannot install facts the student never had. A student that memorized no chemistry will produce a beautifully structured and wrong derivation of a chemical problem, and the structure makes it more convincing, not less. The same argument that made fine-tuning the wrong tool for facts in part 48 applies here with extra force, because a reasoning trace is a more persuasive wrapper around a fabricated intermediate step.
Does not transfer: the teacher's search budget. A teacher that explored five hypotheses internally and produced the trace of the successful one teaches the student the polished path and not the exploration. Students trained this way are notably brittle to adversarial inputs that the teacher would have recovered from, because they never saw the recovery.
The two recipes that mattered
The R1 paper's distillation procedure is deliberately plain: use the teacher to generate reasoning traces on a broad prompt set, filter for correctness where the answer is verifiable, and run supervised fine-tuning on the resulting traces. No reward model, no policy optimization, on the order of eight hundred thousand examples, applied to student models from 1.5B to 70B parameters (DeepSeek-AI). The result was that relatively small dense students scored at levels that previously required much larger models, and — the useful part for engineers — the students were far cheaper to serve than the teacher.
The s1 experiment tightened the budget from the other direction. With a thousand carefully selected questions and their traces, plus a decoding-time intervention the authors call budget forcing — appending an end-of-thinking token to cut deliberation short, or suppressing it to extend deliberation — a small model could be made to outperform much larger closed models on competition mathematics (Muennighoff et al.). Two lessons. First, the selection of training questions dominated the quantity: traces on hard, diverse, traceable problems beat traces on a huge pile of easy ones. Second, thinking length at inference is a controllable knob, and a student trained with traces of varied lengths can be pushed harder on the calls that deserve it.
LIMO pushed the data requirement to its absurd limit and made the same point more loudly: with 817 curated examples and a strong reasoning model as teacher, a 32B student reached high scores on mathematical reasoning benchmarks, the authors arguing that the model's reasoning ability was latent and that the demonstrations merely elicited it (Ye et al.). Whether one believes the strong reading of that claim, the engineering implication is stable: in this regime, curation quality is the lever, and data volume is not a substitute for it.
Where the loop bites back
Length inflation. The most common and expensive failure. Students trained on long traces learn that long traces are expected, and they produce deliberation on questions a one-sentence answer would satisfy. Because output tokens dominate inference cost (see part 50), this can turn a cheap model into an expensive one while looking like an improvement on accuracy benchmarks. The mitigation is to train on length-stratified traces and to evaluate with a token budget, not just an accuracy number.
Brittle reflection. Self-correction learned from traces is imitation, not verification. The literature on self-correction without external feedback found that models frequently revise correct answers to incorrect ones, because the critique step has no ground truth to anchor it (Huang et al.). A distilled student inherits the vocabulary of self-correction along with its unreliability, and it will happily announce that it found an error and then introduce one.
Contamination and provenance. Traces are generated against problems, and it is easy to generate against problems that overlap your evaluation set, or to train on a teacher's traces for a benchmark the teacher was tuned on. Everything part 36 said about contamination applies, with a twist: the traces are a second-hand artifact, so the contamination is inherited and harder to see.
Licensing. If the teacher is a commercial API, the terms of service usually govern whether its outputs may be used to train a competing model, and the distinction between "using outputs" and "distilling a model" is one contract lawyers are actively testing. This was covered in part 37 and it has not gotten simpler.
Field notes
If I were building this today, the pipeline would be: a small, hand-selected set of hard problems from my own domain with verifiable answers; traces generated by the strongest model I am permitted to use, with correct final answers only; length-bucketed, deduplicated, and stripped of any reflection step that cannot be executed; a supervised fine-tune on a small base; and an evaluation that scores accuracy, tokens per correct answer, and p95 latency side by side. Then a second evaluation on questions the student should refuse or escalate, because the trace format makes confidently wrong intermediates much more persuasive than they were before. The whole build is a week, the evidence is unambiguous, and the failure mode is visible in the cost column.
The strategic point is the one that runs through this whole stretch of the series. Reasoning capability arrived as an expensive, inference-time phenomenon and was converted, in about eighteen months, into a trainable artifact that fits on a single accelerator. That conversion is the pattern to watch: whatever the frontier does with a large compute budget this year becomes a small model's competence next year, and the teams that position themselves to capture it — with data, evaluation, and a deployment target — are the ones who get to keep the margin.
What to measure after you distill
The evaluation for a distilled reasoner is not the same evaluation you ran on the base model, because the thing that changed is not only accuracy. Five measurements, and the second through fourth matter more than the first.
Accuracy on your task. Necessary and insufficient, and it should be measured on held-out problems with verifiable answers rather than on a benchmark the teacher may have seen.
Tokens per correct answer. The composite metric that exposes length inflation, and the one number I would put on the dashboard. A student that gains three points of accuracy while tripling output length has made the system more expensive to run and slower to respond, and on many products that is a regression dressed as an improvement.
Latency percentiles. A model that reasons in long chains has a tail-latency profile that a single-shot model does not, and the p95 matters more than the median because it governs both the user experience and the capacity you have to provision for. If you cannot serve that tail, you need a budget mechanism at inference time, and the s1 result suggests that the student will tolerate one if it was trained on traces of varied lengths.
Calibration and abstention. Force the student to answer questions outside its competence and measure whether it escalates. A distillation that improves accuracy while destroying the ability to say "I don't know" has transferred confidence without transferring competence, which is the worst possible trade in any workflow where a human acts on the answer.
Behavior on adversarial and out-of-distribution inputs. Students trained on the teacher's successful traces inherit the polished path and not the exploration, so they fail differently when the problem deviates from the training distribution. Sample a set of perturbed questions — wrong preconditions, contradictory constraints, missing information — and compare the student against the base model and the teacher. This is the evaluation that most often justifies the whole technique, because a small model that escalates cleanly is more useful than a large model that invents.
One deployment note to close: reasoning distillation pairs naturally with the compression work of part 17 and the on-device constraints of part 46, and the pairing is where the real leverage sits. A 4-bit quantized 7B student that produces a medium-length reasoned answer and knows when to stop is, for an enormous class of production tasks, a better system than the frontier model it learned from, at a small fraction of the cost and with the latency of a local call. That is the shape of applied AI in 2025: not the smartest model available, but the cheapest model that is reliably right about your problem, and the techniques in this entry are how it gets built.
Works Cited
DeepSeek-AI. "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning." arXiv, 2025, arxiv.org/abs/2501.12948. Accessed 3 July 2025.
Guha, Neel, et al. "Teaching Small Language Models to Reason." arXiv, 2022, arxiv.org/abs/2212.08410. Accessed 3 July 2025.
Ho, Namgyu, Laura Schmid, and Se-Young Yun. "Large Language Models Are Reasoning Teachers." arXiv, 2022, arxiv.org/abs/2212.10071. Accessed 3 July 2025.
Hsieh, Cheng-Yu, et al. "Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes." arXiv, 2023, arxiv.org/abs/2305.02301. Accessed 3 July 2025.
Huang, Jie, et al. "Large Language Models Cannot Self-Correct Reasoning Yet." arXiv, 2023, arxiv.org/abs/2310.01798. Accessed 3 July 2025.
Muennighoff, Niklas, et al. "s1: Simple Test-Time Scaling." arXiv, 2025, arxiv.org/abs/2501.19393. Accessed 3 July 2025.
Ye, Yixin, et al. "LIMO: Less Is More for Reasoning." arXiv, 2025, arxiv.org/abs/2502.03387. Accessed 3 July 2025.
Zelikman, Eric, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. "STaR: Bootstrapping Reasoning with Reasoning." arXiv, 2022, arxiv.org/abs/2203.14465. Accessed 3 July 2025.