AI Frontiers, part 48: Fine-tuning versus prompting — the decision nobody makes explicitly
Part 48from the AI Frontiers series · 65 parts in all
The question arrives in nearly every applied-AI project I have consulted on, and it almost always arrives late: "should we fine-tune?" By the time someone asks, the team has usually already done one of two things by default — shipped a prompt that works well enough, or started a fine-tuning run because the prompt was not converging. Neither is a decision. The choice between adapting a model in weights and adapting it in context is the most consequential architectural call in an applied system, and it is worth making with evidence.
Part of why the conversation goes badly is that "fine-tuning versus prompting" is a false dichotomy. They modify different things. Prompting changes the model's inputs: it selects a region of behavior the model already supports. Fine-tuning changes the model's weights: it shifts which behaviors are likely in the first place. A model that does not know your product's terminology can be taught it by either route; a model that has never seen a task shape in its pretraining is a much harder proposition for prompting and a much better fit for fine-tuning. Getting the diagnosis right determines the prescription.
What each one actually changes
Prompting — including few-shot examples and retrieval — conditions the model on information supplied at inference time. The mechanistic picture is well supported: in-context learning looks like an implicit, temporary adaptation implemented in the forward pass, and a body of 2022–2023 work showed that transformer layers can implement gradient-descent-like updates over the context (von Oswald et al.; Dai et al.; Shen et al.). Crucially, that adaptation evaporates when the context is gone. Nothing is stored, so there is nothing to drift, nothing to version, and no training pipeline to maintain. The cost is paid on every call: more tokens, more latency, and a prompt that has to be kept short enough to leave room for content.
Supervised fine-tuning instead changes the weights so that a desired output becomes more likely without being shown an example. Its cost profile is the mirror image. You pay once, upfront, in data collection and training compute; afterward the behavior is free to invoke — tokens go down, prompts get shorter, latency improves. What you have bought, though, is a new artifact with all the obligations of an artifact: versioning, evaluation, rollback, and a supply chain for the training data. The teams that regret fine-tuning are almost never the ones who got worse quality; they are the ones who got quality and could no longer remember why the model behaved the way it did.
The evidence, roughly
Four findings from the literature have most shaped how I think about this.
Parameter-efficient fine-tuning beats few-shot prompting on narrow tasks. The T-Few result was the first clean demonstration at scale: with about a thousandth of the parameters trained, a PEFT model outperformed in-context learning on several benchmark families while being cheaper to run (Liu et al.). LoRA and QLoRA made the recipe accessible, so the practical question is no longer whether PEFT works but whether it is worth the pipeline. And on genuinely task-specific work, it usually is.
But few-shot fine-tuning and in-context learning are only comparable when the comparison is fair. Mosbach and colleagues' careful evaluation is the paper I hand to anyone who has read a too-good fine-tuning result: much of the reported advantage of fine-tuning over prompting at small example counts came from unequal budgets — fine-tuning trained on far more data than the prompt could show, and the comparisons ignored the hyperparameter sensitivity that hurts few-shot fine-tuning. When the example budget is matched and tuning is done sensibly, the gap narrows substantially (Mosbach et al.). The takeaway is not "prompting wins"; it is "your A/B test between the two is probably rigged by whoever configured it."
Format and style are easy; new knowledge is hard. LIMA's widely quoted result — that a thousand carefully curated instruction-response pairs produced a well-behaved assistant (Zhou et al.) — supports the position that most "alignment" fine-tuning is teaching surface form and interaction register, not capability. Follow-up work sharpened this: Gekhman and colleagues found that fine-tuning on examples containing knowledge the model did not already possess encourages hallucination, because the training objective rewards producing the unknown fact confidently (Gekhman et al.). If your goal is to teach the model facts about your business, fine-tuning is the wrong tool and retrieval is the right one. If your goal is to teach it your house style, the response schema, and how to behave when it does not know, that is exactly what fine-tuning is for.
Adapters are not full fine-tunes, and the difference is measurable. Work comparing LoRA against full fine-tuning found that low-rank adaptation both learns less and forgets less: it underperforms on genuinely new domains while preserving more of the base model's generality (Biderman et al.). A 2024 study pushed further, showing that LoRA and full fine-tuning find functionally different solutions and that LoRA's rank constrains what it can express (Zhu et al.). Practically: if you need the base model's other abilities intact and you are mostly teaching style or format, LoRA is the right choice. If you are repurposing a model substantially, budget for full fine-tuning or accept a narrower target.
A decision procedure
Five questions, in order, settle most cases.
1. Is the missing information, or the missing behavior? Missing information — facts, documents, current state — is a retrieval and prompting problem. Missing behavior — format, tone, output structure, refusal policy, task shape — is a fine-tuning candidate. Confusing these is the single most common diagnosis error, and it is why so many teams fine-tune on facts and then wonder why the model fabricates.
2. Is the task narrow and high-volume? Fine-tuning's economics are a function of call volume. A thousand calls a day justify an afternoon of pipeline work only if the prompt savings are large; a million calls a day justify it almost immediately, and the latency win probably matters more than the token win. Low-volume, high-variance work should stay on prompts and a frontier model.
3. Do you have a few hundred to a few thousand good examples? Fine-tuning needs demonstrations of the behavior you want, and quality dominates quantity far more than teams expect (LIMA is the extreme version). If you cannot produce the data, you cannot fine-tune — which is convenient, because it also means you cannot evaluate the result.
4. Can you measure the change? A fine-tune without an evaluation set is a coin flip you cannot re-run. If you do not have the private test set described in part 42, build it before training, not after.
5. What breaks when the base model improves? A prompt stack survives a model upgrade; a fine-tune is invalidated by one, because the tuned behavior now sits on a new base. Teams that fine-tune early accept re-tuning on every base-model release. That is a real ongoing cost and it belongs in the plan, not in the surprise column six months later.
The middle ground that keeps winning
The best answer in practice is frequently neither extreme, and 2024's tooling made the middle the default. A base model plus a small adapter trained per task, swapped at request time, gives you most of the token and latency savings of a fine-tune with one training pipeline and many reusable artifacts — the "one base, many adapters" pattern part 9 described. Pair it with retrieval for facts, a structured output contract for format (part 45), and a small rubric-driven evaluation harness, and you have an architecture where each adaptation mechanism is used for the thing it is actually good at.
There is also a newer and quite attractive middle: fine-tune not to change what the model knows but to change what the prompt has to say. Distilling a long, expensive, carefully engineered system prompt into weights — or distilling a frontier model's judgment on a narrow classification task into a small local model — captures the output quality of the expensive configuration at the cost of the cheap one. This is the same distillation dynamic part 9 described for capability transfer, applied to prompt engineering itself, and it is arguably the most economically significant use of fine-tuning in production systems today.
What I would write in the design doc: prompting first, because it is reversible; retrieval for knowledge, always; fine-tuning when a measurable, stable, narrow behavior needs to stop costing tokens on every call — and only once an evaluation set exists to prove the result. The teams I have seen make this decision well treated fine-tuning as a cost-optimization technique with a quality constraint, rather than as a capability upgrade. That framing gets almost every case right.
The data work is the decision
If a fine-tune is on the table, the dataset becomes the most consequential artifact the team produces, and it deserves more design attention than the training configuration. Four points from practice.
The best source of training data is production output you edited. Any workflow with a human review step is generating preference data for free: the draft the model produced and the version the human sent are a natural pair, and the diff between them is a precise, domain-specific statement about what the model got wrong. Teams that capture those pairs and fine-tune on them consistently do better than teams that commission a synthetic dataset, because the corrections target the errors the system actually makes. Capture the pair at the moment of the edit, and store the original alongside the correction.
Rejection sampling is the cheapest way to manufacture demonstrations. Rather than paying someone to write perfect examples, generate many candidates for each input, score them with a verifier or a rubric, and keep the ones that pass. The self-taught reasoner recipe formalized this shape for reasoning tasks (Zelikman et al.), and it generalizes to almost any task with a checkable output. The failure mode to guard against is selection bias: if your verifier only recognizes one style of correct answer, you will train the model to produce only that style, which is fine until it is not.
Hold out data before you look at it. The average applied fine-tune is validated on examples the trainer has already inspected, the base model has already seen in some form, or both. Split first, keep the held-out set out of all tooling including the notebook where you inspect data, and treat contamination in the training corpus as your problem to check rather than the provider's. This is the same discipline that part 36 argued for on the benchmark side, applied to your own house data.
Version the dataset like code. A fine-tuned model without its training data is not reproducible, and the data will be the thing that goes missing. Its own repository, a hash per release, and a short note on provenance for each batch. If you ever have to answer "why does the model behave this way in this case," the answer is in the dataset.
The break-even arithmetic
Fine-tuning is a capital expenditure that buys lower operating cost, so it has a break-even point and it can be estimated in a spreadsheet before any GPU time is booked. Count four numbers. A: the prompt tokens and latency you save per call, in dollars. B: the volume of such calls per month. C: the one-time cost of building the pipeline and the dataset — data collection, annotation, engineering, and training runs, measured in person- days rather than GPU hours, because that is where the cost actually sits. D: the expected useful lifetime of the tuned artifact before a base-model upgrade invalidates it, which in 2025 is a matter of months.
If A times B times D is not materially larger than C, the fine-tune does not pay, and the honest move is to make the prompt better and revisit the question next quarter. When it does pay, the payback is usually visible within weeks, which is why high-volume narrow workloads are where fine-tuning consistently wins. The teams that get burned are those who estimate C from the training run and forget the data work, and estimate D as "forever" because they did not plan for base-model upgrades. Both errors are optimistic in the same direction, and they compound.
Works Cited
Biderman, Dan, et al. "LoRA Learns Less and Forgets Less." arXiv, 2024, arxiv.org/abs/2405.09673. Accessed 6 Feb. 2025.
Dai, Damai, et al. "Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta-Optimizers." arXiv, 2022, arxiv.org/abs/2212.10559. Accessed 6 Feb. 2025.
Gekhman, Zorik, et al. "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" arXiv, 2024, arxiv.org/abs/2405.05904. Accessed 6 Feb. 2025.
Liu, Haokun, et al. "Few-Shot Parameter-Efficient Fine-Tuning Is Better and Cheaper than In-Context Learning." arXiv, 2022, arxiv.org/abs/2205.05638. Accessed 6 Feb. 2025.
Mosbach, Marius, et al. "Few-Shot Fine-Tuning vs. In-Context Learning: A Fair Comparison and Evaluation." arXiv, 2023, arxiv.org/abs/2305.16938. Accessed 6 Feb. 2025.
OpenAI. "Fine-Tuning." OpenAI Platform Documentation, 2024, platform.openai.com/docs/guides/fine-tuning. Accessed 6 Feb. 2025.
Shen, Lingfeng, et al. "Do Pretrained Transformers Learn In-Context by Gradient Descent?" arXiv, 2023, arxiv.org/abs/2310.08540. Accessed 6 Feb. 2025.
von Oswald, Johannes, et al. "Transformers Learn In-Context by Gradient Descent." arXiv, 2022, arxiv.org/abs/2212.07677. Accessed 6 Feb. 2025.
Zhou, Chunting, et al. "LIMA: Less Is More for Alignment." arXiv, 2023, arxiv.org/abs/2305.11206. Accessed 6 Feb. 2025.
Zelikman, Eric, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. "STaR: Bootstrapping Reasoning with Reasoning." arXiv, 2022, arxiv.org/abs/2203.14465. Accessed 6 Feb. 2025.
Zhu, Dan, et al. "LoRA vs Full Fine-Tuning: An Illusion of Equivalence." arXiv, 2024, arxiv.org/abs/2410.21228. Accessed 6 Feb. 2025.