Trust is earned, not given

A different perspective

2023-03-30 · Projects

AI Frontiers, part 25: Instruction tuning — from FLAN to Alpaca

Part 25from the AI Frontiers series · 65 parts in all

On 13 March 2023, a group at Stanford released Alpaca: a 7B model fine-tuned to follow instructions, trained on 52,000 examples generated by a larger model, for a reported cost under $600 (Taori et al.). The response was disproportionate to the artifact and exactly proportionate to what it demonstrated. A model that behaved like an assistant — not a base language model that continues text, but something that answers requests — could be produced by anyone with a GPU rental and a weekend, from another model's output, at a price that made the previous year's assumptions about capital requirements look wrong.

Alpaca was possible because of eighteen months of work on instruction tuning, a research program that had answered a deceptively simple question: given a base model that has read the internet and can complete any text, how do you make it do what you ask? The answers, arriving in stages between 2021 and 2023, are the subject of this entry, and they explain most of what happened to the industry afterwards.

The first answer: demonstrate many tasks in a fixed form

FLAN took the boldest version of the idea. Take a large pretrained model, fine-tune it on a large collection of tasks each phrased as an instruction — dozens of natural language processing datasets, converted into templates — and then measure performance on tasks it was never trained on (Wei et al., "Finetuned Language Models Are Zero-Shot Learners"). The result that made everyone pay attention: a 137B-parameter model that had been instruction-tuned outperformed a 175B-parameter model that had not, on held-out tasks, with no demonstrations in the prompt.

That is a remarkable claim. Instruction tuning did not add knowledge; the base model had all the knowledge. What it added was the behavior of treating an instruction as a request to be fulfilled rather than as text to be continued. The gain came from the format and the diversity of tasks, and the reason it generalized is that the model was not learning individual tasks so much as learning what kind of object an instruction is.

T0 arrived in parallel with the same conclusion from a different design (Sanh et al.): prompted multitask training across a large, deliberately diverse set of tasks produces strong zero-shot generalization. The two papers together established the recipe, and the key word in both is diversity. Fine-tuning on a thousand examples of one task teaches that task. Fine-tuning on a hundred tasks in one format teaches the format, and the format is what transfers.

The second answer: humans, and then a proxy for humans

Contemporaneously, a different group attacked the problem with human preferences. InstructGPT (Ouyang et al.) trained a model on demonstrations of desired behavior, then trained a reward model on human comparisons, then optimized the policy against that reward with reinforcement learning. Two results from that paper shaped everything that followed. First, the resulting model was preferred by human raters over a much larger unaligned model — 1.3B parameters beating 175B on the metrics that mattered. Second, and less quoted, alignment came with a measurable cost on some academic benchmarks, which is the first clear sighting of a tradeoff that has been re-litigated in every generation since.

The RLHF pipeline is expensive because human comparison data is expensive, which is why the field spent 2023 looking for ways to substitute models for people in the feedback loop. That search is the subject of a later entry in this series; here the relevant point is that instruction tuning and preference learning were converging on the same conclusion from opposite directions. Supervised fine-tuning teaches the model what an instruction looks like. Preference learning teaches it which response a person would prefer among several. An assistant needs both, and the second is where the difficulty lives.

The third answer: scale the collection, not the cleverness

By late 2022 the field had hundreds of instruction datasets of varying quality and no principled sense of which parts mattered. The Flan Collection paper (Longpre et al.) did the unglamorous work of finding out. It consolidated a large set of public tuning collections, swept the design choices — how tasks are mixed, whether prompts are templated, how many examples per task, whether inputs are used at all — and reported what actually drove performance.

The findings were deflating in the way good empirical work often is. Task diversity and mixing mattered more than any individual dataset's size. Prompt formatting mattered a great deal, and including the zero-shot form of each task during training mattered more than expected. And a handful of cheap design decisions were worth more than a new architecture. Chung et al. had reached compatible conclusions scaling instruction tuning up to Flan-T5 and Flan-PaLM, and the combination of the two lines of work turned instruction tuning from a craft into a recipe with knobs you could reason about.

Meanwhile, the task-format research had already established that the wording of instructions generalizes across tasks when the form is held constant (Mishra et al.), which is the result that makes a single global tuning run sensible at all. Super-NaturalInstructions (Wang et al.) pushed the same idea to over a thousand tasks with structured task definitions, and OPT-IML (Iyer et al.) applied instruction tuning to a different model family and reproduced the gains. By early 2023 the conclusion was robust: format plus diversity plus scale equals a generalized instruction-follower.

And then the data generated itself

The step that made Alpaca possible was turning the data problem into a generation problem. Self-Instruct (Wang et al.) starts from a small seed set of instructions, has the model generate new instructions and their inputs, filters the results for validity and diversity, and then uses the survivors to fine-tune. The paper's contribution was not just the method but the argument: you do not need a human to write 50,000 instruction pairs, you need a human to write a few hundred and a model to expand them.

Alpaca applied exactly this to a 7B base model, using a larger model's outputs as the teacher, and the cost figure is what travelled. Six hundred dollars of fine-tuning, a few hours on rented GPUs, and an instruction-following model that beat the previous generation's comparably sized models on the evaluations people cared about. The obvious caveats were stated in the release and largely ignored: the evaluation was informal, the safety filtering was minimal, the model hallucinated, and its outputs were released for research rather than for use. The capability claim was modest. The distributional claim was not, and that is why it mattered.

What makes an instruction example good

Once the collection is the product, the question of what belongs in it becomes the central engineering problem, and the empirical answers from this period are unusually consistent.

Diversity beats volume. The Flan Collection's sweeps showed that mixing many task types mattered more than adding more examples of the task types already present, and the mechanism is the one FLAN established: the model is learning the format of an instruction, and it learns the format from variety. A dataset with ten thousand paraphrases of the same task teaches one thing. A dataset with a thousand tasks teaches a general disposition.

Include the zero-shot form. Training each task both with and without examples turned out to matter more than expected, because the deployment condition is usually an instruction with no examples, and a model tuned only on example-bearing prompts is being asked to do something it has not practiced.

Filter, then filter again. Self-Instruct's pipeline rejected generated instructions that duplicated existing ones or failed basic format checks, and the rejection rate was high. The instinct to keep everything because it was expensive to generate is exactly wrong when generation is cheap and judgment is not, and the difference between a mediocre and a good instruction dataset in 2023 was usually the strictness of its filter rather than the size of its generator.

Watch the format, not the eloquence. Downstream evaluations repeatedly showed that consistency of delimiters, role markers and output conventions drove behavior more than refined prose in the instruction itself. This is why templated datasets outperform hand-written ones at equal size, and why the template is worth more attention than the sentence.

Why Alpaca was small on purpose

Alpaca used a 7B base model, and the choice was not a limitation so much as the point. A model that size runs on a single rented GPU, fine-tunes in hours for hundreds of dollars, and can be served by a small team. The claim being made was not that a 7B instruction-tuned model rivals the largest available system. It was that the difference between a base model and an assistant could be crossed at that scale and at that price.

The ceiling is real and worth stating, because a great deal of 2023 disappointment came from ignoring it. Instruction tuning can shape behavior and cannot add knowledge or reasoning ability, so a small model becomes a well-mannered small model. The subsequent history of the field is largely the story of closing that specific gap — distilling stronger models' behavior into small ones, which is the pattern the small-model entry and later the reasoning-model entries both describe.

What Alpaca settled, permanently, is the deployment question. Not every task needs a frontier model, most tasks need one that follows instructions reliably and costs almost nothing, and the tooling to produce such a model had just become available to anyone. The economics of applied AI turned on that realization more than on anything in a benchmark table.

What instruction tuning does not do

It is worth being explicit about the limits, because the 2023 hype cycle blurred all of them together.

It does not add knowledge. Instruction tuning shapes how the model deploys what it learned during pretraining. A model that never saw a fact does not acquire it by being told to be helpful, which is why the limitation shows up as confident fabrication rather than admission of ignorance — the subject of the next entry in this series.

It does not fix reasoning. The gains on reasoning benchmarks from instruction tuning are small relative to the effort, and the gains on formatting, following output conventions and refusing inappropriate requests are large. It teaches the interface, not the intellect, and confusing the two is how teams end up shipping a well-mannered model that is wrong.

It does not remove the need for preference data. A model tuned only on demonstrations imitates them, including their weaknesses. Preference learning is what turns an imitator into something that is reliably helpful rather than merely fluent, and the cost of that data is the reason the pipeline remains the expensive part of the process.

It does not survive badly. Instruction-tuned models forget, and they can be over-tuned into rigid behavior — refusals that fire on benign requests, formatting that cannot be varied, a loss of the base model's breadth. The check on all of this is the same instrument the rest of the series keeps returning to: a held-out evaluation set that covers both what you wanted to improve and what you did not want to break.

The world Alpaca implied

Three things became obvious in the weeks after Alpaca that had not been obvious before.

First, the last mile of model quality was cheap. Pretraining a capable base model required capital that almost nobody had; turning one into an assistant required a few thousand dollars and a dataset. That is an enormous reduction in the barrier to entry, and it is why the open ecosystem exploded over the following months.

Second, the data was the valuable part, and most of it was model output. LIMA would later make this explicit by fine-tuning on a thousand carefully curated examples, and the pattern that emerged — a small amount of high-quality instruction data has enormous leverage — is the same lesson as the Phi line's, arriving from a different direction. If the quality of the demonstrations is what matters, and demonstrations are cheap to generate, then the differentiator is judgment about which ones to keep.

Third, and least comfortably, every instruction-tuned model is downstream of someone else's model and someone else's terms. The distillation question — whether it is permissible to train on a commercial model's outputs, what attribution is owed, what happens when the teacher changes its terms — was raised in the first week and has never been settled. Alpaca's release set the pattern the open ecosystem has followed since: a capability, a cheap reproduction, an unresolved legal question, and a great deal of progress in the meantime.

Works Cited

Chung, Hyung Won, et al. "Scaling Instruction-Finetuned Language Models." arXiv, 2022, arxiv.org/abs/2210.11416. Accessed 30 Mar. 2023.

Iyer, Srinivasan, et al. "OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization." arXiv, 2022, arxiv.org/abs/2212.12017. Accessed 30 Mar. 2023.

Longpre, Shayne, et al. "The Flan Collection: Designing Data and Methods for Effective Instruction Tuning." arXiv, 2023, arxiv.org/abs/2301.13688. Accessed 30 Mar. 2023.

Mishra, Swaroop, et al. "Cross-Task Generalization via Natural Language Crowdsourcing Instructions." arXiv, 2021, arxiv.org/abs/2104.08773. Accessed 30 Mar. 2023.

Ouyang, Long, et al. "Training Language Models to Follow Instructions with Human Feedback." arXiv, 2022, arxiv.org/abs/2203.02155. Accessed 30 Mar. 2023.

Sanh, Victor, et al. "Multitask Prompted Training Enables Zero-Shot Task Generalization." arXiv, 2021, arxiv.org/abs/2110.08207. Accessed 30 Mar. 2023.

Taori, Rohan, et al. "Alpaca: A Strong, Replicable Instruction-Following Model." Stanford Center for Research on Foundation Models, 13 Mar. 2023, crfm.stanford.edu/2023/03/13/alpaca.html. Accessed 30 Mar. 2023.

Wang, Yizhong, et al. "Self-Instruct: Aligning Language Models with Self-Generated Instructions." arXiv, 2022, arxiv.org/abs/2212.10560. Accessed 30 Mar. 2023.

---. "Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks." arXiv, 2022, arxiv.org/abs/2210.10040. Accessed 30 Mar. 2023.

Wei, Jason, et al. "Finetuned Language Models Are Zero-Shot Learners." arXiv, 2021, arxiv.org/abs/2109.01652. Accessed 30 Mar. 2023.