AI Frontiers, part 24: The prompt is not the product — in-context learning and its limits
Part 24from the AI Frontiers series · 65 parts in all
Three months after ChatGPT launched, the job title was already circulating. Companies were hiring prompt engineers, publishing prompt libraries, and treating the phrasing of a request to a language model as a proprietary asset. The intuition behind it was reasonable: if the same model produces a useful answer to one wording and nonsense to another, then the person who knows the right wording owns something valuable. And the finding that started all of this is genuinely one of the most surprising results in the history of the field — GPT-3 could perform tasks it was never trained on, given a handful of examples in the prompt (Brown et al.).
What makes this entry worth writing in February 2023 is that the research community had already spent eighteen months taking that capability apart, and the picture that emerged is much stranger than "the model follows instructions." The examples matter more than the instructions. The labels in the examples barely matter at all. And the whole thing looks less like teaching and more like pointing at a task that the model already knows how to perform, which is a very different mental model with very different implications for what is defensible as a business.
The discovery, stated precisely
In-context learning is the ability to specify a task entirely in the prompt, through demonstrations, and get correct behavior on a new input. GPT-3's paper presented it as the headline result: performance on translation, question answering and word manipulation improved as the number of in-prompt examples grew from zero to a few dozen, without any gradient update. The phrase "few-shot learning" had previously meant meta-learning with careful architecture work; here it meant literally writing examples in a text box.
The mechanism question came immediately and is still not fully settled. One hypothesis is that the model infers a task description from the examples and then executes it. A second is that the transformer's attention layers implement something functionally equivalent to a gradient step over the demonstrations, and papers on both sides of that claim appeared in December 2022 with analyses suggesting that transformer layers can express in-context gradient descent on a linear model (von Oswald et al.; Dai et al.). A third, and the one best supported by subsequent evidence, is that in-context learning is largely about locating and retrieving a skill the pretraining corpus already contained — the demonstrations act as an index rather than an instruction.
The strongest evidence for the third hypothesis is the closest thing this subfield has to a crucial experiment. Min et al. tested what actually matters in a set of demonstrations by replacing the labels with random ones. Random labels barely hurt performance across a wide range of classification and question-answering tasks. The input distribution, the label space and the format mattered; whether the labels were correct did not. A model given four examples of sentiment analysis with labels chosen at random still performed the task.
That result is worth pausing on, because it invalidates a large fraction of the prompt engineering folklore that was being written at the time. If the model is not learning from your labels, then effort spent curating perfect examples is partly wasted, and effort spent communicating what task is being requested — through the format, the phrasing, the label set — is doing the real work. Anthropic's interpretability work on induction heads pointed at the same machinery from the other direction: the specific attention circuits that implement "find this pattern elsewhere and continue it" emerge early in training and appear to be what in-context learning is built on (Olsson et al.).
Why prompt wording is a fragile asset
If the prompt's job is to locate a capability, then the sensitivity of results to prompt wording follows naturally, and the specific sensitivities that were documented are mostly about the model rather than the task.
Order. Zhao et al. showed that few-shot performance varies substantially with the order of the examples and that models have a bias toward the label that appears last in the sequence — a bias that can be corrected by calibrating the output distribution against a content-free input. Lu et al. found the same order sensitivity across tasks and models. A prompt is therefore not a specification; it is a point in a high-dimensional space whose neighborhood you have sampled a few times.
Format, decisively. Webson and Pavlick showed that models often perform a task while ignoring the explicit instruction, and that changing the demonstration format changes results more than changing the instruction's wording. Consistent delimiters, consistent label vocabulary, and consistent whitespace matter more than eloquence.
Pretraining frequency. Razeghi et al. demonstrated that a model's accuracy on numerical reasoning depends on how often the entities in the question appeared in its pretraining data — the same arithmetic performed on numbers the model had seen frequently versus rarely produced measurably different accuracy. This is the least discussed and most important finding, because it means prompt performance is a function of a corpus you cannot inspect.
Stack those three together and the honest description of a hand-tuned prompt is a configuration that happens to work for the inputs you tested, whose performance depends on elapsed time in a training corpus and on an ordering you chose arbitrarily. That is a real engineering artifact. It is not a moat.
What is genuinely valuable, and what is not
The version of this argument that says prompt engineering is worthless is wrong, and it is worth being precise about where the value sits.
Genuinely valuable, in rough order: giving the model the information it needs, which is retrieval and context construction rather than phrasing; giving it a clear output contract, so the result is parseable and checkable; giving it a worked example of the exact format you want, which is what a demonstration is actually good for; and getting the ordering right when the model's biases make it matter — the classic example being to put retrieved evidence before the question, never after. Nobody builds a durable advantage from the wording, and everybody builds a more reliable system by doing these four things well.
Not valuable, despite the effort they attract: elaborate personas, which have small and inconsistent effects; emphatic instructions in capitals and threats of consequences, which mostly change tone; long lists of prohibitions, which are worse than a small number of concrete rules; and the accumulated prompt-library approach of storing thousands of variants of "you are a helpful assistant," which is a symptom of not having an evaluation set.
There is a sharper version of the same test. Take your best prompt and rewrite it in a paragraph of your own words, keeping the examples and the output format identical. If performance holds, the wording was never the asset. If it collapses, you have found something interesting enough to investigate and probably narrow enough not to depend on.
The one place a prompt does become an asset is when it is paired with something the competitor cannot copy — proprietary retrieved context, a tool the model can call, a post-processing pipeline that validates output. Those are the parts that persist when the wording changes, which is the test worth applying: if swapping your prompt for a paraphrase destroys the product, you do not have a product.
Demonstration selection is the real lever
If the demonstrations act as an index rather than an instruction, then the question worth optimizing is which examples to include, and that turns prompt construction into a retrieval problem. Liu et al. tested this directly and found that selecting examples by similarity to the test input beat random selection, and that a simple k-nearest-neighbor search over a pool of candidate demonstrations was often better than a fixed hand-picked set. The reason is intuitive once stated: the examples that most reduce uncertainty about what task is being requested are the ones that resemble the case at hand.
This reframing matters because it changes what you build. A prompt library is a static artifact and cannot adapt. A demonstration index is infrastructure: it can grow, it can be filtered for quality and diversity, and it can be re-ranked per request. Teams that made this switch found that their accuracy on rare input patterns improved without anyone touching the wording of the instruction, which is the tell that the examples were doing the work all along.
There is a constraint that shaped how far anyone could take this in early 2023, and it is worth recording because it explains the tooling of the period. Context windows were small — a few thousand tokens for the models most people could access — so a demonstration-heavy prompt consumed the budget that retrieved context needed. The tradeoff between showing examples and showing evidence was real and painful, and it is the reason the argument about retrieval versus in-context examples in the RAG entry looks so different from the way it looks now.
What the sensitivity results imply for evaluation
A system whose output depends on the arbitrary order of its examples is a system whose performance cannot be characterized by a single number, and this has a direct consequence for anyone trying to measure what they have built.
The minimum defensible practice is to evaluate a prompt across several orderings and report a range rather than a point estimate. If accuracy moves from 71% to 84% when two examples swap positions, the honest summary is not 84%. It is that the prompt is unstable, and that the instability is a bigger risk than the level, because a user will hit the bad ordering eventually.
The stronger practice is to treat the prompt as one component in a pipeline with an evaluation set behind it — the discipline that replaces prompt collecting. Without that instrument, every wording change is an unreviewed deployment, and the person who edited the prompt last has no way to know whether they helped. Which is a strange situation for an industry that would never accept the same standard for a database migration.
The prompt-tuning detour
Meanwhile, a parallel research thread took the opposite approach and made the prompt part of the model. Prompt tuning learns continuous vectors that are prepended to the input (Lester et al.), and prefix tuning does the same inside the layers (Li and Liang). Both were attractive in 2021 for the obvious reason: dozens of parameters per task instead of billions, no change to the base model, and no sensitivity to wording because the prompt is no longer text.
The results were consistent and slightly disappointing in a way that is now standard knowledge: soft prompts match full fine-tuning on very large models but lag badly on small ones, and they inherit the base model's inability to do anything it could not already do. Those are the same limits that LoRA and QLoRA would run into later, for the same reason: parameter-efficient fine-tuning changes behavior, not knowledge. It is worth noting that this detour happened before the API-based prompting wave and reached the opposite conclusion about where the value was, which is a useful reminder that the field's tooling has oscillated between treating the prompt as configuration and treating it as a trainable parameter.
The practical discipline that should have replaced prompt engineering
If I could give February 2023 one piece of advice, it would be to stop collecting prompts and start building an evaluation set. A directory of thirty inputs with expected outputs, covering the cases you care about plus the ones that break, converts every subsequent discussion from opinion into measurement. It tells you when a model upgrade helps, when a wording change is noise, and — critically — when your pipeline is fine and the failure is in the data you retrieved rather than anything the model did.
Everything in this series that follows is a variation on that theme. Prompts are cheap, models churn, and the durable asset is the instrument that tells you whether you are getting better. The people who built that instrument in 2023 found the next two years much easier than the people who built a prompt library, and the difference was not technique. It was deciding which of the two things was actually the product.
Works Cited
Brown, Tom B., et al. "Language Models Are Few-Shot Learners." arXiv, 2020, arxiv.org/abs/2005.14165. Accessed 16 Feb. 2023.
Dai, Damai, et al. "Why Can GPT Learn In-Context? Language Models Secretly Perform Gradient Descent as Meta-Optimizers." arXiv, 2022, arxiv.org/abs/2212.10559. Accessed 16 Feb. 2023.
Honovich, Or, et al. "Instruction Induction: From Few Examples to Natural Language Task Descriptions." arXiv, 2022, arxiv.org/abs/2205.10782. Accessed 16 Feb. 2023.
Kojima, Takeshi, et al. "Large Language Models Are Zero-Shot Reasoners." arXiv, 2022, arxiv.org/abs/2205.11916. Accessed 16 Feb. 2023.
Lester, Brian, et al. "The Power of Scale for Parameter-Efficient Prompt Tuning." arXiv, 2021, arxiv.org/abs/2104.08691. Accessed 16 Feb. 2023.
Li, Xiang Lisa, and Percy Liang. "Prefix-Tuning: Optimizing Continuous Prompts for Generation." arXiv, 2021, arxiv.org/abs/2101.00190. Accessed 16 Feb. 2023.
Liu, Jiachang, et al. "What Makes Good In-Context Examples for GPT-3?" arXiv, 2021, arxiv.org/abs/2101.06804. Accessed 16 Feb. 2023.
Lu, Yao, et al. "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity." arXiv, 2021, arxiv.org/abs/2104.08786. Accessed 16 Feb. 2023.
Min, Sewon, et al. "Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?" arXiv, 2022, arxiv.org/abs/2202.12837. Accessed 16 Feb. 2023.
Olsson, Catherine, et al. "In-Context Learning and Induction Heads." arXiv, 2022, arxiv.org/abs/2209.11895. Accessed 16 Feb. 2023.
Razeghi, Yasaman, et al. "Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning." arXiv, 2022, arxiv.org/abs/2202.07206. Accessed 16 Feb. 2023.
Reynolds, Laria, and Kyle McDonell. "Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm." arXiv, 2021, arxiv.org/abs/2102.07350. Accessed 16 Feb. 2023.
von Oswald, Johannes, et al. "Transformers Learn In-Context by Gradient Descent." arXiv, 2022, arxiv.org/abs/2212.07677. Accessed 16 Feb. 2023.
Wei, Jason, et al. "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv, 2022, arxiv.org/abs/2201.11903. Accessed 16 Feb. 2023.
Webson, Albert, and Ellie Pavlick. "Do Prompt-Based Models Really Understand the Meaning of Their Prompts?" arXiv, 2021, arxiv.org/abs/2109.01247. Accessed 16 Feb. 2023.
Zhao, Zihao, et al. "Calibrate Before Use: Improving Few-Shot Performance of Language Models." arXiv, 2021, arxiv.org/abs/2102.09690. Accessed 16 Feb. 2023.