AI Frontiers, part 42: Evaluation-driven development β building your own test set
Part 42from the AI Frontiers series · 65 parts in all
Half the entries in this series end with the same prescription: measure it against your own data. That is easy to say and specific to do, and the difference between teams that improve a system over eighteen months and teams that churn through models and prompts without converging is almost entirely whether they did it. So this entry is the practical version.
The claim is narrow. A few hundred examples of your actual task, with known outcomes, a written rubric and a cost column, will answer almost every question that arises during development, and it costs about a week of work. Public benchmarks will answer almost none of them, which is the argument of part 36 made concrete.
What it is for, stated as questions
An evaluation set is a tool for answering recurring questions, and it is worth enumerating them because the design follows from the questions.
Did the change I just made help? A model upgrade, a prompt revision, a different chunk size, a retriever swap β each is a hypothesis, and without an instrument every one of them is decided by anecdote from whichever example someone happened to test.
Is this difference noise? Language model outputs vary with sampling; a metric computed on thirty examples has a confidence interval wide enough to swallow most improvements. Knowing the interval prevents the common failure of chasing a two-point gain that is not there.
Which component failed? When an answer is wrong, the cause could be retrieval, the prompt, the model or the post-processing. An evaluation set with per-stage instrumentation localizes the failure, which is the difference between fixing it and re-prompting hopefully.
Is the cheaper configuration good enough? The decision to route a request class to a smaller model or a shorter context is a quality-versus-cost trade, and making it requires knowing how much quality is actually lost. The routing and cost work later in this series depends entirely on being able to answer this.
Would a new model change our answer? Vendors release constantly, and the question of whether to migrate should be a measurement rather than a mood.
Building it, in a week
The construction is more straightforward than the literature implies, provided one accepts that the first version will be imperfect.
Sample from real inputs. The single most important decision. Pull from production logs, support tickets, or the actual queries people typed β never from invented examples, because invented inputs are systematically cleaner, better formed and easier than the real distribution. A test set of imagined cases measures the imaginer.
Cover the distribution, including its edges. Stratify by whatever dimensions matter: input length, topic, language, user type, and β critically β include the failure cases you have already seen, since those are the ones that will recur. A set drawn purely at random from traffic is dominated by easy cases and will report a high score while missing everything that matters.
Write down what correct means. For each example, record the expected outcome and a short note about why. Where the judgment is genuinely subjective, record the rubric that distinguishes acceptable from unacceptable β is it grounded in the source, is it complete, is it actionable β and have a second person apply it independently to a sample. Agreement between two people on the rubric is the evidence that the metric means anything at all.
Separate the tasks from the conversation. Test the retrieval step, the generation step and the end-to-end behavior separately, because they fail for different reasons and a single end-to-end number cannot tell you which one moved.
Hold something back. Keep ten or twenty percent of the examples out of day-to-day use. They will be the honest estimate when a decision matters, because everything you look at regularly becomes a target.
The failure modes of an evaluation set
Four things go wrong, and each is preventable.
Goodharting. The moment a set is the thing people optimize against, they optimize it. This is not misconduct; it is what happens when the metric is visible. The remedies are rotation, a held-back subset, and β most importantly β reading failures rather than only scores, since a failure that a person has read is harder to explain away.
Staleness. Traffic changes, products change, and a set built a year ago measures a distribution that no longer exists. Refresh a fraction each quarter and keep retired examples for regression checks.
One-number averaging. A single accuracy figure hides the fact that the system improved on common cases and got worse on the rare ones that matter commercially. Report per-category numbers and watch the categories that carry the business.
No cost column. The critique of agent benchmarking that arrived later made this point for agents; it applies universally. A configuration that is one point better and four times more expensive is a different decision for a nightly batch job than for an interactive endpoint, and the results table should make that visible rather than leaving it to whoever remembers to ask.
Judge models, and why the obvious approach fails
Manual labeling does not scale, so teams reach for a model as a judge β and introduce a set of biases that the methodology literature documented in detail. Judges prefer longer answers, favor outputs from their own model family, are sensitive to the order in which two candidates are presented, and are unreliable on questions requiring multi-step reasoning (Zheng et al.).
None of that makes model judging useless. It makes it a measurement instrument with known error characteristics, which is a normal engineering situation. The practices that follow are unremarkable: randomize presentation order so position bias averages out; report agreement between the judge and a human-labeled subset, so the judge's error rate is a known quantity rather than an assumption; use detailed rubrics with specific criteria rather than a global preference question, since agreement rises substantially when the judge is asked about one property at a time (Liu et al., G-Eval; Kim et al., Prometheus); and never use the same model to generate and to judge without acknowledging that its preferences are self-favorable.
The honest framing is that a judge model is a cheap approximation of a human rater, and it should be calibrated against human ratings on a sample that is re-measured periodically. A judge whose agreement with people has not been checked in six months is a number generator.
The offline-online gap, and when to trust each
An offline evaluation set measures the system on inputs you collected earlier. Production measures it on inputs people are actually sending now, in an interface where they can see the answer and change their behavior in response. Those are different measurements, and the gap between them is where a lot of confident offline results go to die.
The mechanisms are well understood from a decade of work on counterfactual evaluation in traditional machine learning (Bottou et al.) and from the experimental-design literature in software (Kohavi et al.). Offline evaluation cannot see feedback loops: if the system starts answering a question, users stop asking it. It cannot see interface effects: the same answer presented differently produces different satisfaction. It cannot see distribution shift: prompt patterns change as users learn what works. And it cannot see the long tail of rarely occurring inputs that a long-running system accumulates.
Which produces a division of labor, not a winner. Offline evaluation is the right instrument for technical questions with objective answers β did retrieval find the passage, does the output parse, did the cost go up β and for ruling out regressions before they reach anyone. Online measurement is required for anything involving user preference, engagement or satisfaction, and for verifying that an offline improvement survives contact with real traffic. The order is: fix what is objectively broken offline, then measure what is subjectively better online, and never let an online metric substitute for a correctness check.
Two cautions about the online half. Guardrail metrics matter as much as the metric you are trying to move, because a change can improve the target while degrading something you did not think to watch β trust in the answers, escalation to a human, task completion rather than engagement. And novelty effects are real: a new interface feature often shows a temporary lift that decays, so experiments need to run long enough to see the plateau rather than the spike.
One statistical habit is worth adopting before it bites. On a set of two hundred cases, a two-point difference in accuracy is roughly a handful of examples and is indistinguishable from noise. Report confidence intervals, or at minimum report the number of cases in each category alongside the score, so that the reader can see that a category with nine examples cannot support the conclusion being drawn from it. Most of the false discoveries in applied AI work are differences between five and eight examples that somebody believed because the number had a decimal point in it. The statistical treatment of evaluation in NLP has been written down for years (Dror et al.) and the practical version is simply to say how many cases were involved.
Making it operational
An evaluation set that lives on someone's laptop is a research artifact. Turning it into part of the engineering process takes four habits, none elaborate.
Version it. Cases are added, retired and corrected, and a metric that changes because the test set changed is not a measurement of the system. Store the set as data with a version identifier, and record which version produced each reported result. This is the same discipline as versioning the data pipeline in part 40, applied to the measurement side.
Run it automatically. Wire the evaluation into the deployment path so that a prompt change, a model swap or a retrieval change produces a report before it reaches users. The report should show per-category results, the cost delta, and the specific cases that regressed β not a single aggregate number. A change that fails the guardrail categories should not be mergeable without an explicit decision to accept the tradeoff.
Review failures on a schedule. Once a week, read twenty failures. Not the summary of the failures, the failures. This is the highest-yield half hour in the process, because the aggregate tells you whether something moved and only the individual case tells you why. Every substantive improvement in every system I have worked on came out of someone reading a case and noticing a pattern.
Give it an owner. Evaluation sets decay because they belong to everyone, which means they belong to nobody. A named person responsible for the set's health, its refresh schedule and its honesty is the difference between a maintained instrument and an abandoned script. This is unglamorous organizational advice and it is the actual prerequisite for everything above.
What changes once you have one
The most striking effect is cultural rather than technical. Decisions stop being arguments. The person with the strongest opinion no longer wins; the person who ran the evaluation does, and the disagreements that remain are about what the numbers mean rather than about what happened, which is a much more productive thing to argue about.
Engineering velocity rises because a change can be evaluated in minutes rather than through a week of canary monitoring that mostly measures whether anything catastrophic happened. Regressions get caught before deployment rather than after. And the team develops a shared, accurate mental model of what the system can and cannot do, which is the thing that prevents the two most common failure modes of applied AI projects β overconfidence in a system that produces fluent output, and abandonment of a system that is actually working but fails in a memorable way.
There is a second-order benefit that matters for everything else in this series. Once you have a reliable local instrument, every optimization becomes accessible: you can evaluate quantization, a smaller model, a shorter context, a cheaper reranker, a routing policy β each of which is a trade and none of which can be assessed without a measurement. The teams doing the interesting efficiency work in 2024 were the teams that could measure what it cost them in quality. That is the whole prerequisite, and it is a week of work.
Works Cited
Bavaresco, Anna, et al. "LLMs Instead of Human Judges? A Large Scale Empirical Study Across 20 NLP Evaluation Tasks." arXiv, 2024, arxiv.org/abs/2406.18403. Accessed 1 Aug. 2024.
Bottou, LΓ©on, et al. "Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising." arXiv, 2012, arxiv.org/abs/1209.2355. Accessed 1 Aug. 2024.
Chang, Yupeng, et al. "A Survey on Evaluation of Large Language Models." arXiv, 2023, arxiv.org/abs/2307.03109. Accessed 1 Aug. 2024.
Dror, Rotem, et al. "The Hitchhiker's Guide to Testing Statistical Significance in Natural Language Processing." arXiv, 2017, arxiv.org/abs/1711.01731. Accessed 1 Aug. 2024.
Kohavi, Ron, et al. "Online Controlled Experiments at Large Scale." Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2013. Accessed 1 Aug. 2024.
Kim, Seungone, et al. "Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models." arXiv, 2023, arxiv.org/abs/2310.08491. Accessed 1 Aug. 2024.
Liu, Yang, et al. "G-Eval: NLG Evaluation Using GPT-4 with Better Human Alignment." arXiv, 2023, arxiv.org/abs/2303.16634. Accessed 1 Aug. 2024.
Pineau, Joelle, et al. "Improving Reproducibility in Machine Learning Research." arXiv, 2020, arxiv.org/abs/2011.03395. Accessed 1 Aug. 2024.
Sculley, D., et al. "Hidden Technical Debt in Machine Learning Systems." Advances in Neural Information Processing Systems, 2015. Accessed 1 Aug. 2024.
Zheng, Lianmin, et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv, 2023, arxiv.org/abs/2306.05685. Accessed 1 Aug. 2024.