Trust is earned, not given

A different perspective

2024-02-08 · Projects

AI Frontiers, part 36: Benchmarks and contamination β€” measuring models honestly

Part 36from the AI Frontiers series · 65 parts in all

Every framework in this series depends on the ability to tell whether a change made things better. That is a measurement question, and by early 2024 the field's measurement instruments were in a bad state β€” not because they were badly designed in the first place, but because they had been used so effectively as targets that they stopped measuring what they were built to measure.

The mechanism is simple and unavoidable. A benchmark's value comes from being hard, public and comparable. Its being public means it is in the training data of every model built afterwards. Its being hard means it is exactly what teams optimize for. The result is that a high score is evidence of something, and the field spent 2023 working out what β€” mostly, that the number is a joint measurement of capability and of how much of the test set was in the training corpus.

The instruments, and what each one is good for

It is worth separating the instruments because they fail differently.

Static knowledge benchmarks β€” MMLU (Hendrycks et al.), BIG-bench (Srivastava et al.) and their descendants β€” present multiple-choice questions with known answers. They are cheap, automatic and comparable, and they measure recall-shaped capability well. They are also the most easily contaminated, since a question-answer pair is trivially memorizable, and the most vulnerable to format exploitation, since a model can be tuned to produce the right letter without the right reasoning.

Task-with-verifier benchmarks β€” code on unit tests, mathematics with numeric answers β€” are far more robust, because the answer has structure and the model must produce it rather than recognize it. HumanEval and MATH are the canonical examples, and the reason the field migrated toward verifiable domains is partly that they resist the contamination problem. They are not immune: a model that has memorized a problem's solution generalizes no better than one that guessed.

Human-preference evaluations β€” Chatbot Arena's Elo, MT-Bench style pairwise judging β€” measure something the others cannot, which is whether people like the output. They bring their own distortions, documented candidly in the methodology paper itself (Zheng et al.): judges favor longer answers, prefer their own family's style, are order-sensitive, and are unreliable on hard reasoning. As a complement to automatic benchmarks they are valuable; as a headline number they invite optimization for pleasantness.

Behavioral test suites, following CheckList (Ribeiro et al.), enumerate capabilities with targeted cases rather than averaging over a large set β€” which surfaces specific failures that an aggregate score hides. This is the instrument most teams should build for themselves, and the one almost nobody does.

How contamination actually gets in

The public discussion tends to assume deliberate cheating. The reality is more mundane, and the pathways matter because they determine what detection can catch.

Web-scraped corpora contain test sets as a side effect. MMLU questions appear in slide decks, tutorials, blog posts and forum threads; code benchmarks appear in repositories; the answers travel alongside. Filters catch some of it β€” benchmarks are often excluded by string matching on their distinctive formatting β€” and string filters are easily defeated by paraphrase, by translation, or by the question appearing inside a longer document that does not match the filter's pattern.

Synthetic data introduces a second pathway, and it is the one that compounds. Generating instruction data from a model that was trained on a benchmark can regenerate the benchmark in altered form, which then enters the next model's training set. The generated version no longer matches the filter and is no easier for a model to solve, so the contamination becomes untraceable while remaining real β€” a mechanism worth keeping in mind alongside the model-collapse discussion in part 21.

Fine-tuning introduces the third. A model fine-tuned on a public dataset of instructions that happens to include benchmark items has effectively trained on the test set, and nobody involved made a deliberate decision to do so.

The measurement work of this period established how widespread the effect is (Sainz et al.) and produced detection methods β€” comparing a model's performance on a benchmark with and without access to its own training data, or with perturbed versions of the questions (Golchin and Surdeanu; Oren et al.); checking for verbatim reproduction (Yang et al.); and statistical tests that estimate contamination from the pattern of correct answers on long-tail items (Deng et al.). The consistent finding is that contamination is common, that it inflates scores by amounts that matter, and that publishing the test set makes it worse β€” which is why several benchmark maintainers moved to withholding items and distributing them under agreements.

Why validity is the deeper problem

Contamination is a defect in the instrument's calibration. Construct validity is a defect in the instrument's design, and it is the more serious of the two because it cannot be fixed with better filtering.

The question validity asks is whether a benchmark measures the thing it claims to measure. Accuracy on a set of exam questions is a proxy for competence at tasks that resemble those questions; the proxy assumes that the thing being measured is stable across the shift from answering a question to doing work. Raji et al.'s critique of the everything-benchmark tradition and the measurement-modeling literature from the fairness community (Jacobs and Wallach) make the point precisely: a single number aggregates over a construct that has many dimensions, and the aggregation embeds assumptions nobody stated.

Two practical consequences follow. First, an improvement on a benchmark does not establish an improvement on your task, and the burden of proof is on whoever claims the transfer. Second, and more usefully, the reverse is often true β€” a model that is worse on a benchmark can be better on your task, and you will only discover that if you have your own measurement.

How leaderboards shape behavior

Measurement instruments do not merely observe; they create incentives, and the incentives in this period were visible to anyone watching the release notes.

The pattern was consistent. A model would be published with a table showing the benchmarks on which it led, a footnote for the ones where it did not, and a note that some comparisons used "best-of-n" sampling or a particular prompt format. Each of those choices is defensible individually and the combination is a selection process: the reported number is the maximum over a set of configurations, while the baseline is often reported under a single one. The industry-wide aggregator that emerged β€” the Open LLM Leaderboard β€” helped by standardizing the harness, and it also made the target explicit and public, which is exactly the condition under which contamination flourishes.

A subtler effect is that benchmarks pull research toward whatever is measurable. Question answering over a fixed set of subjects is measurable; whether a system is honest about its uncertainty, whether it degrades gracefully, whether a user can tell when it is wrong β€” those are not, and they accordingly receive less optimization pressure. The parts of a model's behavior that nobody benchmarks are the parts that nobody improves, which is a strong argument for measuring a much wider surface than accuracy.

There is also a genuine market failure. Benchmarking well is expensive and hard to differentiate commercially, because the honest result is usually "the differences are within noise on most tasks." A vendor selling evaluation tooling has an incentive to produce confident rankings, and a lab being evaluated has an incentive to publish favorable ones. The result is a lot of numbers and very little knowledge, which is the situation the field is still in.

What a private evaluation set actually buys

The prescription that follows is unglamorous and, in my experience, the single highest-return investment in any applied AI system: build a small evaluation set from your own traffic, with known correct outcomes, and never train on it.

A few hundred examples is enough to detect a regression and to compare two models. The construction matters more than the size: sample real inputs rather than invented ones, because invented inputs are systematically easier and cleaner than production; include the failure cases you have actually seen, because those are the ones you will see again; label with a written rubric so the labels are reproducible; and hold back a portion that nobody looks at until you are making a final decision.

What this buys is the ability to answer questions that no leaderboard can. Does the new model version help on our data? Did the prompt change help or was that noise? Is the retrieval step or the generation step responsible for this failure? Is the vendor's claim about quality relevant to our workload? Those questions arise weekly, and without a local instrument every one of them is answered by vibes.

What it does not buy is a number you can put on a slide. That is the reason it is underinvested, and it is also the reason the teams with one make better decisions. A public benchmark is a fair way to compare two systems on someone else's task. It is a poor way to compare two systems on yours.

The three practices that make it work

Three habits turn an evaluation set from an artifact into a process.

Refresh it. An evaluation set becomes a target the moment people start optimizing against it, and it becomes stale as traffic changes. Rotate a fraction of the cases every quarter, and keep the retired ones for regression testing. This is exactly what the public benchmark maintainers do, applied locally, and it is the only way to keep a set honest over a long period.

Report cost alongside accuracy. A configuration that is two points better and three times more expensive is a different decision for a nightly batch job than for a chat endpoint. The agent-benchmark critique that arrived later in the year made this argument for agents; it applies just as well to any change you are evaluating, and including a cost column in the results table is the cheapest way to stop making the mistake by omission. The serving optimizations in this series only make sense if you are measuring the thing they optimize.

Make failure analysis routine. Aggregates tell you whether something moved; only reading individual failures tells you why. A weekly sample of twenty failures, read by someone who can change the system, will find more improvement than a month of score-watching. This is the same lesson the retrieval and hallucination entries arrived at from their own directions: the aggregate score is the alarm, and the individual case is the diagnosis.

Capability elicitation, which complicates everything

One more problem deserves mention because it undermines the naive reading of every score. What a model appears capable of depends heavily on the scaffolding around it: the prompting strategy, the number of samples, whether it is allowed tools, whether chain-of-thought is elicited, whether answers are parsed and retried.

This is not a small correction. Work on emergent abilities showed that apparent discontinuities in capability can be artifacts of the metric and the prompt rather than sharp changes in the underlying model (Schaeffer et al.), and the agent literature demonstrated repeatedly that an interface change moved benchmark scores substantially with identical weights. Which means two teams reporting different numbers for the same model are frequently both correct, and a score without the full configuration is not a measurement of the model at all β€” it is a measurement of the model and everything surrounding it, reported without a description of the surroundings.

The practical implication is that when you evaluate a model on your own task, you are measuring your whole system rather than the model's intrinsic ability, which is exactly the quantity you need for a decision. It also means a fair comparison across models requires holding the scaffolding constant, and that a paper claiming a 5-point improvement should be read as a claim about a particular configuration until shown otherwise.

None of this requires a research budget, and all of it requires deciding that measurement is part of the engineering rather than an exercise for the marketing page. The field spent 2023 learning that its public instruments were compromised. The teams that came out of that year in good shape were the ones that had never depended on them.

Works Cited

Bowman, Samuel R., and George E. Dahl. "What Will It Take to Fix Benchmarking in Natural Language Understanding?" arXiv, 2021, arxiv.org/abs/2104.02145. Accessed 8 Feb. 2024.

Deng, Chunyuan, et al. "Investigating Data Contamination in Modern Benchmarks for Large Language Models." arXiv, 2023, arxiv.org/abs/2311.09783. Accessed 8 Feb. 2024.

Golchin, Shahriar, and Mihai Surdeanu. "Time Travel in LLMs: Tracing Data Contamination in Large Language Models." arXiv, 2023, arxiv.org/abs/2308.08493. Accessed 8 Feb. 2024.

Hendrycks, Dan, et al. "Measuring Massive Multitask Language Understanding." arXiv, 2020, arxiv.org/abs/2009.03300. Accessed 8 Feb. 2024.

Jacobs, Abigail Z., and Hanna Wallach. "Measurement Modeling for Fairness in Machine Learning." arXiv, 2019, arxiv.org/abs/1908.05664. Accessed 8 Feb. 2024.

Kiela, Douwe, et al. "Dynabench: Rethinking Benchmarking in NLP." arXiv, 2021, arxiv.org/abs/2104.14337. Accessed 8 Feb. 2024.

Liang, Percy, et al. "Holistic Evaluation of Language Models." arXiv, 2022, arxiv.org/abs/2211.09110. Accessed 8 Feb. 2024.

Magar, Inbal, and Roy Schwartz. "Data Contamination: From Memorization to Exploitation." arXiv, 2022, arxiv.org/abs/2203.08242. Accessed 8 Feb. 2024.

Oren, Yonatan, et al. "Proving Test Set Contamination in Black Box Language Models." arXiv, 2023, arxiv.org/abs/2310.17623. Accessed 8 Feb. 2024.

Raji, Deborah, et al. "AI and the Everything in the Whole Wide World Benchmark." arXiv, 2021, arxiv.org/abs/2111.15366. Accessed 8 Feb. 2024.

Ribeiro, Marco Tulio, et al. "Beyond Accuracy: Behavioral Testing of NLP Models with CheckList." arXiv, 2020, arxiv.org/abs/2005.04118. Accessed 8 Feb. 2024.

Sainz, Oscar, et al. "NLP Evaluation in Trouble: On the Need to Measure LLM Data Contamination for Each Benchmark." arXiv, 2023, arxiv.org/abs/2310.18018. Accessed 8 Feb. 2024.

Srivastava, Aarohi, et al. "Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models." arXiv, 2022, arxiv.org/abs/2206.04615. Accessed 8 Feb. 2024.

Yang, Shuo, et al. "Rethinking Benchmark and Contamination for Language Models with Rephrased Samples." arXiv, 2023, arxiv.org/abs/2311.04850. Accessed 8 Feb. 2024.

Zheng, Lianmin, et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv, 2023, arxiv.org/abs/2306.05685. Accessed 8 Feb. 2024.

"}]