AI fundamentals, part 6: Evaluation — measuring what models can do (and how benchmarks lie)
Part 6from the AI fundamentals series · 8 parts in all
You've built, scaled and served a model (parts 1–5). Now the hardest question in applied AI: is it any good? This part covers what benchmarks measure, how they're gamed, and how to evaluate honestly — the skill that separates teams who ship from teams who demo.
What a benchmark actually is
A benchmark is just a test set with a scorer. The classic kinds:
- Perplexity — from our part-3 loss: how surprised the model is by held-out text. Great for comparing language quality, meaningless across tokenizers, invisible to users.
- Multiple-choice (MMLU-style) — pick A/B/C/D across academic subjects. Cheap, automatable, and saturating: frontier models now score ~90% on tests written to be hard for humans.
- Exact-match generation — math answers, code that passes unit tests (HumanEval-style). Cleaner signal, narrow skills.
- LLM-as-judge — a strong model grades free-form answers (helpfulness, style). Flexible; inherits the judge's biases (longer answers score higher, self-preference, position bias).
Compute the numbers yourself
def evaluate_multiple_choice(model, questions):
# Score MMLU-style items. The correct method: compare the model's *log
# likelihood* of each answer letter, not its sampled chat.
correct = 0
for q in questions:
scores = {}
for letter in "ABCD":
prompt = q["question"] + f"\nAnswer: {letter}"
# log P(full letter | prompt): the fair, order-independent method
scores[letter] = logprob_of(model, prompt)
pick = max(scores, key=scores.get) # model's "choice"
correct += (pick == q["answer"])
return correct / len(questions)
def perplexity(model, held_out_ids):
# exp(mean NLL): 1 = perfect; ~vocabulary-size = random guessing.
nll = loss_and_grads(held_out_ids, model.params)
return float(np.exp(nll))
That log-likelihood detail is not pedantry. Ask the model "respond with just A or D" and you measure its instruction-following, not its knowledge — a different (and gameable) skill.
Three ways benchmarks lie
- Contamination. Benchmark questions crawl into training data (they're on the web). The model may be reciting answers it memorized. Mitigation: canary strings, held-back private test sets, and checking whether the model completes the question before seeing the choices.
- Overfitting-by-selection. Labs tune on the public test, publish the good number, and the number drifts up without real capability gaining. Goodhart's law, in production.
- Format sensitivity. The same model can swing several points on wording, ordering, or whitespace. If an eval can't survive paraphrase, it's measuring the prompt, not the model.
Build your own eval — the only one that matters
For any real application, the benchmark that counts is yours: 50–200 real inputs from your product, expected outputs, scored automatically. Here's the whole harness:
import json, statistics
def run_eval(system_prompt, cases, call_model, judge=None):
# The minimal honest eval loop. cases: [{input, expect}]. Returns accuracy
# plus the failures, which matter more than the score.
results = []
for c in cases:
got = call_model(system_prompt, c["input"]) # your real pipeline
ok = (got.strip() == c["expect"] if judge is None
else judge(c["input"], c["expect"], got)) # or an LLM/regex judge
results.append({"case": c, "got": got, "ok": bool(ok)})
acc = sum(r["ok"] for r in results) / len(results)
fails = [r for r in results if not r["ok"]]
return {"accuracy": acc, "failures": fails}
with open("golden_cases.json") as f:
cases = json.load(f)
report = run_eval(SYSTEM_PROMPT, cases, my_rma_agent_call)
# The discipline: run this on EVERY prompt or model change, before deploy.
print(f"accuracy: {report['accuracy']:.0%}")
for r in report["failures"][:5]:
print("FAIL:", r["case"]["input"][:60], "->", r["got"][:80])
Two rules that keep evals honest: never eval on your dev prompts (build the golden set once, freeze it, extend it), and report the failures, not just the score — five read failures tell you what to fix; "94%" tells you nothing.
The honest stack
In practice, mature teams layer all of it: perplexity for language quality during training; public benchmarks as a smoke test for capability class; LLM-judge for style comparisons; and the private golden eval as the only gate that actually blocks a release. Public scores are marketing; your eval is engineering. Trust the second, and treat surprising public numbers — up or down — as hypotheses to test against your own cases, not facts to adopt.
That closes the core series: tokens → attention → transformer → scaling → inference → evaluation. A final part on how models are aligned to be helpful and safe follows; after that, back to applied topics with these fundamentals under our feet.