Trust is earned, not given

A different perspective

2025-01-22 · Projects

AI fundamentals, part 6: Evaluation — measuring what models can do (and how benchmarks lie)

Part 6from the AI fundamentals series · 8 parts in all

You've built, scaled and served a model (parts 1–5). Now the hardest question in applied AI: is it any good? This part covers what benchmarks measure, how they're gamed, and how to evaluate honestly — the skill that separates teams who ship from teams who demo.

What a benchmark actually is

A benchmark is just a test set with a scorer. The classic kinds:

Compute the numbers yourself

def evaluate_multiple_choice(model, questions):
    # Score MMLU-style items. The correct method: compare the model's *log
    # likelihood* of each answer letter, not its sampled chat.
    correct = 0
    for q in questions:
        scores = {}
        for letter in "ABCD":
            prompt = q["question"] + f"\nAnswer: {letter}"
            # log P(full letter | prompt): the fair, order-independent method
            scores[letter] = logprob_of(model, prompt)
        pick = max(scores, key=scores.get)         # model's "choice"
        correct += (pick == q["answer"])
    return correct / len(questions)

def perplexity(model, held_out_ids):
    # exp(mean NLL): 1 = perfect; ~vocabulary-size = random guessing.
    nll = loss_and_grads(held_out_ids, model.params)
    return float(np.exp(nll))

That log-likelihood detail is not pedantry. Ask the model "respond with just A or D" and you measure its instruction-following, not its knowledge — a different (and gameable) skill.

Three ways benchmarks lie

Build your own eval — the only one that matters

For any real application, the benchmark that counts is yours: 50–200 real inputs from your product, expected outputs, scored automatically. Here's the whole harness:

import json, statistics

def run_eval(system_prompt, cases, call_model, judge=None):
    # The minimal honest eval loop. cases: [{input, expect}]. Returns accuracy
    # plus the failures, which matter more than the score.
    results = []
    for c in cases:
        got = call_model(system_prompt, c["input"])       # your real pipeline
        ok = (got.strip() == c["expect"] if judge is None
              else judge(c["input"], c["expect"], got))   # or an LLM/regex judge
        results.append({"case": c, "got": got, "ok": bool(ok)})
    acc = sum(r["ok"] for r in results) / len(results)
    fails = [r for r in results if not r["ok"]]
    return {"accuracy": acc, "failures": fails}

with open("golden_cases.json") as f:
    cases = json.load(f)
report = run_eval(SYSTEM_PROMPT, cases, my_rma_agent_call)

# The discipline: run this on EVERY prompt or model change, before deploy.
print(f"accuracy: {report['accuracy']:.0%}")
for r in report["failures"][:5]:
    print("FAIL:", r["case"]["input"][:60], "->", r["got"][:80])

Two rules that keep evals honest: never eval on your dev prompts (build the golden set once, freeze it, extend it), and report the failures, not just the score — five read failures tell you what to fix; "94%" tells you nothing.

The honest stack

In practice, mature teams layer all of it: perplexity for language quality during training; public benchmarks as a smoke test for capability class; LLM-judge for style comparisons; and the private golden eval as the only gate that actually blocks a release. Public scores are marketing; your eval is engineering. Trust the second, and treat surprising public numbers — up or down — as hypotheses to test against your own cases, not facts to adopt.

That closes the core series: tokens → attention → transformer → scaling → inference → evaluation. A final part on how models are aligned to be helpful and safe follows; after that, back to applied topics with these fundamentals under our feet.