AI Frontiers, part 51: The evaluation of agents β tasks, traces, and cost-aware scoring
Part 51from the AI Frontiers series · 65 parts in all
Between 2023 and 2025 the agent literature produced a fixture that outlived most of the agents it measured: a leaderboard number that climbed steadily while the systems behind it stayed stubbornly unusable for anything adjacent to the benchmark. The gap is not fraud. It is the predictable result of measuring an expensive, stochastic, environment-dependent process with a metric designed for a single model call.
Part 42 argued that a private evaluation set is the precondition for every improvement in an applied AI system. This entry is the harder version of that argument for agents, where the unit being evaluated is a trajectory in an environment rather than a single response, and where the scoring function has to account for cost, variance, and the possibility that the environment itself is the bottleneck. It closes the loop opened in part 47 and part 49: you cannot evaluate what you did not trace, and you cannot improve what you cannot evaluate.
Why agent evaluation is genuinely harder
Four properties separate it from benchmark evaluation of a model.
Compounding. A ten-step trajectory with 95% per-step reliability succeeds about 60% of the time. The metric is therefore exquisitely sensitive to step count, and two systems with identical per-step competence can score very differently because one takes twice as many steps. Reporting a single pass rate without the step distribution leaves the most important variable unmeasured.
Environment dependence. The agent's score is a function of the task, the tools it can call, the state the environment starts in, and the rate limits of the systems it touches. Benchmarks pin all of this down; production never does. A number from a benchmark is a statement about the benchmark.
Cost as a first-class dimension. An agent can improve its pass rate indefinitely by taking more steps, calling a bigger model, or sampling more attempts. Any evaluation that does not carry a cost column rewards exactly that behavior. The paper that made this argument unavoidable β Kapoor and colleagues' "AI Agents That Matter" β showed that many reported agent gains came from best-of-n sampling and generous budgets, and that joint accuracy-and-cost optimization reorders the rankings substantially.
Stochasticity. Running a task once tells you little. The interesting quantity is not "can it do this" but "how reliably does it do this," which requires repeated runs and a metric that punishes inconsistency rather than averaging it away.
The benchmark landscape, honestly assessed
The 2023β2025 wave produced a genuinely useful set of environments, and it is worth knowing what each one measures. SWE-bench, and its human-validated subset SWE-bench Verified, ask a model to resolve real GitHub issues against real repositories with tests as the arbiter (Jimenez et al.; OpenAI). Its strength is the executable ground truth; its weakness is that the metric conflates localization, editing, and test-passing into one number, and that familiarity with popular repositories is hard to rule out.
WebArena builds a self-hosted replica of realistic web applications and asks agents to complete multi-step tasks in a browser (Zhou et al., "WebArena"), which is the closest existing thing to a fair interactive environment. OSWorld does the analogous job for a desktop operating system with real applications and file systems (Xie et al.). Mind2Web measured generalization to unseen websites (Deng et al.), and AgentBench spanned multiple interactive environments to test whether capability transfers (Liu et al.). GAIA constructed questions that require tool use and multi-step reasoning with unambiguous answers (Mialon et al.). AppWorld generated controllable, stateful API-based applications with programmatic verification (Trivedi et al.), and TheAgentCompany placed agents inside a simulated software company with realistic tasks and partial credit (Xu et al.).
Two things about this landscape deserve saying plainly. First, the benchmarks with executable verification β tests, state assertions, database checks β age far better than rubric-graded ones, because the ground truth cannot drift. Second, benchmark scores have almost no predictive power for your workload. Every domain has its own tool surface, its own error distribution, and its own notion of a good outcome. Use public benchmarks to choose among models on a shortlist; use your own evaluation to ship.
Four measurement sins
1. Accuracy without cost. Report tokens, dollars, and wall-clock per solved task next to every pass rate. If a configuration doubles the score at ten times the cost, that is a decision, not an improvement, and someone has to make it deliberately.
2. pass@k where pass^k is the question. Reporting the best of k attempts measures the model's ceiling and hides its consistency; for anything user-facing, consistency is the product. The Ο-bench authors introduced the cleaner metric β passing all k attempts β precisely because reliability across repetitions is what a deployed system needs, and reported that numbers fall steeply as k grows (Yao et al., "Ο-bench"). A system that solves a task 70% of the time is a system that fails a customer three times in ten, regardless of what any best-of-k figure implies.
3. Test set as training set. Agent evaluation is even more contamination-prone than model evaluation because the agent sees the environment: browsing it, querying it, reading its files. If your evaluation tasks live in a system the agent can explore during development, the score is measuring familiarity. Hold out tasks, and where possible hold out environments.
4. Single-number summaries of multi-cause failures. A 55% pass rate tells you nothing actionable. The failure taxonomy is the actionable output, and the multi-agent failure study from March 2025 provides the most useful one so far: it decomposed failures into specification problems (ambiguous or incomplete task definitions), inter-agent misalignment (information not propagating, conflicting actions), and verification gaps (no check of intermediate results), and found that how you instruct a system about its own responsibilities is a decisive variable (Cemri et al.). Those three categories map onto almost every failure I have debugged. Classify, then count.
Building a private agent evaluation
What works for a production system, in the order I would build it.
Task specifications before tasks. Write down, for each task, the initial state, the allowed tools, the success criterion, and an acceptable cost ceiling. An ambiguous task definition produces ambiguous scores, and no evaluation harness fixes an unclear goal β this is the most common root cause in the taxonomy above.
Reproducible environments. Containerize the world the agent acts in, snapshot the initial state, and reset between runs. Without a reset, runs contaminate each other and the evaluation measures ordering.
Programmatic success checks first. Tests pass, the record exists, the file has the expected content, the API returned the right shape. Rubric-graded evaluation is a fallback for genuinely subjective outcomes, and it needs a human-validated sample to establish that the grader agrees with people β the same discipline as part 42's LLM-as-judge guidance. Agent-as-a-judge approaches, where a model inspects a trajectory against a checklist, are promising for open-ended tasks (Gu et al.), but they are a second grader to validate, not a replacement for one.
Repeated runs, with variance reported. Five runs minimum per task for anything whose variance you care about, and report the pass^k curve rather than its top point. The variance itself is diagnostic: high variance on a task usually means the environment has a race or the task is underspecified.
Traces captured automatically. Every evaluation run should emit the trace format from part 49. Then the same artifact serves three purposes: scoring, failure analysis, and regression fixtures after a change.
A cost column and a budget guard. Cap tokens and steps per task in the harness itself. If a run hits the cap, that is a failure with a distinct label β not a score of "β" that gets quietly dropped, which is how budget overruns turn into apparent accuracy gains.
Scoring, and what to put on the dashboard
The scoring rubric I would sign off on has five columns per task: passed, cost in dollars, steps taken, wall-clock, and failure category if failed. From those, six numbers go on the dashboard: pass^k at k=3 or 5, median cost per task with a p95, median steps with a p95, failure-category distribution, the fraction of runs that hit a budget cap, and the trend of all of the above across releases.
That last one is the point. None of this is a leaderboard. It is an instrument for detecting that a change you made on Tuesday moved reliability in a direction you did not expect β the same purpose as any other test suite, with the added complication that both the code and the world can change underneath it. The teams with good agent evaluations in 2025 were not the teams with the hardest tasks. They were the teams with the fewest ambiguous questions about what success meant, and a habit of rerunning everything on every change.
Evaluate the pipeline, not only the agent
The most common wasted week in agent debugging is a team tuning prompts to fix a failure that was never in the model. Before attributing a failure to the agent's reasoning, establish that each stage it depends on delivered what it should have. That means instrumenting the pipeline with its own scores, independent of the end-to-end result.
Three component metrics catch most misattributions. Retrieval recall: for each task, did the evidence the agent needed appear in the context it was given? If not, no amount of prompt work matters, and the fix belongs in the index, the chunking, or the reranker. Tool call validity: did the agent produce well-formed calls with arguments the tools accepted? A tool schema that is ambiguous produces invalid calls, and the model gets blamed for a contract problem. Environment fidelity: did the tool do what it claimed? Rate limits, stale caches, and permission errors in a test environment are a perennial source of agent failures that reproduce nowhere else.
The cleanest technique for separating the agent's contribution from the pipeline's is an oracle ablation. Run the same tasks with the retrieval step replaced by the ideal context, or the tool replaced by the correct response, and compare against the real pipeline. The gap between the oracle run and the real run is the pipeline's contribution; the oracle run's absolute score is the agent's ceiling. When the gap is small, prompt engineering is the right investment. When it is large, it is the wrong tool entirely, and the team usually finds that out only after weeks of prompt tuning that moves nothing.
Model upgrades, and knowing when a change is real
Once an evaluation exists, its most valuable use is not launch readiness β it is change management. Base models are deprecated on schedules measured in months, providers ship new versions without notice, and a system whose prompt and tooling were tuned against one model will behave differently on its successor. An evaluation suite converts that from a crisis into a routine: run the suite against the candidate, compare per-task outcomes against the incumbent, and ship when the differences are understood.
Two statistical habits make those comparisons trustworthy. First, quantify the noise: repeated runs of an unchanged system produce a spread, and any improvement smaller than that spread is not an improvement. Language-model evaluation has sample sizes that usually make this spread larger than teams assume, and the standard error of an agent pass rate is routinely a few percentage points (Miller). Second, compare per-task rather than in aggregate. If a candidate fixes twenty tasks and breaks eighteen, the aggregate score is honest and useless; the task-level diff is the decision, and the eighteen newly broken tasks are your next three weeks of work.
That is the full argument for treating evaluation as infrastructure rather than as a pre-launch chore. It is the only instrument that tells you whether a new model is free capability or a regression, and it is the only artifact that accumulates value as the system ages. Agents will keep getting more capable on someone else's release schedule. Whether that shows up as progress in your product or as a spike in incident reviews depends entirely on whether you were measuring.
Works Cited
Cemri, Mert, et al. "Why Do Multi-Agent LLM Systems Fail?" arXiv, 2025, arxiv.org/abs/2503.13657. Accessed 24 Apr. 2025.
Deng, Xiang, et al. "Mind2Web: Towards a Generalist Agent for the Web." arXiv, 2023, arxiv.org/abs/2306.06070. Accessed 24 Apr. 2025.
Gu, Jiawei, et al. "Agent-as-a-Judge: Evaluate Agents with Agents." arXiv, 2024, arxiv.org/abs/2410.10934. Accessed 24 Apr. 2025.
Jimenez, Carlos E., et al. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" arXiv, 2023, arxiv.org/abs/2310.06770. Accessed 24 Apr. 2025.
Kapoor, Sayash, et al. "AI Agents That Matter." arXiv, 2024, arxiv.org/abs/2407.01502. Accessed 24 Apr. 2025.
Liu, Xiao, et al. "AgentBench: Evaluating LLMs as Agents." arXiv, 2023, arxiv.org/abs/2308.03688. Accessed 24 Apr. 2025.
Mialon, GrΓ©goire, et al. "GAIA: A Benchmark for General AI Assistants." arXiv, 2023, arxiv.org/abs/2311.12983. Accessed 24 Apr. 2025.
Miller, Evan. "Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations." arXiv, 2024, arxiv.org/abs/2411.00640. Accessed 24 Apr. 2025.
OpenAI. "Introducing SWE-bench Verified." OpenAI, 13 Aug. 2024, openai.com/index/introducing-swe-bench-verified/. Accessed 24 Apr. 2025.
Trivedi, Harsh, et al. "AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents." arXiv, 2024, arxiv.org/abs/2407.18901. Accessed 24 Apr. 2025.
Xie, Tianbao, et al. "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments." arXiv, 2024, arxiv.org/abs/2404.07972. Accessed 24 Apr. 2025.
Xu, Frank F., et al. "TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks." arXiv, 2024, arxiv.org/abs/2412.14161. Accessed 24 Apr. 2025.
Yao, Shunyu, et al. "Ο-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains." arXiv, 2024, arxiv.org/abs/2406.12045. Accessed 24 Apr. 2025.
Zhou, Shuyan, et al. "WebArena: A Realistic Web Environment for Building Autonomous Agents." arXiv, 2023, arxiv.org/abs/2307.13854. Accessed 24 Apr. 2025.