Trust is earned, not given

A different perspective

2025-03-06 · Projects

AI Frontiers, part 49: Agent observability — tracing, replay, and debugging loops

Part 49from the AI Frontiers series · 65 parts in all

A support ticket I have seen a dozen times in different clothes: "the agent gave a customer the wrong refund amount yesterday and we cannot work out why." The team has an application log. It contains a request ID, a latency, a 200, and the final string the agent returned. Everything that mattered — which documents were retrieved, what the tool returned, whether the model misread a number or the retrieval returned the wrong policy, whether a retry silently swapped the answer — is gone, because none of it was recorded in a form anyone thought to look at.

This is the same lesson distributed systems learned in the previous decade, arriving in a new domain with a few genuinely new twists. Agents are non-deterministic, their behavior depends on text that is expensive to store, and their failure modes include loops and incoherence that no exception ever fires for. Observability is not a nice-to-have here; it is the only way the system gets better after it ships. And the artifact at the center of it — the trajectory, or trace — is arguably the single most valuable thing an AI system produces, because it is simultaneously your debugging tool, your evaluation data, and your training data.

Why conventional observability is necessary and insufficient

Request logs, metrics, and distributed tracing were designed for systems whose behavior is determined by code. Their unit of explanation is the call: a function was invoked with these arguments and returned this value. Dapper established the shape of it — annotate the request tree, propagate a trace ID, sample — and OpenTelemetry has since standardized it across languages (Sigelman et al.; OpenTelemetry). Every AI system should still have that: traces that show the request, the database calls, the cache hits, the external HTTP requests.

What conventional tracing does not capture is the part that is now deciding the outcome. Between the inbound HTTP request and the outbound response, an AI system makes semantic decisions: which of forty retrieved chunks entered the context, what the model was told, what the model said before and after a tool call, how many times it looped, what it cost. These are not function calls with typed arguments; they are, functionally, the program. Traces that stop at the HTTP boundary describe the shell and not the machine.

The second difference is directionality of failure. Traditional systems fail by throwing. Agents fail by being plausibly wrong — the subclass of failure part 26 spent an entire entry on. A 200 response containing a fluent, confident, incorrect answer is indistinguishable from a success in the metrics, which is why the discipline around traces has to be designed for inspection rather than alerting alone.

The trace as the primary artifact

A useful agent trace is a tree, and its nodes are typed. What I put in one, and would recommend as a minimum:

Per model call: the exact provider and model identifier including version, the full message list as sent (system prompt included), the full response including any tool-call arguments, the finish reason, token counts for input and output, the latency, and the computed cost. Full prompts, not hashes of prompts. The single most common reason a team cannot debug a bad output is that they hashed the prompt to save space.

Per retrieval step: the query as issued, the retriever configuration and the index version, the top-k candidate IDs with scores before any filtering or reranking, the IDs that survived, and the final context assembly with token counts. That before/after view is what distinguishes "the reranker dropped the right document" from "the retriever never had it," and those two bugs have nothing in common.

Per tool call: the tool name and version, the arguments after any schema validation, a redacted view of the result, the latency, the error if any, and whether the result was truncated before entering the context. Truncation is a chronic silent failure — the model reasons over half a document and nobody records that it did.

Per step: the decision the orchestrator made, if the orchestration is deterministic (which part 47 argued it should be), and the state hash after the step. That gives you loop detection for free: two identical state hashes with no progress is a bug, not a model quirk.

The point of typing nodes rather than dumping a JSON blob per request is that typed traces can be queried. "Show me all traces last week where a retrieval returned zero candidates and the model answered anyway" is a query over typed fields. Over blob logs it is a research project.

Replay: the hard part, and why it is worth the trouble

Replay is the property that separates a debuggable agent from a mysterious one. The requirement is that you can take a recorded trace and re-execute it, either exactly or with one thing changed.

Exact replay is impossible for the model call itself — sampling is not reproducible across provider versions, and providers deprecate models on their own schedule. What is achievable, and what I recommend, is decision-level replay with recorded responses. Store each model call's full request and full response. A replay mode in your runtime serves recorded responses from a cassette when the request matches; when it does not match — because you changed the prompt, the retriever, or the code — it either calls the live model or fails loudly, depending on the mode. This gives you three things at once: a deterministic regression test derived from production (replay a hundred real traces and assert the behavior), a counterfactual instrument ("what would this trace have done with the new prompt, holding all later decisions fixed"), and a cheap local debugging loop that does not burn API budget.

The engineering cost is real. It requires that every non-deterministic input — model responses, tool results, the clock, random seeds — be funneled through an injectable seam, and that the runtime have an explicit mode. It is exactly the discipline that makes a system testable, and the teams that build it early end up with what is effectively a video recorder for their agent. Given that model deprecation windows are measured in months, I would now treat recorded traces as the durable specification of expected behavior, because hand-built evaluation sets go stale while a recorded trace does not.

What to alert on

Metrics that have earned their place, all computed over traces: cost per completed task, because cost per call hides retry storms and long loops; steps per task, whose distribution widening is an early warning of a model or prompt regression; tool error rate and tool-argument-validation failure rate, which catch schema drift; loop rate, as the fraction of traces that repeat a state without progress; context utilization, the fraction of the window a typical trace consumes, which predicts truncation failures before they happen; and the rate at which the system says "I don't know", which should be non-zero — a system that never abstains is not well-calibrated, it is bluffing.

And one human-in-the-loop metric: the rate at which returned answers are edited or rejected downstream. If a human touches the output before it goes anywhere, that edit is a free label, and it is the single highest-signal number available to an applied team. The machine-learning practitioners' literature has made this point for years — that the failures worth studying are the ones in production, and that the organizational gap between who builds the model and who feels the consequences is where reliability goes to die (Sambasivan et al.; Paleyes et al.).

Traces are personal data

One warning that belongs in the design doc rather than the appendices: a full agent trace contains the user's input, retrieved documents (which may belong to other users, or be licensed material), the model's full reasoning, and often credentials or identifiers that appeared in a tool result. That is a rich and sensitive corpus with a retention policy you have to write.

What works in practice: redact at the boundary, with a per-field policy, before the trace is persisted; store a full-fidelity trace only for a sampled fraction of traffic, with the rest truncated or hashed; set retention explicitly and enforce it, because trace stores grow faster than any other data store in an AI system; encrypt at rest; and keep traces inside the same data-residency boundary as the rest of the system, a constraint that becomes a design input the moment you have European traffic. None of this is exotic — it is the normal discipline of handling user data — but it needs to be decided when the tracing is built, not after the first incident review.

The reason all of this is worth the effort is not compliance. It is that the loop from observation to improvement is the entire game in applied AI. Models are getting better on someone else's schedule; your system gets better on yours, and the only fuel it runs on is a faithful record of what it actually did.

Volume, sampling, and what a trace store costs

The cost of tracing is not the instrumentation work, it is the storage, and the number is larger than teams expect. A full-fidelity trace with prompts, retrieved documents, and tool results runs somewhere between fifty and several hundred kilobytes depending on how much context the system uses. A million requests a month at that size is tens to hundreds of gigabytes per month, growing forever, in a store that is queried rarely and read in anger during incidents. That shape needs a policy before it needs a schema.

The policy that works in practice is stratified sampling with a guaranteed error slice. Keep one hundred percent of traces that ended in a failure, a budget cap, a user correction, or a downstream rejection. Keep a small random sample of successes, a percent or two, as a baseline for comparison and as a source of regression fixtures. Keep newly deployed configurations at full retention for their first days, because that window is where regressions surface. Then drop the rest on a rolling schedule measured in weeks, with an archive to cold storage if a compliance requirement says you must. Teams that skip the failure slice and sample uniformly end up with a store that is mostly successful traffic and cannot answer the one question that matters.

One measurement note: the sampling decision has to be made after the outcome is known, which means the trace is buffered until the request completes and the retention decision is applied then. Architectures that decide at request start cannot implement this, and the difference is worth designing for early.

From traces to training data

The reason I argue traces are the most valuable artifact an AI system produces is that they are the raw material for every other improvement loop in this series. The chain is short. A trace with a human correction is a supervised example. A trace with the outcome label is an evaluation case. A trace that failed is a test that must keep failing after the fix. A trace that succeeded slowly is a cost-optimization candidate. Assemble a few thousand of those and you have the dataset for fine-tuning, and the evidence for whether it was worth it.

Two cautions come with that chain, and both are easy to get wrong. The first is selection bias: the traffic you logged is not the traffic you want, and if the system's output fed back into user behavior, your dataset encodes the system's own prior preferences. Training on your own accepted outputs is the fastest way to make a system more confidently itself, which is only an improvement if you are also keeping an adversarial slice in the mix. The machine-learning engineering literature has described this feedback pathology since well before language models (Sculley et al.; Amershi et al.), and nothing about the new stack exempts it.

The second is that trace-derived data inherits the trace's privacy properties. A dataset built from real user interactions carries the same obligations as the traces themselves, plus the additional one that training may make some of it memorizable. Redact before persisting and the problem stays bounded; redact after and it never does.

Built carefully, though, this is the loop that separates systems that improve from systems that merely get replaced. Every incident review should end with a new evaluation case. Every week of production should add a handful of corrected pairs to the dataset. Every model upgrade should be gated on the regression suite assembled from real failures. Observability is not the cost of running an AI system; it is the mechanism by which the system gets better.

Works Cited

Amershi, Saleema, et al. "Software Engineering for Machine Learning: A Case Study." Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice, 2019. Accessed 6 Mar. 2025.

Beyer, Betsy, et al., editors. Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media, 2016. Accessed 6 Mar. 2025.

Breck, Eric, et al. "The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction." Proceedings of the IEEE International Conference on Big Data, 2017. Accessed 6 Mar. 2025.

Kreps, Jay. "The Log: What Every Software Engineer Should Know about Real-Time Data's Unifying Abstraction." LinkedIn Engineering Blog, 2013, engineering.linkedin.com/distributed-systems/log-what-every-software-engineer-should-know-about-real-time-datas-unifying. Accessed 6 Mar. 2025.

OpenTelemetry. "Semantic Conventions for Generative AI Systems." OpenTelemetry Documentation, 2024, opentelemetry.io/docs/specs/semconv/gen-ai/. Accessed 6 Mar. 2025.

Paleyes, Andrei, Raoul-Gabriel Urma, and Neil D. Lawrence. "Challenges in Deploying Machine Learning: A Survey of Case Studies." arXiv, 2020, arxiv.org/abs/2011.09926. Accessed 6 Mar. 2025.

Sambasivan, Nithya, et al. "Who's Responsible for Failures in Machine Learning?" arXiv, 2021, arxiv.org/abs/2112.13143. Accessed 6 Mar. 2025.

Sculley, D., et al. "Hidden Technical Debt in Machine Learning Systems." Advances in Neural Information Processing Systems, 2015. Accessed 6 Mar. 2025.

Sigelman, Benjamin H., et al. "Dapper, a Large-Scale Distributed Systems Tracing Infrastructure." Google Technical Report, 2010, research.google/pubs/pub36356/. Accessed 6 Mar. 2025.

Zaharia, Matei, et al. "Accelerating the Machine Learning Lifecycle with MLflow." IEEE Data Engineering Bulletin, 2018. Accessed 6 Mar. 2025.

Zhang, Xiaoyu, et al. "AgentOps: Enabling Observability of LLM Agents." arXiv, 2024, arxiv.org/abs/2411.05285. Accessed 6 Mar. 2025.