Trust is earned, not given

A different perspective

2026-02-05 · Projects

AI Frontiers, part 60: Continuous evaluation in production — drift, regressions, rollbacks

Part 60from the AI Frontiers series · 65 parts in all

There is a version of AI engineering that ends at launch and a version that begins there, and the difference is not philosophical. It is whether the system has a measurement loop that keeps running after the first real user arrives. A pre-launch evaluation set reflects the traffic someone imagined; production traffic is the only data that describes the traffic that exists. The gap between the two is why systems that looked excellent in staging produce incident reviews six weeks later, and why the teams that compound improvement over years are the ones with the least glamorous monitoring.

This entry is about the loop. It picks up part 49's traces and part 51's measurement design and asks what has to be true for evaluation to be continuous rather than episodic: cheap labels, drift detection that does not drown in false alarms, a way to notice that a model changed underneath you, and a rollback that takes seconds.

What changes the day you launch

Four things make the post-launch problem genuinely different.

Your test set starts aging immediately. It was built from the requests you could imagine, sampled at a moment when your product had a particular set of users doing a particular set of things. Two quarters later, the interesting traffic is elsewhere: a new integration, a new user segment, a new document format, a prompt template copied from a blog. Every one of those is a distribution shift that no amount of care during design anticipates.

The world moves independently of your releases. New products get announced, prices change, regulations take effect, upstream systems change their APIs, and the content your retrieval layer indexes is rewritten. Concept drift in the machine learning sense — the relationship between inputs and correct outputs changing over time — is the norm, not the exception, and the literature on detecting and adapting to it has been developing for two decades (Gama et al.; Webb et al.).

The model layer is not under your control. A hosted model can be updated, its defaults changed, its safety behavior adjusted, or its deprecation announced on a schedule you did not choose. Your evaluation suite is the only instrument that will tell you, and if it only runs on your own release cadence, you will find out from a customer.

Failure becomes quieter. A system that throws errors is easy to monitor. A system that returns fluent, plausible, subtly wrong answers has no exception to catch, which is the ongoing lesson of part 26.

Cheap labels: the whole game

Continuous evaluation is a labeling problem in disguise, and the teams that solve it well use four sources, all of which they had already.

Human edits. Anywhere a person reviews, edits, or overrides output, that pair is a label. A support agent rewriting a suggested reply is telling you precisely what the model got wrong. Capture the original and the correction at the moment of the edit, and you have a stream of domain-specific supervision for the cost of the pipeline that already exists. In every system I have looked at closely, this was the highest-value signal available and the least instrumented.

Downstream outcomes. Did the ticket reopen, the order get refunded, the invoice get rejected, the user retry within a minute? Proxy outcomes are noisy and enormously better than nothing, and they are objective in a way that model judges are not.

Sampled human review. A small, stratified, ongoing sample — fifty interactions a week, weighted toward the failure-prone categories — reviewed by someone with domain knowledge and scored against a short rubric. This is the only source that detects quality changes in cases where nothing downstream visibly breaks, and it is also what keeps your automated judge honest. The discipline of validating an LLM judge against human ratings was covered in part 42, and it has to be repeated periodically because your traffic is not the traffic you validated against.

Explicit feedback. Thumbs, ratings, escalations, complaints. Sparse and biased, and useful mainly as a leading indicator of a category problem rather than as a quality measure.

Drift detection that does not cry wolf

The failure mode of drift monitoring is not missing an event; it is alerting so often that nobody looks. Four measurements, in order of usefulness, with the discipline that makes them actionable.

Input distribution. Track the distribution of a small set of stable features: request length, language, category, retrieval hit count, tool-call types, tenant. Compare windows with a test appropriate to the feature — a Kolmogorov-Smirnov or population-stability measure for numeric features, a divergence measure for categorical. Alert on sustained movement over a trailing baseline rather than on a single window, and set thresholds from your own historical variance, not from a textbook.

Embedding-space drift. Embed the incoming requests with the same encoder described in part 43 and track the centroid's movement and the emergence of new clusters. This catches semantic changes that feature-level monitoring misses — a new class of request, a new jargon, a new competitor's product name — and new clusters are usually the most informative signals the system produces. It is also cheap, since only a sample needs embedding.

Output distribution and abstention rate. Length distribution, structured-output validator failures, refusal rate, and the rate at which the system says it does not know. A falling abstention rate is a warning sign, not an improvement: it usually means the system has started guessing.

Cost and latency distribution. Tracked per task, as part 50 argued, because retries and runaway loops show up here first and because a change in the p95 is often the earliest symptom of a model behavior change.

Attach every alert to a decision. If an alert fires and the documented response is "look at it," the alert is decoration. Each threshold should name the action: page, open a ticket, add a case to the evaluation set, or nothing.

Regressions from a layer you do not own

The model provider's release calendar is now part of your release calendar, and treating it that way removes most of the surprise. Three practices.

Pin everything and upgrade deliberately. Pin model identifiers including version where the provider supports it, pin the client library, pin the evaluation suite version. When a pin expires, that is a scheduled project, not an incident — the deprecation notice is the advance warning, and the evaluation suite is how you decide whether the upgrade candidate is better.

Shadow-evaluate before switching. Run the candidate model against your last few weeks of recorded traces, and against your held-out evaluation set, and compare per-task rather than in aggregate. Recall from part 51 that a candidate that fixes twenty cases and breaks eighteen is a decision, not a win, and that the spread of repeated runs is the threshold below which no difference is real (Miller).

Canary anything that changes behavior. Route one percent of traffic to the new configuration, watch the outcome proxies and the failure categories, and only then proceed. This is ordinary deployment practice from the infrastructure world, and it applies unchanged to model changes because a model change is a code change with an unreadable diff.

Rollback, and the reason it comes first

The property that makes all of this tractable is the ability to reverse a change in one action. Everything that is configurable and versioned is reversible; everything embedded in a prompt string or hard-coded in a handler is not. Practically, that means: model selection, prompt templates, retrieval parameters, routing thresholds, and tool permissions all live in versioned configuration with a single source of truth; the application reads them at request time or on a signal; and rolling back is one change, deployed in the ordinary way.

Systems that can do this recover from bad releases in minutes and can therefore be aggressive about trying things. Systems that cannot spend their energy on caution and still get surprises. The observational literature on production machine learning has described this dynamic for years — that the failures that matter are usually integration and drift failures rather than algorithmic ones, and that the organizations which build the release machinery around the model are the ones that can iterate (Sculley et al.; Breck et al.; Paleyes et al.). The practitioner literature on rules of thumb for machine learning systems and the general engineering guidance on error budgets and canaries both point the same way (Zinkevich; Beyer et al.).

Who is on call for a language model

The organizational question is the one teams skip, and it determines whether any of this happens. Someone has to own the evaluation dashboard, triage its alerts, and decide when a quality drop warrants a rollback. In practice this ends up as an extension of existing on-call for AI-adjacent services, with two additions: an incident review step that turns every production failure into a permanent evaluation case, and an explicit cadence — weekly review of sampled interactions, monthly recalibration of the automated judge, quarterly refresh of the held-out set from new traffic.

The reward for doing this is the compounding effect that runs through the entire series. A system with a live measurement loop gets better every month on the strength of its own traffic: new evaluation cases come from real failures, new training data comes from real corrections, new prompt and routing decisions come from real measurements. A system without one gets replaced when the next model generation makes someone ask why the old thing feels stale. The models will keep improving on someone else's schedule. The loop is the only part of the advantage that is actually yours.

A weekly operating rhythm

The difference between teams that maintain a measurement loop and teams that built one during a launch sprint is almost entirely cadence. What has worked, in systems I have seen sustain it for years, is a small set of recurring activities with named owners — deliberately small, because a review process that takes a day a week gets skipped the first busy month and never restarts.

Weekly: read twenty sampled interactions. Not a dashboard, actual cases, including failures. An engineer and a domain expert, thirty minutes. The output is not a report; it is zero to three concrete changes, which may be a prompt edit, a retrieval fix, a new evaluation case, or a note that everything is fine. This single meeting is the highest-value item on the list, because it is the only activity that reliably detects the failure mode where every metric looks healthy and the output is subtly worse.

Weekly: triage the alerts that fired. Which thresholds were crossed, whether each was real, and whether the threshold needs changing. Alert fatigue is self-inflicted, and the only cure is to adjust thresholds as evidence accumulates rather than leaving them at whatever value was chosen during a launch.

Monthly: refresh the evaluation set. Add cases from real failures, retire cases that no longer reflect the product, and re-validate the automated judge against a fresh human-labeled sample. The set should grow slowly and change constantly; a set that has not changed in six months is measuring last year's product.

Monthly: review the drift panels. Input distribution, embedding drift, abstention rate, cost per task. The purpose is to notice slow movement, which is invisible in any single week and obvious across three months — and slow movement is what makes a system feel stale before anyone can point to a regression.

Quarterly: upgrade and re-baseline. Evaluate the candidate models and provider versions, decide which to adopt, and re-baseline every threshold against the new configuration. Doing this on a schedule converts a series of crises into a routine.

Quarterly: exercise the rollback. Actually perform it in a staging environment, or better, in production during a low-traffic window. The one time you need it is the worst time to discover that the model version was pinned in three places and only two of them are in configuration.

Total cost: a few hours a week across a team, plus the discipline to keep showing up. The return is that quality problems are found by the team rather than by a customer, which is the entire point of operating rather than merely deploying.

Works Cited

Beyer, Betsy, et al., editors. Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media, 2016. Accessed 5 Feb. 2026.

Breck, Eric, et al. "The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction." Proceedings of the IEEE International Conference on Big Data, 2017. Accessed 5 Feb. 2026.

Gama, João, et al. "A Survey on Concept Drift Adaptation." ACM Computing Surveys, vol. 46, no. 4, 2014, pp. 1–37. Accessed 5 Feb. 2026.

Miller, Evan. "Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations." arXiv, 2024, arxiv.org/abs/2411.00640. Accessed 5 Feb. 2026.

Paleyes, Andrei, Raoul-Gabriel Urma, and Neil D. Lawrence. "Challenges in Deploying Machine Learning: A Survey of Case Studies." arXiv, 2020, arxiv.org/abs/2011.09926. Accessed 5 Feb. 2026.

Sculley, D., et al. "Hidden Technical Debt in Machine Learning Systems." Advances in Neural Information Processing Systems, 2015. Accessed 5 Feb. 2026.

Webb, Geoffrey I., et al. "Characterizing Concept Drift." Data Mining and Knowledge Discovery, vol. 30, 2016, pp. 964–994. Accessed 5 Feb. 2026.

Zinkevich, Martin. "Rules of Machine Learning: Best Practices for ML Engineering." Google Developers, 2017, developers.google.com/machine-learning/guides/rules-of-ml. Accessed 5 Feb. 2026.