AI Frontiers, part 62: Human-in-the-loop design — review queues and trust calibration
Part 62from the AI Frontiers series · 65 parts in all
Almost every serious deployment of these systems has a human somewhere in the loop, and in almost every one of those deployments the human was added as an afterthought: a confirmation dialog, a review queue bolted onto the side, an escalation path defined by whoever was on call that week. Then three months later someone notices that the reviewers approve ninety-nine percent of what they see, that approval takes four seconds per item, and that the human in the loop is a signature rather than a control.
That outcome is not a personnel problem. It is a design problem with a well-documented history in the automation literature, and it is avoidable if the human is treated as what they are: a component in the system with a latency budget, a cost per operation, an error distribution, and a fatigue curve. This entry is about designing that component properly — the interface, the queue, the escalation policy, and the metrics — and it follows from the security argument in part 57 that containment, not detection, is what the human provides.
What the human is actually for
Conflating three different jobs is the root of most bad human-in-the-loop designs, because each has a different cost structure and a different failure mode.
Labeling. The human produces training and evaluation data. This is the job that compounds — it feeds the loops in part 60 — and it is the job companies most often fail to instrument, because it looks like ordinary work rather than like data collection.
Gating. The human authorizes an irreversible or high-consequence action. The design question is not accuracy; it is what the human needs to see to make a genuine decision in a bounded time, and whether the decision is actually theirs to make.
Exception handling. The human resolves cases the system cannot: ambiguity, missing information, novel situations, and anything requiring judgment about consequences. This is where human skill is most valuable and where the system's job is to route well rather than to answer.
Mix them and you get the classic pathology. A reviewer who is nominally gating, but who receives a thousand items a day with a two-second budget each, is not gating; they are rubber-stamping with extra steps, and the organization believes it has a control it does not have. Naming which job a given review station does, and measuring it against that job, is the first fix.
Trust calibration, or why appropriate reliance is the goal
The human factors literature settled this question long before language models existed in a usable form. Trust in automation should be calibrated rather than maximized: reliance that matches the system's actual competence in the specific situation, so that the human neither over-trusts (accepting wrong answers) nor under-trusts (ignoring good ones) (Lee and See). Parasuraman and Riley named the four pathologies — use, misuse, disuse, abuse — that follow from miscalibration, and the medical and aviation literatures catalogued automation bias decades ago: people who monitor an automated decision tend to accept it, particularly when the decision is difficult, when they are under time pressure, and when the interface presents the output as a recommendation rather than as one input among several (Goddard et al.; Skitka et al.).
Three findings from the human-AI collaboration literature are worth building around because they are counterintuitive.
Explanations do not reliably improve team performance. The experiments on complementary team performance found that model explanations could increase reliance without increasing accuracy — people learned to copy the model more, not to supervise it better — and in some setups explanations actively harmed accuracy while raising acceptance (Bansal et al., "Does the Whole Exceed Its Parts?"). The practical conclusion is that an explanation should be designed to help the human disagree — showing the evidence, the source, the alternatives considered — rather than to justify the answer.
Mental models predict performance better than accuracy does. Studies of human-AI teams found that participants performed better when their mental model of the AI's strengths and weaknesses matched its actual behavior, independent of their overall trust level (Bansal et al., "Beyond Accuracy"). That suggests investing in training and interface cues about where the system is likely to be wrong, which is a different investment from improving the system's average score.
People under-use good algorithmic advice when they can see it fail. Studies of decision support found that a small number of visible errors produced lasting disuse, even when the system was better than the human overall (Green and Chen). The design implication: never present accuracy as a single number, and make the failure modes legible so that disuse is proportional to the actual error rate rather than to the memorability of a bad case.
Interface design
Beyond the general guidance for human-AI interaction, which was distilled into a set of well-tested design guidelines by a large cross-industry effort (Amershi et al., "Guidelines for Human-AI Interaction"), four patterns have earned their place in review interfaces.
Show the basis, not just the answer. The extracted field with a link to the source span; the recommendation with the retrieved policy it came from; the generated summary with the count of source documents. When a reviewer can check a claim in one click, they do, and the review is real. When checking requires opening another system, they do not, and the review is theater.
Separate confidence from correctness. A model's self-reported confidence is weakly calibrated at best. Interface elements that convey uncertainty should be derived from something observable — retrieval scores, agreement across samples, whether an executable check passed, whether the case resembles the training distribution — rather than from the model's own claim of certainty.
Make rejection cheap and specific. A single "reject" button produces a boolean. A small set of structured reasons produces a dataset. The review interface is your best labeling instrument; design it to emit labels, and the correction loop from part 60 pays for itself.
Order the queue by consequence, not by arrival. A queue sorted by risk — dollar amount, irreversibility, tenant tier, model uncertainty — lets a reviewer with limited attention spend it where the damage would be worst. Arrival-order queues are an implementation convenience masquerading as a policy.
Queue design is capacity planning
The queue is where the abstractions meet arithmetic, and the arithmetic is unforgiving. If the system produces a thousand items a day that nominally require review, and a meaningful review takes two minutes, that is thirty-three reviewer-hours per day. Those are real people with real costs, and if they do not exist, the review does not happen — it just gets recorded as having happened.
Four design choices reduce the required capacity without reducing the control.
Sample where the stakes are low. Uniform review of low-consequence outputs is a waste of the scarcest resource in the system. Review a stratified sample, use it for measurement rather than gating, and reserve full review for the categories where a mistake is expensive.
Defer rather than block where the workflow allows. A draft sent for review in parallel with other work costs latency; a blocking modal costs attention. The interaction design literature has argued for mixed-initiative systems that interrupt only when the expected value of the interruption exceeds its cost (Horvitz), and that calculus is worth doing explicitly at design time rather than discovering by watching reviewers close dialogs.
Budget the escalation rate as a monitored metric. The fraction of items escalated to humans should be a number someone owns, with a target range. Rising escalation is an early signal of model regression; falling escalation with falling quality is a signal of rubber- stamping. Neither is visible if nobody tracks it.
Instrument the human's contribution. Acceptance rate, override rate, override accuracy (measured against later ground truth), time per item, and — the most informative and most neglected — the rate at which the human catches something the system would have gotten wrong. A review process that never catches anything is either unnecessary or not working, and those two possibilities require different responses.
Learning to defer, statistically
There is a formal version of the question "when should the system hand off?" and it is worth knowing because it gives you a principled objective instead of a threshold chosen by feel. The learning-to-defer literature learns a policy that trades off the system's prediction against a human's, accounting for the human's accuracy and the cost of asking, with variants that incorporate fairness constraints and partial human availability (Madras et al.; Mozannar and Sontag). You do not need the machinery to benefit from the framing: the decision rule should minimize expected cost across both actors, which means knowing the human's error rate by category and the cost of a consultation, not just the model's confidence.
Two practical consequences. First, measure the human's accuracy by category, not in aggregate, because the value of deferral is category-specific — a reviewer who is excellent on ambiguous policy questions and no better than the model on arithmetic should only see the former. Second, price the consultation. A review that costs two minutes of a specialist's time is not free, and a deferral policy that ignores that cost will escalate far more than the economics support.
The organizational part
The failure mode that kills human-in-the-loop programs is not technical. It is that the reviewers are treated as a source of capacity rather than as a source of evidence, so nobody asks them what they are seeing until the incident review. The teams that get this right run a standing channel between the review queue and the model team, with the reviewers' structured rejections feeding the evaluation set weekly and their unusual cases read as a leading indicator. In those organizations the loop runs the other way as well: the model team tells the reviewers what changed and what to watch, which is the mental-model investment the studies above point to.
The uncomfortable but liberating conclusion is that a human-in-the-loop system is not a system with a safety net. It is a team, and teams have interfaces, capacity limits, training needs, and their own error modes. Design it as a team and the human adds real safety and real data. Design it as a dialog box, and you have purchased the appearance of oversight at the cost of the attention you will need when the model is finally wrong in a way that matters.
Training and recalibrating the reviewer
The mental-model finding — that team performance tracks whether the human understands where the system is weak, more than how much they trust it — implies something most teams never do: train the reviewer on the model.
What that looks like in practice is modest. An onboarding session that shows the system's known failure categories with real examples, so the reviewer has a concrete picture of what to look for rather than a general sense of caution. A short reference of the systematic biases: the model is worse on the long tail of rare formats, worse when the source document is contradictory, better on well-templated requests than on unusual ones. A standing channel where reviewers report patterns and the model team reports changes, so the two groups are not reasoning about different systems. And a periodic calibration exercise where several reviewers score the same batch and compare — not to impose consensus, but to surface the cases where the rubric is ambiguous, which is usually a rubric problem rather than a reviewer problem.
Two effects make this worth the effort beyond accuracy. Reviewers who understand the failure modes generate better labels, because their rejections are specific. And reviewers who know the system is sometimes wrong stay engaged, which is what prevents the drift into acceptance that makes the whole control illusory. The cost is a few hours of onboarding and a recurring thirty minutes; against the cost of the review capacity itself, it is a rounding error.
Works Cited
Amershi, Saleema, et al. "Guidelines for Human-AI Interaction." Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019. Accessed 9 Apr. 2026.
Bansal, Gagan, et al. "Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance." Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, 2019. Accessed 9 Apr. 2026.
Bansal, Gagan, et al. "Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance." Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 2021. Accessed 9 Apr. 2026.
Goddard, Kate, Abdul Roudsari, and Jeremy C. Wyatt. "Automation Bias: A Systematic Review of Frequency, Effect Mediators, and Mitigators." Journal of the American Medical Informatics Association, vol. 19, no. 1, 2012, pp. 121–127. Accessed 9 Apr. 2026.
Green, Ben, and Yiling Chen. "The Principles and Limits of Algorithm-in-the-Loop Decision Making." Proceedings of the ACM on Human-Computer Interaction, vol. 3, CSCW, 2019. Accessed 9 Apr. 2026.
Horvitz, Eric. "Principles of Mixed-Initiative User Interfaces." Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 1999, pp. 159–166. Accessed 9 Apr. 2026.
Lee, John D., and Katrina A. See. "Trust in Automation: Designing for Appropriate Reliance." Human Factors, vol. 46, no. 1, 2004, pp. 50–80. Accessed 9 Apr. 2026.
Madras, David, Toniann Pitassi, and Richard Zemel. "Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer." Advances in Neural Information Processing Systems, 2018. Accessed 9 Apr. 2026.
Mozannar, Hussein, and David Sontag. "Consistent Estimators for Learning to Defer to an Expert." Proceedings of the 37th International Conference on Machine Learning, 2020. Accessed 9 Apr. 2026.
Parasuraman, Raja, and Victor Riley. "Humans and Automation: Use, Misuse, Disuse, Abuse." Human Factors, vol. 39, no. 2, 1997, pp. 230–253. Accessed 9 Apr. 2026.
Skitka, Linda J., Kathleen Mosier, and Mark Burdick. "Accountability and Automation Bias." International Journal of Human-Computer Studies, vol. 52, no. 4, 2000, pp. 701–717. Accessed 9 Apr. 2026.