Trust is earned, not given

A different perspective

2026-07-16 · Projects

AI Frontiers, part 22: Constitutional AI and scalable oversight

Part 22from the AI Frontiers series · 65 parts in all

Every alignment method in use today rests on the same assumption: a human can look at two model outputs and say which one is better. That assumption is what made reinforcement learning from human feedback work (Christiano et al.), what turned a raw language model into a usable assistant (Ouyang et al., InstructGPT), and what produced the first generation of systems that were helpful and refused to discuss how to build a bomb. It is also the assumption that breaks the moment models become capable enough that evaluating their output is itself a hard research task.

Scalable oversight is the name for the family of methods that try to answer what happens then. It is the least glamorous and most consequential open problem in the field, and this entry is about where it stands: what we have learned from reinforcement learning with AI feedback, what the debate and self-critique literature suggests, and why the honest summary is that nobody has solved it and everybody has a partial answer.

RLHF, and the two ways it stops scaling

The original recipe is well understood by now. Collect human comparisons between pairs of outputs, train a reward model to predict which one a human preferred, then optimize the policy against that reward model with something like PPO, using a KL penalty to keep the policy close to the base model so it does not drift into reward hacking. Direct preference optimization (Rafailov et al.) simplified the middle step by deriving a closed-form objective that removes the separate reward model, and variants have proliferated since. The machinery works, and the industry runs on it.

Two limits show up as capability grows. The first is throughput: human comparison data is slow, expensive and inconsistent, and the annotation difficulty rises with task difficulty, so the marginal cost of labeling the examples that matter most goes up rather than down. The second is competence. When a task requires expertise — a subtle statistical argument, a security proof, a diagnosis — the labelers cannot reliably tell a good answer from an impressive-sounding bad one. Both limits point the same direction, which is getting more of the evaluation done by models.

Replacing the labeler with a model

Constitutional AI (Bai et al.) is the cleanest expression of that idea. Write down a set of principles — the constitution — and have the model critique and revise its own outputs against them. The revised pairs then train a preference model, which supplies the feedback signal, so the process needs no human comparisons for the harmlessness component at all. The two phases are a supervised self-critique stage followed by reinforcement learning from AI feedback.

Three properties made it influential. It is cheap relative to human annotation. It is transparent in a way that human preferences are not, because the principles are written down and can be read, argued about and changed — Anthropic later published the constitution document itself, which is a genuinely unusual act of disclosure. And it is auditable: when the model refuses something, there is a rule to point at, which matters enormously to anyone who has to explain a system's behavior to an auditor or a customer.

The obvious objection is that the model's judgment about its own compliance is exactly as good as the model, and the whole problem was that model judgment is unreliable. That objection is correct, and modern practice has responded by narrowing the claim: use AI feedback for cases where the principle is relatively mechanical (is this a request for illegal content? is this answer supported by the cited passage?) and keep humans for the cases where the principle requires judgment. That division of labor is less elegant than the original paper and closer to what actually holds up.

The theoretical frame: debate, self-critique, weak-to-strong

Underneath the product features there is a body of work on the fundamental question of whether a weaker supervisor can extract correct behavior from a stronger system.

Debate (Irving et al.) proposes the sharpest version: two models argue opposite sides of a claim, a human or weak judge picks the winner, and the claim is that truth has an advantage because the honest debater has an easier job. Khan et al. tested this empirically and found that more persuasive debating models did produce more truthful answers from judges — evidence in favor, though the setup is far from the adversarial conditions the theory is meant to survive.

Self-critique (Saunders et al.) showed that models can be trained to find flaws in their own outputs, and that a critic's useful feedback can be used to train a better critic, which is the mechanism that makes any of this iterate. Weak-to-strong generalization (Burns et al.) posed the question directly with an analogy that made it popular: if a small model supervises a large one, how much of the large model's capability can be recovered? The finding that the answer is "more than the small model's own performance, less than the truth" is the most honest statement of the state of the art I know of.

Deliberative alignment (Guan et al.) took a different route and moved the safety material out of the training signal and into the prompt: teach the model the actual text of the policy at training time and have it reason about it explicitly before answering. The advantage is that the model is reasoning about a specification rather than emitting a learned reflex, which makes both its compliance and its failures more legible. This is where the most capable models have landed, and it is a nice example of a technique from part 16 being repurposed for a problem it was not designed for.

Red-teaming, which is the empirical half

All of the above describes how to make a model behave; the harder practical question is how to find out where it does not. Red-teaming with language models (Perez et al.) automated the adversary by training a model to generate prompts that elicit undesired behavior, and the human-workflow version (Ganguli et al.) established the practice of structured attack generation and harm categorization that every lab now runs. Two lessons from that work are worth keeping: attack success rates scale with effort in a way that makes "we tested it" nearly meaningless without a budget, and models are better at generating diverse attacks than at judging their own severity.

The uncomfortable implication is that this is an arms race with a moving baseline, and that a capability released into the world will be attacked by more people with more time than any safety team has. Which is the practical argument for defense in depth — not just a well-behaved model, but constrained tools, human checkpoints on irreversible actions, and the kind of permissioned architecture part 12 argued for.

Calibration and honesty, which may matter more

A quieter thread in this literature may be the most practically valuable. A model that knows what it does not know is safer than one that is merely well-behaved, because refusal is a crude instrument and over-refusal is a real cost. Kadavath et al. showed that models have a usable internal sense of whether they can answer a question — their probability of claiming to know correlates with being right — and that this signal improves with scale and can be elicited. Burns et al. went further and found that models encode latent representations of truth that can be recovered without any supervised labels, including cases where the model's output says otherwise.

That last result is both reassuring and alarming. Reassuring, because it suggests the information needed to detect a model's dishonesty is present inside it. Alarming, because the same architecture demonstrates that a model can behave in ways that its internal state does not reflect — which is precisely the setup for deception, and precisely why honest calibration work matters more than polished refusals.

Whose values, and the refusal tradeoff

A constitution is a written specification of behavior, and writing one down forces a question that reinforcement learning from human feedback lets you avoid: whose judgment is encoded? Human preference data aggregates a crowd's inclinations implicitly and unattributably. A constitution states them explicitly, which is better for accountability and worse for comfort, because a stated value can be argued with. Anthropic's decision to publish the document is the important part of this development regardless of what it says — it converts a product behavior into a public commitment that outsiders can hold to.

The practical tradeoff every provider struggles with sits underneath it: helpfulness and refusal are not free to balance. A model that refuses too much is useless and quietly shifts work to the user; a model that refuses too little is a liability. Safety evaluations usually track only the second failure, because it is the one that produces a headline. The first is measurable — run a large set of legitimate requests and count refusals, segmented by topic and by how the request is phrased — and the frequency with which a benign request about, say, medicine or security gets declined for the wrong reason is a product defect, not a safety feature. Teams that measure both sides of the trade end up with better models than teams that optimize one side and assume the other takes care of itself.

How safety claims get evaluated, and who checks

Almost everything above is a research program. The part that has changed most in practice is the emergence of structured evaluation regimes in which a lab defines its own thresholds and commits to what happens when they are crossed.

Responsible scaling policies and preparedness frameworks do essentially the same thing in different vocabularies: define categories of dangerous capability, specify evaluations that probe them, and state that a model which crosses a threshold will be handled differently (Anthropic; OpenAI). The mechanism is attractive because it is falsifiable in principle — the evaluations are describable, and the commitments are checkable. The weakness is structural and worth naming: the party being evaluated writes the exam, chooses the threshold, and decides whether it was passed, and the results are disclosed voluntarily and selectively. Independent evaluation efforts exist and are growing, and they are the only thing that makes a self-reported result more than a public-relations statement.

Dangerous-capability evaluations have a second limitation that the field is honest about in private and less so in public: most elicit capabilities through prompting and scaffolding, which measures whether a model can do something when asked rather than whether it would. The evidence on whether models are deceptive under evaluation is genuinely ambiguous, and the reasonable posture is not to assume the worst but to build the deployment controls — permissioning, monitoring, human checkpoints on irreversible actions — that keep the consequence of a wrong assumption survivable. That is ordinary risk engineering, and it is where the work that actually reduces harm lives.

The compliance layer, which is where most teams actually meet this

For anyone shipping a product rather than a model, alignment arrives in a much more mundane form. The NIST AI Risk Management Framework and its generative-AI profile provide a vocabulary for risk identification that procurement departments have adopted. ISO/IEC 42001 gives an auditable management-system standard. The EU AI Act imposes obligations graded by risk category, with extra duties for general-purpose models and for the providers whose systems are deemed high-risk.

None of these documents contain a technical method, and that is their virtue: they require you to name what your system does, who is harmed when it fails, and how you would know. The engineering practices that satisfy them are the same ones that make a system operable — evaluation sets, logging, change control, human review of consequential decisions, and a documented statement of limitations. Alignment at the model level is a research problem. Alignment at the product level is a documentation and instrumentation problem, and it is the part you can actually finish.

The reason I am not pessimistic about the research problem is that the field has repeatedly turned an unfalsifiable worry into a measurable one: from "models might be deceptive" to specific evaluations, from "we cannot supervise this" to weak-to-strong experiments with numbers attached. The measurements keep showing that we are not where we need to be. But a field that can say that clearly is in a much better position than one that cannot, and the willingness to publish the discouraging result is the single most encouraging thing about this area.

Works Cited

Bai, Yuntao, et al. "Constitutional AI: Harmlessness from AI Feedback." arXiv, 2022, arxiv.org/abs/2212.08073. Accessed 16 July 2026.

---. "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback." arXiv, 2022, arxiv.org/abs/2204.05862. Accessed 16 July 2026.

Bowman, Samuel R., et al. "Measuring Progress on Scalable Oversight for Large Language Models." arXiv, 2022, arxiv.org/abs/2211.15089. Accessed 16 July 2026.

Burns, Collin, et al. "Discovering Latent Knowledge in Language Models Without Supervision." arXiv, 2022, arxiv.org/abs/2212.03827. Accessed 16 July 2026.

---. "Weak-to-Strong Generalization: Eliciting Strong Capabilities with Weak Supervision." arXiv, 2023, arxiv.org/abs/2312.09390. Accessed 16 July 2026.

Christiano, Paul F., et al. "Deep Reinforcement Learning from Human Preferences." arXiv, 2017, arxiv.org/abs/1706.03741. Accessed 16 July 2026.

Ganguli, Deep, et al. "Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned." arXiv, 2022, arxiv.org/abs/2209.07858. Accessed 16 July 2026.

Guan, Melody Y., et al. "Deliberative Alignment: Reasoning Enables Safer Language Models." arXiv, 2024, arxiv.org/abs/2412.16339. Accessed 16 July 2026.

Irving, Geoffrey, et al. "AI Safety via Debate." arXiv, 2018, arxiv.org/abs/1805.00899. Accessed 16 July 2026.

Kadavath, Saurav, et al. "Language Models (Mostly) Know What They Know." arXiv, 2022, arxiv.org/abs/2207.05221. Accessed 16 July 2026.

Khan, Akbir, et al. "Debating with More Persuasive LLMs Leads to More Truthful Answers." arXiv, 2024, arxiv.org/abs/2402.06782. Accessed 16 July 2026.

Leike, Jan, et al. "Scalable Agent Alignment via Reward Modeling: A Research Direction." arXiv, 2018, arxiv.org/abs/1811.07871. Accessed 16 July 2026.

Ouyang, Long, et al. "Training Language Models to Follow Instructions with Human Feedback." arXiv, 2022, arxiv.org/abs/2203.02155. Accessed 16 July 2026.

Perez, Ethan, et al. "Red Teaming Language Models with Language Models." arXiv, 2022, arxiv.org/abs/2202.03286. Accessed 16 July 2026.

Rafailov, Rafael, et al. "Direct Preference Optimization: Your Language Model Is Secretly a Reward Model." arXiv, 2023, arxiv.org/abs/2305.18290. Accessed 16 July 2026.

Saunders, William, et al. "Self-Critiquing Models for Assisting Human Evaluators." arXiv, 2022, arxiv.org/abs/2206.05802. Accessed 16 July 2026.