Trust is earned, not given

A different perspective

2024-10-10 · Projects

AI Frontiers, part 44: Synthetic preference data and the RLAIF flywheel

Part 44from the AI Frontiers series · 65 parts in all

Part 27 walked through RLHF as it was practiced in 2022 and 2023: humans rank pairs of model responses, a reward model learns the ranking, and policy optimization pushes the model toward high-reward outputs. The uncomfortable part of that pipeline is the first step. Human preference labels cost roughly a dollar to a few dollars each at commercial rates, take minutes to hours of wall-clock time, and carry an inter-annotator agreement that hovers around seventy percent on subjective questions. A useful alignment dataset is tens or hundreds of thousands of comparisons. The arithmetic never worked, and everyone in the field knew it.

So the flywheel turned inward. If a language model can write instructions, it can write responses; if it can write responses, a capable judge — human-written criteria in hand, or another model — can rank them. By late 2024, the majority of the preference signal used to align open models was model-generated, and a subfield had formed around making that substitution work rather than collapse. This entry is about that substitution: what it buys, where the loop bites back, and how to run one without producing a model that is confidently wrong in a new and more expensive way.

The three substitutions

The synthetic-alignment industry is really three separate substitutions, and conflating them causes bad engineering decisions because they fail differently.

Instructions. Self-Instruct established the pattern in late 2022: seed the model with a handful of hand-written tasks, ask it to generate more, filter for diversity against the existing pool, and generate the responses (Wang et al., "Self-Instruct"). Alpaca applied this at small scale with a strong teacher and got a chat-shaped model from a few hundred dollars of API spend — the moment, covered in part 25, when the long tail of open models started.

Responses. Given an instruction, generate several candidate answers. This is cheap and embarrassingly parallel; its limitation is diversity, since samples from one model cluster.

Judgments. Rank the candidates. This is the expensive psychological step, because it is where we stop asking a model to produce text and start asking it to have taste. The RLAIF paper did the controlled comparison: swap the human labeler for an off-the-shelf LLM given the same instructions, and measure (Lee et al.). On summarization and dialogue tasks the result was close to parity with human feedback — sometimes better on the human-evaluated axes, and dramatically cheaper. Constitutional AI had already shown a structured version of the same idea, using a written list of principles to critique and revise outputs before any preference ranking happened (Bai et al.) — the subject of part 22.

The pipeline that actually shipped

The most influential artifact of this era is not a paper but a dataset and a pipeline. UltraFeedback collected roughly 64,000 prompts, generated four responses to each from a mixture of models, and had a strong model grade each response on separate axes — helpfulness, honesty, instruction-following, truthfulness — then trained a reward model on the fine-grained scores (Cui et al.). Two design choices are worth stealing. First, decomposing the judgment into named criteria: a judge asked for a single "which is better" verdict is doing an unstable summarization of several independent considerations, while a judge asked for scores on four axes can be audited per axis, and the axes fail independently. Second, keeping the generator pool heterogeneous, because ranking responses from four different models produces a much better-conditioned signal than ranking four samples from one.

Zephyr showed the end-to-end version: distilled supervision from a large teacher, then DPO on synthetic preferences, and a 7B model that punched well above its class (Tunstall et al.). The DPO detail matters for cost. Direct preference optimization removes the separate reward model and the RL loop, turning alignment into a classification-style loss over preference pairs (Rafailov et al.) — which means the whole pipeline from synthetic judgment to trained model is a sequence of fine-tuning runs, no reinforcement-learning infrastructure required. That is the reason RLAIF-style data did not stay in the labs: it fits in the same tooling as supervised fine-tuning.

Closing the loop: when the model grades itself

Once judgments can be generated, the tempting move is to remove the teacher entirely. Self-Rewarding Language Models did exactly that: the same model acts as instruction follower and as LLM-as-a-judge, iteratively generating its own preference data across rounds and improving on both axes (Yuan et al.). SPIN replaced comparisons with a self-play objective in which the model's own previous-round outputs serve as the dispreferred examples (Chen et al., "Self-Play Fine-Tuning"). Both papers report real gains across iterations; both are also the clearest window into the failure mode, because a system that scores its own work has no external source of truth to correct a mistaken criterion.

The empirical check on this came from the critical evaluation of AI feedback by Sharma and colleagues: RLAIF can beat RLHF when the judge is a stronger model than the policy, and can lose to it when judgments come from a model of similar capability that is given no ground truth. Reward models trained on AI feedback are systematically over-optimistic about fluent-but-wrong answers, because fluency correlates with the judge's preferences and correctness does not correlate nearly as strongly. The paper also found that AI feedback tolerates reward-model over-optimization more gracefully than human feedback does — the degradation is slower — which is a genuine argument in its favor and not a license to skip verification.

Four ways the flywheel breaks

1. Criterion monoculture. If every generated judgment reflects one model family's notions of good writing, you optimize for that style. Symptoms are a distinctive register in all outputs, systematic verbosity, and — the dangerous one — hedging on questions where a direct answer was available. The fix is deliberate: multiple teachers, human ratings on a stratified sample of the synthetic data, and per-axis audits rather than a single quality score.

2. Judge leakage. The judge model has memorized benchmark answers, so preferences measuring "correctness" partly measure "is this in the judge's training data." Sharma et al. call this out directly; the practical mitigation is to hold out a human-judged slice and verify that gains on it track gains on the synthetic set. When they diverge, the synthetic metric is measuring the judge, not the model.

3. Entropy collapse. Repeated optimization against a self-generated signal narrows the output distribution. The model stops sampling the rare correct answer, because the rare correct answer scores no better than the common plausible one under a judge that cannot tell them apart. If you track nothing else while running this loop, track output diversity — distinct n-grams, response-length variance, self-BLEU — and stop when it falls.

4. Recursive degradation. Training on model output and training judges on model output are both compounding operations, and part 21 covered the collapse debate around the first. The distinguishing feature of preference data is that it does not have to compound: judgments are typically attached to human-written or human-selected prompts, which anchors the input distribution. The flywheel loses its anchor the moment prompts are also generated and then filtered by the same judge.

Operating a flywheel

What I would actually build, if I were aligning a model or a production policy: a stratified prompt pool with a human-curated core that never leaves the mix; three or more generators for candidate responses; a judge with an explicit rubric, required to justify before it scores, and audited on a fixed sample against human raters; a DPO or similar offline objective so training stays reproducible; and a reporting sheet that carries, next to every quality metric, the human-agreement rate of the synthetic set that produced it. That last column is the one that catches trouble, because it is the only number that can degrade while every other number improves.

The honest summary is that synthetic preference data solved a cost problem and created a measurement problem. Human labels were scarcer than we wanted; model judgments are abundant, fast, and circular unless you deliberately keep an outside reference in the loop. The teams that got this right in 2024 were not the ones with the biggest teacher — they were the ones that spent a few percent of the budget keeping their own scoreboard honest.

The economics, with numbers

The cost asymmetry is worth spelling out because it is the entire reason this shift happened, and it is bigger than people expect. Human preference collection through a commercial vendor runs roughly one to three dollars per comparison at the quality level needed for alignment work, with a cheap tail well below that and a long tail far above it for expert domains. Inter-annotator agreement on subjective quality judgments typically lands in the sixty-to-eighty-percent range, which means even perfect execution gives you a reward signal with substantial noise. Producing a hundred thousand comparisons is a six-figure, multi-week project before anyone trains anything.

The synthetic version of the same dataset is a token bill. Generating four candidate responses to a hundred thousand prompts and grading each of them with a strong judge is on the order of a few million tokens in and a few million out — hundreds of dollars at 2024 prices, not hundreds of thousands, and hours rather than weeks because the work parallelizes without limits. The training run that consumes the data remains the expensive part; the data generation has stopped being the bottleneck.

That arithmetic has a specific organizational consequence. When labels were expensive, dataset quality was a resource-allocation problem and there was a natural limit on how many experiments you could run. When labels are nearly free, the limiting factor moves to evaluation — because you can now generate twenty candidate preference datasets and you have no cheap way to tell which of them trained a better model. Teams that scaled their synthetic data generation without scaling their human-verified evaluation capacity simply converted an old bottleneck into a new one, and the new one is harder to see because the symptoms are subtle: everything improves on every metric you are tracking.

Choosing the teacher, and knowing when not to use one

Three properties of the teacher determine whether the flywheel produces a better model or a more assertive one.

Capability gap. The Sharma et al. result is worth restating as a design rule: a judge stronger than the policy gives useful signal, and a judge comparable to the policy mostly amplifies the policy's existing preferences. If your policy model and your judge model are the same size, you are not aligning; you are sharpening. The common failure is a team using a local 7B judge to grade a local 7B policy and slowly teaching it to sound more confident.

Rubric specificity. A judge asked for a holistic preference ranks by fluency and length. A judge asked for scores on named criteria — does the answer satisfy the request, is every claim supported, does it decline where it should — ranks by those criteria and can be audited per criterion. The published recipes that held up all decomposed judgment this way, including the mixture-of-teachers pipeline behind Nemotron-4 (Adler et al.).

Whose values. A single teacher's preferences encode that vendor's editorial choices, and deploying those choices as an invisible default is a governance decision that nobody signed. Anthropic's Collective Constitutional AI experiment is interesting precisely because it made the choice explicit, deriving constitutional principles from a public process rather than from a room of researchers (Huang et al.). Whatever process you use, the principles should be written down somewhere a customer could read them, because in a regulated domain they will eventually be asked for.

Where a teacher should not be used is where correctness is checkable. If a test suite, a type checker, a database, or an authoritative document can adjudicate the answer, that adjudication is a better label than any model's opinion, and it costs nothing to trust. The most robust alignment pipelines I have seen put executable verification first, model judgment second, and human review only on the residue.

Works Cited

Adler, Bo, et al. "Nemotron-4 340B Technical Report." arXiv, 2024, arxiv.org/abs/2406.11704. Accessed 10 Oct. 2024.

Bai, Yuntao, et al. "Constitutional AI: Harmlessness from AI Feedback." arXiv, 2022, arxiv.org/abs/2212.08073. Accessed 10 Oct. 2024.

Chen, Zixiang, et al. "Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models." arXiv, 2024, arxiv.org/abs/2401.01335. Accessed 10 Oct. 2024.

Cui, Ganqu, et al. "UltraFeedback: Boosting Language Models with Scaled AI Feedback." arXiv, 2023, arxiv.org/abs/2310.01377. Accessed 10 Oct. 2024.

Huang, Saffron, et al. "Collective Constitutional AI: Aligning a Language Model with Public Input." arXiv, 2024, arxiv.org/abs/2406.07814. Accessed 10 Oct. 2024.

Lee, Harrison, et al. "RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback." arXiv, 2023, arxiv.org/abs/2309.00267. Accessed 10 Oct. 2024.

Rafailov, Rafael, et al. "Direct Preference Optimization: Your Language Model Is Secretly a Reward Model." arXiv, 2023, arxiv.org/abs/2305.18290. Accessed 10 Oct. 2024.

Sharma, Archit, et al. "A Critical Evaluation of AI Feedback for Aligning Large Language Models." arXiv, 2024, arxiv.org/abs/2402.12366. Accessed 10 Oct. 2024.

Tunstall, Lewis, et al. "Zephyr: Direct Distillation of LM Alignment." arXiv, 2023, arxiv.org/abs/2310.16944. Accessed 10 Oct. 2024.

Wang, Yizhong, et al. "Self-Instruct: Aligning Language Models with Self-Generated Instructions." arXiv, 2022, arxiv.org/abs/2212.10560. Accessed 10 Oct. 2024.

Yuan, Weizhe, et al. "Self-Rewarding Language Models." arXiv, 2024, arxiv.org/abs/2401.10020. Accessed 10 Oct. 2024.