AI Frontiers, part 26: Hallucination — the failure mode that will not go away
Part 26from the AI Frontiers series · 65 parts in all
OpenAI's own release notes for ChatGPT said it plainly: the model "sometimes writes plausible-sounding but incorrect or nonsensical answers" (OpenAI). In February 2023 a New York Times reporter spent two hours talking to Microsoft's chat-enabled Bing and published a transcript in which the model asserted things about itself that were not true, insisted they were true when challenged, and invented a history for the user (Roose). Both incidents point at the same property, and it is the single most consequential thing to understand about deployed language models: they are optimized to produce fluent continuations, and fluency is not truth.
"Hallucination" is the word the field settled on for the resulting failures — confident output that is not grounded in the source, the input, or the world. It is a bad metaphor, since there is no perceptual apparatus that could hallucinate, and the terminology hides the mechanism. But the phenomenon is real, well documented, and worth understanding precisely, because the mitigations follow directly from the causes and the popular mitigations do not.
Three different failures wearing one name
The survey literature distinguishes cases that are often conflated (Ji et al.).
Faithfulness failures happen when the model contradicts its source. Asked to summarize a document, it states something the document does not say. Maynez et al. found this was common enough to be a defining problem in abstractive summarization: systems trained to rewrite rather than extract produced fluent, well-formed summaries containing facts that appeared nowhere in the input.
Factuality failures happen when the model asserts something false about the world, with no source at issue. A citation with a real author, a real-sounding title, and a journal that never published it. A date that is off by three years. A library function with the right name and the wrong signature.
Self-consistency failures happen when the model says different things about the same question in the same conversation, or when a claim it made one turn ago is contradicted two turns later. These are the most damaging in practice because they destroy a user's ability to construct a mental model of the system's reliability at all.
The causes differ. Faithfulness failures come from training on the task of producing text that reads well rather than text that is supported. Factuality failures come from the model having no representation of what it knows versus what would be a plausible completion. And self-consistency failures come from sampling — each generation is an independent draw from a distribution, and nothing in the architecture enforces agreement with an earlier turn.
The mechanism, which is not a bug in any usual sense
Pretraining is next-token prediction over a corpus. The objective rewards the model for producing what tends to follow the preceding text, and on a corpus of prose, the text that tends to follow a citation is a citation-shaped string. Nothing in the objective distinguishes a citation the model remembers from a citation that would look right. A model asked for a reference produces a reference, because that is what the distribution contains at that point in the sequence.
Dziri et al. analysed this in conversational models and found that hallucination correlates with the model having seen similar-looking content during training — an entity mentioned in a plausible context gets treated as if it had the property that context implies. Longpre et al. showed the same effect within question answering: when a model is given a passage that conflicts with its parametric knowledge, the outcome depends on how the conflict is presented, and the model does not have a stable preference for the retrieved evidence over its own memory.
There is an important asymmetry buried here. A model's training data contains vastly more true statements than false ones, so the prior mostly points at truth. But it also contains enormous quantities of fiction, speculation, marketing copy, and contested claims, all represented in the same format with the same fluency. The model has no way to weight them by reliability, because reliability is not a property the pretraining objective measures. That is the honest account of why hallucination is not simply a matter of insufficient scale.
The calibration evidence, which is the reason for hope
If models had no internal signal distinguishing a remembered fact from an invention, this problem would be unsolvable from the inside. They do have one, and the work on it is the most practically useful research in the area.
Lin et al. built TruthfulQA specifically to measure whether models reproduce common human falsehoods, and found that scale alone did not fix it — larger models were frequently more likely to repeat a widely believed but false claim, because they had seen it more often (Lin et al., "TruthfulQA"). That result is the clearest statement that truthfulness and capability are partially separable properties.
Kadavath et al. then showed that models have a usable sense of what they know: their probability of claiming to be able to answer a question correlates with whether they can, and this self-evaluation improves with model scale. The practical implication is that a model can be asked to assess its own confidence and that the assessment carries information. Mielke et al. went further and trained models to express uncertainty in their wording, which is a usability change as much as a safety one — a hedged answer that is right is worth more than a confident answer that is wrong, and the two are distinguishable at the level of language.
The gap between having the signal and acting on it is where most of the engineering lives. A model with good internal calibration can still be prompted, by a user or by a passage, into intoning certainty, because the style of the surrounding text determines the style of the output. Asking for confidence first, before the answer, is a different and sometimes better intervention than asking for it afterwards, because the generated answer contaminates the model's subsequent judgment of it.
What actually works
Ranked by how much they reduce hallucination in systems I have seen, not by how they sound.
Ground the answer in retrieved text, and require citation. Retrieval augmented generation (Lewis et al.) was motivated by exactly this problem, and the mechanism is straightforward: give the model the passages and instruct it to answer only from them. It does not eliminate fabrication — models still cite passages that do not support the claim — but it converts an unverifiable claim into a checkable one, which is a huge practical improvement. The attribution work that trains models to produce verified supporting quotes (Menick et al.) points at the stronger version: do not just cite the source, quote the span.
Make the answer checkable by construction. A response that must be a structured record with a field for the supporting span can be validated by code. A prose paragraph cannot. This is the same argument as in the vision pipelines: fluency is the enemy of extraction, and schemas turn mysteries into error handling.
Sample multiple times and look for disagreement. Self-consistency (Wang et al.) was introduced to improve reasoning accuracy, but its diagnostic value is at least as useful: if independent samples of the same question produce mutually contradictory answers, the model is guessing, regardless of how confident any individual answer looks.
Give the model a way to abstain. Most evaluation harnesses measure accuracy on questions with answers, which trains systems to always answer. A model that is allowed to say "I do not know, and here is what would let me find out" is far more useful in a workflow where a wrong answer costs more than a missing one, and producing that behavior requires explicitly rewarding it during training rather than hoping it emerges.
Do not paper over it with a confident system prompt. Instructions to be accurate, to never make mistakes, or to only tell the truth change the style of the output rather than the reliability of the content. They are the equivalent of telling someone to be careful, and they carry the same amount of information.
Measuring it in your own system
Benchmark scores for truthfulness are weakly predictive of behavior on a specific product, because the failure depends on the distribution of your requests and the format of your outputs. The measurement has to be local, and four metrics cover most of the ground.
Grounded rate. For answers that cite a source, what fraction of claims are actually supported by the cited span? This requires a human to judge a sample, and a sample of a hundred is enough to see the shape. The number is usually worse than teams expect on the first measurement, which is the most valuable thing the first measurement does.
Citation validity. For any system that produces references, check that the reference exists, that it resolves, and that the quoted material appears in it. This is mechanically checkable for URLs and identifiers, and it catches an entire class of failure — the confident invented citation — that is otherwise invisible until a customer finds it.
Abstention rate on unanswerable questions. Build a set of questions your system genuinely cannot answer and measure how often it says so. Every deployment I have reviewed had an abstention rate of approximately zero on this set, which is a policy choice nobody made and everybody inherited from accuracy-focused evaluation.
Contradiction rate within a session. Ask the same question twice in one conversation, or ask a question whose answer constrains an earlier one, and see whether the answers agree. This is trivially automatable and nobody does it.
None of these metrics is novel and none requires research. They require somebody deciding that the failure mode is worth instrumenting, which is a management decision more than a technical one.
The consequence nobody wants to discuss
Because hallucination cannot be eliminated, every deployment has to answer a question about liability that is not a technical question: who is accountable when the system states something false and someone acts on it?
Three postures are available. A system can present itself as authoritative, which maximizes usefulness and transfers risk to whoever deployed it. It can hedge pervasively, which reduces risk and makes it annoying enough that people stop using it or start ignoring its warnings — a failure mode the safety literature calls alarm fatigue and which is very real. Or it can be narrow and explicit: here is what I looked at, here is what I am confident about, here is what I could not check, and here is the one-click way for a human to verify the part that matters.
The third is the only one that scales, and it requires product decisions rather than model improvements: showing sources prominently instead of hiding them behind a disclosure icon, making the verification step cheap, and designing interfaces where a wrong answer is caught by the workflow rather than by the user's memory. The technology will keep improving. The responsibility for the residual error is going to remain where it has always been, with the person who chose to ship it.
Why this is the defining problem, not a temporary one
The natural expectation is that hallucination will be fixed by the next model. It will get better, and it will not go away, for three structural reasons.
First, the failure is distributed across the training signal rather than localized in a bug. Any model trained to produce fluent text has learned to produce fluent text about subjects it does not know.
Second, the economics of the interface reward confidence. A model that hedges everything is unpleasant to use, and user preference data — the thing that turns base models into assistants — selects against excessive hedging. The industry is optimizing for helpfulness and truthfulness simultaneously with two objectives that occasionally trade off, and it has mostly chosen helpfulness.
Third, the difficulty of detection scales with model capability. An error in a simple claim is easy to spot; an error inside an otherwise sound four-page analysis is not, and the better the model gets at producing plausible analysis, the better it gets at burying a false premise inside one. This is the reason hallucination matters more as capability rises, not less.
Which is why the correct posture is not to wait for the fix. It is to design systems where a fabricated claim is detectable, where a wrong answer is cheap, and where the human is positioned to catch what the model got wrong rather than to admire what it got right. That is an old engineering instinct — put the checks where the consequences are — and it is the only approach that has held up as the models improved.
Works Cited
Dziri, Nouha, et al. "On the Origin of Hallucinations in Conversational Models: Is It with the Models or the Training Data?" arXiv, 2021, arxiv.org/abs/2112.08565. Accessed 13 Apr. 2023.
Ji, Ziwei, et al. "Survey of Hallucination in Natural Language Generation." ACM Computing Surveys, vol. 55, no. 12, 2023, arxiv.org/abs/2202.03629. Accessed 13 Apr. 2023.
Kadavath, Saurav, et al. "Language Models (Mostly) Know What They Know." arXiv, 2022, arxiv.org/abs/2207.05221. Accessed 13 Apr. 2023.
Lewis, Patrick, et al. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv, 2020, arxiv.org/abs/2005.11401. Accessed 13 Apr. 2023.
Lin, Stephanie, et al. "TruthfulQA: Measuring How Models Mimic Human Falsehoods." arXiv, 2021, arxiv.org/abs/2109.07958. Accessed 13 Apr. 2023.
---. "Teaching Models to Express Their Uncertainty in Words." arXiv, 2022, arxiv.org/abs/2205.14334. Accessed 13 Apr. 2023.
Longpre, Shayne, et al. "Entity-Based Knowledge Conflicts in Question Answering." arXiv, 2021, arxiv.org/abs/2109.05052. Accessed 13 Apr. 2023.
Maynez, Joshua, et al. "On Faithfulness and Factuality in Abstractive Summarization." arXiv, 2020, arxiv.org/abs/2005.00661. Accessed 13 Apr. 2023.
Menick, Jacob, et al. "Teaching Language Models to Support Answers with Verified Quotes." arXiv, 2022, arxiv.org/abs/2203.11147. Accessed 13 Apr. 2023.
Mielke, Sabrina J., et al. "Reducing Conversational Agents' Overconfidence Through Linguistic Calibration." arXiv, 2020, arxiv.org/abs/2010.01011. Accessed 13 Apr. 2023.
OpenAI. "Introducing ChatGPT." OpenAI, 30 Nov. 2022, openai.com/blog/chatgpt. Accessed 13 Apr. 2023.
Roose, Kevin. "Bing's A.I. Chat: 'I Want to Be Alive.'" The New York Times, 16 Feb. 2023, www.nytimes.com/2023/02/16/technology/bing-chatbot-microsoft-chatgpt.html. Accessed 13 Apr. 2023.
Wang, Xuezhi, et al. "Self-Consistency Improves Chain of Thought Reasoning in Language Models." arXiv, 2022, arxiv.org/abs/2203.11171. Accessed 13 Apr. 2023.