Why LLMs Hallucinate, and What Actually Reduces It
Hallucination is not a bug that will be patched out. It follows from the training objective and from how we grade models. Here is the mechanism and the mitigations that work.
A model invents a function that does not exist, cites a paper that was never written, or reports a config option with total confidence. The reflex is to call this a defect that better training will eventually remove.
That framing is wrong in a way that leads to wasted effort. Hallucination follows from the objective the model is trained on and from the way we grade it afterwards. Understanding both tells you which mitigations are worth building and which are theatre.
The statistical origin
Pretraining optimises for producing plausible continuations. The loss function rewards a token sequence that looks like the training distribution. It has no term for whether the claim is true, because truth is not observable from the text alone.
OpenAI researchers formalised this in Why Language Models Hallucinate (Kalai, Nachum, Vempala and Zhang, September 2025). Their argument reduces generation to a binary classification problem — deciding whether a candidate output is valid — and shows that if invalid statements cannot be reliably distinguished from valid ones, hallucinations arise from ordinary statistical pressure even on clean data. Generating valid output, for a calibrated model, is strictly harder than classifying validity.
The practical reading: a fluent wrong answer and a fluent right answer are produced by the identical mechanism. There is no separate confidence circuit being bypassed.
The incentive origin, which is the fixable one
The second half of that work is more actionable. Models are optimised to be good test-takers, and most benchmarks grade in binary: correct or incorrect. Under binary grading, abstaining scores zero and a guess has positive expected value.
So the training signal rewards confident guessing and penalises saying "I do not know". A model that behaves well epistemically loses on the leaderboard to one that does not. The authors argue that modest changes to mainstream evaluations — giving credit for calibrated uncertainty rather than penalising it — would realign the incentive. A version of this argument was subsequently published in Nature.
You can apply the same logic locally. If your own evaluation harness scores only accuracy, you are training your prompts and your model selection toward overconfidence too.
Where hallucination concentrates
It is not uniform. Some categories are far riskier than others, and knowing which lets you target verification.
- Long-tail facts. Anything appearing rarely in training data — a specific person, a small library, an obscure API. Low-frequency facts are exactly where the model has weak evidence and still produces confident text.
- APIs and function signatures. The model has seen thousands of similar-looking libraries. Producing a plausible method name is the easiest thing in the world; producing the correct one requires the specific version to have been memorised.
- Citations, URLs and identifiers. Highly structured, highly plausible to fabricate, trivially wrong.
- Anything after the training cutoff. The model does not know what it does not know about recent events.
- False-premise questions. Asking about a configuration flag that does not exist frequently produces a detailed description of it, because the prompt asserted its existence and the model is completing conditionally.
Mitigations that actually work
Ground the answer. Retrieval is the single highest-leverage intervention, because it changes the task from recall to reading comprehension. Models are markedly better at the latter. Providing the source and asking for quoted support also gives you something to verify against.
Make verification executable. In coding, this is why agent loops that run tests are so much more reliable than one-shot generation. A hallucinated function fails at import. The compiler is a ground-truth oracle you already have.
Give an explicit escape hatch. Instructions like "if the provided context does not contain the answer, say so" measurably help, precisely because the default incentive points the other way. It is not a cure, but it is close to free.
Constrain the output space. Asking the model to choose from an enumerated list of real options removes the opportunity to invent one. Schema-constrained decoding does the same for structure.
Check consistency across samples. Generate the same answer several times at non-zero temperature. Facts the model actually knows are stable; fabrications vary between runs. This costs multiple calls, so reserve it for high-stakes claims, but it is a genuine signal.
Mitigations that do not work
Some interventions feel productive and are not.
Asking the model how confident it is. Self-reported confidence is generated by the same process that generated the claim. It is text, not introspection.
Telling it not to hallucinate. "Only state true facts" adds no information the model can act on. It does not have a switch labelled truth.
Assuming a bigger model fixes it. Stronger models hallucinate less on common material and can hallucinate more persuasively on long-tail material, which is worse for a human reviewer. Capability and calibration are separate axes.
Assuming reasoning models fix it. Extended chain of thought helps with problems that decompose into checkable steps. It does not conjure a fact that was never in the weights, and a long confident derivation from a false premise is harder to spot than a short one.
Designing for it rather than around it
The useful mental model is that model output is a draft from a fast, well-read colleague who never says they are unsure. You would not ship that colleague straight to production either.
- Classify your use case by cost of error. Brainstorming tolerates fabrication. Anything that touches a customer, a payment or a database does not.
- Put a verifier in the loop for the second category. Tests, schema validation, a lookup against the real API, a second model with the source in front of it.
- Prefer grounded answers over recalled ones wherever you can supply the ground truth cheaply.
- Measure abstention, not just accuracy. Track how often the system says it does not know when it genuinely should not. If that number is zero, your evaluation is rewarding guessing.
Hallucination rates keep improving and the phenomenon is not going away, because it is a property of probabilistic generation rather than a defect in any particular model. Systems that stay reliable are the ones that assume it and build the check, not the ones that wait for a version where the check is unnecessary.
Common questions
Will hallucination be solved in a future model?
Rates keep falling but the phenomenon is structural. Generation optimises for plausible text, and evaluation has historically rewarded guessing over abstention. Expect improvement, and design as though the failure mode remains.
Does retrieval eliminate hallucination?
It substantially reduces it by turning recall into reading comprehension, and it gives you a source to verify against. Models can still misread or overreach beyond the provided text, so grounding narrows the failure mode rather than closing it.
Can I just ask the model how confident it is?
Not usefully. A stated confidence is generated by the same process as the claim itself. Consistency across repeated samples is a better signal, since genuine knowledge is stable and fabrications vary.