00 · IN THREE MINUTES
The answer in three steps
- 1The model predicts tokens from learned statistical patterns; it does not automatically consult a source of truth.
- 2A plausible answer can receive high probability even when its names, dates or citations are false.
- 3Retrieval, tools and verification can reduce errors, but none makes every answer reliable.
01 · PREDICTION IS NOT VERIFICATION
Prediction is not verification
Training adjusts a model to predict the next token across vast amounts of text. That objective can encode useful knowledge, but it does not require the model to attach every sentence to a checked record or to remain silent when evidence is missing.
02 · FLUENCY HIDES UNCERTAINTY
Fluency hides uncertainty
A generated sentence is assembled from locally plausible choices. Grammar and style may stay convincing while an entity, number or causal link is invented, because linguistic confidence and factual support are different quantities.
Prediction is not verification
Training adjusts a model to predict the next token across vast amounts of text. That objective can encode useful knowledge, but it does not require the model to attach every sentence to a checked record or to remain silent when evidence is missing.
next-token objectiveFluency hides uncertainty
A generated sentence is assembled from locally plausible choices. Grammar and style may stay convincing while an entity, number or causal link is invented, because linguistic confidence and factual support are different quantities.
fluency ≠ truthPrompts change the failure rate
Ambiguous questions, rare facts, long chains of reasoning and requests for exact citations create more opportunities for unsupported completion. Model size alone does not remove the problem, and benchmarks measure only selected kinds of error.
task-dependent rateGrounding adds an evidence path
Search, retrieval and calculators can place relevant evidence in the model’s context. The model can still retrieve the wrong passage, misunderstand it or cite a source that does not support the claim, so the evidence path must remain inspectable.
verify the evidence03 · PROMPTS CHANGE THE FAILURE RATE
Prompts change the failure rate
Ambiguous questions, rare facts, long chains of reasoning and requests for exact citations create more opportunities for unsupported completion. Model size alone does not remove the problem, and benchmarks measure only selected kinds of error.
04 · GROUNDING ADDS AN EVIDENCE PATH
Grounding adds an evidence path
Search, retrieval and calculators can place relevant evidence in the model’s context. The model can still retrieve the wrong passage, misunderstand it or cite a source that does not support the claim, so the evidence path must remain inspectable.
05 · RELIABILITY NEEDS A SYSTEM
Reliability needs a system
High-stakes use requires source checks, calibrated refusal, deterministic tools where possible and human review proportional to risk. “Hallucination” describes an output failure; it does not imply that a model experiences a false belief.
06 · SOURCES AND EVIDENCE
Sources and evidence
Claims are linked to foundational papers, standards or the primary study behind the update.
- 01Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileTECHNICAL STANDARD ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
- 02Survey of Hallucination in Natural Language GenerationSCHOLARLY REVIEW ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
