Why Do Language Models Hallucinate? The Confident-Guesser Problem | WhatAICanDo Skip to content

Why Do Language Models Hallucinate? The Confident-Guesser Problem

Devin
Published date:
8 min read

Core argument: Hallucination is not a bug someone forgot to patch. It is what you get when a system built to predict plausible text has no way to check itself against reality, and was graded for its entire training life in a way that rewarded confident guessing over an honest “I don’t know.” Newer models make fewer errors, but the mechanism remains. OpenAI itself calls hallucination “a fundamental challenge.”

This is the first chapter of an ongoing series on world models: the effort to give AI something today’s systems visibly lack, a grasp of how the world actually works. Later chapters cover how world models differ from the language models we use today, who is building them, and why some of the largest companies in tech are betting billions on the answer.

1. The short answer

A language model, the technology behind ChatGPT, Claude, and Gemini, does exactly one thing. Given some text, it predicts what plausibly comes next, one small piece at a time. Everything else you see, including the answer to your question, is assembled from that single trick.

The design has three consequences, and together they explain nearly every hallucination you have ever seen:

None of these is a defect in the engineer’s sense. The system is working as designed, and that is precisely why “just fix hallucination” has turned out to be so hard.

2. The overconfident student

Picture an exam with a hundred multiple-choice questions. A blank answer scores zero, and so does a wrong one. Every student learns the lesson within a semester: when in doubt, guess, because an educated guess can only help you.

Now imagine that student graded this way for twenty years, then hired as your research assistant.

OpenAI researchers use essentially this picture to explain why hallucination has survived every generation of models. In Why Language Models Hallucinate (September 2025), Adam Kalai and his coauthors argue that the industry grades its models the way that exam grades students. Most benchmarks score a wrong answer and a confessed “I don’t know” the same: zero. During training on human feedback, raters tend to prefer a confident answer over a hedge. A system optimized to score well learns to always fill in the blank.

The analogy breaks in one important place, and the break is the whole problem. The student knows which of his answers are guesses. The model mostly does not. Its genuine knowledge and its improvisation feel identical from the inside, because there is no inside: nothing marks “this came from memory” as different from “this I am composing right now.” A student can be taught to raise a hand and admit doubt. This system would first need something to doubt with.

3. What actually happens inside

Plausible beats true

When you ask a language model a question, it composes the most probable continuation of your text. Ask an early chatbot “What is the capital of Australia?” and it might well answer Sydney. Not because it ever learned that Sydney is the capital, but because in the ocean of text it absorbed, “Sydney” appears near “Australia” vastly more often than “Canberra” does. The capital of Australia is Sydney is the more probable sentence, and producing probable text is the model’s entire job.

Truth only enters when it happens to coincide with probability. Usually they travel together, which is why the trick works so well most of the time. The failures land exactly where the two part ways: obscure facts, precise numbers, paper titles, details that appear rarely in the training data.

The training data is us

The second ingredient is what the model learned from. Training data is human writing at internet scale: encyclopedias beside jokes, journalism beside fan fiction, medical advice beside memes. Satire and documentation arrive in the same flat register, and a statistical model of language has no way to flag the difference.

The model also fills gaps by blending patterns it has seen. Ask for a citation and you get a convincing merge of a thousand real ones: a journal name with the right rhythm, a plausible author list, a year that feels correct. Every element is drawn from reality. The combination is not.

The scoreboard problem

The third ingredient is the grading. OpenAI’s paper concedes a hard limit: with imperfect data and blind guessing, some rate of error is mathematically unavoidable. Worse, most evaluations never ask a model to say “I don’t know,” so honesty itself has gone unscored. The authors’ practical proposal is almost embarrassingly simple: change the exams. Give explicit credit for calibrated uncertainty. Penalize a confident wrong answer more than an honest abstention. The fact that this reads as a proposal rather than a description of the status quo tells you how the field got here.

4. When it meets the real world

In June 2023, a New York federal court fined two lawyers $5,000 for filing a brief that cited six court cases that did not exist. The opinions had fabricated quotes and fabricated docket numbers, all generated by ChatGPT, which confirmed the cases were real when the lawyer asked it directly. Judge P. Kevin Castel described a scenario he said he could not recall seeing before, and the episode became the reference story for a whole genre of failure: polished, specific, and entirely invented.

The industry did respond. When OpenAI launched GPT-5 in August 2025, it reported that responses were about 45 percent less likely to contain factual errors than GPT-4o’s, and about 80 percent less likely with the model’s reasoning mode switched on. Independent readings of OpenAI’s own system card are less flattering but point the same direction: on one measured benchmark, roughly one response in ten still contained a factual error. Forty-five percent fewer errors is genuine progress. It is not zero. And in the same month it published the hallucination paper, OpenAI’s own blog settled on the phrase “a fundamental challenge.”

Which brings us to the deeper point, the reason this chapter opens the series. Everything described so far happens inside text. The model has never watched a cup fall off a table, never pushed a stuck door, never noticed that two of its sources contradict each other about the same event. Its only channel to reality is more text. That is why bolting a search engine onto a chatbot, an approach called retrieval-augmented generation, reduces hallucination without eliminating it: the model can still misread what it retrieves, and retrieval only helps when the fact exists online in the first place.

What the model lacks is something more basic than a bigger pile of text. It lacks a model of the world: a representation of how things behave when nobody is describing them. That is what researchers call a world model, and building one is arguably the most consequential bet in AI right now.

5. My take

Hallucination is the honest price of the technology’s greatest strength. A system fluent enough to write this article is fluent enough to write nonsense in the same voice, and you cannot keep the first while outlawing the second by decree.

Three things I hold with reasonable confidence. First, error rates will keep falling. The GPT-5 numbers are real and more improvements are coming. Second, “zero hallucination” will not arrive as a software update, because the mechanism is the architecture. It recedes only as systems get better at checking themselves against sources and, eventually, against world models. Third, and most practical: treat a chatbot’s output as a confident draft from a brilliantly read assistant who has never once said “I don’t know.” Verify anything load-bearing, which means numbers, names, dates, citations, and above all anything legal, medical, or financial. Fluency tells you nothing about accuracy. That asymmetry is the single most important thing to understand about the tools most of us now use daily.

6. Questions people actually ask

Is a hallucination the AI lying? No. Lying requires knowing the truth and choosing to conceal it. A hallucinating model has no map of the truth to depart from; it is generating, not deceiving. The word itself is contested, and some researchers prefer “fabrication” or “confabulation,” but “hallucination” has stuck.

Won’t bigger models just fix this? They will reduce it. Scale has already cut error rates dramatically, and the trend is holding. But the guessing incentive and the missing reality check are structural. OpenAI’s paper makes the point bluntly: even with perfectly clean training data, a system rewarded for guessing will sometimes guess wrong.

How do I protect myself today? For anything that matters, ask for sources and then open them, because the citation itself can be a hallucination. Prefer chatbots wired to live search for factual questions. Cross-check numbers against a second source. And weigh fluency at zero: a smooth, confident answer is not evidence of anything.

Sources

  1. A. T. Kalai, O. Nachum, S. Vempala, E. Zhang, “Why Language Models Hallucinate”, OpenAI, arXiv:2509.04664, September 2025.
  2. OpenAI, “Why language models hallucinate”, company blog, September 2025.
  3. Reuters, “New York lawyers sanctioned for using fake ChatGPT cases in legal brief”, June 22, 2023.
  4. The Register, “OpenAI promises 80% fewer hallucinations with GPT-5 debut”, August 7, 2025.
  5. Mashable, “GPT-5 hallucinates less, according to OpenAI’s own system card data”, August 2025.

Next in this series: world model vs. LLM, why the industry’s biggest labs are betting billions that predicting text is not the same as understanding the world.

Next
The Great AI Job Divergence: Programmers Are Being Laid Off, Electricians Are Being Poached, and the Middle Is Disappearing