Why Language Models Invent Sources, and What Actually Stops It
By Pooja Goenka ยท 2026-09-29
In 2023 a lawyer in New York needed case law for a filing and asked ChatGPT to find it. It came back with six decisions. He was suspicious enough to check, so he did the thing that feels like checking: he asked the chatbot whether the cases were real. It said yes. It added that he could find them on Westlaw and LexisNexis.
He filed them. Opposing counsel could not locate the cases. Neither could the judge. None of the six existed. The court sanctioned the lawyer, a colleague and their firm five thousand dollars, and the opinion described the reasoning inside one of the invented decisions as gibberish. The case is Mata v. Avianca, decided in the Southern District of New York on 22 June 2023, and you can read the opinion yourself.
The story usually gets told as a warning about trusting AI. I think that reading misses the useful part. Nothing malfunctioned that day. The model did exactly what it is built to do, and a citation is close to the worst possible thing to ask it for. That is worth understanding properly, because the same mechanism is sitting underneath every product any of us are building right now.
The explanation that sounds right and is not
The common account is that the model lied, or hallucinated, or was overconfident. Those words all smuggle in a mind that knew something and chose to misreport it.
There was no such mind. A language model does not hold a set of facts and consult them. It holds a very large statistical picture of which token tends to follow which other tokens, and at every step it produces a probability for every token in its vocabulary and picks from that distribution. That is the whole mechanism, and once you have it in view the invented citations stop looking like a malfunction and start looking like arithmetic.
Sharp distributions and flat ones
Give a model an unambiguous context and the distribution has one obvious winner. For "The capital of France is ___", Paris takes about 94 percent and the rest of the vocabulary splits a few points between them. The model looks like it knows a fact. What it has is a very lopsided distribution, which for practical purposes is the same thing.
Now give it an open context. "After the meeting, Anjali was feeling ___" produces nothing like a winner. Tired might be 12 percent, frustrated 10, optimistic 9, overwhelmed 8, and nearly half the mass scattered across thousands of other plausible words. There is no fact to retrieve here, and the model is not in any difficulty. It samples something reasonable and moves on, exactly as designed.
The trouble is that from the outside those two situations produce identical looking output. Same fluent sentence, same steady tone. The shape of the distribution is invisible in the result.
A case citation is the second situation wearing the clothes of the first. "Varghese v. China Southern Airlines Co., 925 F.3d 1339 (11th Cir. 2019)" is a highly constrained format. Two party names, a volume number, a reporter, a page, a circuit, a year. The model has seen tens of thousands of citations and has a precise sense of what one looks like. What it does not have, unless someone put it there, is the actual reporter. So it fills a well understood shape with plausible contents, and the result is a string that passes every test a reader applies by eye.
Why asking it to double check cannot work
The lawyer did try to verify. He asked the model directly whether the case was real, and it confirmed it and named the databases where he would find it. Most retellings skip that detail, and it is the most useful thing in the whole story.
That confirmation was generated the same way the citation was. "Is Varghese a real case?" is a context, and given everything that came before it in the conversation, "Yes, it is a real case" is an extremely likely continuation. The model was not consulting anything. It had no more access to Westlaw in that second message than it had in the first. Asking a model whether its own output is true is asking it to predict what a confirmation would sound like, and it is very good at that.
So the lesson is not that he failed to check. It is that the check he performed was not capable of being a check. There was no oracle in the room, only the same generator being asked a second question.
What retrieval actually changes
The fix is unglamorous and it is mostly not about the model. Before the model answers, you search a real source you control, pull back the specific documents that bear on the question, and put them into the prompt. Then you ask the model to answer from those. That is retrieval augmented generation, and the important word in it is retrieval.
The model has not become more truthful. It is a next-token predictor before and after. What changed is the context it is predicting from. When the relevant passage is sitting in the prompt, the continuation that reproduces that passage is now the high probability one, and the invented alternative is competing against a real document instead of against nothing.
That also explains the shape of the failures you get afterwards, which is a good sanity check on whether you have followed the argument. A grounded system still gets things wrong, in more tractable ways: retrieval returns the wrong passage, or returns nothing and the model answers anyway, or returns the right passage and the model summarises it badly. Each of those has a name, a test and a fix. "The model invented a source" has none of the three, because there is nothing there to repair.
Where the model's own memory is still the right tool
I do not want to leave you with the idea that everything needs a retrieval pipeline. Plenty of things do not.
For anything the model absorbed in enormous volume and that has not changed since, parametric memory is fine and much cheaper than a search: language itself, common algorithms, and the ordinary work of summarising, rewriting, translating and classifying text you have already handed it. In all of those the model is either working on what is in front of it or on patterns so widely attested that the distribution really is sharp.
The rule I use is this: if being wrong about a specific detail matters, and that detail lives in a source you could look up, look it up. Identifiers, prices, policies, dates, names, anything from your own systems, and anything from after the training cutoff. If you could not check the claim yourself without opening a document, the model could not either.
What to take from it
Assume a specific, checkable claim from a bare model is unverified until you have verified it against the source yourself. Not by asking the model again.
Notice when you are requesting a fact and when you are requesting prose. Those feel like the same request and are not, and the output will not tell you which one you made.
If you are building something where the facts come from your own data, put retrieval in from the start. Retrofitting it after someone notices the invented answer is the expensive order to do it in.
Where you can watch this happen
We built the LogicWiz Generative AI course around this problem, because it is the difference between using these models and shipping something that runs on them.
One lesson puts the probability distribution on screen so you can switch prompts and watch it go from a single 94 percent spike to a flat spread across a dozen candidates. Another runs the same question through the same model twice, once with nothing and once with a single retrieved support ticket, and you read the two answers side by side. Chapter five goes through retrieval properly: embeddings, chunking, indexing, and what to do when the retrieved passage is the wrong one.
Everything in it is free. Chapters one to three open with no account at all, and from chapter four a free account keeps you going. There is no card at any point.
The court opinion is worth twenty minutes of your time too. It is a clear account of what happens when a very capable writer of plausible text is mistaken for a source.