This is the shortest possible explanation of these tools, and the only technical section on this site. There is no math in it. Everything practical on the other tabs comes out of it.
What Westlaw does
You type a query. Westlaw searches an index of documents that exist, ranks them, and hands them back. That is retrieval. If it returns a case, the case is real, because the system had to find it somewhere in order to return it. Provenance comes free.
Retrieval has a characteristic failure: it misses things. The case you needed was worded differently and never surfaced. That failure is familiar, and every judge and clerk has developed habits for managing it. You search again, from a different angle.
What a large language model does
Nothing gets looked up. The model was trained on an enormous quantity of text, and from that training it learned, in effect, what words tend to follow what other words in what contexts. When you give it a prompt, it produces an answer one piece at a time, each piece being the most plausible continuation of everything so far. That is processing, and it is not a lookup with extra steps. It is a different operation.
Nothing in that process consults a library. The output is not the result of a search. It is a construction, assembled to be plausible.
Why it invents citations
Now the important part. A fabricated citation is not a malfunction. It is the system working exactly as designed. Asked for authority, the model produces text that has the shape of authority: a plausible case name, a plausible reporter, a plausible year, a plausible pin cite. Producing that shape is the job. Whether the case exists is not a question the architecture ever asks, because the architecture has no mechanism for asking it.
This is why the industry term, hallucination, is misleading. It implies a lapse from an otherwise sound faculty, as though the model normally knows and occasionally slips. There is no faculty to lapse from. The model is doing the same thing when it is right and when it is wrong. You just happen to like one result better.
Which means: the question is not when this gets fixed. Models get better, and error rates fall. But making the text more plausible is what improvement means here, and more plausible fabrications are harder to catch, not easier. The kind of tool it is does not change.
Why bias is built in the same way
Same architecture, same consequence. The model produces what is probable given the text it was trained on, and that text is the accumulated written output of a society with all of its assumptions in it. Probable is not the same as fair, or accurate, or representative. Patterns that nobody would defend if they were stated out loud are still patterns, and the model reproduces them because reproducing patterns is what it does.
Vendors work hard to suppress the obvious cases, and they largely succeed at the obvious cases. The subtle ones are the ones that matter to a court, and they are not detectable by looking at any single output. This is a structural feature, not a settings problem.
Four things this predicts
- Give the model the source and reliability rises sharply. When the relevant document is sitting in the prompt, the most plausible continuation is anchored to text the model can see. Errors drop a great deal. They do not drop to zero, and the model will still smooth over a distinction that matters, but this is the single largest lever you have.
- Ask it to work from memory and reliability is worst. No anchor, and fluent output regardless. This is precisely the mode in which a model invents a case, and it is why asking one to draft an opinion out of its own knowledge is the wrong use.
- Confidence carries no information. The tone of the answer is generated by the same process as the content of the answer. A wrong answer arrives in exactly the same voice as a right one. Hedging, when it appears, is a stylistic feature, not a reliability signal.
- It does not know what it does not know. Asking a model whether it is sure, or asking it to check its own work, produces more plausible text. It is not verification. Verification means going to the source.
A word about the “grounded” legal tools
Westlaw and Lexis have built AI layers that put a retrieval step in front of the model: find the real documents first, then have the model summarize and synthesize them. That genuinely helps, and it is a meaningfully better risk profile than a general chatbot answering a legal question from memory.
It is not a cure. The retrieval step can miss, and the model still writes the summary, which means it can still mischaracterize what it found, and still smooth two cases into one proposition that neither supports. Grounding narrows the failure mode. It does not eliminate the need to read the case.