DIZON.LAW

Insights · Tools

Why grounded (RAG) legal AI hallucinates less, and why you still verify

A system that answers from a defined set of documents hallucinates less than a raw chatbot, but a lower error rate is still an error rate, and the duty to verify does not change.

There is a useful distinction hiding inside the phrase "legal AI," and most of the confusion about hallucination comes from missing it. A raw chatbot answers from memory. A grounded system answers from documents you can name. The difference is not cosmetic, and it is the single most important thing for a managing partner to understand before letting any tool near a client matter.

Two different machines that look the same

When you type a question into a general-purpose chatbot, it does not look anything up. It predicts the next words based on patterns absorbed during training. That is why it can produce a citation that has the cadence, the reporter, the page number, and the confident tone of a real case, and yet describe an opinion that was never written. The model is not lying in any human sense. It is generating fluent text, and a plausible fake is, to the model, just as easy to produce as the truth. This is the failure that put two lawyers in front of Judge Castel in Mata v. Avianca (S.D.N.Y. 2023), where a brief built on cases invented by ChatGPT drew a $5,000 sanction and an order to notify the real judges falsely named as authors.

Retrieval-augmented generation, usually shortened to RAG, is a different machine. Before the model writes anything, the system runs a search against a defined body of documents: a commercial case database, a statute set, or a firm's own files. It pulls back the passages that appear most relevant, places them in front of the model, and instructs it to answer from that material. The model still writes the prose, but now it is summarizing text it has actually been handed rather than reaching into a foggy memory of the entire internet. That architecture is the heart of nearly every serious retrieval augmented generation law product on the market, and it is what people mean when they talk about grounded legal AI.

Why grounding lowers the hallucination rate

The intuition is straightforward. If the right answer is sitting in the model's context window, quoted from a real source, the model has far less reason to invent one. Grounding narrows the model's job from "recall the law of the United States" to "read these six passages and tell me what they say." The second task is much harder to get catastrophically wrong.

This is not just theory. The most rigorous independent look at the question is the Stanford RegLab and Human-Centered AI study "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools" (Magesh, Surani, Dahl, Suzgun, Manning, and Ho), released as a preprint in 2024 and later published in the Journal of Empirical Legal Studies. The researchers ran more than 200 legal research queries through the major grounded tools and hand-scored every answer. Two findings matter for a firm leader.

First, grounding works. The retrieval-based legal research tools hallucinated substantially less than a general model like GPT-4 used on its own. Grounding is a real engineering improvement, not marketing.

Second, grounding does not eliminate the problem. The study found these specialized, document-grounded legal tools still produced incorrect or misgrounded answers a meaningful share of the time, on the order of one in six for the better performer and roughly one in three for another. That was measured against vendor marketing that had described, in one instance, "hallucination-free" linked citations. The gap between the claim and the measured result is the whole lesson.

How a grounded system still gets it wrong

It helps to know the specific ways a RAG system fails, because they are different from how a raw chatbot fails, and they are sneakier.

  • Retrieval misses. If the search step does not surface the controlling authority, the model answers confidently from whatever it did retrieve. The omission is invisible in the output. The answer looks complete because nothing flags what is absent.
  • Misgrounding. The model cites a real, retrieved case but characterizes its holding incorrectly, or stitches a citation to a proposition the source does not support. The citation checks out; the claim attached to it does not. This is the RAG hallucination pattern that fools a quick reviewer, because the case is real and the quote may even be verbatim while standing for the opposite of what the brief says.
  • Stale or wrong corpus. A system grounded on a database that is missing recent decisions, or that includes overruled authority without flagging it, will ground its answer in genuine but no-longer-good law. Grounding guarantees the source exists. It does not guarantee the source is still controlling.
  • Confident synthesis across sources. Asked to reconcile several passages, the model may produce a clean rule that none of the sources actually states. The output reads like a holding. It is the model's own composition.

None of these is exotic. Each is a normal, recurring behavior of grounded tools, which is why the technology reduces risk rather than removing it.

The firm precedent engine, done correctly

The most valuable version of this technology for a 3-to-30-lawyer firm is not a better way to search published case law. The commercial databases already do that. The opportunity is to ground a system on the firm's own work product: its briefs, motions, contracts, prior research memos, and closed-matter files. Call it a firm precedent engine. Ask it how the firm argued a particular motion to dismiss two years ago, and it retrieves the actual brief and answers from it, rather than from a generic notion of how such motions are written.

This is where grounding pays off most, for a reason worth naming. The firm's own documents are a closed, knowable, high-quality corpus. You wrote them. You can verify them against the matter file. A system grounded on your own filings is both more useful and more checkable than one reaching across the open web, because the source of every answer is a document already in your possession. It turns the institutional knowledge that currently lives in a few partners' heads, and walks out the door when they retire, into something the whole firm can query.

The same caution still applies. A precedent engine can retrieve the wrong prior brief, or summarize a settlement term incorrectly, or surface a clause from a deal that fell through. Grounding makes the answer traceable. It does not make it true. The advantage is that traceability makes verification fast, because the answer points you straight to the source document instead of leaving you to confirm a citation from scratch.

Why you still verify everything

Here the engineering question becomes an ethics question, and the answer is settled. ABA Formal Opinion 512 (July 29, 2024), the ABA's first formal ethics opinion on generative AI, is explicit that a lawyer must independently verify the output of these tools, with scrutiny proportional to the stakes. That duty does not soften because a system is grounded. A lower hallucination rate is still a hallucination rate. A tool that is wrong one time in six is wrong often enough to end a career if the wrong one reaches a court, and the courts have shown no patience for "the software cited it."

What grounding changes is not whether you verify but how cheaply you can. With a raw chatbot, a citation is a lead you have to run down from nothing. With a well-built grounded tool, every assertion should link to the passage it came from, so verification becomes reading the cited source and confirming it says what the tool claims. That is the right way to evaluate any vendor: not "does it hallucinate," because every one of them does, but "how fast does it let me catch the hallucination."

A short standard for any grounded tool, whether you buy it or build it:

  • Ask what corpus it is grounded on, and how current that corpus is.
  • Require pinpoint citations back to specific passages, not just a list of sources at the end.
  • Open the cited source and read it. Confirm it exists and that it supports the exact proposition, every time, before anything leaves the firm.
  • Watch for what is missing, not only what is wrong. Retrieval failures hide in omissions.
  • Keep a human accountable for every filed or sent work product, in writing, as firm policy.

The plain takeaway

Grounding is a genuine advance and the right architecture for legal work. A system that answers from a defined set of documents, your firm's files or a real legal database, hallucinates meaningfully less than a chatbot answering from memory, and it makes its few remaining errors easier to catch because it shows its sources. That is a strong foundation to build a practice on. It is not a substitute for a lawyer reading the case. The firms that do well with this technology are not the ones who trust it. They are the ones who use grounded tools precisely because grounded output is faster to verify, and then they verify it.

This is general information for lawyers and law-firm leaders, not legal advice, and it does not create an attorney-client relationship. The authorities are cited so you can read them yourself.

The longer argument continues in AI in the Defender’s Office, a national field guide now in production.

Read about the book