Insights · legal AI
AI Legal Research Tools Are Useful and Unreliable: What the Stanford Studies and ABA Opinion 512 Require
Two Stanford studies put numbers on legal AI error rates, and the sanctions cases and ABA Formal Opinion 512 make verification a professional duty, not an option.
Two Stanford RegLab studies, taken together, settle a question many firms still treat as open: AI legal research tools are useful, but none of them is reliable enough to file without reading the underlying authority yourself. The first study showed that general chatbots get the law wrong most of the time. The second showed that the purpose-built, citation-grounded tools sold by the largest legal publishers get it wrong a meaningful fraction of the time. Neither finding is a reason to avoid these tools. Both are reasons to verify.
What the Stanford research actually found
In January 2024, researchers at Stanford's RegLab and Institute for Human-Centered AI (Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E. Ho) published "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models" (later published in the Journal of Legal Analysis). Testing general-purpose models then in wide use (GPT-3.5, Llama 2, and PaLM 2) against large volumes of legal queries, they found hallucination rates ranging from 69% to 88% on specific legal questions. Performance was worst exactly where lawyers need precision: when asked about a court's central holding, the models hallucinated at least 75% of the time. The models also tended to be confidently wrong, reinforcing a flawed premise in the question rather than correcting it.
That study measured raw chatbots with no connection to a legal database. The obvious rebuttal from vendors was that their products are different because they retrieve real authority before generating an answer. So the same group tested that claim. In May 2024, Magesh, Surani, Dahl, Suzgun, Christopher Manning, and Ho released "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools," evaluating the retrieval-grounded products from the two dominant publishers. The results were better than the chatbots and still sobering. Lexis+ AI and Thomson Reuters' Ask Practical Law AI each hallucinated on roughly 17% of queries. Westlaw's AI-Assisted Research hallucinated on roughly a third of queries (about 34%). For comparison, the study put a general model's error rate on legal queries in the range of 58% to 82%.
The headline for AI legal research accuracy is therefore not "these tools are broken." It is that even the best legal-specific tools, marketed at the time with language about being "hallucination-free," produced an unsupported or incorrect answer on roughly one in six to one in three queries. A tool that is wrong one time in six is genuinely valuable and absolutely cannot be trusted unchecked.
What "hallucination" means here, and why "misgrounded" is the dangerous one
The legal AI hallucination rate is easy to misread if you picture only invented case names. The Stanford researchers used a broader and more practical definition with two failure modes:
- Incorrect: the response misstates the law or the facts of a cited authority.
- Misgrounded: the response cites a real, existing source, but that source does not actually support the proposition the tool attached to it.
The second category is the one that should keep managing partners up at night, because it survives the most common quick check. An associate who confirms that every cited case exists, that the reporter citation is real, and that the case is good law has caught the crude error and missed the subtle one. A misgrounded answer looks fully clothed: real case, real citation, plausible parenthetical, wrong proposition. You only catch it by reading the cited authority and confirming it says what the tool claims. That is the entire verification lesson in one sentence.
Why retrieval grounding helps but does not solve the problem
Retrieval-augmented generation works by fetching relevant documents from a curated database and instructing the model to answer from them. It meaningfully reduces fabricated citations, which is why the legal tools outperformed the bare chatbots. It does not eliminate error, for reasons that are structural rather than fixable by a software update:
- The retrieval step can surface the wrong documents, or miss the controlling one, so the model reasons from an incomplete record.
- Even given the right documents, the model can summarize or characterize them incorrectly, which produces a misgrounded answer.
- Legal questions often turn on synthesis across authorities, jurisdiction, procedural posture, and whether a case remains good law, and these are precisely the tasks where the Stanford work found accuracy degrading.
The practical takeaway is that grounding changes the shape of the risk without removing it. The fabricated-case problem becomes rarer; the plausible-but-wrong-characterization problem remains. Verification has to be aimed at the second.
When a software flaw becomes a sanctions problem
Courts have already converted this reliability gap into professional discipline, and they have not been forgiving of the "the computer did it" defense.
The canonical case is Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023). Counsel submitted a brief containing cases ChatGPT had invented, then doubled down when opposing counsel and the court could not find them. On June 22, 2023, Judge P. Kevin Castel sanctioned the attorneys $5,000 under Rule 11, finding they had acted with subjective bad faith by continuing to stand behind fabricated authority after its existence was questioned. The lesson the court drew was not that using AI is sanctionable. It was that filing unverified output, and failing to correct it once warned, is.
The problem reached the appellate level in Park v. Kim, 91 F.4th 610 (2d Cir. 2024). Counsel's reply brief cited a case that did not exist, which she conceded she had generated with ChatGPT. The Second Circuit referred her to its Grievance Panel, framing the failure as a breach of the basic obligation to confirm that the authority she cited was real before putting it before the court. Two years on, trial and appellate courts across the country have issued comparable orders, and several have adopted standing orders requiring disclosure or certification of AI use. The pattern is consistent: the sanction attaches to the lack of verification, not to the tool.
What the rules already require: ABA Formal Opinion 512
None of this requires new ethics rules. On July 29, 2024, the ABA Standing Committee on Ethics and Professional Responsibility issued Formal Opinion 512, "Generative Artificial Intelligence Tools," mapping existing Model Rules onto AI use. The opinion is the cleanest statement of what a small or mid-sized firm must do, and it tracks the Stanford findings closely.
- Competence (Model Rule 1.1). Lawyers need not become AI experts, but must have a reasonable understanding of the capabilities and limitations of any tool they use. A 17% to 34% error rate is a limitation you are now on notice of.
- Candor and meritorious claims (Model Rules 3.1 and 3.3). Output verification is not optional, and the opinion ties the required level of independent review to the task. Generating an idea needs less scrutiny; relying on a tool's legal analysis or citations needs a lawyer to confirm it against the source.
- Confidentiality (Model Rule 1.6). Entering client information into a tool can implicate the duty of confidentiality, which matters when a "research" tool also trains on or retains inputs.
- Communication (Model Rule 1.4). Depending on circumstances, clients may need to be told how AI is being used in their matter.
- Supervision (Model Rules 5.1 and 5.3). Partners are responsible for establishing that associates and nonlawyer staff verify AI output. In a 3-to-30-attorney firm, this is a management obligation, not an individual one.
- Reasonable fees (Model Rule 1.5). You generally cannot bill for the hours a tool saved as though you spent them, and the time spent verifying is part of competent work, not a separate luxury.
The opinion does not tell you to avoid these tools. It tells you that the duty of competent representation now includes understanding that the tool can be wrong and building review around that fact.
A verification protocol that scales to a small firm
The goal is a habit cheap enough to follow every time, because a verification step you skip under deadline pressure is the one that ends up in a sanctions order. A workable minimum:
- Pull every cited authority. Confirm the case, statute, or regulation exists and is reported as cited. This catches fabrication.
- Read the cited portion, not the headnote. Confirm the source actually stands for the proposition the tool attached to it. This catches the misgrounded answer, which is the one verification of mere existence will miss.
- Confirm it is still good law. Check subsequent history and treatment in a conventional citator. A tool's training or index may predate a reversal.
- Check jurisdiction and posture. Confirm the authority governs your court and procedural setting, not merely a similar question elsewhere.
- Treat the tool's summary as an argument to test, not a fact to adopt. Use it to find leads and draft language, then verify as if a junior associate with a known error rate wrote it, because functionally one did.
- Set a firm policy and a named owner. Document which tools are approved, what may be entered into them given confidentiality duties, and who confirms verification before anything is filed. Formal Opinion 512 makes this a supervisory expectation.
For a firm in this size range, the cost of this protocol is minutes per brief. The cost of skipping it is the Rule 11 motion, the grievance referral, the client's lost confidence, and the published opinion with your name in it.
The bottom line
The Stanford study and its sequel give lawyers something the marketing did not: a number. General chatbots get the law wrong on a majority of legal queries. The best legal-specific tools get it wrong somewhere between one in six and one in three times, often by misgrounding a real citation rather than inventing a fake one. That is good enough to make these tools a real productivity gain and nowhere near good enough to file unread. The professional standard, under Formal Opinion 512 and the sanctions cases enforcing it, is not whether you used AI. It is whether you verified what it produced. Adopt the tools. Keep the duty to check.
This is general information for lawyers and law-firm leaders, not legal advice, and it does not create an attorney-client relationship. The authorities are cited so you can read them yourself.
The longer argument continues in AI in the Defender’s Office, a national field guide now in production.
Read about the book →