AI & Data Privacy
AI Hallucinations in Legal Work: What an ADGM Costs Order Should Teach Every Firm
In December 2025 an ADGM court ordered a law firm to pay AED 282,508 in costs after its pleadings cited authorities that did not exist, or did not say what the pleadings claimed. The mechanism behind that is worth understanding precisely.
On 18 December 2025, the ADGM Court of First Instance ordered a law firm to pay AED 282,508 — roughly US$77,000 — in costs, after pleadings filed on behalf of its client were found to contain authorities that were fabricated, wrongly cited, or that did not support the proposition they were cited for. The court linked the errors to AI-assisted research used without adequate verification.
It is worth being precise about what that decision was, because the precision is the lesson. It was not a regulatory fine and it was not a finding about AI being impermissible. It was a costs order: the other side and the court had their time wasted, and somebody had to pay for it. The judge's point was narrow and hard to argue with — authorities produced with AI assistance must be treated sceptically and verified before they are put in front of a court, and that duty sits with the lawyer.
Key takeaways
- The expensive failure mode is not the invented case. It is the real case cited for something it does not say — which survives a surface check.
- This is not a prompting problem. Purpose-built legal research tools measure badly on it too.
- Every professional-conduct framework already answers the question: verification is the lawyer's duty, and it does not transfer to a vendor.
- Open-domain generation and closed-domain extraction carry structurally different risk. Ask which one a tool is doing before you ask about its accuracy.
Three failures wearing one name
“Hallucination” is used loosely for three quite different problems, and they are not equally dangerous.
The fabricated authority
A case, statute, or clause that does not exist. This is the one that makes headlines, and it is the least dangerous, because the first person to look it up finds nothing. It is caught by an ordinary check.
The misattributed citation
A real authority, cited with the wrong reference — right principle, wrong case number or paragraph. Annoying, embarrassing, and usually caught by a careful reader.
The misapplied authority
A real case, correctly cited, advanced for a proposition it does not support. This is the expensive one. Every surface check passes: the case exists, the citation resolves, the court is real. Only someone who reads the judgment finds the problem — and under deadline, that is exactly the check that gets skipped.
Watch out
The ADGM pleadings reportedly contained all three categories. That combination is the signature of unverified AI-assisted research, and it is why “we checked the citations exist” is not a sufficient answer.
The measurements are worse than most people assume
The instinct is to treat this as a discipline problem — a careless associate, a bad prompt. The measurements do not support that reading.
- A 300-task benchmark of ten frontier models published by the legal AI vendor HAQQ found that 24% of answers cited law that did not actually support the proposition advanced. Not invented law — real law, misapplied. Roughly one answer in four.
- Stanford's RegLab assessed leading purpose-built legal research tools in a study pointedly titled Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, and found hallucination rates in the range of roughly 17% to 33% despite vendor claims to the contrary.
Read those together and a conclusion follows that is uncomfortable for the whole category: retrieval and legal specialisation reduce the rate, but neither eliminates it, and the marketing word “hallucination-free” has already been tested and found wanting more than once. Any vendor using that phrase — including any vendor writing a blog post like this one — should be asked what it means operationally rather than as an adjective.
The duty does not move
Whatever the tool, the obligation to verify stays with the lawyer. That is not a novel principle invented for AI; it is the ordinary competence duty applied to a new instrument, and every professional-conduct framework in the region reaches it by one route or another. The practical consequence is that a tool's value depends less on how often it is right than on how cheaply you can tell whether it is right this time.
A tool that is right 95% of the time and gives you no way to check is worse in practice than one that is right 90% of the time and shows its source for every claim — because the first shifts the entire verification burden onto you, and the second removes most of it.
Generation and extraction are not the same risk
The cases above share a shape: a model was asked an open-domain question — what does the law say, what authority supports this — and answered from what it had learned. There is no bounded source to check the answer against, so plausibility is the only thing constraining the output. That is the condition under which hallucination is not a bug but the expected behaviour.
Closed-domain extraction is a different problem. The question is not “what does the law say” but “what does this document say, and where”. The source is supplied, bounded, and known. Every output can carry a pointer back into it. The failure modes do not disappear — extraction can still miss things, and a missed clause matters — but they change character: an omission you can test for, rather than an assertion you cannot trace.
That is why the useful question to put to a vendor is not “does it hallucinate?”, to which the answer is always no. It is: what is the bounded source, and does every output point back into it?
What to require, whatever you buy
- 01A reference on every assertion — page, paragraph, or timestamp — that resolves in one click. Anything without one is a claim you must re-derive.
- 02A stated scope. Is the tool answering from a supplied document set, or from a model's training? Those are different products and should be evaluated differently.
- 03The ability to say nothing. A tool that returns an empty field and a note is more trustworthy than one that always produces an answer.
- 04Reproducibility. Run the same document twice. If the output differs, you cannot audit it, and neither can anyone reviewing your work.
- 05A verification step in the workflow, not in the policy document. “Lawyers must check AI output” is not a control. A review pass with the source one click away is.
The ADGM order is a useful marker precisely because it is regional, recent, and undramatic. Nobody was struck off; a firm paid the other side's costs because work went out unverified. That is the ordinary shape of this risk, and it is entirely avoidable at the point where a tool is chosen.
The other half of choosing a tool is what it does with your material once you upload it — covered in why pasting case files into ChatGPT could violate client confidentiality.
This article is general information about professional practice and technology, not legal advice. Data protection and professional conduct rules differ by jurisdiction and change often — the position described here is as at August 2026. Check the rules that bind you, and your firm's own policies, before changing how you handle client material.
Keep reading
AI & Data Privacy
Why Pasting Case Files Into ChatGPT Could Violate Client Confidentiality
A general-purpose chatbot is a third party. The moment a client document goes into one, you have made a disclosure decision — whether or not you thought of it that way.
Workflow
5 Ways Paralegals Can Cut Case File Review Time Without Cutting Corners
Speed in file review comes from doing fewer passes over the same pages, not from reading faster. Five techniques that hold up when the file is 4,000 pages.