The research note reads well. The proposition is stated crisply, the authority is on point, the citation is properly formatted, and the paragraph number is right there. It took forty seconds to produce. The only difficulty is that the judgment does not exist, and nothing on the page will tell you that.
The problem is not that AI tools are sometimes wrong. Every research method is: a junior misreads a judgment, a headnote misleads, a digest goes stale. The problem is that when an AI tool is wrong in this particular way, the output carries none of the usual signals of being wrong. It is fluent, confident, correctly formatted, and indistinguishable from the output that is right. AI legal research verification is a discipline precisely because the artefact you are handed offers no evidence of its own reliability, and the person who signs the filing carries the consequence either way.
Fluency Is Not Accuracy, and They Are Not Even Related
Professional judgement about written work rests on an assumption so basic we rarely notice it: that fluency correlates with competence. It is a reliable assumption, because in humans fluency is expensive. Producing a perfectly formed citation to a judgment you have never read is more work than producing a correct one, so nobody does it. A language model breaks that completely. For a model, fluency is free. It costs exactly as much to produce a flawless citation to a judgment that does not exist as to produce one that does. Fluency and accuracy are independent properties of model output, and every instinct you have for reading professional work assumes they are not.
The signal you are used to reading is simply not present
When a colleague hands you a confidently written note, the confidence is evidence, because a person who is unsure usually writes like a person who is unsure. When a model hands you the same note, the confidence is not evidence of anything. It is a property of the generation process, not a report on the underlying facts.
The Mechanism: Why a Model Invents a Citation

It is tempting to call this lying, or guessing. Both mislead, because both imply the model has access to a truth it is choosing to depart from. It has none. Understanding what actually happens tells you which citations to distrust, and why.
A model predicts the next token, and nothing else
A language model is trained to do one thing: given a sequence of text, predict what is statistically likely to come next. That is the entire objective. At every point the model answers the question, what would plausibly come next here, and never the question, is what I am about to say true. There is no term in that objective for whether a case exists, and no register being consulted. The model is not checking anything, because nothing in the architecture has checking as its job.
A citation is the most predictable string in legal writing
Consider what a citation is from the model's point of view: two party names joined by a versus, a year in brackets, a volume number, a reporter abbreviation from a small fixed vocabulary, a page number, a paragraph reference. The model has seen this pattern hundreds of thousands of times. So when generation reaches a point where a citation belongs, the pattern completes itself almost effortlessly, and every component is individually plausible, because every component is drawn from a distribution the model learned accurately. What the model learned is the shape of a citation, not the existence of a case. Format is learnable from text. Existence is not derivable from format.
A language model does not know whether a judgment exists. It knows what a citation to an existing judgment tends to look like. Those are not the same knowledge, and only one of them is in the model.
This also disposes of the hope that the model's own confidence might serve as a filter. Because the format is so heavily constrained, a model can generate the tokens of a fabricated citation with very high internal probability. It is confident about the shape. The shape is what it learned. The confidence attaches to precisely the part that was never in question.
It fails hardest exactly where you needed it most
A model's accuracy on a proposition tracks how heavily that proposition and its authority co-occur in what it was trained on. Ask for the leading authority on a foundational point under the Indian Contract Act 1872 and you will probably get something real, because that pairing appears constantly in legal writing. Ask about a narrow, unusual point, the kind you are researching precisely because you could not find it yourself, and the training data is thin. Thin data is where the model interpolates, and interpolation is where fabrication lives. The failure rate is not spread evenly: it concentrates in the queries you were least able to check, which is the same set that made you reach for the tool. The tool is least reliable exactly where it is most useful.
No retrieval step
A bare model looks nothing up. It answers from patterns compressed into its parameters, not from a document in front of it. Nothing ever asks whether the judgment named is on any court's record.
Format without referent
The format is learned from millions of examples and reproduced faithfully. The case it points at is generated like the rest of the sentence. Well-formed and real are different claims, and the model makes only the first.
Helpfulness pulls towards answering
Asked to supply authority, the path of least resistance is to supply something shaped like authority. Reporting that none was found is harder to produce than completing the pattern.
Three Failure Modes, Not One
The invented case gets the attention because it is the most vivid. It is also the easiest of the three to catch. The other two are quieter and considerably more dangerous, because they survive a superficial check.
| Failure mode | Why a quick check misses it | What catches it |
|---|---|---|
| The judgment does not exist | Right shape, real reporter, plausible volume. Nothing on its face is wrong. | Search the primary source by party name and by citation, independently. A fabrication fails at least one route. |
| It exists but does not say it | The citation resolves. The bench and parties are real, so the usual check passes. | Read the paragraph. The proposition is often counsel's recorded submission, an aside, or the dissent. |
| It says it, but is no longer good law | The judgment itself is correct. The defect is in what later benches did to it. | Trace subsequent treatment. An authority is only as current as the last bench that dealt with it. |
Note what all three have in common. Every one is caught by going to the primary source and reading it. None is caught by reading the AI output more carefully. You cannot verify the output from inside the output, however sceptically you read it, because the defect is not in the text you are looking at. It is in the relationship between that text and a document you have not opened.
Why It Fails Under Pressure, and Why That Is No Defence
Producing a fabricated citation takes the model a second. Disproving one takes a search, a portal, a failed lookup, a second search under a different spelling, and a judgement call about whether to keep looking, repeated for every authority in the note. Generation and verification are separated by orders of magnitude, and the whole burden of the imbalance falls on the one person in the chain with a deadline. So the failure clusters predictably. Nobody fabricates a citation into a filing on a quiet Tuesday. It happens at eleven at night before a listing, when the note arrived looking finished and the last authority on the list felt like the safest thing in the world to skip. The output arriving in a finished state is itself part of the problem.
None of which is a defence. An advocate is an officer of the court before being an agent of the client: that is the premise of the Advocates Act 1961 and of the Bar Council of India's rules on standards of professional conduct and etiquette. Nothing in either contemplates a representation made to a bench for which the advocate is not answerable. When you cite an authority you represent that the judgment exists, that it says what you say it says, and that it is good law. That representation was yours when the research came from a junior, and it is yours when it came from a model. The chain of production is your internal arrangement. The court sees a signature. Courts in several jurisdictions have now dealt with fabricated citations in filings, and the pattern has been consistent: the tool has not been accepted as an answer.
What is actually at stake
The immediate exposure is an adverse order and costs, possibly personal. The larger exposure is quieter. Once a bench has found one fabricated citation in your filing, every other authority in it is read differently, and so is your next filing. Credibility with a bench is built slowly and spent all at once. Whether the matter is a Section 138 complaint or a several-crore arbitration, the duty is identical, and so is the damage.
A Verification Workflow That Works With Any Tool

Nothing here depends on which tool you use. It catches all three failure modes, and it is ordered deliberately: each step is cheap, and each is a gate, so a citation that fails at step one costs you almost nothing.
Establish that the document exists, before you engage with the reasoning
Everyone skips this, because the reasoning is the interesting part. Do not evaluate the argument until the judgment is in front of you as a document. Search the primary source twice, independently: once by party name, once by citation. A real judgment converges, and both routes land on the same document. The check costs seconds, and it saves the time you would otherwise spend on reasoning with no judgment underneath it.
Read the paragraph, not the characterisation
The AI's summary is a claim about the text. It is not the text. Open the judgment, find the passage, and ask three things. Is this the court speaking, or counsel's submission as recorded? Is it the ratio, or an observation in passing? Is it the majority, or the dissent? A real judgment attached to a proposition it does not support is more dangerous than an outright fabrication, because it passes step one and keeps passing until someone opposite reads it.
Trace what happened to it afterwards
A correct citation to a dead authority is still a defective filing, and arguably a worse look, because it suggests you opened the judgment and stopped there. Check whether it has been followed, distinguished, doubted, referred to a larger bench, or overruled, and whether the provision it construed still reads the same way. A model trained to a cutoff has no idea what a bench did last month.
Test whether the proposition survives your facts
The step no tool performs for you, and the one that is actually legal work. Retrieval matches on semantic similarity, which is not legal applicability. A judgment can be real, correctly characterised and good law, and still not travel to your matter: the facts may be distinguishable, the bench coordinate rather than binding on the point, the observation per incuriam. The tool finds the authority. Whether it applies is your judgement, and that is not delegable.
The check before it leaves your desk
Before a filing goes to the Registry, for every authority in it:
- You opened the judgment itself. Not a summary, not a headnote, not an AI extract.
- You read the paragraph you rely on, in its surrounding context.
- You know whether the passage is ratio or obiter, and whether it is the majority.
- You checked its subsequent treatment, and know the date to which you checked.
- You can say why it applies to these facts without reference to anything the tool said.
- If the bench asked you to take it to the passage, you could, from your own reading.
The last is the real test, and worth asking while drafting rather than while being asked. If you could not take the bench to the passage from your own reading, you have not verified the citation. You have relayed it.
Signals That Should Slow You Down
None of these proves fabrication, and none substitutes for the workflow. They are triage: which authority in a long note to check first.
Instability under re-asking is diagnostic. Ask the same question three times in fresh sessions. Three different citations for the same proposition means you are watching generation, not retrieval, because a retrieved fact is stable and a generated one is resampled. The logic runs one way only: instability is strong evidence of fabrication, but stability is not evidence of truth.
A citation without a link is a citation the tool did not retrieve. This is the sharpest triage available. If a system can name a judgment but cannot take you to it, then at no point in producing that answer did it have the judgment. It produced the name from its parameters, and everything downstream of that name, including the paragraph and the quoted text, was produced the same way. The link is not a convenience feature. It is the evidence that a document was involved at all.
How Retrieval-Grounded Search Changes the Economics
That is the hinge, and where the architecture of a tool matters more than the quality of its prose. There are two fundamentally different things a system can do with a legal question. It can answer from what it absorbed during training, which is the pattern-completion described above. Or it can search a corpus of actual judgments first, retrieve the documents that are genuinely there, and answer from those, carrying the reference forward. In a chat window these look identical. They are not the same operation, and they fail in entirely different ways.
What grounding fixes, and what it does not
Grounding does not make an AI infallible. A grounded system can still retrieve the wrong judgment, over-read a passage, or miss a distinction that matters. What grounding changes is not whether you verify. It is what verification costs. When every assertion arrives attached to the judgment it came from, checking is a click and a read, not a search, a doubt and a hunt. The claim worth making is that verification is fast, not that it is unnecessary.
| What you are checking | Ungrounded generation | Retrieval-grounded search |
|---|---|---|
| Does the judgment exist? | Unknown from the output. You search the portals yourself. | Settled before the answer was written. The document was retrieved. |
| Does it actually say that? | A claim about a text the system never had in front of it. | A claim about a text you can open and read in place. |
| Is it still good law? | Silent, and bounded by a cutoff you cannot see. | Answerable from tracked treatment. |
| Cost per authority | A search, a doubt and a hunt. Minutes, sometimes tens of them. | One click to the primary source. |
| Where the burden falls | Entirely on you, at the worst possible time. | On the system to show its source, and on you to read it. |
What this looks like in practice
CourtMesh is built on that second architecture. The corpus is roughly 310 million cases: the Supreme Court, all 25 High Courts, the district judiciary and tribunals including the NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and DRT, taken directly from the official government portals with no third-party intermediary. That sourcing is not a technicality. It determines what you are verifying against: the record as the court published it, rather than a version that has passed through somebody else's editorial process.
Search works two ways: keyword, when you know the phrase, the provision or the party; and AI semantic search, when you know the proposition but not the words the bench happened to use, which is the situation that sends most people to an AI tool in the first place. Every result links to the actual judgment. That is the property doing the real work here: not the semantic search, which is convenience, but the link, which is evidence. Where a judgment carries deep analysis, citation relationships are tracked as relationships rather than a flat list of mentions: whether it was followed, distinguished, overruled or referred. That is step three above, the step most commonly skipped, and it gets skipped because doing it by hand is tedious, not because anyone thinks it unimportant.
One limitation, stated plainly, because an article about fabricated claims is a poor place to make loose ones: only a subset of the corpus carries full AI-derived analysis, including the citation relationships described above. Retrieval covers the corpus. Deep analysis covers part of it, and it would be wrong to let the first quietly imply the second. On infrastructure, since client data sits on top of research: CourtMesh is designed to support DPDP Act 2023 readiness, hosted in AWS Mumbai, with AES-256 encryption at rest and TLS 1.3 in transit. Relevant here mainly because the matters you research belong to clients who never consented to their disputes being typed into a general-purpose chatbot.
What to ask any vendor, including this one
- Does every assertion link to a primary source? Not a case name and a summary: the judgment itself, one click away.
- Where does the corpus come from? Directly from the courts, or through an intermediary? Every layer between bench and screen can introduce error, and is one you cannot inspect.
- What is the coverage, precisely? A corpus that stops at the High Courts is a different product from one reaching the District Courts and tribunals.
- How current is it, and how would I know? Ask the lag between pronouncement and availability, and whether the interface shows the date to which it runs.
- Is subsequent treatment tracked? A tool that finds judgments but not what happened to them has automated the easy half and left you the half that gets people into trouble.
- What happens to what I type? Your research queries describe your client's matter.
- What does the vendor say it gets wrong? A vendor with no articulable failure modes has either not looked or is not telling you. The most diagnostic question on the list.
Make This a Habit of the Firm, Not of the Individual
The most common route by which a fabricated citation reaches a bench is not that somebody decided not to check. It is that everybody assumed somebody else had. A junior runs the research through a tool and drafts the note, treating the authority as settled because the tool presented it as settled, on the reasonable assumption that a legal research product returns legal research. The senior reads the note, finds it well drafted and coherent, and assumes the junior pulled the judgments, because pulling the judgments is what doing the research has always meant. The signature goes on. Nobody was careless. Nobody verified. The gap is in the space between two reasonable assumptions.
Name the verifier, in writing
For every authority, one person is responsible for having opened the judgment. Not the team, not the process: a person. Responsibility distributed across a workflow lives nowhere in particular.
Record the check, not just the citation
The note should carry, beside each authority, who opened it, on what date, and to what date treatment was checked. Seconds of work when the source is a click away, and it converts an assumption into a fact.
Keep retrieval and analysis apart
Separate what the tool found from what you concluded. Blended into one fluent prose, a reader cannot tell which parts have a judgment behind them, and nor, a month later, can the author.
None of this is exotic. It is the discipline the profession already applies to limitation, where nobody would rely on a computed date without opening the file, because the consequence of being wrong is total and the check is cheap. Fabricated citations deserve the same reflex, for exactly the same reason.
The Position Worth Holding
The reasonable view is neither of the extremes. It is not that these tools are a menace to be kept out of chambers: they compress genuine hours of genuine drudgery. Nor is it that they can be trusted and the verification problem is a transitional wrinkle the next model will iron out. It will not, because the fabrication is not a bug in the model. It is what the model does, applied to a domain that demands something the model was never built to provide.
Use the tools. Verify the output. Choose tools that make verifying cheap. Those three are compatible, in that order, and the third is what makes the second survive a bad week: a discipline that costs an hour will be honoured on quiet days and quietly abandoned at eleven at night before a listing, while a discipline that costs a click survives contact with the cause list. The citation that did not exist was never really a technology story. It is an old professional question in new clothing: what are you prepared to put your name to, and how do you know?
Verify in a click, not in an hour
CourtMesh is search grounded in the record: roughly 310 million cases across the Supreme Court, all 25 High Courts, the district judiciary and tribunals, taken directly from the official government portals with no intermediary. Keyword and AI semantic search, every result linked to the actual judgment, and citation relationships tracked as followed, distinguished, overruled or referred across the subset that carries deep analysis. We do not claim our AI never errs, and we do not claim the corpus or the citation graph is complete. We claim something more useful: every answer is grounded in a retrieved judgment that links to its source, so checking it takes a click. Verification is still your job. It should just stop taking your evening.
Explore CourtMesh


