A general purpose AI assistant is usually reliable on stable doctrine and structure: what a writ petition is, how the appellate hierarchy works, the difference between cognizable and non-cognizable offences, what limitation does. It is unreliable on citations, section numbers, current statutory text and anything decided recently. The critical property is that it fails silently: a fabricated case citation is formatted exactly like a real one, delivered in exactly the same confident register, in the same paragraph as three correct statements. There is nothing in the output that marks the false part.
So the practical answer to the question in the title is: trust it as a starting point, never as a source. Use it to understand a concept, to structure a research plan, to summarise a document you supplied, or to draft a first pass. Do not use it as the origin of any citation, section number or proposition that will go into a pleading, an opinion or a client email without opening the primary source yourself.
The rest of this piece explains why this happens mechanically, what makes it worse in India specifically, why grounding is the variable that actually matters, and gives you a verification workflow that takes minutes. It is written by a company that sells AI legal research, which is a reason to be sceptical of it and also a reason it is unusually specific about the limits.
Why It Fabricates, in Terms a Lawyer Can Use
A language model generates text by predicting what comes next, one token at a time, based on patterns learned from an enormous quantity of text. It is not looking anything up. When it produces an answer from its own parameters, it is producing text that looks like the kind of text that answers your question.
For explaining a doctrine, that works remarkably well, because the pattern of a correct explanation is closely tied to the substance of a correct explanation. There are many ways to say what a caveat does and they mostly say the same thing.
For a citation, it fails badly, and for a specific reason. A citation is a highly patterned string with almost no semantic content. A party name, a versus, another party name, a year, a reporter abbreviation, a volume, a page. Every element of that pattern is easy to generate and impossible to derive. The model has learned exactly what a plausible Indian citation looks like. It has not, in that moment, learned which ones exist.
A fabricated citation is not a mistake in the ordinary sense. It is the model doing precisely what it does, correctly, in a situation where pattern-shaped output happens to require a fact. Nothing in the process distinguishes the two.
Two consequences follow, and both matter more than the general warning. First, fluency is not evidence of accuracy, and in fact confident fluency is what makes fabrication hard to catch. Second, the failure is not detectable by proofreading. You cannot read a citation carefully enough to determine whether it exists. Only looking it up does that.
The Failure Modes That Are Specific to Indian Law
Every jurisdiction has the general problem. India has several additional ones, and they compound.
The 2023 recodification
The BNS, BNSS and BSA replaced the IPC, CrPC and Evidence Act with effect from 1 July 2024. An assistant may answer with the old numbering when the new applies, or with new numbering when the matter is governed by the old, and both errors read as authoritative. This is currently the single largest source of wrong section numbers in Indian AI output.
Wrong old-to-new mapping
Worse than using the wrong code is mapping between them incorrectly. Asked what Section 154 CrPC corresponds to, a model can produce a confident and wrong BNSS number. The mapping is factual, and factual mappings are exactly what parametric memory is bad at.
Party names repeat endlessly
Indian citations are full of common surnames, State of X as a party, and long running litigation producing several rounds with near identical names. A model assembling a plausible-looking name has a very high chance of assembling one that resembles something real without being it.
Reporter citations are trivially fabricable
A volume number, a reporter abbreviation and a page number carry no meaning that can be checked by reading. They are the easiest thing in law to invent and among the hardest to spot as invented.
High Courts differ
A confident sentence beginning the law in India is that can be true in one High Court and wrong in another, with nothing overruled anywhere. Assistants tend to flatten this into a single national position, which is the shape of answer users reward.
The training cutoff
The most recent development is often the most important one and is exactly the one missing. A judgment delivered after the cutoff, or an amendment brought into force after it, is invisible, and the model has no way to tell you that it is answering about a world that has moved.
What Assistants Are Genuinely Good At
A piece that only listed the failures would be dishonest in the other direction. Used for the right tasks, these tools are extremely useful, and the pattern of what works is consistent.
| Task | How much to trust an ungrounded assistant | Why |
|---|---|---|
| Explaining a legal concept | High, with a check on anything specific | Stable doctrine is heavily represented in the training data and the explanation pattern tracks the substance |
| Structuring a research plan | High | It is generating a plan, not a fact. The output is a set of questions for you to answer, which is exactly the right shape |
| Summarising a document you supply | High, within the document | The text is in front of it. This is comprehension rather than recall, and it is the strongest use case there is |
| Drafting a first pass | Moderate, as a starting point | Structure and language are pattern tasks. Every substantive assertion in the draft still has to be independently established |
| Translating legalese into plain language | High | Again a transformation of text you supplied rather than a retrieval of facts it must remember |
| Naming a case for a proposition | Very low | This is the fabrication zone. Treat every unverified citation as a hypothesis about a case that may not exist |
| Stating a section number | Very low in India right now | The recodification makes this worse than it would otherwise be, in both directions |
| Telling you the current position on a recent question | Very low | Training cutoff, and no awareness of what it does not know |
The Variable That Actually Matters Is Grounding
The reliability question is usually framed as being about how clever the model is. It is not. It is about where the answer came from.
An ungrounded answer is generated from the model's parameters: the compressed statistical residue of its training. Nothing was consulted at the time of answering. There is no document behind the sentence.
A grounded answer is generated after retrieving actual documents from a real corpus, with the answer constructed from what those documents say and linked back to them. The model is still doing the writing, and the substance now has provenance. It can still misread a document, over generalise from it, or miss a better one. What it is much less likely to do is invent a judgment that does not exist, because the judgment it is describing was retrieved rather than recalled.
That difference is why the same underlying model can be dangerous in one product and useful in another, and why the useful question to ask about any legal AI tool is not how good the model is but what corpus it searched and whether it will show you the source.
A link is not proof, and this catches out careful people
Grounded systems attach sources, which is good, and users then stop checking, which is not. A link can be real and irrelevant. A citation can resolve to a genuine judgment that does not say what the answer claims. A paragraph reference can be to the wrong paragraph. Open the link, find the passage, and read the sentences around it. The presence of a citation tells you the system had a document in mind. It tells you nothing about whether the document supports the sentence.
A Verification Workflow You Can Run in Minutes
Separate the answer into claims and explanation
Explanation of a doctrine is low risk. Specific claims, case names, section numbers, time limits and dates, are the high risk part. Mark them. Everything below applies to the marked items only, which is what keeps this fast.
Verify every citation exists, independently
Search for the case in a source that is not the assistant: the court's own judgment search, the e-SCR for reported Supreme Court judgments, or a research platform. A citation you cannot find is a citation that does not exist until proven otherwise. Do not conclude that your search was inadequate.
Open the judgment and find the proposition
Existence is not enough. Locate the actual passage supporting the claim, in the numbered paragraphs. If the proposition is nowhere in the judgment, the citation was attached to a sentence it does not support, which is a distinct and very common failure.
Check the statute as it stands today
Read the provision in the current text, and where the matter is governed by an earlier version, in the version that applied on the relevant date. For criminal matters, establish whether the BNS, BNSS and BSA apply or whether the earlier codes govern, and never take a mapping between them on trust.
Ask the forward question
Has the judgment been overruled, doubted, or referred to a larger bench? A judgment is a finished document that answers only backwards, and no assistant can tell you from the text what happened to it afterwards.
Check the jurisdiction
Is the authority from the Supreme Court, from the High Court where you are appearing, or from a different High Court? A proposition settled in one High Court and rejected in another is not the law in India; it is a conflict, and which side of it you are on depends on where you are appearing.
Check the date against the cutoff
If the question turns on anything recent, assume the assistant does not know about it and go and look. This is the failure mode with no warning signal at all.
Record what you verified and when
In the file: which citations were checked, against what source, on what date. This is what makes the research auditable, and it is what protects you if the position moves between the opinion and the hearing.
The Professional Dimension
There is a version of this discussion that treats fabricated citations as an embarrassing technical quirk. It is not. Filing a fabricated authority in a court is a serious professional matter, and courts in several jurisdictions have taken action against lawyers who have done so. The most widely reported instance is the United States case of Mata v. Avianca (2023), in which counsel filed a brief containing citations generated by an AI tool that did not exist, and were sanctioned.
The principle underneath it is not jurisdiction specific and does not depend on any particular rule. The duty to the court is personal. An advocate who cites an authority is representing to the court that the authority exists and says what it is being cited for. That representation cannot be delegated to a tool, and no software is enrolled under the Advocates Act 1961.
There is a second, quieter risk that deserves naming because it is less discussed than fabrication. A clean, confident, well organised answer stops you looking. A researcher without a tool knows the check is outstanding. A researcher looking at a tidy answer with five citations believes the work is done. Over trusting a research tool fails worse than not having one, because it removes the sense that anything remains to be done.
Where CourtMesh Fits, Stated Honestly
This is where a vendor usually tells you their product solves the problem. It does not, and the useful thing is to be exact about what it changes.
CourtMesh covers roughly 310 million cases sourced only from official government portals, spanning the Supreme Court, all 25 High Courts, the district judiciary and tribunals including the NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and DRT. AI research runs over a semantic index that covers a subset of that corpus rather than all of it, and citation relationships, typed as followed, distinguished, overruled and referred, are likewise available only across a subset. Those subsets are expanding. They are not complete and they will not be.
What retrieval grounding changes is the fabrication risk. An answer constructed from judgments actually retrieved from that corpus, with the judgment linked, is far less likely to describe a case that does not exist. That is a genuine and significant reduction in the most dangerous failure mode.
What it does not change: you still have to open the judgment. Retrieval can surface the wrong document, the answer can over generalise from a document that was retrieved correctly, and coverage gaps are invisible from inside the result. And the most important sentence in this article about any research tool, ours included: a clean result means the tool found nothing. It does not mean there is nothing.
General information, not legal advice
This article discusses the reliability of AI generated legal information. It is not legal advice, and neither is any output from any AI system. The record of the court concerned is the authority on any judgment, the statute as in force is the authority on any provision, and the professional judgement about what applies to a set of facts belongs to a practitioner.
The right question is not how clever, it is grounded in what
Ask any legal AI tool two things: what corpus did you search, and will you show me the judgment. An answer from parametric memory and an answer retrieved from a real corpus look identical on the screen and are not remotely the same thing. CourtMesh searches roughly 310 million cases from official government portals, with semantic retrieval and citation analysis over a subset it does not pretend is the whole. Opening the judgment is still your job, and it should be.
Explore CourtMesh


