Skip to main content
    All articles

    Should You Trust a Chatbot's Answer About Indian Law?

    8 August 202611 min readCourtMesh Team
    Cover card headed Not How Clever, Grounded in What, with the line: ask for the source

    A general purpose AI assistant is usually reliable on stable doctrine and structure: what a writ petition is, how the appellate hierarchy works, the difference between cognizable and non-cognizable offences, what limitation does. It is unreliable on citations, section numbers, current statutory text and anything decided recently. The critical property is that it fails silently: a fabricated case citation is formatted exactly like a real one, delivered in exactly the same confident register, in the same paragraph as three correct statements. There is nothing in the output that marks the false part.

    So the practical answer to the question in the title is: trust it as a starting point, never as a source. Use it to understand a concept, to structure a research plan, to summarise a document you supplied, or to draft a first pass. Do not use it as the origin of any citation, section number or proposition that will go into a pleading, an opinion or a client email without opening the primary source yourself.

    The rest of this piece explains why this happens mechanically, what makes it worse in India specifically, why grounding is the variable that actually matters, and gives you a verification workflow that takes minutes. It is written by a company that sells AI legal research, which is a reason to be sceptical of it and also a reason it is unusually specific about the limits.

    Why It Fabricates, in Terms a Lawyer Can Use

    A language model generates text by predicting what comes next, one token at a time, based on patterns learned from an enormous quantity of text. It is not looking anything up. When it produces an answer from its own parameters, it is producing text that looks like the kind of text that answers your question.

    For explaining a doctrine, that works remarkably well, because the pattern of a correct explanation is closely tied to the substance of a correct explanation. There are many ways to say what a caveat does and they mostly say the same thing.

    For a citation, it fails badly, and for a specific reason. A citation is a highly patterned string with almost no semantic content. A party name, a versus, another party name, a year, a reporter abbreviation, a volume, a page. Every element of that pattern is easy to generate and impossible to derive. The model has learned exactly what a plausible Indian citation looks like. It has not, in that moment, learned which ones exist.

    A fabricated citation is not a mistake in the ordinary sense. It is the model doing precisely what it does, correctly, in a situation where pattern-shaped output happens to require a fact. Nothing in the process distinguishes the two.

    Two consequences follow, and both matter more than the general warning. First, fluency is not evidence of accuracy, and in fact confident fluency is what makes fabrication hard to catch. Second, the failure is not detectable by proofreading. You cannot read a citation carefully enough to determine whether it exists. Only looking it up does that.

    The Failure Modes That Are Specific to Indian Law

    Every jurisdiction has the general problem. India has several additional ones, and they compound.

    The 2023 recodification

    The BNS, BNSS and BSA replaced the IPC, CrPC and Evidence Act with effect from 1 July 2024. An assistant may answer with the old numbering when the new applies, or with new numbering when the matter is governed by the old, and both errors read as authoritative. This is currently the single largest source of wrong section numbers in Indian AI output.

    Wrong old-to-new mapping

    Worse than using the wrong code is mapping between them incorrectly. Asked what Section 154 CrPC corresponds to, a model can produce a confident and wrong BNSS number. The mapping is factual, and factual mappings are exactly what parametric memory is bad at.

    Party names repeat endlessly

    Indian citations are full of common surnames, State of X as a party, and long running litigation producing several rounds with near identical names. A model assembling a plausible-looking name has a very high chance of assembling one that resembles something real without being it.

    Reporter citations are trivially fabricable

    A volume number, a reporter abbreviation and a page number carry no meaning that can be checked by reading. They are the easiest thing in law to invent and among the hardest to spot as invented.

    High Courts differ

    A confident sentence beginning the law in India is that can be true in one High Court and wrong in another, with nothing overruled anywhere. Assistants tend to flatten this into a single national position, which is the shape of answer users reward.

    The training cutoff

    The most recent development is often the most important one and is exactly the one missing. A judgment delivered after the cutoff, or an amendment brought into force after it, is invisible, and the model has no way to tell you that it is answering about a world that has moved.

    What Assistants Are Genuinely Good At

    A piece that only listed the failures would be dishonest in the other direction. Used for the right tasks, these tools are extremely useful, and the pattern of what works is consistent.

    TaskHow much to trust an ungrounded assistantWhy
    Explaining a legal conceptHigh, with a check on anything specificStable doctrine is heavily represented in the training data and the explanation pattern tracks the substance
    Structuring a research planHighIt is generating a plan, not a fact. The output is a set of questions for you to answer, which is exactly the right shape
    Summarising a document you supplyHigh, within the documentThe text is in front of it. This is comprehension rather than recall, and it is the strongest use case there is
    Drafting a first passModerate, as a starting pointStructure and language are pattern tasks. Every substantive assertion in the draft still has to be independently established
    Translating legalese into plain languageHighAgain a transformation of text you supplied rather than a retrieval of facts it must remember
    Naming a case for a propositionVery lowThis is the fabrication zone. Treat every unverified citation as a hypothesis about a case that may not exist
    Stating a section numberVery low in India right nowThe recodification makes this worse than it would otherwise be, in both directions
    Telling you the current position on a recent questionVery lowTraining cutoff, and no awareness of what it does not know

    The Variable That Actually Matters Is Grounding

    The reliability question is usually framed as being about how clever the model is. It is not. It is about where the answer came from.

    An ungrounded answer is generated from the model's parameters: the compressed statistical residue of its training. Nothing was consulted at the time of answering. There is no document behind the sentence.

    A grounded answer is generated after retrieving actual documents from a real corpus, with the answer constructed from what those documents say and linked back to them. The model is still doing the writing, and the substance now has provenance. It can still misread a document, over generalise from it, or miss a better one. What it is much less likely to do is invent a judgment that does not exist, because the judgment it is describing was retrieved rather than recalled.

    That difference is why the same underlying model can be dangerous in one product and useful in another, and why the useful question to ask about any legal AI tool is not how good the model is but what corpus it searched and whether it will show you the source.

    A link is not proof, and this catches out careful people

    Grounded systems attach sources, which is good, and users then stop checking, which is not. A link can be real and irrelevant. A citation can resolve to a genuine judgment that does not say what the answer claims. A paragraph reference can be to the wrong paragraph. Open the link, find the passage, and read the sentences around it. The presence of a citation tells you the system had a document in mind. It tells you nothing about whether the document supports the sentence.

    A Verification Workflow You Can Run in Minutes

    1

    Separate the answer into claims and explanation

    Explanation of a doctrine is low risk. Specific claims, case names, section numbers, time limits and dates, are the high risk part. Mark them. Everything below applies to the marked items only, which is what keeps this fast.

    2

    Verify every citation exists, independently

    Search for the case in a source that is not the assistant: the court's own judgment search, the e-SCR for reported Supreme Court judgments, or a research platform. A citation you cannot find is a citation that does not exist until proven otherwise. Do not conclude that your search was inadequate.

    3

    Open the judgment and find the proposition

    Existence is not enough. Locate the actual passage supporting the claim, in the numbered paragraphs. If the proposition is nowhere in the judgment, the citation was attached to a sentence it does not support, which is a distinct and very common failure.

    4

    Check the statute as it stands today

    Read the provision in the current text, and where the matter is governed by an earlier version, in the version that applied on the relevant date. For criminal matters, establish whether the BNS, BNSS and BSA apply or whether the earlier codes govern, and never take a mapping between them on trust.

    5

    Ask the forward question

    Has the judgment been overruled, doubted, or referred to a larger bench? A judgment is a finished document that answers only backwards, and no assistant can tell you from the text what happened to it afterwards.

    6

    Check the jurisdiction

    Is the authority from the Supreme Court, from the High Court where you are appearing, or from a different High Court? A proposition settled in one High Court and rejected in another is not the law in India; it is a conflict, and which side of it you are on depends on where you are appearing.

    7

    Check the date against the cutoff

    If the question turns on anything recent, assume the assistant does not know about it and go and look. This is the failure mode with no warning signal at all.

    8

    Record what you verified and when

    In the file: which citations were checked, against what source, on what date. This is what makes the research auditable, and it is what protects you if the position moves between the opinion and the hearing.

    The Professional Dimension

    There is a version of this discussion that treats fabricated citations as an embarrassing technical quirk. It is not. Filing a fabricated authority in a court is a serious professional matter, and courts in several jurisdictions have taken action against lawyers who have done so. The most widely reported instance is the United States case of Mata v. Avianca (2023), in which counsel filed a brief containing citations generated by an AI tool that did not exist, and were sanctioned.

    The principle underneath it is not jurisdiction specific and does not depend on any particular rule. The duty to the court is personal. An advocate who cites an authority is representing to the court that the authority exists and says what it is being cited for. That representation cannot be delegated to a tool, and no software is enrolled under the Advocates Act 1961.

    There is a second, quieter risk that deserves naming because it is less discussed than fabrication. A clean, confident, well organised answer stops you looking. A researcher without a tool knows the check is outstanding. A researcher looking at a tidy answer with five citations believes the work is done. Over trusting a research tool fails worse than not having one, because it removes the sense that anything remains to be done.

    Citing a case that does not exist, because the citation looked exactly like a real one
    Citing a real case that does not say what it was cited for
    Using an IPC or CrPC section number where the BNS or BNSS now governs, or the reverse
    Relying on a mapping between the old and new codes that the assistant generated rather than looked up
    Treating a proposition as national law when it is the position in one High Court and contested in another
    Missing a development after the training cutoff, with no signal that anything was missing
    Believing a clean result means there is nothing to find, when it means the tool found nothing

    Where CourtMesh Fits, Stated Honestly

    This is where a vendor usually tells you their product solves the problem. It does not, and the useful thing is to be exact about what it changes.

    CourtMesh covers roughly 310 million cases sourced only from official government portals, spanning the Supreme Court, all 25 High Courts, the district judiciary and tribunals including the NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and DRT. AI research runs over a semantic index that covers a subset of that corpus rather than all of it, and citation relationships, typed as followed, distinguished, overruled and referred, are likewise available only across a subset. Those subsets are expanding. They are not complete and they will not be.

    What retrieval grounding changes is the fabrication risk. An answer constructed from judgments actually retrieved from that corpus, with the judgment linked, is far less likely to describe a case that does not exist. That is a genuine and significant reduction in the most dangerous failure mode.

    What it does not change: you still have to open the judgment. Retrieval can surface the wrong document, the answer can over generalise from a document that was retrieved correctly, and coverage gaps are invisible from inside the result. And the most important sentence in this article about any research tool, ours included: a clean result means the tool found nothing. It does not mean there is nothing.

    General information, not legal advice

    This article discusses the reliability of AI generated legal information. It is not legal advice, and neither is any output from any AI system. The record of the court concerned is the authority on any judgment, the statute as in force is the authority on any provision, and the professional judgement about what applies to a set of facts belongs to a practitioner.

    The right question is not how clever, it is grounded in what

    Ask any legal AI tool two things: what corpus did you search, and will you show me the judgment. An answer from parametric memory and an answer retrieved from a real corpus look identical on the screen and are not remotely the same thing. CourtMesh searches roughly 310 million cases from official government portals, with semantic retrieval and citation analysis over a subset it does not pretend is the whole. Opening the judgment is still your job, and it should be.

    Explore CourtMesh
    AIChatbotsLegal InformationHallucinationTrust
    X LinkedIn