Legal Research Tools for Judicial Officers | CourtMesh
    Skip to main content
    All articles

    Jurisprudential Consistency: The Research Problem Nobody Builds For

    16 July 202617 min readCourtMesh Team
    CourtMesh cover card headed "Jurisprudential Consistency" and "The problem nobody builds for", with the line "research tooling is built for advocates, not for the bench"

    Almost every legal research product in India is built around a single question: find me authority that supports my client's position. It is a good question, and the tools that answer it answer it well. It is also not the question a judge is asking, and the difference is not one of degree. It is a different problem, with a different success criterion and a different failure mode, and almost nobody is building for it.

    Every Research Tool Is Built for the Advocate

    Legal research, as a product category, has one archetypal user in mind: the advocate preparing a matter. Every design decision follows. Results are ranked by relevance to a query, and the query encodes a position. Headnotes exist so you can assess whether a judgment helps you. Citators exist so you can confirm your authority is good before you stake an argument on it. Similar-case discovery means cases that would help you in the same way.

    None of this is a criticism: the bench is a far smaller population than the bar, and building for the larger group is rational. But the design has a consequence that is easy to miss. A tool optimised to find the best case for a proposition is, by construction, optimised to stop looking once it has found one. Ranking is a filter. It promotes the strongest supporting material and pushes the rest down the page, and what it pushes down is not marked as suppressed. It simply is not there.

    The bench inherits these tools unchanged. Same search box, same ranking, same citator. And it is asked to do a completely different job with them.

    The Judge's Problem Is the Inverse

    An advocate's research succeeds when it finds the best authority for a position. A judge's research succeeds when nothing relevant has been missed. Those sound like one activity described from two angles. They are two tasks, and they pull in opposite directions.

    Start with the failure modes, because that is where the difference is sharpest. The advocate's failure mode is being outgunned: opposing counsel produces a stronger line and you had not seen it. It costs the client. But the system is built to absorb that failure, because catching it is precisely what the other side is for.

    The judge's failure mode has no such mechanism. It is deciding a question in a way that does not sit with the settled position, because the settled position was never in front of the bench. Nobody in the room has the job of noticing. The decision goes into the reports and becomes part of the law, and the gap that produced it is invisible in the judgment, because a judgment can only reflect what was before the court.

    There is a second, subtler inversion, and it is the one that makes conventional tools genuinely ill-suited rather than merely imperfect. The material a judge most needs is the material that cuts against the view they are forming. That is what tests the view. But relevance ranking computes relevance against the query, and the query encodes where you already are. The system is structurally most likely to bury the exact thing you most needed to see. It is not being unhelpful. It is being helpful at the wrong job.

    This is not a soft preference for even-handedness. Under Article 141 of the Constitution, the law declared by the Supreme Court is binding on all courts within the territory of India: consistency with binding authority is an obligation, not an aspiration. And the doctrine of per incuriam exists precisely because the failure has been recognised and named, a decision rendered in ignorance of binding authority or a statutory provision carrying less weight for it. But notice what that doctrine does. It assigns a consequence after the fact. It offers the bench nothing to help avoid the situation in the first place. The doctrine is downstream. The problem is upstream, and it is retrieval.

    DimensionThe advocate's research problemThe judge's research problem
    GoalFind the strongest authority supporting a position already chosen.Establish what the settled position actually is, before a view has hardened.
    Success criterionThe best available case for the client's argument has been found.Nothing material has been missed, including what neither side thought to cite.
    Failure modeBeing outgunned by authority the other side found first. Caught in the room.Deciding without the field in view. Nobody in the room is assigned to catch it.
    What completeness meansEnough to argue the point persuasively. Sharp diminishing returns past the strongest line.The whole line of authority on the question, including the parts that cut the other way.
    Contrary authorityDistinguish it, or find a reason the matter turns on something else.Must be confronted on the merits, whether or not either side raised it.
    Who checks the workOpposing counsel immediately, then the bench.Appellate scrutiny, eventually, in the fraction of matters carried up.
    The binding constraintThe brief, and what the client will pay for.The cause list.

    The line this article does not cross

    Nothing here suggests that software should decide anything, recommend an outcome, predict a result, or reason toward a conclusion. It should not, and a tool that offered to would be worth refusing. The whole claim in this article is about retrieval: finding the relevant material faster, so that more of the field is in front of the bench when the bench decides. What that material means, what weight it carries, whether it binds and what follows from it are judicial questions, and they are the judge's alone. A retrieval tool has finished its entire job the moment it hands over what it found. Judicial independence is not a constraint this kind of tooling has to work around. It is the reason the tooling must stay firmly on its own side of the line.

    What Was Never Put Before the Bench

    The adversarial system rests on a design assumption, and it is a good one: that between them, two motivated and competent sides will place the relevant field before the court. Each searches exhaustively for what helps it. The union of two exhaustive searches approximates the whole. Most of the time this works well enough. But it has a structural gap, and the gap is not bad luck in a particular matter. It is built into the shape of the thing.

    Aligned omission

    The authority that hurts the appellant may also, for entirely different reasons, hurt the respondent. Both sides have the same incentive and it points the same way. Nobody is assigned to bring it, so it is not brought.

    Asymmetric preparation

    The union of two searches only approximates the field when both are thorough. A large team on one side and a sole practitioner on the other produces one thorough search and one thin one. The party-in-person produces one.

    The unanticipated point

    Sometimes the question that decides the matter is not the one either side prepared for. It emerges during argument. Neither party has researched it, because neither knew it was coming.

    So the bench sees a curated subset of the field and has no systematic way of knowing what the subset leaves out. Not because anyone has misled the court, but because completeness was nobody's assignment. The court has real mechanisms here: it can call for further assistance, appoint an amicus, direct counsel to address a specific authority, reserve the matter and look for itself. All exist and all are used. But every one of them requires the bench to already suspect that something is missing. The gap is invisible from inside it, and that is the whole difficulty.

    The court can direct counsel to address any authority it wishes. It cannot direct counsel to address an authority that nobody has mentioned.

    The Hardest Precedent to Recall Is Your Own

    A judicial officer with a long tenure has decided thousands of matters. Over years. On a board that rotates. Across subject matter that changes with each assignment. Somewhere in that body of work is a narrow question of law that arose four years ago, in a matter mostly about something else, argued for ten minutes and disposed of in three paragraphs. Those three paragraphs are the officer's own reasoning on the point, and in a real sense the officer's position on it.

    When the question comes back, only one outcome is easy: the officer remembers deciding it, remembers what they held, and can find it. Otherwise they may recall the conclusion but not the reasoning, or the reasoning but not which matter it was in, which makes it effectively unfindable. Or, the difficult case, the question arrives dressed differently: different counsel, different vocabulary, a matter about an entirely different subject. It does not register as the same question at all. There is no moment of deciding to check, because the question never announces itself.

    This is not a failure of memory or of care. It is what a large docket does to anybody. The relevant observation is not about the officer. It is that the tools are not helping.

    Why searching your own judgments does not work

    Consider what the officer would have to do to check. Search their own prior judgments, by what? Not the name of the matter, because that is precisely what has been forgotten. Not the party, for the same reason. The only handle is the point of law itself. And the point of law, as recorded in the earlier judgment, is expressed in the vocabulary of the matter it arose in: that statute, those facts, that counsel's framing. Which is not the vocabulary of the matter now on the board.

    So keyword search over one's own judgments fails in exactly the case where it is most needed. It works when you already remember enough to find the thing, which is when you did not need it. It fails when the point is the same but nothing on the surface says so, which is the only case that matters. The tool is reliable in precise proportion to how little you need it.

    And the cost is real. Two decisions from the same officer, each defensible on its own facts and reasoning, reaching different conclusions on the same narrow question, are corrosive in a way neither is on its own. The litigant who lost cannot be told the position is settled, because visibly it is not. The advocate who advised on the strength of the earlier decision looks unreliable, and was not. Nobody wanted this, and no individual decision caused it. It is an artefact of the earlier judgment not being findable at the moment it was needed.

    Coordinate Benches, Parallel Lines

    How coordinate benches diverge on the same question over time

    Now take the same problem and scale it from one officer to a whole court. Two benches of the same High Court, of the same strength, decide the same question and take different views. Neither is bound by the other. A coordinate bench decision carries persuasive weight, and judicial discipline expects that a bench disagreeing with a coordinate bench should refer the question to a larger bench rather than take the contrary view and leave two conflicting decisions on the books. That convention is sound and it works. But look at what it presupposes: it only operates once the bench knows the coordinate decision exists.

    If it does not, there is no disagreement to manage. There are just two lines of authority running in parallel, each internally consistent, each perfectly coherent read from the inside, neither aware of the other.

    And here is the part that is easy to underestimate: divergence is self-reinforcing. Once two lines exist, each is cited by the advocates it benefits, which means each is put before benches by people motivated to present it as the position. Each attracts subsequent decisions that follow it. Each accumulates its own vocabulary and its own body of practitioners who have never had occasion to encounter the other line. After a few years there are two settled positions. In the same court. On the same question.

    This can persist for a long time. It surfaces when somebody happens to notice: an advocate arguing one line appears before a bench familiar with the other, or a judgment finally does the survey nobody had done. The question then goes to a larger bench and is resolved, and that resolution is the system working exactly as designed. But everything decided in the interval was decided under one of two incompatible positions, and which one you got depended on the board your matter was listed before. That is a real cost, borne by litigants who had no way to know.

    What catches this earlier is not better judging. The judging was never the problem. What catches it is a view of the whole line of authority on the question, rather than the fraction the parties chose to put in the paper book. That is a retrieval problem. It has always been a retrieval problem, and it is the specific kind that is close to impossible by hand and quite tractable with the right index. The costs of the gap are specific, and they compound:

    A binding authority not cited by either side, and so never considered
    Two decisions of the same officer on the same narrow question, reaching different conclusions
    Coordinate benches building parallel lines of authority, neither aware of the other
    A decision open to challenge as per incuriam for reasons that had nothing to do with its reasoning
    Divergence that persists for years because no single matter ever put the whole line before a bench
    Litigants getting different answers to the same question depending on the board they were listed before

    A Throughput Problem, Not a Competence Problem

    It would be easy to read all of this as a complaint about how judicial officers work. It is not, and the misreading is worth heading off. The officer deciding a matter is very often the person in the room who knows the area best, having seen the question arise in far more contexts than either advocate arguing it. Whatever the constraint is, it is not knowledge and it is not care.

    The constraint is time, and specifically the shape of the time. A board of matters. Each gets what it gets. Most are routine, some are heavy, and the heavy ones are not reliably the ones that looked heavy when the board was published. The question that turns out to need a full survey of authority is often not flagged in advance at all: it emerges during argument, in a matter listed for something else, at a point when the survey would have to happen now.

    Exhaustive manual research on every question that might need it is not something anybody can do. Not for want of trying or of ability. Arithmetic. If a proper survey takes three hours and the board does not have three hours to give a matter, the survey does not happen. It is replaced by what counsel cited, plus what the officer remembers, plus whatever a quick search turns up. That is a rational allocation of a scarce resource and any sensible person would make it. It is also, precisely, where the gaps come from.

    This is the exact class of problem tooling is supposed to absorb. Not the judging. The looking. If the survey that took three hours takes twenty minutes, it starts happening in matters where it currently cannot, and the bench decides with more of the field in view. Nothing about the deciding has changed. The bench has simply stopped having to choose between thoroughness and the cause list, which was never a choice anybody should have had to make.

    The constraint on the bench has never been judgement. It is the hours it takes to find the material that the judgement is exercised upon.

    What Retrieval Built for the Bench Would Do

    Retrieving your own prior judgments and lines of authority

    If you took the judicial research problem seriously as a distinct problem rather than as a variant of the advocate's, a few things would follow. None are exotic. Most are simply the consequence of asking a different question of the same corpus.

    1

    Start from the issue, not the vocabulary

    The real question is whether this fact pattern has been decided, not which judgments contain these particular words. Semantic search over a description of the issue retrieves conceptually relevant judgments regardless of the terminology a given bench or draftsman used. That matters most where keyword search collapses: when the same question is framed in language you would never have thought to search for.

    2

    Ask what you yourself have held

    Filter by judge, and a question that was unanswerable becomes a lookup: have I decided this, and what did I say. This is retrieval, not inference. It is the cheapest check on the list, and without a judge filter across a full corpus it is close to impossible to perform at all.

    3

    Widen to the coordinate benches

    The same question, the same court, other benches, across the years. Divergence becomes visible only if you are looking at the line of authority rather than at the citations in the paper book.

    4

    Trace the line up and down

    Where citation relationships have been derived, a judgment can be followed through what followed it, what distinguished it, what referred it and what overruled it. That traversal gives you the shape of the line. Note the conditional in that first clause: it matters, and the next section is why.

    5

    Then read the judgment

    Every step above produces candidates. Not one produces a holding. The judgment is read at source, in full, and it is read by the judge. There is no version of this in which a summary substitutes for the text.

    Where This Falls Down, Stated Plainly

    A tool aimed at the bench that oversells itself is worse than no tool, because the entire proposition is that what comes back can be relied on. So here are the limits, stated as plainly as we know how.

    The corpus is broad. The analysis is not as broad.

    CourtMesh indexes roughly 310 million cases: the Supreme Court, all 25 High Courts, the district judiciary and tribunals including the NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and DRT, sourced only from official government portals. That is the corpus. It is not the number of judgments carrying full AI-derived analysis, which is a subset: roughly a couple of million judgments, not 310 million. The citation relationships described above exist over that subset.

    Absence of a flag is not evidence of anything

    This is the easiest thing on this page to get wrong. If a judgment carries no overruled marker against it, that does not mean it has not been overruled. It may simply sit outside the analysed subset. The citation graph covers that subset and not the full corpus, so a clean record is the absence of a finding and never a finding of absence. Treat every flag as a lead worth following, and never treat the lack of one as a clearance. That distinction is the whole difference between a tool that helps and a tool that misleads.

    AI output can be fluent and wrong

    CourtMesh AI can produce output that is confidently phrased and incorrect. Fluency is not accuracy, and the two come apart in a way that reading human writing does not prepare anybody for: with a person, confident phrasing usually carries some information about how sure they are. With a model it should not be read as carrying anything at all. Any AI-surfaced material must be read at source before it is relied on. A mischaracterised authority in a written submission is a bad submission. A mischaracterised authority in a judgment is a different order of problem entirely.

    Better retrieval is not completeness

    A good index raises the floor. It does not guarantee that nothing has been missed, and no system can, ours included. What changes is the probability, and the price in hours the improvement costs. That is worthwhile. It is not the same thing as a solved problem, and we would rather say so.

    What happens to a search query

    For a judicial user this is not an idle question, and it is worth spelling out why. The queries an officer runs while a matter is reserved are a live record of what that officer is thinking about a pending matter. That is about as sensitive as a search log gets. CourtMesh is designed to support DPDP Act 2023 readiness, hosted in AWS Mumbai, with AES-256 encryption at rest and TLS 1.3 in transit, and data is not used to train AI models. Where the AI features are involved, a query is processed in order to answer it, and our AI use terms set out what that processing involves. Anyone for whom this matters should read those terms rather than take a summary of them on trust, here or anywhere else.

    The Tooling Failed the User, Not the Other Way Round

    The argument here reduces to one observation. The judiciary has a research problem that is structurally different from the advocate's, harder in the dimension that matters most, and served for decades by tools designed for somebody else and handed over unmodified. That is a failure of the tooling, not of the bench, and framing it the other way round gets both the diagnosis and the remedy wrong. The bench has been doing a harder job with instruments built for an easier one, under time constraints that made the difference impossible to close by effort. Effort was never going to close it. The gap is arithmetic.

    What closes it, partially and honestly, is retrieval designed for the question actually being asked: what is the settled position, what did I hold last time, what have my colleagues held, and what is in this field that nobody has put in front of me. Those are answerable questions. Nobody was asking them, because nobody was building for the person who asks them.

    If This Is a Problem You Recognise

    CourtMesh indexes roughly 310 million cases from the Supreme Court, all 25 High Courts, the district judiciary and tribunals, drawn only from official government portals. You can search by issue rather than keyword, filter by judge, court, case type, year and date range, and trace citation relationships across the subset where they have been derived. It is a retrieval tool and nothing more: it finds material and hands it over, and everything that matters after that point is yours alone. If the problem described here is one you recognise, the corpus is there to be looked at, and we would rather it be judged on whether it finds what you need than on anything we have said about it.

    Explore CourtMesh
    JudiciaryJudicial OfficersConsistencyLegal ResearchIndia
    X LinkedIn