Skip to main content
    All articles

    Open Court Data: What India Publishes, and What It Should

    7 August 202610 min readCourtMesh Team
    Cover card headed Published but Not Machine Readable, with the line: available is not open

    India is, by the standards of most legal systems, unusually generous with judicial information. Judgments of the Supreme Court and the High Courts are published free of charge on official portals. District court orders are published through a national judgment search. Cause lists are available daily. National pendency statistics are open to anybody with a browser. The Supreme Court's own reports are freely accessible, and the Court has adopted neutral citation for its judgments.

    It is also true that almost none of this is published in a form designed to be used at scale. The material is available and it is not accessible in the sense that matters for research, for accountability or for building anything. That distinction, between publication and usability, is the subject of this article, and closing it is probably the highest-return transparency reform available to the Indian judiciary per rupee spent.

    What India Actually Publishes Today

    LayerWhat is publicFormPractical usability
    Supreme CourtJudgments and orders, the Court's own reports, case status, cause lists, and aggregate statisticsDocuments on official portals, with neutral citation adopted for judgmentsGood by Indian standards. Free, comprehensive at the judgment level, and increasingly well identified.
    High CourtsJudgments and orders through a consolidated national judgment search and through each court's own site, plus cause lists and case statusDocuments, with metadata that varies by courtReasonable for lookup, poor for systematic work. Twenty-five separate publication practices and no common metadata standard.
    District judiciaryJudgments and orders where uploaded, case status, cause lists, and national statisticsDocuments of highly variable quality, with hand-entered metadataWeakest layer. Publication is uneven, formats vary, and party name and act tagging are inconsistent.
    TribunalsOrders on each tribunal's own portalBench-wise, with each tribunal's own numbering and search grammarFragmented. Each portal is its own island with no shared identifiers.
    StatisticsNational pendency, disposal and institution data through the judicial data grid, with an access route for institutional usersDashboards, with defined access arrangements beyond themGenuinely good for system-level questions. Not a substitute for case-level data.
    Revenue and other quasi-judicial forumsVery little, and what exists is State-specificState portals where they exist at allThe largest gap in the entire picture, and the one that affects the most common category of dispute.

    The pattern is consistent and it is not accidental. Publication is best where the output is citable precedent and worst where it is volume adjudication, which is precisely the inverse of where the litigation actually is.

    Available Is Not the Same as Accessible

    Here is the distinction that matters. A judgment published as a document on a website is available to anybody who knows it exists and can find it. It is not accessible in the sense that a research question, an accountability question or a public interest question requires.

    No bulk access

    Court data is published for individual retrieval. There is no general facility for obtaining a corpus, which means any systematic study of a court's output begins with an infrastructure problem rather than a research question.

    No common metadata standard

    Party names, case types, act and section tags, judge names and disposal reasons follow local conventions. Nothing joins across courts, so reconciling a party or a statute nationally requires normalisation work that everybody redoes independently.

    Documents, not data

    A scanned or unstructured document contains the reasoning and none of the structure. The holding, the provisions applied, the outcome and the citation relationships are all in the text and none of them are fields.

    Access friction by design

    Public search interfaces are built for a person looking up one matter, with the controls appropriate to that use. That is a defensible design for the purpose, and it is why usability for research is poor even where publication is complete.

    Unclear reuse terms

    Whether and on what terms published judicial material may be reused is not stated consistently. Judgments themselves are not the kind of material anybody should be claiming proprietary rights over, and the absence of an explicit position creates avoidable uncertainty.

    No stable identifiers everywhere

    Neutral citation is a genuine advance and its adoption is uneven across courts. Without a stable identifier, the same judgment is referred to differently in every system that holds it.

    The gap is not publication, it is structure

    India publishes an enormous quantity of judicial material and publishes almost none of it as data. A judgment released as an unstructured document with inconsistent metadata has been made public without being made usable, and the difference decides whether anybody can answer questions about how courts actually behave. Fixing this requires standards and stable identifiers rather than new portals.

    The Anonymisation Tension, Taken Seriously

    Any argument for more open court data has to engage with the counter-argument properly, because it is a real one. Judgments contain personal information: names, addresses, financial details, medical history, allegations of crime, family circumstances. Making that material easily searchable at national scale has consequences for individuals that did not exist when a judgment sat in a physical volume in a court library.

    Indian law already draws several firm lines, and they should be stated because they are frequently forgotten in this debate.

    • Victims of sexual offences. The identity of a victim of specified sexual offences may not be disclosed, and the prohibition applies to publication in any form. Courts anonymise accordingly.
    • Children. The juvenile justice legislation prohibits disclosure of the name, address, school or any particular that may lead to the identification of a child in conflict with law or in need of care and protection. The child protection legislation contains parallel restrictions.
    • In camera proceedings. Family and matrimonial matters may be heard in camera, which limits what forms part of the public record at all.
    • Sealed material. Courts direct sealing of specific material in appropriate cases, and that direction binds anybody handling the record.
    • Data protection. The Digital Personal Data Protection Act, 2023 brings a general framework for personal data with exemptions relevant to legal proceedings and the enforcement of legal rights. How that framework interacts with published judicial records is still being worked out, and it would be dishonest to state the position more firmly than that.

    Beyond these statutory lines sits the harder question, which is whether an individual has any claim to have their name removed from a published judgment once the matter is over. Indian High Courts have not spoken with one voice on it. Some have granted interim relief directing masking of a petitioner's name in a judgment where continued searchability was causing demonstrable harm. Others have declined, holding that judicial records are public and that a general right to erasure from them is not recognised, while accepting that redaction is appropriate in the categories where statute or established practice requires it, notably matrimonial matters and offences where identity is protected.

    The unresolved state of that question is itself an argument for a policy rather than a series of individual orders. What is needed is a rule about what is published, applied uniformly, rather than an outcome that depends on which High Court a person happens to approach.

    Anonymisation is not the enemy of open court data. Inconsistent anonymisation is, because a rule applied unevenly protects nobody reliably and obstructs everybody unpredictably.

    What Should Be Published, and How

    Here is the constructive part. A serious open court data policy would not require new institutions or large expenditure. It would require decisions.

    1

    A stable identifier for every judgment, everywhere

    Neutral citation across all courts, applied consistently, so that a decision has one name that every system can use. This is the single highest-value item on the list and the Supreme Court has already shown it can be done.

    2

    A common metadata schema

    An agreed set of fields for every published order: court, bench, judges, date, case type, parties, acts and sections in issue, disposition. Agreement on the field names matters more than the sophistication of the schema.

    3

    Machine-readable text as the default output

    Judgments produced as text rather than as images, with searchable structure. This is already the practice in the higher judiciary and is uneven below.

    4

    A clear and permissive reuse position

    An explicit statement that published judicial material may be reproduced and reused, with the conditions that apply. Ambiguity here helps nobody and deters exactly the public interest work that open data is meant to enable.

    5

    A published anonymisation standard

    A uniform rule about which categories are anonymised and how, applied at the point of publication by the court rather than left to downstream handlers to guess at.

    6

    A route to bulk access for legitimate use

    A defined mechanism under which researchers, institutions and public interest organisations can obtain corpora, on terms, rather than an environment in which the only options are individual retrieval or nothing.

    Notice that none of these is a technology project. They are standards decisions, and the institutions that would have to make them, the e-Committee, the High Courts and the National Informatics Centre, already exist and already work together.

    Why This Is the Highest-Return Reform Available

    Compare open court data with the other reforms on the table. Filling judicial vacancies requires appointments and posts. Building courtrooms requires capital expenditure. Adding case management capacity requires judicial time that does not exist. All are necessary and all are slow and expensive.

    Standardising identifiers and metadata on material that is already being published costs comparatively little and unlocks a great deal. It makes empirical study of the judiciary possible, which is the precondition for evidence-based reform of everything else. It makes inconsistency detectable, which is a precondition for addressing it. It lets the judiciary see its own behaviour in ways that aggregate dashboards cannot show. And it makes it possible to answer the questions this whole subject keeps stumbling on: whether special courts work, whether commercial court timelines hold, whether comparable matters get comparable outcomes.

    The honest disclosure of our own interest

    We build a legal search product. Better open court data makes our work easier and our product better, and it would be dishonest to argue for it without saying so. The argument stands independently of that interest, and the test is straightforward: every reform on the list above benefits researchers, journalists, litigants, courts and competing products equally. If we were arguing for privileged access rather than for standards, that would be a different and much less defensible article.

    Where Public and Private Interests Actually Align

    There is a lazy framing in which open data pits public transparency against private legal publishers who profit from friction. In India that framing gets the incentives wrong.

    The historical business of Indian legal publishing was curation: selecting judgments worth reporting, writing headnotes, checking text against the record and assigning citations. That work was valuable because raw judgments were hard to obtain. Now they are not. What is scarce today is not the judgment but the ability to find the right one among hundreds of millions, to relate it to others, and to know whether it still stands. Those are retrieval and analysis problems, and they get easier, not harder, when the underlying corpus is well identified and consistently structured.

    So a serious open data policy does not threaten the products built on this material. It raises the floor for all of them and shifts competition to where it belongs, which is quality of retrieval rather than privileged access to documents that should never have been hard to get.

    That is the position we build from. CourtMesh sources everything from official government portals, with no third-party intermediaries, and indexes the Supreme Court, all twenty-five High Courts, the district judiciary and tribunals including NCLT, NCLAT, ITAT and CESTAT, with roughly 310 million records keyword-searchable and roughly 2 million carrying deeper semantic indexing. Every improvement in how courts publish is an improvement we inherit directly, and so does everybody else, which is exactly the point.

    Treating availability of a judgment on a portal as evidence that the material is usable for research
    Assuming judicial material carries a clear reuse permission when the position is not stated
    Publishing or republishing material that identifies a person whose identity is statutorily protected
    Relying on party name matching across courts without normalisation, when no common identifier exists
    Reading an absence from a portal as an absence of a proceeding
    Arguing for open court data without engaging with the anonymisation questions, which discredits the argument

    The fair summary is that India has done the expensive part. It built the systems, it publishes at national scale, and it did so while the courts were running. What remains is a set of decisions about identifiers, fields, formats and a uniform anonymisation rule, none of which requires new money and all of which would transform what anybody can learn about how Indian justice actually works. Reforms that cheap are rare enough to be worth naming.

    Search the published record as one corpus

    Indian courts publish an extraordinary amount of material across twenty-five High Courts, the district judiciary, the Supreme Court and a dozen tribunals, in a dozen different grammars. CourtMesh puts all of it behind one search, sourced only from official government portals, with filters for court, judge, year, case type and date range. Ask the question once, see where the answer came from, and verify it against the issuing court's own record.

    Explore CourtMesh
    Open DataTransparencyJudiciaryPublic AccessPolicy
    X LinkedIn