Skip to main content
    All articles

    How to Choose a Legal Research API in India

    5 June 202617 min readCourtMesh Team
    Cover card headed No Coverage Claim Survives Your Queries, with the line: ask where it thins

    You are about to make a decision that will be extremely expensive to reverse. Whichever legal research API you pick, within six months your product will have its identifiers, its field names, its silences and its rate limit baked into your schema, your cache layer and your users' expectations. Migration is not a sprint. It is a quarter. So the evaluation deserves considerably more rigour than a demo call and a features grid.

    The difficulty is that every vendor in this market says the same words. Comprehensive coverage. Powerful AI. Real time updates. Enterprise ready. None of those phrases carries information, and the ones that sound most reassuring are usually the least verifiable. What follows is a framework for cutting through that: five dimensions that actually differ between vendors, a scorecard with weights you can copy into a spreadsheet, an evaluation protocol that puts real queries before pricing, and a list of questions to put in writing before you sign anything.

    One argument runs through all of it. Free tier generosity and prepaid credit expiry are the two most predictive signals of a vendor you can build on, and neither has anything to do with technology. They tell you whether the vendor expects to win by being useful or by making it awkward to leave.

    The Five Dimensions That Actually Differ

    Almost every real difference between Indian legal research APIs collapses into five questions. Everything else on a comparison table is either downstream of these or is noise.

    Corpus coverage

    How many records are indexed, and where did they come from. A vendor indexing reported judgments has a different product from one indexing the full registry record. The gap is measured in orders of magnitude, and it is the gap your users fall into.

    Court breadth

    Supreme Court and a few High Courts is not India. Indian commercial litigation runs through NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and the DRT, and Indian volume litigation runs through the district judiciary. Ask which forums, and which years within each.

    Pricing transparency

    Whether the price is on a page or behind a sales call. Published pricing is a commitment to treat a two person firm the same as a large one. Contact us pricing is a statement that the price depends on what they think you can pay.

    Credit and quota terms

    What happens to what you paid for and did not use. Expiry terms, rollover, refill behaviour and whether unused balance survives a plan change. This is where the money quietly leaks.

    AI capability, honestly scoped

    Not whether they have AI, but which subset of the corpus it has actually touched. Indexed is not analysed, and the ratio between the two is the single most informative number in the entire evaluation.

    Operational reality

    Rate limits, error semantics, response consistency, and what the API does when a registry never published a field. This is the dimension you cannot assess from a brochure and can assess in an afternoon of real calls.

    Corpus Coverage and the Question Nobody Wants Asked

    Start with size, because size is falsifiable. Ask for a number. Then ask what the number counts: documents, cases, orders, or all three added together in a way that flatters the total.

    Then ask the harder question, which is provenance. Indian court data originates with the registries and reaches the public through eCourts and the NJDG. A vendor either goes to those sources or buys from someone who did. That distinction matters for three practical reasons: latency, because every intermediary adds a lag between the registry publishing and you seeing it; continuity, because a reseller's access can be withdrawn by whoever they buy from; and correctness, because every hop is an opportunity for a transformation you cannot inspect. For reference, the CourtMesh corpus is roughly 310 million cases taken from official government portals with no third party intermediary in the chain.

    Then ask about the shape of the coverage, which is where most vendors get vague. Coverage is never uniform. Metadata completeness varies by court and by year: District Court records are consistently thinner than High Court records, and records from 1995 are thinner than records from 2024 everywhere. A vendor who tells you coverage is complete and even is either not looking or not telling. A vendor who tells you exactly where it thins out is describing a system they actually understand.

    The registry is the authority

    Whatever you buy, you are buying a view of what registries published, not the record itself. Any responsible vendor will say this out loud, and any responsible product will say it to your users too. If a vendor positions their aggregate as authoritative, they have either not thought about it or they are hoping you will not. Before advising a client or acting on a status, the position is confirmed against the official record of the court concerned. That does not make the aggregate less useful. It makes it correctly scoped.

    Court Breadth: Which Forums, Which Years

    Court breadth is the dimension most likely to sink a product after launch, because it fails asymmetrically. A demo built around Supreme Court constitutional matters will look magnificent on every platform in the market. The failure arrives when a real user searches for a section 138 Negotiable Instruments Act 1881 complaint pending before a magistrate, an O.M.P. under section 34 of the Arbitration and Conciliation Act 1996 in a High Court's original side, or an insolvency petition before the NCLT, and gets nothing.

    So enumerate. The Supreme Court, all 25 High Courts, District Courts, and tribunals including NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and the DRT. For each, ask for a start year, because a vendor may technically cover a High Court while holding only the last four years of it. Ask specifically about the courts your users actually litigate in, not the ones the demo used.

    Ask about case type codes too, because they are the practical handle for filtering. If your users think in CRL.A., W.P.(C), CS(COMM), O.M.P., SLP and Arb.P., an API that cannot filter on registry case types is going to force you to build a mapping layer nobody budgeted for. Some APIs expand a primary type into the underlying registry codes for you, so a request for Appeal resolves server side into CA, CRA, LPA and the rest. That is a small feature which quietly removes a whole category of work.

    Pricing Transparency and the Credit Expiry Question

    Here is the claim, stated plainly: how a vendor prices tells you more about how they will behave in year two than any technical assessment you can run in week one.

    Published pricing is a constraint the vendor has accepted. It means a solo advocate in Indore pays the published rate and so does a large firm in Mumbai, and it means you can model your unit economics without a procurement cycle. Pricing behind a mandatory sales conversation means the number is a function of what the seller estimates you can bear. That is a legitimate business model. It is also a signal about what renewal will feel like once your product depends on them.

    Credit expiry is the sharper test. Consider the arithmetic in the abstract, with no real numbers attached. Suppose you prepay for a block of credits and use sixty percent of them in the term. If unused credits expire, you paid full price for sixty percent of the value, and your effective rate is a third higher than the published rate. Now suppose your usage is lumpy, which it will be, because litigation support work follows filing seasons and client mandates rather than a smooth curve. Expiry converts every quiet month into a permanent loss and every busy month into an overage conversation.

    There is a second effect that is worse than the money. Expiring credits change user behaviour: teams start rationing calls near the end of a term, or burning them on low value work to avoid waste. Either way the vendor's incentive and yours have separated. A vendor whose credits do not expire, or that rolls unused balance forward, has aligned itself with you using the product when you need it.

    Why the free tier tells you the most

    A generous free tier is expensive for the vendor and produces no revenue on its own. A vendor offers one because they believe the product wins on merit once you have used it properly. A thin free tier, or one that gates the endpoints that actually matter, tells you the vendor's confidence sits in the contract rather than in the product. Judge the free tier by whether it lets you run your real workload, not by the size of the number. A thousand calls that exclude the AI endpoints is a smaller offer than a hundred that include them.

    AI Capability, Scoped Honestly

    Every vendor in this market now has AI on the homepage. The question that separates them is not whether a model is involved but what it has actually processed.

    There are two corpora in any legal AI product, and their sizes are not close. There is the corpus indexed for keyword search, which can plausibly run to hundreds of millions of records, because indexing text is cheap. And there is the corpus that has been embedded for semantic search and processed for structured analysis, which is expensive per document and is therefore far smaller. On CourtMesh the keyword corpus is roughly 310 million records and the semantically embedded, analysed corpus is in the low millions. That ratio is not a defect. It is arithmetic, and every vendor faces the same arithmetic whether or not they mention it.

    So ask the ratio. If a vendor implies that every record in a nine figure corpus carries model derived analysis, they are describing a compute bill nobody in this market is paying. Ask instead which courts and which years the analysed subset covers, because that is what determines whether AI features work for your users or only for your demo.

    Then ask what the analysis actually contains, because the word covers everything from a two line abstract to a structured breakdown of issues, holdings, reasoning, cited cases and precedent relationships. Ask whether the analysis carries citations back into the source document. Analysis that points at the passage it came from is checkable by a lawyer in seconds. Analysis that does not is something you have to take on faith, and faith is not a professional standard.

    A Scorecard You Can Copy

    Weights are arguable and should be argued about, inside your own team, before you speak to any vendor. What follows is a defensible default for a product that serves practising lawyers. Score each vendor 1 to 5 on each criterion, multiply by the weight, and total. The discipline is not the arithmetic; it is being forced to justify a score of 4 on coverage when the vendor would not give you a number.

    CriterionWeightWhat a 5 looks likeWhat a 1 looks like
    Corpus size and provenance20A specific record count, direct sourcing from eCourts, NJDG and registries, and a clear statement of what the count countsComprehensive with no number, or a refusal to say where the data comes from
    Court and forum breadth15Supreme Court, all 25 High Courts, the district judiciary and the major tribunals, with a start year given per forumSupreme Court and selected High Courts, years unspecified
    Analysed subset size and depth15An explicit ratio of analysed to indexed, structured analysis fields, and citations pointing back into the sourceAI powered as a claim, with no scope, no ratio and no grounding
    Pricing transparency12Published rates, published credit consumption per endpoint, no mandatory sales call to see a numberContact us, with the rate determined after they have seen your funding announcement
    Credit expiry and rollover12Unused credits carry forward, terms stated on the pricing page, plan changes do not destroy balanceAggressive expiry, silent forfeiture on plan change, terms only in the contract
    Free tier usefulness8Enough allowance on the endpoints that matter, including AI endpoints, to run a real workload before payingA trial that excludes exactly the endpoints you need to evaluate
    Rate limits and operational fit8Published limits, standard rate limit headers on every response, a retryAfter value you can honourUndocumented limits discovered in production, or 429 with no guidance
    Response consistency and errors6One success and error envelope across every endpoint, per field validation messagesDifferent shapes per endpoint, errors as HTML or bare strings
    Documentation quality4Every endpoint, every parameter, every error, with the absent field problem addressed honestlyA quickstart, a sales deck, and silence about edge cases

    Notice that documentation carries the lowest weight and corpus the highest. That ordering is intentional and it is the opposite of how most evaluations run in practice, because documentation quality is visible in five minutes and corpus quality takes a day to test. Weight by consequence, not by how easy something is to assess. You can work around thin documentation. You cannot work around a corpus that does not contain your users' cases.

    The Protocol: Ten Real Queries Before You Look at Price

    This is the part that changes decisions. Do not evaluate with the vendor's demo queries. Do not evaluate with queries you invented for the evaluation. Take ten real research questions from matters your firm or your users have actually run in the last year, write them down before you speak to anyone, and put the identical ten to every candidate.

    1

    Write the ten queries first, and freeze them

    Pull them from real files. A mix: two identifier lookups where you know the exact case number, two named party searches, two settled terms of art, two conceptual questions where you can state the principle but not the label, one fact pattern as a client described it, and one exhaustive listing such as every judgment of a given High Court under section 138 of the Negotiable Instruments Act 1881 since a given year. Freeze the list. Adding a query mid evaluation is how you end up rationalising a preferred vendor.

    2

    Record the answer you expect

    For each query, note the judgment or the set you already know should come back. This is the only way to measure recall. Without a known answer you are grading vendors on how confident their results look, which is precisely the failure mode AI search encourages.

    3

    Run all ten against every candidate on the free tier

    Same queries, same filters, same day. Do this before any pricing conversation, because knowing the price contaminates how generously you read the results. If a free tier will not let you run this, that is itself a finding and it belongs in the scorecard.

    4

    Score recall and precision separately

    Did the known answer come back, and where in the ranking. Then, how much of the top ten was actually relevant. Vendors optimise for precision because it looks good in a demo. Recall is what determines whether your user misses the judgment that decides their matter.

    5

    Deliberately test the thin edges

    Search a District Court in a state nobody demos, in a year before 2010. Look at which fields come back empty. This tells you more about the underlying pipeline than any Supreme Court query will, because everyone's Supreme Court data is good.

    6

    Break it on purpose

    Send a malformed date, an out of range year, an empty query, a deactivated key. Read the errors. An API that returns a field level validation message telling you exactly which parameter was wrong will save your team weeks over an API that returns a generic 400.

    7

    Push through the rate limit

    Find the limit, breach it, and read what comes back. You want a 429 carrying a retryAfter in seconds and a reset timestamp, plus rate limit headers on every ordinary response so your client can steer before it hits the wall. Discovering an undocumented limit in production is a bad way to learn it exists.

    8

    Only now, open the pricing page

    With the ten query results in a table in front of you. The vendor that lost on recall does not become acceptable because it is cheaper, and the expensive one does not become correct because it is expensive. Price is the last input, not the first.

    Questions to Put in Writing

    Ask these by email, and keep the reply. Not because you expect to litigate it, but because a vendor answers a written question differently from a spoken one, and the difference between the two answers is information.

    1. How many records are indexed for keyword search, and how many of those have been embedded for semantic search and processed for AI analysis. Please give both numbers.
    2. Which courts and forums are covered, and from which year for each: Supreme Court, each High Court, District Courts, and each tribunal including NCLT, NCLAT, ITAT, CESTAT, SAT, TDSAT and DRT.
    3. Does your data come directly from eCourts, the NJDG and court registries, or through a third party data supplier? If a supplier is involved, name them and state what happens to our access if that relationship ends.
    4. What is the rate limit per API key, expressed per minute and per day, and is it documented publicly? Do responses carry rate limit headers, and does a 429 include a retry interval?
    5. What happens to prepaid credits that are unused at the end of a term, and what happens to the balance if we change plan mid term?
    6. Which endpoints consume credits, at what rate, and is that consumption published?
    7. For a District Court record from 2012, which fields are typically populated and which are typically empty? A sample response would be ideal.
    8. Is analysis grounded in the source document, meaning does it carry citations pointing back at the passages it was derived from?
    9. How are API keys stored, and can your support team recover the plaintext of a key we have lost? The correct answer is no.
    10. What is logged per API key, how long is it retained, and are sensitive headers and body values redacted before storage?
    11. Under the DPDP Act 2023, what is your role in respect of data we send you, and what are your obligations to us on breach notification?
    12. If we stop paying, what happens to data we have already retrieved and stored? Are there restrictions on retention, redistribution or derived works?

    The last two are asked least often and matter most at the point of exit. A licence that quietly forbids retaining what you already pulled turns migration from a quarter of engineering into a legal problem.

    Signals Worth Walking Away From

    Some answers should end the evaluation rather than lower a score.

    A refusal to state corpus size, or a number that changes between the website, the deck and the email
    AI analysis implied across an entire nine figure corpus, which is arithmetically not happening
    No published rate limit, so your capacity planning is a guess and their enforcement is discretionary
    Prepaid credits that expire aggressively, combined with no rollover and forfeiture on plan change
    Support able to retrieve the plaintext of your API key, which means keys are not stored as hashes
    Data sourced from an unnamed third party, leaving your continuity dependent on a contract you cannot see
    A free tier that excludes the exact endpoints the product is sold on, so the evaluation is impossible by design
    Claims of complete and uniform coverage across all Indian courts and all years, which nobody has

    A vendor who tells you precisely where their coverage thins out has understood their own pipeline. A vendor who tells you coverage is complete has either not looked or is hoping you will not.

    What the Framework Adds Up To

    Run the protocol and the scorecard and something slightly uncomfortable usually happens: the vendor that looked best in the demo is not the one that wins. Demos are optimised for precision on familiar Supreme Court material, and precision on familiar material is the easiest thing in this entire domain. Ten real queries, including the awkward District Court one and the exhaustive listing, produce a different ranking almost every time.

    The other thing the framework does is force you to price the commercial terms alongside the technology, where they belong. An API with slightly worse recall and non expiring credits, published rates and a free tier that lets your team actually experiment may be a better place to build than a marginally better corpus behind a sales process and a term that quietly forfeits what you did not use. You are not buying a query response. You are buying a relationship whose terms you will be living inside while your product depends on it.

    For reference against your own scorecard, the CourtMesh API exposes twelve endpoints over roughly 310 million cases spanning the Supreme Court, all 25 High Courts, the district judiciary and tribunals, sourced directly from official government portals. The endpoint reference, including every parameter and every documented error, is at the API documentation. Credit consumption and plan terms are published at API pricing rather than quoted on a call. Run your ten queries against it and against everyone else, and let the table decide.

    Run your own ten queries

    No comparison table survives contact with real research questions from real files. Take ten queries off your own matters, freeze the list, and put the same ten to every candidate before anyone shows you a price. The CourtMesh corpus covers roughly 310 million cases from official government portals across the Supreme Court, all 25 High Courts, the district judiciary and tribunals. Endpoint reference and error semantics are at the API documentation, plan and credit terms are published at API pricing, and an overview with access details is at the API overview.

    Explore CourtMesh
    APILegal ResearchDevelopersEvaluationIndia
    X LinkedIn