Ask any AI chatbot to translate a clean, typed English paragraph into Hindi, and it will do a competent job in seconds. That is not, however, the document Indian legal teams actually spend their days on.
The real file is a photocopied sanad from 1974, faded in exactly the spot where the survey number sits. It is a land record scanned at low resolution, in Kannada. It is a departmental note typed on a manual typewriter decades ago, in Tamil, with a handwritten figure in the margin that may or may not be an assessment value. This is not an unusual file. For most Indian legal teams working on property, recovery, or compliance matters, it is a routine one.
We built Lexi for exactly this kind of document. To see how it performs, we ran a full benchmark: Lexi against Claude, ChatGPT, Gemini and Grok, on genuinely degraded Indian legal documents, and published every number.
What This Benchmark Shows
Three findings, in order of significance:
1. Lexi is the only system that gets a legal notice filing-ready in every language. A legal notice carries a list of attachments that has to stay as a checkable list. Lexi produced this correctly in all six languages tested. Three of the four other systems did not manage it in any language.
2. Lexi leads the field on reading degraded, decades-old scans. Averaged across every scanned document, Lexi recovered 98.5% of the critical information, ahead of Claude (96.6%), Gemini (94.4%), Grok (91.4%), and ChatGPT (87.5%).
3. Factual accuracy no longer separates the field. All five systems preserved every one of the 16 legally operative details, in all six languages. Lexi clears that baseline cleanly, but so does everyone else, which is why the two results above are the ones that decide anything.
What Is Benchmarking, and Why Should a Legal Team Care?
Benchmarking simply means giving several tools the exact same task and comparing how each one performs: a fair, apples-to-apples test rather than a marketing claim. Every system here read the same documents and translated the same notice. Nothing was tuned per tool, and nothing was cherry-picked to flatter one system over another.
It matters because a legal team choosing an AI tool isn't choosing based on a demo. They're choosing based on whether the tool holds up on the worst document already sitting in their own filing cabinet, and a benchmark is the closest thing to that test that can be run and published at once, for everyone to check.
What We Tested
This was a substantial test: 55 separate outputs, produced by five AI systems, measured against six real source documents. Two kinds of work were tested, because Indian legal teams routinely need both directions.
Both were chosen for the same reason: they are the two points where an error is expensive and invisible. Reading a damaged scan wrong loses a fact nobody knows is missing. Translating a notice wrong produces a document that looks correct and fails when it is served. Tasks where mistakes announce themselves are not worth benchmarking. These are.
Reading damaged scans (Indic → English). Each system was given genuinely poor-quality source material, including a badly degraded Kannada land record and a Tamil departmental file note, and asked to extract the numbers that carry legal weight: survey numbers, order dates, file references, and assessment figures. 89 such numeric entities were checked across the scanned documents. This direction is unforgiving by design: there is no room to paraphrase a survey number, only to get it right or wrong.
Translating a legal notice (English → Indic). A formal statutory notice, the kind sent under Section 138 read with Sections 141 to 142 of the Negotiable Instruments Act, common in cheque-bounce matters, was translated into six Indian languages: Gujarati, Hindi, Kannada, Marathi, Punjabi, and Tamil. Sixteen legally critical details were tracked across every language: a cheque number, an amount, five dates, two invoice references, three statutory section citations, the Act year, the notice period, and both PIN codes. None of these sixteen is a minor formatting choice. Each one does real legal work, from establishing the statutory window to fixing the address that determines whether service was good.

Translation inside Lexi: language and quality mode selected before the document is processed.
What We Found
1. Lexi Gets the Document Filing-Ready
This is the single clearest result in the entire study, and arguably the most practically important one.
Why this test exists: An enclosure schedule is the one part of a notice that gets handled physically. If it arrives as a paragraph rather than a list, someone has to retype it into a checkable form before service, by hand, every time. That is slow, and worse, it is the point where an item quietly goes missing. A notice served without an annexure it claims to carry is a defect the other side can raise, so this is not a formatting preference. It decides whether the document works on the day it is used.
A legal notice like the one tested here comes with a schedule of enclosures: a cheque scan, a return memo, invoices, delivery challans, a ledger extract, and screenshots, seven items in total. At the point of service, a clerk checks that list against the physical attachments in hand. For that check to work at all, the seven items need to survive as discrete, numbered entries. A single run-on paragraph cannot be checked off item by item.

Lexi was the only system to itemise the enclosure schedule in every language tested.
Against three of the four systems tested, the gap is absolute: 100% against zero, in every single language. Only Grok came close, at 83.3%.. This is a document-understanding result, not a language result. Knowing that an annexure schedule has to survive as a checkable list is knowledge about how legal documents actually get used, not knowledge about how to translate a sentence. That is the difference between a translated file and a document a clerk can actually work from.
2. Lexi Leads on Reading the Hardest Documents
Reading a faded, decades-old scan in a regional script is far more demanding than translating clean, typed text, and it's the part of the job that matters most in practice, since a title search or a recovery matter often comes down to a handful of numbers buried in exactly this kind of document.
Why this test exists: A missed number in a scanned record does not announce itself. The translation still reads fluently, the sentences still make sense, and nothing in the output tells a reviewing lawyer that a survey number or an order date was dropped on the way through. That is what makes recall the right thing to measure, and it is also why performance on the worst document in a set matters more than an average across easy ones. A tool is only as useful as its behaviour on the file nobody wants to be assigned.
Averaged across every scanned document in the study, Lexi led the field outright:

Readings of critical entities across all scanned documents in the benchmark.
And on the single hardest document in the study, a heavily degraded Kannada land record with 26 critical numbers to extract, Lexi led there too:

Detection rate on the most degraded document in the corpus.
Land records this old are frequently the only surviving evidence of a right, with no clean digital original to fall back on if a number is misread. On a separate departmental file note in Tamil, Lexi also surfaced a figure that no other system in the study recovered at all. It is the kind of catch that, in title and recovery work, is the difference between a complete record and a silent gap that surfaces years later.
3. Every Model was Able to Preserve the Legal Facts
Why this test exists: This is the catastrophic-failure check. A wrong statutory period, a dropped section citation or an altered PIN code can make a notice unservable, and like a missed OCR entity, the failure is silent: the document still reads correctly. Every system clearing it is itself the finding. Factual accuracy is now table stakes rather than a differentiator, which is precisely why the two results above are the ones that separate the field.
On the translation task, every system tested, Lexi included, preserved all 16 operative legal facts correctly, across all six languages, without exception.

All sixteen operative facts preserved across every target language.
This is the floor every serious legal AI tool needs to clear, and the whole field clears it. Lexi included, without a single exception across six scripts. For a notice that has to survive service and later scrutiny, that floor is the foundation the two results above are built on. It is also the reason those two results, not this one, are where a legal team should be looking.
Why the Real Advantage Isn't in the Numbers Alone
General-purpose AI tools take in one document and hand back one file. They don't know what matter that document belongs to, can't check it against the other documents already in the file, and can'troute it to the right person for review. Every one of those steps, reconciling terminology, tracing a figure back to its place on the source scan, review by someone qualified to sign it, and assembly into a filing-ready bundle, has to happen manually, by hand, every single time.
Lexi is built around a different premise: translation is one governed step inside a chain that already holds the entire matter.

The Lexi pipeline.
● Scan intake- the document enters the system already linked to its matter: parties, jurisdiction, practice area, key dates.
● OCR and translation- the step this benchmark measured, now informed by the surrounding case context rather than treated as an isolated file.
● Citation check- statutory references are verified as existing, current, and supporting the proposition made.
● Filing pack- the reviewed document is assembled into an indexed, paginated, filing-ready bundle.
● Matter update- status writes back automatically.
None of this shows up in a recall percentage. It's the entire difference between a translated file sitting in an inbox and a matter actually moving toward being closed.
How We Scored This
Getting a benchmark like this right takes real care, and our scoring approach went through more than one correction along the way. Sharing that process is part of what makes the results above worth trusting.
Numeral scripts can create false failures. An early pass of scoring flagged what looked like a dropped reference number in a Marathi output. It had not been dropped. It had been rendered correctly in Devanagari numerals (such as २७४३२५) rather than Western digits. Any evaluation that doesn't normalise numerals across scripts before comparing them will systematically, and wrongly, penalise the systems doing the most thorough native-script localisation.
Exact string matching can reward reproducing a source's own errors. One early pass credited a system with uniquely finding an order date that others had supposedly missed. In fact, no date had beenmissed. The source document itself rendered the date inconsistently, and only one system had
reproduced that exact quirk. When the reference text is itself imperfect, exact-string matching measures which system best replicates the original's flaws, not which system found the correct value.
Consensus scoring can penalise a system for reading more carefully. Treating agreement between systems as a proxy for truth rewards conformity, and marks down whichever system surfaces something genuine that the others simply missed. In legal work specifically, that overlooked fact is often exactly the one that matters.
Methodology & Sources
Every figure in this guide is drawn from an ongoing benchmark evaluation of five AI systems, Lexi, Claude, ChatGPT, Gemini and Grok, across 55 provider outputs and 6 source documents, with OCR figures reflecting our latest analysis.
The OCR task used a consensus-based scoring method rather than a single fixed "ground truth," since the source scans were too degraded in places to serve as a clean reference. As set out above, consensus scoring carries a known bias: it can mark down a system that correctly surfaces something the others missed. We used it because no cleaner reference existed, andmitigated it by cross-checking uniquely recovered figures against each document's own embedded text layer wherever one survived. The full methodology and underlying dataset are available on request, we would rather a reader ask us for the data than take any figure here on faith.
Reach out at hi@getlexi.io.
Frequently Asked Questions
What makes Lexi different from a general-purpose AI model for Indian legal documents?
A general-purpose AI model receives a single document, has no knowledge of the wider matter it belongs to, and the interaction ends once it returns a file. Lexi keeps the entire matter in context (parties, jurisdiction, practice area, prior filings, and the original scan) and carries the translated document through citation verification, advocate review, and filing pack assembly automatically.
Does a legal notice need more than accurate translation?
Yes. A notice also needs to stay structurally usable once served. For example, its schedule of enclosures needs to remain a checkable, itemised list rather than collapse into a paragraph. In this benchmark, Lexi was the only system that preserved this structure consistently across all six languages tested.
Can AI accurately translate Indian legal documents into regional languages?
Yes. Every system in this benchmark, Lexi included, preserved all 16 legally operative facts correctly across all six Indian languages tested. Accurate preservation of legal substance is achievable with today's AI tools; the more differentiating question is what happens to a document's usability beyond that baseline, which is where Lexi pulls ahead.
How does Lexi compare to Claude and Gemini for Indian legal documents?
Lexi led both on OCR recall: 98.5% mean against Claude's 96.6% and Gemini's 94.4%, and 96.2% against 92.3% and 88.5% on the single hardest document tested. On the enclosure schedule, Lexi itemised it correctly in all six languages; neither Claude nor Gemini managed this in any language.
How does Lexi compare to ChatGPT for Indian legal documents?
Lexi's mean OCR recall (98.5%) was well ahead of ChatGPT's (87.5%), and on the hardest scanned document in the study the gap was the widest of any matchup tested, at 96.2% against 57.7%. Lexi also itemised the enclosure schedule correctly in all six languages; ChatGPT did not manage this in any of them.
How does Lexi compare to Grok for Indian legal documents?
Lexi led Grok on both OCR measures (98.5% versus 91.4% mean; 96.2% versus 76.9% on the hardest document). On the enclosure schedule, Lexi itemised all six languages correctly; Grok was the closest competitor of the four, itemising five of six.
Is a benchmark result enough to choose an AI tool for legal work?
Not on its own. A benchmark is a useful, verifiable signal, but the more durable test is to take the single most degraded document already sitting in your own files, and run it through whatever tool you're considering. That's the test that predicts what a team will actually experience day to day.
The Bottom Line
Lexi was the only system in this study that produced a legal notice ready for real-world use, a properly itemised enclosure schedule, in every language tested. It led the field on reading degraded, decades-old scans, both on average across every document and on the single hardest document in the study. And it preserved every one of sixteen legally critical facts without exception, matching the full field.
The deeper reason to choose a tool like this goes beyond any single number. Translation is only one link inside a system that carries the whole matter forward, from a degraded scan at the bottom of a file, to a document a lawyer can put in front of a court or a client.
Lexi is a legal operating system built for Indian legal teams: matter management, document processing, drafting, review, and filing, all in one place. To see how it handles your own documents, reach out at hi@getlexi.io.
