Sarvam Vision 2.1: Indian Language OCR Benchmark Review

Sarvam Vision 2.1 is the first OCR model to handle both Indian language scripts and English document structure without sacrificing either.

Sarvam Vision 2.1: Indian Language OCR Benchmark Review

For the first time, an OCR model reads Indian languages as fluently as English documents. Sarvam Vision 2.1 breaks a long-standing trade-off. Indian businesses had to choose between structural document parsing and script recognition — never both. Finance teams automating invoices faced this. Insurance companies processing handwritten claims in multiple states faced this. Organizations digitizing regional records faced this. All had to pick accuracy or coverage. This model does both. Below is a look at what changed, where it wins, where it struggles, and how the benchmarks break down.

Key Features of Sarvam Vision 2.1

Value Extraction From Tables

Pulling named fields out of a table rather than transcribing the entire grid. This works for statements, ledgers, or anything where a value only makes sense in relation to its row and column headers.

Value Extraction From Forms

The same idea applied to form fields. The label and value sit in separate boxes. The pairing has to be inferred from layout, not reading order.

Indic Handwritten Extraction

Handwriting in Indian scripts extracted into structured fields rather than transcribed as a block. This is the hardest capability. Sarvam trained it partly on video sources to capture real handwriting variety, not just synthetic samples.

What the Numbers Actually Show

Until now, if a model was good at English document parsing it was usually mediocre at Indian languages, and the models that handled Indic scripts well were not competitive on general document structure. Those were separate tools with separate failure modes.

Sarvam Vision 2.1 Benchmark

Each point is one system scored on both axes. Upper right is better on both at once.

Look at where the other points sit. Infinity-Parser2 Pro scores 86.1 on English, second only to Sarvam, and then collapses to 49.83 on Indic. Google Cloud Vision does the reverse: 39.6 on English document structure, but 71.76 on Indian languages, because it has had Indic OCR for years without the layout intelligence. Sarvam 2.1 is the only point in the upper right.

Where It Wins and Where It Doesn’t

The Indic Benchmark

Sarvam also released the benchmark itself: 6,909 samples, 6,609 spanning all 22 official Indian languages and 300 in English, drawn from newspapers, brochures, textbooks, and historical writing dated from 1800 to the present.

The benchmark is published on Hugging Face, which matters. A vendor-built benchmark that the vendor wins is worth little on its own. A vendor-built benchmark released publicly so others can run it is a different proposition, and it is the right way to do this.

olmOCR-Bench

The community benchmark for English document parsing. It runs pass-fail checks across eight categories: arXiv maths, base text, headers and footers, tiny text, multi-column pages, degraded old scans, old-scan maths, and tables. It tests whether facts are present or absent rather than scoring subtle differences.

Sarvam notes that this benchmark is officially English-only but contains some contaminant samples in Chinese and other scripts. Their previous release reported on a filtered English-only set; this time they report on the official set for parity with competitors, which is the more conservative choice.

ModelMathTablesOldScanMultColOverall
Sarvam Vision 2.190.591.955.382.187.3
Infinity-Parser2 Pro87.488.958.083.386.1
Opus 590.089.554.085.885.1
Chandra-OCR286.587.549.282.484.5
Mistral OCR483.788.648.985.783.1
Gemini 3.6 Flash86.585.948.178.682.4
GPT 6 Astra82.690.947.077.881.8

OmniDocBench v1.6

A different measure: structural fidelity rather than fact presence. It is a composite of text edit distance, table structure scored with TEDS, formula recognition scored with CDM, and reading order, run over newspapers, textbooks, magazines, and financial reports.

ModelText edit dist (lower better)Formula CDMTable TEDSOverall
PaddleOCR-VL 1.60.03560.9850.93196.01
Sarvam Vision 2.10.02890.9880.89094.97
GLM-OCR0.03740.9840.89594.71
GPT 6 Astra0.04600.9670.89193.74
Gemini 3.6 Flash0.03710.9760.86993.58
Opus 50.04710.9670.85692.51

Read that table across rather than down. Sarvam has the best text edit distance of any model at 0.0289 and the best formula score at 0.988. It loses the top spot purely on table structure, where PaddleOCR-VL scores 0.931 against Sarvam’s 0.890. Sarvam calls both benchmarks arguably saturated, which is fair when the top twelve models sit inside four points of each other.

On the Global Benchmarks

Sarvam leads olmOCR-Bench at 87.3 overall. It does not lead every category:

CategorySarvam 2.1Best scoreHeld by
Old scans55.358.0Infinity-Parser2 Pro
Multi-column82.185.8Opus 5
Tiny text92.593.5Opus 5
Tables91.991.9Sarvam 2.1
Math90.590.5Sarvam 2.1

On OmniDocBench v1.6, Sarvam is second, not first: 94.97 against PaddleOCR-VL 1.6 at 96.01. Sarvam wins on text edit distance and formula recognition; PaddleOCR wins on table structure with a TEDS of 0.931 against 0.890.

Santhali is the clear loss. Sarvam scores 53.91 and Bodhan Indic-OCR scores 68.30, a gap of more than fourteen points. Odia is a narrower loss to Gemini 3.6 Flash, 80.01 against 81.01. Kashmiri is not a loss but it is weak in absolute terms at 54.82, the best score any model manages on that language.

Sarvam Vision 2.1 Architecture

Sarvam Vision 2.1 Architecture

  • The vision-language model doesn’t read the page directly
  • Two harnesses sit in front of it: a semantic layout parser that segments the page into regions, and a pointer network that establishes reading order
  • The VLM can work at page level alone, but Sarvam found the accuracy trade-off makes harnessing worthwhile
  • This is why multi-column newspapers and merged-cell tables are the hard test cases: if the harness segments wrongly, the VLM transcribes correct text in the wrong order — fluent and wrong, which is harder to catch than garbled output
  • Post-training combined supervised fine-tuning with RLVR (reinforcement learning with verifiable rewards), which fits OCR well since correctness against a known transcription is programmatically checkable

Conclusion

Sarvam Vision 2.1 is a meaningful step forward for document AI applied to Indian languages. It is the first model to occupy the upper-right quadrant on both English document structure and Indic script recognition simultaneously, and the public release of the benchmark dataset allows the broader community to verify those claims independently. Gaps remain — particularly on Santhali and in old-scan parsing — but the core trade-off that has constrained Indian document automation for years appears, for the first time, to be broken.