Artificial Intelligence

India’s Sarvam AI Beats Google Gemini, ChatGPT on OCR Benchmarks

BENGALURU: Bengaluru-based startup Sarvam AI claims its models outperformed Google Gemini and ChatGPT on optical character recognition and text-to-speech benchmarks focused on Indian languages, marking a milestone for domestic AI development.

Sarvam Vision achieved 84.3% accuracy on olmOCR-Bench, surpassing Gemini 3 Pro and DeepSeek OCR v2, while ChatGPT ranked significantly lower. On OmniDocBench v1.5, Sarvam Vision scored 93.28% overall, excelling in complex formulas and layout parsing.

Co-founder Pratyush Kumar shared benchmark results on X, stating “On Indian languages, Sarvam Vision is the best model by far, while supporting all 22 scheduled Indian languages.”

Read more: Gnani AI Raises $10M to Build Sovereign Voice AI for Enterprise Markets

The Vision series includes a 3-billion-parameter state-space model capable of image captioning, scene text recognition, chart interpretation, and complex table parsing. The model handles messy layouts, tables, mathematical formulas, and technical documents where traditional OCR tools struggle.

Read more: Meta Plans Facial Recognition for Ray-Ban Smart Glasses This Year

Alongside Vision, Sarvam launched Bulbul V3, a text-to-speech model supporting 35 voices across all 22 official Indian languages. Bulbul V3 handles smooth language switching between Tamil and English or Hindi and English without disruption.

Tech commentator Deedy Das acknowledged changing his earlier skepticism: “I was wrong about Sarvam. When I wrote about them a year ago, I felt the direction to train small Indic language models was wrong. But they have the best text-to-speech, speech-to-text, and OCR models for Indic languages.”

Union IT Minister Ashwini Vaishnaw said the work reflects success of India’s AI mission.

Sarvam made its Document Intelligence API free through February 2026. The startup positions itself as building “sovereign AI” developed within India for government projects, public infrastructure, and BFSI sector applications, alongside innovations like indigenous AI smart glasses.

Anurag Shukla

Anurag Shukla is a Senior Journalist with over two decades of experience across television, digital, and print media. He has worked with leading national news organisations and has also served as a Research Officer in the Prime Minister’s Office (PMO), contributing to media research and policy-level content. A former journalism academic, Anurag brings strong editorial depth and a keen understanding of how technology, governance, and society intersect at Tea4Tech.

Recent Posts

Amid Backlash, Meta Goes Camera-Free With New AI Glasses

California: Meta introduced its first pair of camera-free AI glasses on Wednesday at its Connect…

11 hours ago

Anthropic Brings Claude Chat and Cowork Together Into a Single Window

Bengaluru: Anthropic has combined Claude's chat and Cowork into one interface, so users no longer…

1 week ago

OpenAI Goes All In on Hardware, Buys Glass Imaging for $300M+

California: OpenAI has reportedly acquired Glass Imaging, a startup specializing in AI-enhanced smartphone camera technology,…

1 week ago

OpenAI Launches GPT-6 Astra With Advanced Computer and AI Capabilities

New Delhi: OpenAI has launched GPT-6 Astra, its latest and most advanced AI model, with…

2 weeks ago

India’s Rilo Joins Adobe in AI Marketing Automation Deal

New Delhi: Adobe has acquired Rilo, an India based marketing intelligence startup, in a deal…

3 weeks ago

Anthropic Levels Up Claude Fable 5.1 With More Power, Less Cost

New York: Anthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, its latest AI…

3 weeks ago