Sarvam expands its multilingual AI stack
Indian AI company Sarvam AI is pushing deeper into speech technology with Saaras V4, its latest automatic speech-recognition model.
The model supports all 22 constitutionally recognised Indian languages, alongside English.
Sarvam says the system is designed for noisy audio, code-mixed speech and dialect variation.
Those capabilities target some of the hardest problems in building voice AI for India. (Sarvam AI)
The company has positioned Saaras V4 as a foundation for voice agents and multilingual applications.
That makes the release relevant beyond transcription.
Speech recognition can become the input layer for customer-service agents, government services, healthcare tools, education platforms and enterprise software.
India needs more than English-first AI
Most global AI systems were initially developed around high-resource languages.
English benefited from enormous training datasets.
Indian languages present a different challenge.
Many have less digitised training data.
Users also frequently mix languages within the same conversation.
Speech varies across regions.
Background noise can further reduce recognition quality.
Therefore, an Indian voice model needs to handle more than vocabulary.
It needs to understand context.
Saaras V4 is designed specifically around these conditions.
Sarvam says the model achieved state-of-the-art performance across all 22 Indian languages in its evaluations. These are company-reported benchmark claims, so they should be distinguished from independently reproduced results. (Sarvam AI)
A different architecture under the hood
Saaras V4 combines an audio encoder with a 3-billion-parameter hybrid state-space language-model decoder trained by Sarvam.
The architecture is designed to handle multiple output formats.
Users can request standard transcription, translation, verbatim text, transliteration or code-mixed output. (Sarvam AI)
That flexibility matters for India.
A Hindi speaker may want Devanagari text.
Another user may want Romanised Hindi.
An enterprise may need English translation.
A voice agent may need code-mixed text.
One model supporting multiple outputs can simplify application development.

Keyterm prompting targets specialised vocabulary
Sarvam has also added keyterm prompting to Saaras V4.
The API allows developers to provide up to 50 domain-specific terms.
Those can include names, places, brands or technical terminology.
The system then uses those terms to improve recognition.
This feature could be useful in specialised environments.
Consider a hospital.
A transcription system may need to recognise drug names.
A financial application may need company names.
A government application may need local place names.
Generic speech recognition can struggle with such vocabulary.
Domain-specific prompting gives developers another way to improve accuracy without rebuilding the model.
Sarvam’s API documentation confirms that keyterm prompting became available for Saaras V4 through its REST and batch APIs. (Sarvam AI Developer Documentation)
Real-time voice is becoming more important
Speech AI is also moving toward real-time applications.
Sarvam made Saaras V4 available through its realtime API earlier this month.
The company says the model is designed to handle noisy audio and code-mixed speech while supporting multilingual recognition. (Sarvam AI Developer Documentation)
That could become important as AI agents become more conversational.
Traditional voice assistants generally follow a simple pattern.
Listen.
Transcribe.
Generate a response.
Speak.
The quality of transcription therefore directly influences the entire interaction.
If the first stage fails, later reasoning cannot completely recover.
Consequently, better speech recognition can improve the reliability of the entire voice-AI pipeline.
The commercial opportunity is large
India has hundreds of millions of internet users who communicate in languages other than English.
That creates a large potential market for voice interfaces.
However, building a good model is only one part of the challenge.
Developers need APIs.
They need predictable pricing.
They need latency low enough for real-time applications.
They also need privacy and reliability.
Sarvam’s product approach is therefore focused on making the technology accessible to developers rather than keeping it solely as a research system.
India’s AI competition is moving into specialised layers
The AI market is increasingly fragmented.
Some companies build foundation models.
Others focus on chips.
Some build agents.
Others specialise in speech, vision, healthcare or enterprise workflows.
Sarvam is concentrating heavily on the language and speech layer.
That can create a specialised competitive position.
The company does not need to dominate every part of AI.
Instead, it can focus on making voice interfaces work across India’s linguistic diversity.
What Saaras V4 represents
Saaras V4 is ultimately part of a larger shift.
India’s AI opportunity is not limited to building another general chatbot.
It also involves adapting AI to the country’s languages, accents, institutions and everyday use cases.
Speech could be one of the most important interfaces in that transition.
If voice systems become reliable across Indian languages, AI applications can reach users who may not prefer English-first interfaces.
That makes multilingual speech recognition more than a technical feature.
It can become infrastructure for the next generation of Indian AI applications.
Tags: Sarvam AI, Saaras V4, Indian languages, speech AI, voice AI, artificial intelligence, Indian AI startups
Flairius CTA: Follow Flairius News — sharp takes on AI, business, and India’s startup economy — flairiusnews.com

