What is Saaras V4?
Saaras V4 is Sarvam AI's speech-to-text model, released on 25 September 2026. It transcribes the 22 scheduled Indian languages and English, now including non-Indian English accents, and it works out the language itself. Sarvam reports a wrong-language rate of 5.22% across all 22 languages and 2.9% across the ten most spoken.
Inside, an audio encoder feeds a 3-billion-parameter language model that Sarvam trained from scratch. That design lets one model write the same call out five ways:

Two other features matter for real calls. Keyterm prompting takes names, brands and product terms before transcription starts; on AI4Bharat's IndicContextEval benchmark, with the terms supplied, Sarvam reports 16.03% word error, the lowest it found. And streaming starts returning words in under 150 milliseconds, by Sarvam's figure. The docs say it handles 8 kHz phone audio.
Sarvam's own benchmark claims are strong. It reports the best results on all 22 Indian languages and the lowest average word error across seven English test sets, from meetings and earnings calls to Indian-accented English. On noisy Indian speech, it reports less than half the error of Deepgram Nova-3 and GPT-4o Transcribe. All of that is the vendor's own measurement.

How does it score on independent tests?
Saaras V4 has not been independently tested yet. Sarvam's earlier models have, and the gap they show is wide.
Voice of India, built by AI4Bharat at IIT Madras and Josh Talks, is 536 hours of unscripted phone conversations from 36,691 speakers in 15 languages. Its scoring accepts legitimate spelling variants, so a model is not punished for writing a word two valid ways. In the paper, Sarvam's audio model had the lowest error rate in 13 of the 15 languages. Google's Gemini 3 Pro won on the Hindi dialects Bhojpuri and Chhattisgarhi. No system stayed under 20% word error in every language.
Tamil shows how far apart the models are:
Word error rate on Tamil phone calls, Voice of India benchmark
Hindi is closer at the top. On the same Voice of India table, Sarvam's audio model scored 4.95%, Gemini 3 Pro just under 6% and Amazon Transcribe 6.78%, against 12.97% for Deepgram Nova-3 and 33.95% for GPT-4o Transcribe. A separate live Hindi leaderboard from the same team, updated 25 July 2026, puts Saaras v3 second at 3.78%, behind Bodhan's Indic Transcribe Core at 3.48%, with Gemini 3 Pro third at 4.70%.
Two notes on method. Voice of India scored only models that list the language, so Deepgram's 68% on Tamil is a supported language, not a model asked to guess. And "Sarvam Audio", the paper's winner, is a separate Sarvam model from Saaras.
In plain terms, 14% word error is about one word in seven, and 68% is about two in three. On a three-minute call at a typical 150 words a minute, that is roughly 60 wrong words against 300, and the ones that go wrong are often names and amounts.
So the receipt for the claim is real for Sarvam's previous generation. The global model that keeps up is Gemini 3 Pro, not the APIs most teams reach for first. When V4 is added to the leaderboard, check it in your languages before you trust the launch post.
Does the price decide it?
Not on its own. For recorded calls, Saaras is fourth of five, a few dollars above Deepgram. For live streaming, it is the cheapest of the five we priced. Here is 10,000 hours of call audio a month at list prices, using AWS Marketplace's dollar rate for Sarvam:
Sarvam charges the same ₹30 an hour to stream as for a batch file, which is why it leads the streaming column, even against Deepgram's current promotional rate; for a voice agent that is the number that counts. Speaker labels cost extra almost everywhere; on Sarvam the rate goes to ₹45 an hour, about $4,740 for 10,000 hours. Gemini 3 Pro, the strongest global model on Voice of India, is billed per token on Google's API, so price it on a sample of your own audio rather than from a rate card.
The bigger cost sits after the transcript. A 14% error rate and a 68% error rate on Tamil do not cost the same once a person has to re-listen to the call, fix the CRM note or redo the claim. Take 3.5 minutes of a reviewer's time at $48 an hour, the assumptions we used in our gate simulation: one re-listened call costs $2.80, the same as about nine hours of Saaras transcription. At a tenth of that hourly rate, it still costs about an hour of transcription. Use your own loaded rate. So the spread in the table above is small next to the share of calls a person has to touch. Price the review minutes per call in each language, not the transcription hour. The job picks the stack, and here the job is the language mix of your callers.
Which mode and which API for which job?
Saaras has three ways in, all at the same base price per hour:

- REST takes up to 30 seconds of audio. Use it for voice notes, commands and short WhatsApp clips.
- Batch takes files up to 2 hours, 20 files a job, and is where speaker labels live. Use it for recorded calls, meetings and claims.
- WebSocket streaming is for live calls and voice agents. Check the model on this path: Sarvam's newer realtime endpoint, with partial transcripts as the caller speaks, accepts only
saaras:v3-realtime, so a live agent there runs v3, not V4. Sarvam ships integrations for LiveKit, Pipecat and the Vercel AI SDK, and SDKs for Python and Node.
Pick the mode by where the text goes next. A live agent reads transcribe or codemix. A regulator or a dispute needs verbatim. An English-only analytics or LLM step takes translate. Anything a customer reads back in chat suits translit. Each request takes one mode, so the compliance copy and the CRM copy are two requests on the same audio, both from one model with no separate translation step. Confirm with Sarvam how a second mode on the same audio is billed.
Know the limits before you design around them. The docs say word-level timestamps are not supported, only chunk-level, so subtitles aligned to each word need another tool. For card numbers and one-time codes, a pause on the call recorder does not stop the live audio going to the speech service, and a filter on the transcript only acts after the provider has heard it. Keep that audio out of the stream instead: pause the speech-to-text feed during card capture, or take the card details through a separate keypad or secure payment channel. Speaker labels come only in batch, so live agent assist cannot split agent from customer from Saaras alone; a stereo recording with one channel per side does it at the telephony layer.
Where does it fit, industry by industry?
Where the caller speaks an Indian language and the audio is a phone line, Saaras has the strongest case. Four patterns cover most of it:
- Live voice agent. Phone or web call, Saaras streaming, then your LLM, then a check before anything is written to the CRM or core system, then speech back to the caller.
- Post-call analytics. Recordings go to batch with speaker labels in
verbatimfor the record andtranslatefor English analytics, then an LLM pulls out intent, promises and complaints into the case record. - CRM notes from every call.
transcribetext on the activity, the English summary on the case, product and branch names passed as keyterms. - Compliance archive.
verbatimonly, a set retention period, and keys you hold.
A global model fits better in four cases. Your calls are non-Indian English only. You need a language Sarvam does not list, Arabic among them. You need word-level timestamps. Or callers speak a Hindi dialect such as Bhojpuri, where Gemini 3 Pro scored best on Voice of India.
Can it work for the UAE?
Partly. It can carry the Indian-language share of a UAE contact centre, which is large, but it cannot carry Arabic, and the data path needs a decision.
The Consul General of India in Dubai put the Indian community at about 4.36 million at the end of 2024, roughly 38% of the UAE's 11.35 million people, with Kerala the largest group. A Dubai real estate desk, remittance line, telecom shop or clinic hears Malayalam, Tamil, Telugu, Hindi and Urdu every day, often mixed with English. That is the audio Saaras was built for.
Arabic is a different job. Saaras lists no Arabic, and Gulf dialects are hard for every model: on the open Arabic speech leaderboard (updated 6 May 2026), the best open model, Meta's omniASR LLM 7B, averages 28.32% word error, and 56.46% on the Casablanca set, which includes Emirati speech. Commercial APIs from ElevenLabs, Google and Microsoft list Arabic, but we found no independent test of them on Gulf phone calls, so put two of them in your test on your own recordings.
Calls that switch mid-sentence between Arabic and Malayalam or Hindi need one more step. Saaras detects only its own 23 languages, so it will misread an Arabic stretch rather than flag it. Split the call into segments, detect the language on each with the Arabic provider or a separate language-ID step, and send each segment to the model that covers it. Test how often that split lands in the wrong place.

Then the data. Sarvam's managed service runs in Azure's Central India region, so audio from a mainland UAE company leaves the country. The UAE PDPL (Federal Decree-Law No. 45 of 2021) allows transfers abroad under Articles 22 and 23. The UAE Data Office has published no adequacy list, so a transfer rests on a contract that binds the recipient to the law's protections, or on the caller's explicit consent. Under Article 29, companies have six months to comply from the day the law's Executive Regulations are issued, extendable once, and those regulations are still unpublished. Health data, under Federal Law No. 2 of 2019, and banking data, under Central Bank rules, sit outside that law under their own rules, which generally keep them in the UAE. DIFC and ADGM run their own regimes.
The way through for regulated calls is Sarvam's single-tenant deployment in your own cloud account, which its trust center offers on AWS, Azure or Google Cloud, placed in a UAE region. Ask Sarvam in writing which Saaras version that deployment runs: its self-hosting docs still describe v3.

What do you check before you sign?
Four things in the paperwork, before any audio moves.
- Retention. Sarvam's product pages say data is deleted after processing by default. Its retention docs say a workspace keeps data indefinitely until an owner sets a period, and that zero retention is available for Model APIs but not yet for Voice Agents. Set the period on day one and get the default in writing.
- Where it runs. Managed service in Central India; single-tenant, on-premise or air-gapped on request. Sarvam lists SOC 2 Type II and ISO 27001 and says it does not train on customer data.
- Which version. V4 is the default on the speech-to-text endpoint. The newer realtime streaming endpoint runs
saaras:v3-realtime, and a private deployment may be v3 until Sarvam says otherwise. Confirm which model each of your paths runs. - The DPDP contract in India. Under the Digital Personal Data Protection Act, 2023, you are the data fiduciary and Sarvam processes on your behalf, so the contract has to carry your purpose, retention and deletion terms. Sarvam's trust center offers a standard data processing agreement covering its role as a processor under DPDP. Your notice and consent to callers stay your job, and so do any sector rules from your regulator.
What would we do first?
Run a two-week test on your own calls before you change anything.
- 01Set the retention period in the Sarvam workspace before any audio goes in, and record the transfer basis if you are in the UAE.
- 02Pull 200 to 300 real calls per language, including the noisy and code-mixed ones.
- 03Run Saaras V4, Gemini 3 Pro and your current provider on the same calls, in the mode each downstream system needs.
- 04Score word error with spelling variants accepted, and separately score names, amounts and dates, which break CRM records.
- 05Price each provider per resolved call, including the minutes a person spends fixing transcripts in each language.

If you want a second pair of hands, send us your language mix, monthly call hours and the systems the transcript feeds, using the brief below. You hear back within 24 hours from the person who signs the work.
If the transcript feeds a voice agent, the next decision is the channel. Our page on voice agents built on Retell AI covers the call side, and our customer service agent page covers the write into the case. For UAE teams, our AI agents in the UAE page covers the local constraints.

