Speaking twice at Dreamforce · Sept 15-17 →

Voice AI

India · UAE

Sarvam AI speech to text: where Saaras V4 beats global models, and where it does not

By Shivanath DevinarayananPublished 13 min read
Written forCTO, CXO, or VP or director of engineering who owns the agent stack and the model bill“Which model should run which step, and what does each step cost?”

The short answer

When should a company use Sarvam's Saaras V4 instead of a global speech-to-text model?

Use Saaras V4 when your callers speak Indian languages beyond standard Hindi, the audio is phone-grade or code-mixed, or the data must stay in India. Price matters less than the error rate in your callers' languages. At Mindcat, we route each call by language and test on real calls first.

Telephone exchange with coloured cables into one board: Saaras V4 against global speech-to-text, no Arabic

Takeaways

  1. 01Saaras V4 covers 22 Indian languages and English, with five transcript formats from one call.
  2. 02On Voice of India's Tamil phone calls, Sarvam's audio model scored 14% word error. Deepgram Nova-3 scored 68%.
  3. 03For batch files, three of the five services we priced cost less than Saaras. For live streaming, Saaras is the cheapest.
  4. 04Saaras has no Arabic. In the UAE, pair it with an Arabic model and record the basis for sending audio to India.
  5. 05Sarvam's docs keep workspace data until an owner sets a retention period. Set it before the first call.

What is Saaras V4?

Saaras V4 is Sarvam AI's speech-to-text model, released on 25 September 2026. It transcribes the 22 scheduled Indian languages and English, now including non-Indian English accents, and it works out the language itself. Sarvam reports a wrong-language rate of 5.22% across all 22 languages and 2.9% across the ten most spoken.

Inside, an audio encoder feeds a 3-billion-parameter language model that Sarvam trained from scratch. That design lets one model write the same call out five ways:

ModeWhat you getUse it for
verbatimEvery word as spoken, in the native scriptCompliance records and disputes
transcribeNative script, with numbers and dates normalisedCRM notes and live agents
codemixNative script, with English words left in EnglishHinglish and other mixed speech
translitThe whole call in Roman scriptChat, WhatsApp and search
translateAn English translationAnalytics and English-only LLM steps
One call, five transcripts: verbatim, transcribe, codemix, translit and translate, all from Saaras V4
One call, five transcripts.

Two other features matter for real calls. Keyterm prompting takes names, brands and product terms before transcription starts; on AI4Bharat's IndicContextEval benchmark, with the terms supplied, Sarvam reports 16.03% word error, the lowest it found. And streaming starts returning words in under 150 milliseconds, by Sarvam's figure. The docs say it handles 8 kHz phone audio.

Sarvam's own benchmark claims are strong. It reports the best results on all 22 Indian languages and the lowest average word error across seven English test sets, from meetings and earnings calls to Indian-accented English. On noisy Indian speech, it reports less than half the error of Deepgram Nova-3 and GPT-4o Transcribe. All of that is the vendor's own measurement.

If you are building the voice agent around the transcript, our page on the experience plane covers where speech, channel and agent meet.
A railway junction at dawn where one line splits into many tracks, each lit by a lantern of a different colour
One line in, a track for each language.

How does it score on independent tests?

Saaras V4 has not been independently tested yet. Sarvam's earlier models have, and the gap they show is wide.

Voice of India, built by AI4Bharat at IIT Madras and Josh Talks, is 536 hours of unscripted phone conversations from 36,691 speakers in 15 languages. Its scoring accepts legitimate spelling variants, so a model is not punished for writing a word two valid ways. In the paper, Sarvam's audio model had the lowest error rate in 13 of the 15 languages. Google's Gemini 3 Pro won on the Hindi dialects Bhojpuri and Chhattisgarhi. No system stayed under 20% word error in every language.

Tamil shows how far apart the models are:

Word error rate on Tamil phone calls, Voice of India benchmark

Sarvam Audio14.16%
Sarvam Saarika 2.514.86%
Gemini 3 Pro15.72%
Amazon Transcribe19.32%
GPT-4o Transcribe64.22%
Deepgram Nova-367.77%
Source: Voice of India per-language OI-WER table, Josh Talks evaluation report, read 7 October 2026. Sarvam Audio and Saarika 2.5 are Sarvam's earlier models, not Saaras V4

Hindi is closer at the top. On the same Voice of India table, Sarvam's audio model scored 4.95%, Gemini 3 Pro just under 6% and Amazon Transcribe 6.78%, against 12.97% for Deepgram Nova-3 and 33.95% for GPT-4o Transcribe. A separate live Hindi leaderboard from the same team, updated 25 July 2026, puts Saaras v3 second at 3.78%, behind Bodhan's Indic Transcribe Core at 3.48%, with Gemini 3 Pro third at 4.70%.

Two notes on method. Voice of India scored only models that list the language, so Deepgram's 68% on Tamil is a supported language, not a model asked to guess. And "Sarvam Audio", the paper's winner, is a separate Sarvam model from Saaras.

In plain terms, 14% word error is about one word in seven, and 68% is about two in three. On a three-minute call at a typical 150 words a minute, that is roughly 60 wrong words against 300, and the ones that go wrong are often names and amounts.

So the receipt for the claim is real for Sarvam's previous generation. The global model that keeps up is Gemini 3 Pro, not the APIs most teams reach for first. When V4 is added to the leaderboard, check it in your languages before you trust the launch post.

Does the price decide it?

Not on its own. For recorded calls, Saaras is fourth of five, a few dollars above Deepgram. For live streaming, it is the cheapest of the five we priced. Here is 10,000 hours of call audio a month at list prices, using AWS Marketplace's dollar rate for Sarvam:

ProviderBatch rate10,000 hours, batch10,000 hours, streamingIndian languages covered
Sarvam Saaras V4₹30 an hour (about $0.316 on AWS Marketplace)₹3,00,000 (about $3,160)
Deepgram Nova-3 Multilingual$0.0052 a minute$3,120
ElevenLabs Scribe v2$0.22 an hour$2,200
AssemblyAI Universal-3.5 Pro$0.21 an hour$2,100
Google Speech-to-Text v2, standard$0.016 a minute$9,600

Sarvam charges the same ₹30 an hour to stream as for a batch file, which is why it leads the streaming column, even against Deepgram's current promotional rate; for a voice agent that is the number that counts. Speaker labels cost extra almost everywhere; on Sarvam the rate goes to ₹45 an hour, about $4,740 for 10,000 hours. Gemini 3 Pro, the strongest global model on Voice of India, is billed per token on Google's API, so price it on a sample of your own audio rather than from a rate card.

The bigger cost sits after the transcript. A 14% error rate and a 68% error rate on Tamil do not cost the same once a person has to re-listen to the call, fix the CRM note or redo the claim. Take 3.5 minutes of a reviewer's time at $48 an hour, the assumptions we used in our gate simulation: one re-listened call costs $2.80, the same as about nine hours of Saaras transcription. At a tenth of that hourly rate, it still costs about an hour of transcription. Use your own loaded rate. So the spread in the table above is small next to the share of calls a person has to touch. Price the review minutes per call in each language, not the transcription hour. The job picks the stack, and here the job is the language mix of your callers.

Which mode and which API for which job?

Saaras has three ways in, all at the same base price per hour:

Which Saaras API to call: REST for clips up to 30 seconds, Batch for files up to 2 hours, WebSocket for live calls
Three APIs, one base price per hour. Speaker labels, in batch only, cost ₹45 an hour.
  • REST takes up to 30 seconds of audio. Use it for voice notes, commands and short WhatsApp clips.
  • Batch takes files up to 2 hours, 20 files a job, and is where speaker labels live. Use it for recorded calls, meetings and claims.
  • WebSocket streaming is for live calls and voice agents. Check the model on this path: Sarvam's newer realtime endpoint, with partial transcripts as the caller speaks, accepts only saaras:v3-realtime, so a live agent there runs v3, not V4. Sarvam ships integrations for LiveKit, Pipecat and the Vercel AI SDK, and SDKs for Python and Node.

Pick the mode by where the text goes next. A live agent reads transcribe or codemix. A regulator or a dispute needs verbatim. An English-only analytics or LLM step takes translate. Anything a customer reads back in chat suits translit. Each request takes one mode, so the compliance copy and the CRM copy are two requests on the same audio, both from one model with no separate translation step. Confirm with Sarvam how a second mode on the same audio is billed.

Know the limits before you design around them. The docs say word-level timestamps are not supported, only chunk-level, so subtitles aligned to each word need another tool. For card numbers and one-time codes, a pause on the call recorder does not stop the live audio going to the speech service, and a filter on the transcript only acts after the provider has heard it. Keep that audio out of the stream instead: pause the speech-to-text feed during card capture, or take the card details through a separate keypad or secure payment channel. Speaker labels come only in batch, so live agent assist cannot split agent from customer from Saaras alone; a stereo recording with one channel per side does it at the telephony layer.

Where does it fit, industry by industry?

Where the caller speaks an Indian language and the audio is a phone line, Saaras has the strongest case. Four patterns cover most of it:

  • Live voice agent. Phone or web call, Saaras streaming, then your LLM, then a check before anything is written to the CRM or core system, then speech back to the caller.
  • Post-call analytics. Recordings go to batch with speaker labels in verbatim for the record and translate for English analytics, then an LLM pulls out intent, promises and complaints into the case record.
  • CRM notes from every call. transcribe text on the activity, the English summary on the case, product and branch names passed as keyterms.
  • Compliance archive. verbatim only, a set retention period, and keys you hold.
IndustryJobMode and APIWhy Saaras
Banking and lendingCollections and onboarding agents; mis-selling checksStreaming transcribe; batch verbatim with speaker labels
InsuranceClaim intake and renewal calls with an audit trailBatch verbatim, plus translate for the claim summary
Telecom and contact centresTag and score every call; regional-language IVRBatch codemix into analytics; streaming for the IVR
HealthcareDoctor dictation and telemedicine notesStreaming with keyterms for drug names; translate for the record
Government servicesCitizen helplines and grievance intakeOn-premise or air-gapped deployment
Retail and quick commerceVoice orders and WhatsApp voice notes to ticketsREST with translit
Real estateLead qualification calls and site-visit notesStreaming agent; batch with speaker labels
Field service and logisticsVisit notes and proof of delivery by voiceBatch, or Sarvam's offline on-device models

A global model fits better in four cases. Your calls are non-Indian English only. You need a language Sarvam does not list, Arabic among them. You need word-level timestamps. Or callers speak a Hindi dialect such as Bhojpuri, where Gemini 3 Pro scored best on Voice of India.

Can it work for the UAE?

Partly. It can carry the Indian-language share of a UAE contact centre, which is large, but it cannot carry Arabic, and the data path needs a decision.

The Consul General of India in Dubai put the Indian community at about 4.36 million at the end of 2024, roughly 38% of the UAE's 11.35 million people, with Kerala the largest group. A Dubai real estate desk, remittance line, telecom shop or clinic hears Malayalam, Tamil, Telugu, Hindi and Urdu every day, often mixed with English. That is the audio Saaras was built for.

Arabic is a different job. Saaras lists no Arabic, and Gulf dialects are hard for every model: on the open Arabic speech leaderboard (updated 6 May 2026), the best open model, Meta's omniASR LLM 7B, averages 28.32% word error, and 56.46% on the Casablanca set, which includes Emirati speech. Commercial APIs from ElevenLabs, Google and Microsoft list Arabic, but we found no independent test of them on Gulf phone calls, so put two of them in your test on your own recordings.

Calls that switch mid-sentence between Arabic and Malayalam or Hindi need one more step. Saaras detects only its own 23 languages, so it will misread an Arabic stretch rather than flag it. Split the call into segments, detect the language on each with the Arabic provider or a separate language-ID step, and send each segment to the model that covers it. Test how often that split lands in the wrong place.

Route each call by language: detect it, send Indian languages to Saaras V4 and Arabic to a tested model, then the CRM
Route each call by language.

Then the data. Sarvam's managed service runs in Azure's Central India region, so audio from a mainland UAE company leaves the country. The UAE PDPL (Federal Decree-Law No. 45 of 2021) allows transfers abroad under Articles 22 and 23. The UAE Data Office has published no adequacy list, so a transfer rests on a contract that binds the recipient to the law's protections, or on the caller's explicit consent. Under Article 29, companies have six months to comply from the day the law's Executive Regulations are issued, extendable once, and those regulations are still unpublished. Health data, under Federal Law No. 2 of 2019, and banking data, under Central Bank rules, sit outside that law under their own rules, which generally keep them in the UAE. DIFC and ADGM run their own regimes.

The way through for regulated calls is Sarvam's single-tenant deployment in your own cloud account, which its trust center offers on AWS, Azure or Google Cloud, placed in a UAE region. Ask Sarvam in writing which Saaras version that deployment runs: its self-hosting docs still describe v3.

Using Saaras from the UAE: Saaras for Indian languages, a separate Arabic model, a transfer basis, UAE deployment
Using Saaras from the UAE.

What do you check before you sign?

Four things in the paperwork, before any audio moves.

  • Retention. Sarvam's product pages say data is deleted after processing by default. Its retention docs say a workspace keeps data indefinitely until an owner sets a period, and that zero retention is available for Model APIs but not yet for Voice Agents. Set the period on day one and get the default in writing.
  • Where it runs. Managed service in Central India; single-tenant, on-premise or air-gapped on request. Sarvam lists SOC 2 Type II and ISO 27001 and says it does not train on customer data.
  • Which version. V4 is the default on the speech-to-text endpoint. The newer realtime streaming endpoint runs saaras:v3-realtime, and a private deployment may be v3 until Sarvam says otherwise. Confirm which model each of your paths runs.
  • The DPDP contract in India. Under the Digital Personal Data Protection Act, 2023, you are the data fiduciary and Sarvam processes on your behalf, so the contract has to carry your purpose, retention and deletion terms. Sarvam's trust center offers a standard data processing agreement covering its role as a processor under DPDP. Your notice and consent to callers stay your job, and so do any sector rules from your regulator.

What would we do first?

Run a two-week test on your own calls before you change anything.

  1. 01Set the retention period in the Sarvam workspace before any audio goes in, and record the transfer basis if you are in the UAE.
  2. 02Pull 200 to 300 real calls per language, including the noisy and code-mixed ones.
  3. 03Run Saaras V4, Gemini 3 Pro and your current provider on the same calls, in the mode each downstream system needs.
  4. 04Score word error with spelling variants accepted, and separately score names, amounts and dates, which break CRM records.
  5. 05Price each provider per resolved call, including the minutes a person spends fixing transcripts in each language.
Before you switch: set retention, pull 200 calls per language, test against today, score names, price review time
Five checks before you switch.

If you want a second pair of hands, send us your language mix, monthly call hours and the systems the transcript feeds, using the brief below. You hear back within 24 hours from the person who signs the work.

If the transcript feeds a voice agent, the next decision is the channel. Our page on voice agents built on Retell AI covers the call side, and our customer service agent page covers the write into the case. For UAE teams, our AI agents in the UAE page covers the local constraints.

Filed under

Voice AIsarvam aisarvam ai apisarvam api pricinghindi speech to texttamil speech to textspeech to text api
Useful to someone on your team?

FAQ

Questions teams ask next

How much does Sarvam AI speech to text cost?
Sarvam lists speech-to-text at ₹30 an hour, billed per second, for its real-time, streaming and batch APIs, and ₹45 an hour with speaker labels. AWS Marketplace lists the same service at about $0.32 an hour. Enterprise volume, VPC and on-premise pricing are quoted separately.
Does Saaras V4 support Arabic?
No. Sarvam lists 23 languages for Saaras V4: the 22 scheduled Indian languages and English, including non-Indian English accents. A UAE contact centre needs a separate model for Arabic calls.
Is Saaras V4 better than Deepgram or OpenAI for Hindi?
Sarvam's own tests say so, and independent tests back the gap for Sarvam's earlier models. On the Voice of India benchmark, Sarvam's audio model scored 4.95% word error on Hindi phone calls, Deepgram Nova-3 12.97% and GPT-4o Transcribe 33.95%. Saaras V4 itself has not been independently tested yet.
Where does Sarvam process audio?
Sarvam's trust center says its managed service runs on Microsoft Azure's Central India region, and that it also deploys single-tenant in your own cloud account, on-premises, or air-gapped. Its self-hosting docs still describe Saaras v3, so ask which version a private deployment gets.

Sources

  1. 01Sarvam AI: Introducing Saaras V4 (25 September 2026)Accessed 07 Oct 2026
  2. 02Sarvam API Docs: Saaras model pageAccessed 07 Oct 2026
  3. 03Sarvam API Docs: Realtime Streaming Speech-to-Text APIAccessed 07 Oct 2026
  4. 04Sarvam API Docs: PricingAccessed 07 Oct 2026
  5. 05Sarvam API Docs: Speech-to-text FAQsAccessed 07 Oct 2026
  6. 06Sarvam API Docs: Data retentionAccessed 07 Oct 2026
  7. 07Sarvam Trust CenterAccessed 07 Oct 2026
  8. 08AWS Marketplace: Sarvam Models APIsAccessed 07 Oct 2026
  9. 09Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India (arXiv 2604.19151)Accessed 07 Oct 2026
  10. 10Voice of India evaluation report with per-language OI-WER table (Josh Talks)Accessed 07 Oct 2026
  11. 11voiceofindia.ai live speech-to-text leaderboard, Hindi (updated 25 July 2026)Accessed 07 Oct 2026
  12. 12Deepgram pricingAccessed 07 Oct 2026
  13. 13AssemblyAI: Speech-to-Text API Pricing (rates as of July 2026)Accessed 07 Oct 2026
  14. 14IA Decider: Speech-to-Text Cost and Accuracy Reference 2026 (rates checked 19 September 2026)Accessed 07 Oct 2026
  15. 15MarkTechPost: Sarvam AI Releases Saaras V4 (26 September 2026)Accessed 07 Oct 2026
  16. 16UAE Federal Decree-Law No. 45 of 2021 on the Protection of Personal DataAccessed 07 Oct 2026
  17. 17Kayrouz and Associates: Cross-Border Data Transfers Under UAE Law in 2026Accessed 07 Oct 2026
  18. 18MIS Legal: Cross-Border Data Transfers after the UAE Data Protection Law (May 2026)Accessed 07 Oct 2026
  19. 19Telangana Tribune: Indian population in the UAE rises to 4.36 million (Consul General of India, Dubai)Accessed 07 Oct 2026
  20. 20Open Universal Arabic ASR Leaderboard results (updated 6 May 2026)Accessed 07 Oct 2026
Shivanath Devinarayanan, founder of Mindcat

Written by

Shivanath Devinarayanan

Founder, Mindcat Consulting · Salesforce MVP Hall of Fame

Runs 200+ agents in production. Reads every brief that comes in and signs the work that goes out.

About Shivanath

First Agent in Production · the written assessment

We work in real estate and right now, the AI pilot never left the sandbox.

Industry
Problem

Brief: Real estate. Leads from portals and WhatsApp. Follow-up has no owner. Stalled AI or Salesforce pilot. Need a finish, cut, or rebuild call. From: /blog/sarvam-saaras-v4-speech-to-text.