What does Agentforce Voice speak in India today?
Hindi and Hinglish. Salesforce launched Agentforce Voice in Hindi on 19 May 2026 and said it understands "natural code-switching between Hindi and Hinglish." The press release does not call it generally available or beta. Salesforce's own Help page for Agentforce Voice lists Hindi with a "(Beta)" tag and limited support for pronunciation dictionaries.
Three more languages are close. On 11 June 2026, Arun Kumar Parameswaran, who runs sales for Salesforce in South Asia, told Business Today that Tamil, Marathi and Telugu voice were in beta testing and expected "in the coming weeks". They are three of seven Indian languages Salesforce plans to add. When we checked on 10 October, we found no later Salesforce announcement that they had left beta, so treat them as beta until your org's voice language table says otherwise. He said Hindi and English together with those languages would cover 95 to 97% of customer language needs in India. That is Salesforce's estimate. Your caller list may disagree, and in Kerala, Bengal or Gujarat it will.
Here is where each large Indian language stands for voice, on what Salesforce and Sarvam have published:
Indian languages for voice: Agentforce Voice and Sarvam, checked 10 October 2026

Text is further ahead than voice. Salesforce Help lists Malayalam, Marathi, Tamil and Telugu as beta languages for text agents, and adds that "language availability for voice-enabled agents can differ." So a chat agent on WhatsApp may answer in Malayalam while the phone agent for the same service cannot.
Why does one language per agent matter?
Because your phone line is not one language. Salesforce Help is plain about it: "Agentforce Voice agents are monolingual. Each voice agent can only listen and respond in its configured default language." The Allowed Languages setting applies to text agents only, and adaptive language mode, which answers in the caller's language, is not available for voice.
A service line in Chennai or Kochi gets callers who open in Tamil, switch to English for the policy number and finish in Malayalam. The Hindi agent handles Hinglish inside one call. It will not follow a caller from Hindi into Tamil.
So the language decision happens before the agent picks up. There are three ways to make it, and they cost different amounts of time on the call:
- The language on the record. The caller's number matches a contact whose preferred language is stored. Nothing is added to the call.
- An IVR menu. "Press 1 for Hindi." A few seconds, once, before anyone speaks.
- Detection on the first words. The most natural for the caller, and the only one that can guess wrong.
Each language then gets its own agent. For Hindi and English that agent can be Agentforce Voice. For Malayalam, Bengali, Gujarati, Kannada, Punjabi and Odia, Salesforce has published no voice date, so you need a second speech stack or a person.

Where does Sarvam fill in?
In the languages Agentforce Voice has not reached. Sarvam AI, a Bengaluru model company, hears and reads 22 Indian languages and English, and speaks 10 of them plus English. It publishes three pieces a voice agent needs:
- Bulbul v3 turns text into speech in Hindi, Bengali, Tamil, Telugu, Gujarati, Kannada, Malayalam, Marathi, Punjabi, Odia and English, with 8 kHz output, the sample rate phone lines carry. List price: ₹3 per 1,000 characters.
- Saaras turns speech into text in 22 Indian languages and English, including code-mixed speech. List price: ₹30 per audio hour.
- Sarvam Voice Agents packages speech in and out for phone, web and WhatsApp calls in the languages Bulbul speaks.
Bulbul v4 Flash, Sarvam's newer voice model, lists no Malayalam or Odia voices yet, so a Malayalam line stays on v3 for now.
Sarvam is one of several options for the second stack. This post prices it beside ElevenLabs, because both publish list prices and both took part in the same independent test. The Voice of India benchmark, built by AI4Bharat at IIT Madras and Josh Talks, scored speech-to-text models on unscripted phone conversations. Malayalam and Bengali are the two languages this post keeps naming, so here they are:
The Malayalam set is 45.2 hours from 1,090 speakers, and the Bengali set 50.3 hours from 6,014 speakers (Table 1 of the paper). The paper treats 20% word error as a common bar for practical use, and in Malayalam only Sarvam's two models clear it, by about a point. These are Sarvam's earlier models. Saaras V4 has no independent score yet, so its Malayalam and Bengali figures do not exist in public. Our Saaras V4 post covers the Hindi and Tamil results.
That gives a rule for each language: shortlist the stacks whose public score clears your bar, run your own test calls through each, and route the language to the cheapest stack per resolved call that meets your pass mark. If none meets it, that language goes to a person.
Sarvam has two gaps that matter here. Bulbul does not speak Urdu or Arabic, so a Hyderabad Urdu line or a Gulf contact centre needs another voice. And there is no Salesforce connector. Sarvam's docs say its agents reach your CRM, core banking and payment systems through tools. For Salesforce, that tool is something you build: an API user with narrow permissions, the calls that read the case or the loan, and the call that writes the outcome back.
Then the data. Sarvam's managed service is India-hosted by default, on Microsoft Azure's Central India region according to its trust center. The same page lists single-tenant, on-premises and air-gapped deployments, SOC 2 Type II, ISO 27001, a data processing agreement for DPDP, and no training on customer data. Its retention docs say a workspace keeps data until an owner sets a period, and that zero retention is available for Model APIs but not yet for Voice Agents. Set the period before the first call. On the Salesforce side, ask your account team where Agentforce Voice processes and keeps audio for your org, and get both answers in writing.
That integration is most of the project. The speech is the part Sarvam already ships. Our Sarvam page lists the models, prices and limits in one place, and the Agentforce plus Indian-language voice use case shows the fields this pattern writes.
How do the two connect to Salesforce?
Through the record. Agentforce Voice and the Sarvam agent never talk to each other. Both read the same Case, Account or loan record and both write their result to it, so a supervisor sees one history whatever language the caller spoke.

Three rules keep it honest. The Sarvam agent writes only through a dedicated integration user, so its actions show up with their own name in the audit trail. It writes a fixed set of fields, the same ones the Agentforce agent writes, so reports do not split by stack. And it hands off to the same Omni-Channel queue with the transcript attached, so the person who picks up does not ask the caller to start again. To Salesforce, the Sarvam agent is one more system calling its APIs through a named user, inside your org's API limits.
Telephony is part of the build. Calls in the new languages reach the Sarvam agent through your telephony provider, so number routing and the transfer back to a person are written and tested before go-live. How your Service Cloud Voice or contact-centre setup hands a live call to an outside agent and back varies by setup, so it is one of the first questions we put to your Salesforce account team.
What does each language cost, and how fast is a turn?
Two bills, priced in different units. Salesforce's Help Agent pricing bills $2 (400 Flex Credits) for a resolved session and caps a voice session at ten minutes per resolution; on Flex Credits a standard action is $0.10. Our Agentforce pricing post works through both. A second stack bills per audio hour, per character, per token and per phone minute.
The table prices one 1.6-minute call at list prices, with Sarvam's speech beside ElevenLabs' speech. The reasoning model and the phone line are the same in both columns, so the gap is the speech alone:
Three notes on those numbers. Sarvam cut Sarvam 105B to ₹15 in and ₹60 out per million tokens on 10 October 2026, by its changelog, and our Sarvam pricing post uses the same rates. ElevenLabs is running a launch discount on v4 Turbo until 12 October, and the table uses the regular $0.04. And the Voice of India score above is for Scribe v2, not the realtime model priced here, so test the one you would run. If you consider ElevenLabs, ask where it processes and keeps audio, as you would ask Salesforce.
Why Sarvam 105B for the think step? Sarvam lists its sarvam-105b-conversations variant, at the same price, for real-time dialogue and voice agents. The next model down on Sarvam's list, Gemma 4 31B (in beta), costs more per token (₹35 in, ₹90 out), so a smaller model is no saving there. No public test shows either one handling a claim-status call with tool use, so your shadow run has to.
Compare either column with Salesforce only after you add what each leaves out. The Salesforce price includes the reasoning and the actions inside Salesforce, but not your telephony. The columns above leave out the integration build and the person who runs the second stack. Divide each by resolved calls in that language, not by minutes. At 20,000 calls a month, the 62 paise gap between the two columns comes to ₹12,400 (our arithmetic).
Speed is the other number. In conversation, people usually answer each other within about 200 milliseconds, across the ten languages Stivers and colleagues measured. A voice agent will not match that. Vendors publish figures for some steps of a turn and leave the rest to you:
The published parts add up to about 250 milliseconds. The pause before the stack decides the caller has finished, the model and the network are unknown until you measure them, so the turn budget is a requirement you write, not a figure you can look up. We put one in the spec before the shadow run, for example 1,200 milliseconds at the median and 2,000 at the slowest 5% of turns, from the caller's last word to the agent's first sound. Those are starting targets for a test, not results we have measured. Tighten them if callers talk over the agent or hang up. Measure it this way:
- 01Make 50 live test calls per language from ordinary mobiles on two networks, reading a script with names, amounts and dates.
- 02Record both sides in stereo, and log a timestamp at each step in the trace: end of speech detected, final transcript, first model token, first audio sent, and first audio heard on the recording.
- 03Report the median and the slowest 5% for each step and each language. The step that breaks the budget is the one to change.
Language detection adds its own error. Sarvam reports that Saaras V4 picks the wrong language on 2.9% of verified IndicVoices clips across the ten most spoken Indian languages, and 5.22% across all 22. Code-mixed openings on a phone line are a different test, so count misroutes in the shadow run. If detection on the first words adds a delay or misroutes callers, move that language to the record or the IVR menu.
What do you own, and what do you rent?
Own what makes the agent yours. Rent what is easy to replace.
This is what keeps the exit cheap. When Salesforce ships Malayalam voice, you run the same Malayalam test calls through Agentforce Voice, compare against the pass mark, and change one row in the routing table if it wins. When Sarvam changes a price, the same applies to the next vendor. The spec and the test calls are the work you keep.
Name an owner for the second stack in the spec: a person on your team who answers for its prompts, its language list, its incidents and its monthly bill. During the build that is shared with us. After go-live it is yours, with the Sarvam account in your company's name.
When does a language move from shadow to live?
When it passes a mark you wrote down before the shadow run started. We use three stages, each with its own test:
- Shadow. The Sarvam agent hears recorded calls and writes nothing. For each call, compare its outcome code and the values it would write (names, amounts, dates) with what your team logged. Count misroutes from the language step.
- Gated writes. The agent writes, but a person approves each write. Count how many approvals were changed.
- Live. The agent writes on its own for that job and language. A sample of calls is still reviewed each week.
Set three numbers in the spec before stage 1: the outcome match you need, the misroute rate you accept, and the turn time you accept. Then size the sample so the result means something. With 200 calls in a language and 190 matches, the true match rate could sit anywhere from about 91% to 97% (a Wilson 95% interval, our arithmetic). Read the low end against your mark. Any wrong amount or date in a field the agent would write keeps that language in shadow, whatever the overall score.
Two failures to plan for. Beta behaviour: Salesforce Help warns that beta languages "can have some inconsistencies when retrieving information," so the Hindi agent gets the same shadow test before it changes accounts. Spec drift: two agents built in two tools drift apart when one gets a new refund rule and the other does not. Run the same test calls through both agents after every spec change.

What would we do in week one?
Read before we build. We can't name clients, so this post shows the method and the tests rather than a case study. Week one ends with a written assessment you keep whether or not we do the build:
- 01Pull a month of call records and count calls by language and by job, so the first language pair is chosen from your data.
- 02Check the voice language table in your org, and ask your account team how a live call passes to an outside agent and back.
- 03Write the spec for one job, for example Malayalam claim-status calls: what the agent reads, the fields it writes, when it hands off, what it never says.
- 04Pick 200 recorded calls in that language for the shadow run, and write the three pass marks.
- 05Price both stacks per resolved call on your volumes, and settle the hosting route and retention period in writing.
From there our target is one job in one language pair live in about 30 days, shadow period included, in our method: spec, shadow, gate. Add the next language only when the first one holds, with the same spec and the same fields.
The use-case list has more jobs this pattern fits, and our India page covers the RBI, TRAI and DPDP rules a calling agent has to follow. If you want a second pair of hands, send us the job, the languages your callers speak and the Salesforce cloud you run, using the brief below. We put one agent on it with a person on the exceptions, as in First Agent in Production. You hear back within 24 hours from the person who signs the work.

