What does Sarvam charge?
Sarvam AI publishes rupee list prices for each model, pay as you go. These are the ones a voice agent or a document workflow uses, as we read them on Sarvam's pricing page on 10 October 2026:
Prices updated on 10 October 2026. That day Sarvam cut Sarvam 105B from ₹29.28 in, ₹10.98 cached and ₹73.20 out per million tokens to ₹15, ₹5 and ₹60, according to its docs changelog. Every figure in this post now uses the new rates. When we checked, the marketing pricing page still showed the old ones, and the docs pricing page showed the new.
Two things are missing from that page. Telephony is not on it, because Sarvam's voice agents run on a carrier you bring or a number you rent. And there is no per-minute price for Sarvam's packaged Voice Agents product. If you are offered one, get it in writing with what it includes, then compare it with the arithmetic below.
A note on who is doing the arithmetic. We are a Salesforce consulting partner, not a Sarvam partner, and you would hold the Sarvam account in your own name. We can't name clients, so every figure here is a list price, a cited source or an assumption we label. The method works the same for any speech vendor.
What does one connected call cost in models?
About ₹3.60, on the assumptions below. We built the call from Sarvam's list prices and a typical Indian outbound call, and we show every input so you can swap in your own.

Speech output is the biggest line, about two thirds of the total. The language model is about a ninth.
For comparison, India Fintech puts a cascaded voice agent at about ₹4.68 a minute at list prices, and a fully wrapped vendor product at about ₹10.3. Those figures cover a fuller stack than our three lines. The gap between ₹2.25 and ₹4.68 is roughly what telephony, hosting and the orchestration layer add, and you should price those from your own quotes.
Why is the minute the wrong unit?
Because nobody buys minutes. A collections head buys promises to pay. A service head buys closed cases. A minute that ends in "please call back later" costs the same as one that ends in a payment date.
So divide by what the call achieved. For a reminder, we count two things: a promise captured (a date and amount on the loan record, known the same day) and a promise kept (the payment arrived, known after the due date). The table below uses promises captured, by the agent or by the telecaller it handed to. Here is the same worked example for 1,000 connected calls:
1,000 connected calls: AI agent with hand-offs against telecallers only (our arithmetic)
Read the last two agent columns. If the agent resolves 40 calls in 100 instead of 60, cost per resolved call rises by half. If, as usually happens, the weaker agent also hands off twice as many calls, it lands at ₹16.50 to ₹21: above the telecaller team at ₹5 a minute and level with it at ₹8. Nothing on Sarvam's rate card changed. The 60% and 40% are assumptions to show the swing, not measurements; your own resolution rate is the number to find first.
Kept promises divide again. If 70 in 100 captured promises are paid (our assumption), the ₹8.50 to ₹10 becomes about ₹12 to ₹14 per kept promise. Apply the same divisor to the telecaller column only when your data shows the same kept rate for both.

How does it compare with a telecaller?
Cheaper, by less than vendor decks suggest. India Fintech puts a human telecaller at ₹5 to ₹8 per connected minute and the saving from voice AI in India at 1.5 to 3 times. In its words, "In India the wedge is not 10x." Exotel, citing Caller.digital, puts a human-resolved contact at ₹40 to ₹120 and an AI-resolved one at ₹12 to ₹25.
Exotel's per-contact figures overlap the top of our range and run higher because they carry more: telephony, platform fees, supervision and the cost of calls that need several attempts. Our table counts models and hand-off minutes only. Use it to see what moves, and use your own full cost to decide.
Replace the ₹5 to ₹8 with your own loaded figure before you compare. Add a month's salaries, incentives, the supervisors' share, seats and telephony for the calling team, then divide by the connected minutes they logged that month. That is your cost per connected minute, and it already carries the hours your team cannot call.
What does it cost a month at your volume?
Multiply, then add what the unit leaves out. On the same assumptions (₹3.60 in models and 15% hand-offs at ₹5 to ₹8 a minute per connected call):
At every volume the people on hand-offs cost more than a third as much as the models, so the hand-off rate is a budget line, not a quality metric. Outbound adds the dials: India Fintech reports connect rates of 46 to 48%, and your carrier bills for the calls that ring out. The models cost nothing on a call nobody answers, but the telephony does.

How do you measure resolution before you trust the number?
On recorded calls, per language, with the sample size written next to the result. This is the test we run before any cost model goes to finance:
- 01Write the outcome codes first: promise captured, already paid, dispute, hardship, wrong person, call back, handed off. One code per call.
- 02Pull 200 recorded calls per language, including the noisy and code-mixed ones, and have your team code each call.
- 03Run the agent in shadow on the same calls. It writes nothing. Code its outcomes the same way, and have a second person check every disagreement.
- 04Report the match rate and the resolution rate with a range, not a point.
The range matters. If the agent resolves 120 of 200 calls, the true rate could be anywhere from about 53% to 67% (a Wilson 95% interval, our arithmetic). On the first column of the table, that moves cost per resolved call from ₹8.50 to ₹10 to anywhere between about ₹7.60 and ₹11.30. Budget on the low end of resolution. The same width limits model comparisons: at 200 calls each, two models less than about ten points apart cannot be told apart (our arithmetic), so compare them on more calls or treat them as equal and pick on price.
Language is where resolution splits. On the Voice of India benchmark, 536 hours of unscripted phone calls from 36,691 speakers in 15 languages, Sarvam's audio model scored 4.95% word error on Hindi and 14.16% on Tamil. That is roughly one word in twenty against one in seven, and the wrong words are often names and amounts. These are Sarvam's earlier models; our Saaras V4 post has the full table. Expect a separate resolution rate, and so a separate cost, per language.
When does the choice of model matter?
Less than the hand-off rate, on these numbers. A hand-off costs ₹10 to ₹16 in telecaller time. The whole language-model line is ₹0.40 a call. Suppose a cheaper model halves it (our assumption): you save about ₹0.20 a call, and one extra hand-off every 50 to 80 calls wipes that out. So we pick the model on resolution in the shadow run, then on price.
The line does grow on longer calls. If each turn adds 400 tokens of history instead of staying flat at 3,000, the think line rises from ₹0.40 to ₹0.57, about 15% of the call. Sarvam lists cached input at ₹5 per million tokens; if 2,000 of each turn's tokens are the same instructions and the cache serves them, that ₹0.57 falls to about ₹0.41. Check the cost per turn on your longest calls, not your average one.
Own the parts you would need to swap a vendor: the outcome codes, the 200 test calls per language, the prompts and the Salesforce integration. Rent the models. Sarvam 105B is open weights under Apache 2.0, so self-hosting is an option, but price the servers and the people against ₹0.40 a call before you take it.
What else goes on the scorecard?
Three more numbers sit beside the cost.

- Agent speech. Bulbul bills per character, so long replies cost the most. Render fixed lines (the greeting, the recorded-call notice, the closing) once and replay the audio. In our example, taking 300 fixed characters out of each call's 800 drops speech from ₹2.40 to ₹1.50.
- Turn time. Sarvam reports streaming speech-to-text returning words in under 150 milliseconds, but a turn adds the model, the speech back and the phone line. In the shadow run, time the gap from the borrower's last word to the agent's first sound, at the median and at the slowest 5% of turns. Long pauses show up as hang-ups, and hang-ups show up as lower resolution.
- Data terms. Sarvam's managed service is India-hosted by default, on Azure's Central India region according to its trust center, which lists SOC 2 Type II, ISO 27001, a DPDP data processing agreement and no training on customer data. Its retention docs say a workspace keeps data until an owner sets a period, and zero retention is not yet available for Voice Agents. Set the retention period before the first call.
Documents follow the same logic. Sarvam's Document AI extracts fields at ₹1 a page, so a 12-page bank statement costs ₹12. The real cost is the reviewer's minutes on low-confidence fields; our loan document extraction use case shows which fields we would send to a person.
What would we do in week one?
Week one builds your number. It ends with a written assessment you keep:
- 01Collect a month of call logs: dials, connects, minutes, outcomes by language, and your loaded telecaller cost per connected minute.
- 02Write what counts as resolved for one job, and the fields the agent writes to the loan record: outcome, promise date and amount, recording link, next attempt or a task for a person.
- 03Pick the 200 calls per language for the shadow run, price the reviewer hours to code them, and set the resolution and hand-off rates the business case needs to hold.
- 04Fill in this post's tables with your figures and quotes for telephony, and agree the hosting route and retention period in writing.
From there our target is one job in one language pair live in about 30 days, shadow period included, in our method: spec, shadow, gate. The EMI reminder use case is one job where "resolved" is easy to define, and the use-case list has others. For a reminder, the agent calls between 8 a.m. and 7 p.m. (RBI's window for recovery calls, which we apply to reminders by choice), from a 1600-series number if RBI regulates you, and hands disputes and hardship to a person the same day. Our India page covers the full set of calling rules. If you want a second pair of hands, send us the job, the languages and your monthly call volume, using the brief below. We put one agent on it with a person on the exceptions, as in First Agent in Production. You hear back within 24 hours from the person who signs the work.

