Speaking twice at Dreamforce · Sept 15-17 →

OpenAI and other models

US · Global

LLM market share on OpenRouter, read for the models your company runs

By Shivanath DevinarayananPublished 18 min read
Written forCTO, CXO, or VP or director of engineering who owns the agent stack and the model bill“Which model should run which step, and what does each step cost?”

The short answer

What does OpenRouter's usage data say about which models a company should run?

OpenRouter counts requests and tokens on its own traffic, not users or spend, so it shows what builders run, not enterprise share. Cheap Flash models carry the volume, and decision models are new. At Mindcat, we put a frontier model on hard cases, a workhorse on the repeated step, and a decision model at each gate.

Canal lock gates at dusk, small boats through the small gate: price each agent step before you pick a model

Takeaways

  1. 01OpenRouter says its rankings count requests and tokens routed through it, and none of them measures users or spend.
  2. 02Anthropic had 2.5% of OpenRouter requests in the week of 28 September. A 2025 enterprise survey put it first, at 32%.
  3. 03Five of OpenRouter's ten most-used models that week carry Flash in their name. Cheap models carry the volume.
  4. 04In OpenRouter's 50-question test, a checked cheap draft cost 7% of the frontier model, and two answerable questions went to a person.
  5. 05Free variants can log prompts for training. Set a data policy before any customer text goes near one.

What do OpenRouter's rankings measure?

Traffic through OpenRouter, counted two ways, and nothing else. Whether your agents are still in a pilot or already in production, the next decision is the same: which model does which step, and what checks the step before it writes. OpenRouter's public rankings are a free view of what other builders run instead. Its own page says how to read them: the model charts rank by tokens processed, the market share chart counts requests, and "none of them measures users or spend."

That caveat decides how you use the numbers. In the week beginning 28 September, Anthropic had 2.5% of text requests on OpenRouter. A Menlo Ventures survey of 150 technical decision-makers in July 2025 put Anthropic first, at 32% of production workloads, weighted by each application's scale. The survey is 15 months older, so read the gap as a difference in method, not a trend. Both can be true. Enterprise contracts that go direct to a lab never pass through OpenRouter, and OpenRouter's own State of AI study says enterprise usage and closed internal systems "remain outside the scope of our data."

MeasureWhat it countsUse it for
OpenRouter rankingsTokens and requests routed through OpenRouter. Not users, not spendWhat builders and small teams run, week by week
Enterprise survey (Menlo, July 2025)Share of production workloads, reported by 150 decision-makers, weighted by scaleWhat large companies run in production
Your own invoicesWhat your team pays for, by model and taskThe only number that sets your budget
OpenRouter counts tokens and requests, a survey counts workloads, invoices count spend: Anthropic is 2.5% and 32%
Three ways to count LLM market share.
If you already know which step in your agent should go to which model, our page on routing LLM calls with a decision model shows the gate in practice.
A lit signal box above a switchyard where one track splits into three, a short train and a long freight train waiting
One track, three routes. The gate decides.

Who leads OpenRouter right now?

DeepSeek, Google and OpenAI, with 61.5% of requests between them. In the week beginning 28 September, OpenRouter's market share chart had DeepSeek at 21.8% of text requests, Google at 21.5% and OpenAI at 18.2%. An anonymous stealth provider came fourth at 10.5%, and its requests rose 296% on the week before.

Share of text requests on OpenRouter, week of 28 September 2026

DeepSeek21.8%
Google21.5%
OpenAI18.2%
Stealth10.5%
Z.ai6.6%
Qwen6.2%
Anthropic2.5%
Source: OpenRouter rankings, Market Share chart, read 7 October 2026. Requests, not users or spend

By tokens, the week to 5 October looked like this:

RankModelAuthorTokensChange on the week
1Space Bunny AlphaStealth
2DeepSeek V4.1 FlashDeepSeek
3GLM 5.3 FlashZ.ai
4MiMo-V2.6-FlashXiaomi
5Hy4 previewTencent
6GPT-6 LunaOpenAI
7DeepSeek V4 Flash 0731DeepSeek
8Nemotron 3 Ultra (free)Nvidia
9GPT-5.6 LunaOpenAI
10DeepSeek V4 Flash 0423DeepSeek

OpenRouter's 30-day table marks DeepSeek V4.1 Flash as new, with no usage in the 30 days before, so all of its 72.3 trillion tokens arrived in the last 30 days. OpenRouter's page for Space Bunny Alpha says a third-party provider operates it and "has chosen to remain anonymous during this preview." Anthropic's newest models were climbing fast from a small base: Claude Sonnet 5.5 was on OpenRouter's trending list at 926 billion tokens for the week, up from almost nothing the week before.

Why do cheap and free models carry the volume?

Because most agent work is the same step done thousands of times, and builders send that step to a cheap model. Five of the ten models in the table above carry Flash in their name, and six come from DeepSeek, Z.ai, Xiaomi or Tencent. The free variant of Nvidia's Nemotron 3 Ultra was eighth, at 5.53 trillion tokens.

OpenRouter's own research adds a caution. Its State of AI study, published with a16z in December 2025, found demand "relatively price-inelastic" across the platform: a 10% price cut moved usage only about 0.5 to 0.7%. So a price cut alone does not pull work to a model. Our reading of the table is that builders pick a model per step: a cheap one where the step repeats, and they keep paying for an expensive one where the step is hard.

Free comes with terms. OpenRouter's free router warns that "prompts and completions may be logged by the provider and used to improve the model." Free variants are capped at 20 requests a minute, and at 50 a day until an account has bought 10 credits, then 1,000. For a company, two settings matter more than the cap. Zero data retention is not data residency, so for UAE or India data, check where the provider serves the request as well. You can refuse providers that store data, account-wide or per request, and you can require zero data retention. Space Bunny Alpha's page says prompts "may be retained by the provider but are not used for training", from a provider you cannot name. Neither is a place for customer text.

Free routes: prompts may be logged and used to improve the model, 20 requests a minute, 50 or 1,000 a day
Free tokens come with terms.

What is a decision model, and where does it sit?

A model that returns a typed answer with probabilities instead of prose, priced for the gates in your agent. OpenRouter now lists a Decisions category. Jev 1.13 from TypeSafe went up on 18 September at $0.042 per million input tokens, with output free. Perplexity's Decider V1 27B costs $0.04 per million input tokens, and OpenAI added GPT-6 Luna Decisions on 6 October at $0.10. Liquid AI's d1 uses the same schema.

ModelMakerInput price per million tokensOutput
Jev 1.13TypeSafe$0.042
Decider V1 27BPerplexity$0.04
GPT-6 Luna DecisionsOpenAI$0.10
d1Liquid AINot shown in the category text we read
Decision models on OpenRouter: Jev 1.13, Decider V1 27B, GPT-6 Luna Decisions and d1, priced on input, output free
Decision models on OpenRouter, 7 October 2026.

Each one answers three kinds of question about the state you send: yes or no, one choice from a list you define, or a score on an ordered scale. Your code reads the probability and acts. Nothing needs parsing.

The cost per decision is small enough to put one at every gate. OpenRouter's classification cookbook labelled 550 posts for $0.014, about $0.026 per 1,000 items at 610 input tokens each. That is the job a chat model does badly at volume: it writes a paragraph, you parse it, and you pay for every output token.

Disclosure: Mindcat implements Jev for clients, and we have a page on it. The four decision models above come from four makers. The pattern does not depend on Jev, but the published tests we cite all use it, and we have not tested the other three here.

This is the belief we hold hardest: not every hop is a model call. A router, a policy check or a "should a person see this?" question wants a fast, typed, predictable answer. A chat model gives you prose and a bill for the words.

Where does the gate sit before a write?

Right before the write, with a person behind it. A decision model is cheap enough to ask "is this action what the user asked for?" before every write to your system of record. Above a threshold the write goes through. Below it, the request goes to a person, in Slack or wherever your approvers already work. OpenRouter publishes a cookbook for gating agent tool calls this way: safe calls run, unsupported ones are refused, and only the unclear ones wait for a person.

The gate model sees the same customer text as the rest of the agent, so the data policy above applies to it too. Check its provider's retention before customer data reaches it.

The gate does not replace identity. In Salesforce, the agent should write as its own registered identity with a narrower permission set, which is how Salesforce describes registering agents. And the approval has to bind to the exact action, or an edited request can ride on an old yes, as our test of human approvals shows. When a rep triggers the write, our rule is to bound it by the narrower of the agent's permissions and the rep's. Log the gate's question, its probability, the route it took and the approver against the record, so the audit trail answers "who acted" without a search. Name the person who owns the exceptions queue before the gate goes live, not after. The same gates apply whether the agent runs in Agentforce or in a custom stack. How much model choice you get differs, and that is part of how we pick.

One agent, four tiers: decision model at every gate, workhorse for the repeated step, frontier for hard cases, a person
One agent, four kinds of compute.

What do published gate tests show?

Fast and cheap in every test, accurate in most, and every sample is small. We have no production table of our own to publish, so here are the four public measurements we could find, with their sizes and who ran them. A simulation of our own comes after the worked example.

Test, publisher, sizeAccuracySent to a personLatencyCost per check
Tool-call gate recipe, Agency Enterprise (jev-recipes), 104 cases, recorded 4 October 202698 of 104 graded correct by the recipe, handoffs included (94%); all 76 it decided alone were right28 of 104 (27%)
agent-autoguard, open-source package, 20 held-out calls8/8 safe allowed, 8/8 dangerous caught, 4/4 ambiguous sent to a person4 of the 4 ambiguous
Same 20 calls, Claude Sonnet 5 as the guard7/8 safe allowed, 5/8 dangerous caught, 1/4 ambiguous sent to a person1 of the 4 ambiguous
AmpUp, Claude Code permission gate, 8 situations run 5 times, 18 September 202615 of 40 exact labels (Claude Haiku 4.5: 25 of 40). Neither model allowed a dangerous actionBoth cautious on harmless cases

Read the rows for what they are. The first recipe's labels are mostly AI-written synthetic cases without independent human review, and its makers build on Jev. The only independent test, AmpUp's, found Jev less exact than a small chat model on its labels, though safer than an accidental allow. OpenRouter's own refund-gate cookbook adds that each check cost $0.00003 to $0.00004 and came back in under 600 ms, and that repeating the same request moved the probabilities by up to 0.08 without changing the outcome.

The threshold is where you choose your escalation rate. The recipe publishes the trade-off: at a confidence threshold of 0.5, 6% of calls went to a person and 96% of the rest were right; at 0.8, 27% went to a person and all of the rest were right. That is the same number that decided the bill in the worked example below. A stricter gate is safer and sends more work to people.

Median time per gate decision, in milliseconds

Jev, jev-recipes96 ms
Jev, OpenRouter p50180 ms
Jev, AmpUp228 ms
Jev, autoguard428 ms
Haiku 4.5, AmpUp733 ms
Sonnet 5, autoguard1992 ms
Source: jev-recipes tool-call gate (4 Oct 2026); OpenRouter Jev 1.13 model page p50; AmpUp (18 Sep 2026); agent-autoguard holdout set. Different tests and networks, so compare within a test

What does a verified cascade save?

In OpenRouter's own test, zero wrong answers in both runs at about 7% of the frontier model's cost, with two answerable questions handed to a person instead. The cookbook drafts each answer with a cheap model, asks Jev whether the retrieved text supports the draft, and sends it to the frontier model only when the check fails. On 50 support questions, 27 answerable from the help-center excerpts and 23 not, the results were:

OpenRouter's verified-cascade test, 50 questions

CriterionWrong answers, reported runWrong answers, earlier runCost, reported run
Frontier model on every question20$0.175
Cheap model on every question02$0.004
Cheap draft, Jev check, escalate on fail00$0.012
From OpenRouter's cookbook 'Cut LLM Cost with a Jev-Verified Cascade', read 7 October 2026. Cheap model: GPT Luna. Frontier model: GPT Astra. Graded against hand-written answers.
Cheap model drafts, Jev checks, frontier rewrites on fail, a person takes the rest: 0 wrong at $0.012
The verified cascade, from OpenRouter's cookbook.

The cookbook gives costs only for the reported run, and it does not say which questions the frontier model got wrong. The cheap model alone was cheapest, and in the reported run it was also right every time. In the earlier run it shipped 2 wrong answers. The cookbook calls a 2-answer swing noise, and the cascade was the only setup with zero wrong answers in both runs. That is the case for the check: you pay a little more than the cheap model to stop the bad run from reaching a customer. The cascade sent 2 questions to the frontier model and 2 answerable ones to a person. In one, both tiers were right but Jev's confidence stayed under its 0.8 threshold. In the other, the draft declined to answer. Those two handoffs are a person's time, and the $0.012 does not include it. Fifty questions is a small test, so treat the 7% as a direction to measure on your own traffic.

What does one step cost, worked through?

Here is the sum for a support-reply agent, at OpenRouter list prices read on 7 October and with token counts we made up for the example. Swap in your own counts from a week of logs. The prices: Jev 1.13 at $0.042 per million input tokens with output free, DeepSeek V4.1 Flash at $0.045 in and $1.20 out, and Claude Sonnet 5.5 at $2 in and $10 out. Each run uses 3,000 input tokens and 400 output tokens to draft, 600 tokens for the routing gate and 3,500 for the check. OpenRouter's model pages also show the median time each model takes: 0.18 seconds for Jev 1.13, 1.25 seconds for DeepSeek V4.1 Flash and 2.88 seconds for Claude Sonnet 5.5.

StepModelCost per 1,000 runs
Route the requestJev 1.13$0.025
Draft the replyDeepSeek V4.1 Flash$0.615
Check the draftJev 1.13$0.147
Rewrite the 10% that fail, and check againClaude Sonnet 5.5, Jev 1.13$1.015
Cascade total$1.802
Frontier model on every runClaude Sonnet 5.5$10.000

Using Decider V1 27B at $0.04 for the two gates would change the total by less than a cent. The cascade comes to 18% of the frontier-only bill here, and more than half of it is the 10% that escalate. That rate is our assumption: in OpenRouter's 50-question test, 2 of 50 went to the frontier model and 2 more to a person, so 4 of 50 (8%) left the cheap path, and in the gate recipe above a person saw 6% to 27% depending on the threshold. That is the number to watch. At a 5% escalation rate the cascade is about 13% of the frontier bill, and at 20% it is about 28%. Prices vary by provider on OpenRouter, so read the model page on the day you budget. The gates cost almost nothing. The escalation rate decides the bill.

Two things this example leaves out. It uses list prices with no prompt caching or batch discount; if most of your 3,000-token context is static and cached, the frontier bill drops and the cascade's share rises, so price both at your cache hit rate. And the hard cases get slower: at the median times above, a run that passes the check takes about 1.6 seconds, and one that escalates and is checked again takes about 4.7 seconds against 2.9 for the frontier model alone. If a person waits on the reply, set the escalation path's time limit before you set the threshold.

What does the bill look like once people are counted?

The model bill is the small part. We do not yet measure escalation rate or cost per resolved write on a live agent, so we built a seeded simulation of one and published the script. Anyone can re-run it and get the same numbers. These are simulated numbers, not a client and not a benchmark.

It models a CRM-update agent over 28 days. DeepSeek V4.1 Flash extracts the fields and drafts each write to Salesforce. Jev scores whether the write is correct: 0.88 or above writes, a middling score retries on Claude Sonnet 5 and is scored again, and a low score goes to a person. Every input is an assumption we chose: a 7% draft error rate, rising on harder records, and 3.5 minutes of review at $48 an hour, plus a 2-minute spot check on 2% of automatic writes. Prices are DeepSeek's direct peak rate ($0.30 in, $1.20 out per million tokens, higher than the OpenRouter rate above), Jev at $0.042 in and Sonnet 5 at $2 in and $10 out, with no caching. We left two messy events in: a provider outage on day 9 that sent 40 of 96 runs straight to Sonnet, and a stricter person threshold from day 15 (0.70 raised to 0.75).

WeekWritesTo a personRetried on SonnetWrong ones caughtWrong ones writtenAll-in cost per write
150757 (11.2%)
253059 (11.1%)
348747 (9.7%)
450867 (13.2%)
Total2,032230 (11.3%)

Of the 230 writes sent to a person, 119 would have been wrong and 111 were fine, each costing a person 3.5 minutes. Model spend came to $0.0046 per write, and $0.34 once review time is counted, so the model bill was 1.3% of the cost of running the agent. Sonnet ran on 13.5% of writes, counting retries and the outage, and took 51% of the model spend. Running every step on Sonnet with no gate would cost $0.0184 per write in model spend, four times as much, with nothing to catch its errors. With no gate at all, our assumed error rate puts about 145 wrong writes into Salesforce in the month instead of 8.

The 1.3% comes from our review assumptions. Halve the review time and the model bill is still under 3% of the total. Week 3 shows how little one month tells you: a stricter threshold should send more writes to a person, yet the share went down and the most wrong writes got through. At this volume, chance moves the weekly numbers more than the threshold does. Picking a model per step still matters, because it sets the 8 that slipped through and the 111 false alarms. But the bill is decided by how many runs reach a person and how long they stay there.

Which coding agents are builders running?

A mix, and the top of it is open source, in OpenRouter's apps ranking. Its top coding agents on 7 October were Hermes Agent at 1.52 trillion tokens, Claude Code at 1.25 trillion, Kilo Code at 569 billion, pi at 528 billion and Codex at 508 billion. The first and third, Hermes Agent and Kilo Code, describe themselves as open source. These are OpenRouter tokens only. A team on a direct Anthropic plan runs Claude Code without touching OpenRouter, so the ranking says nothing about Claude Code's total use.

The pattern is the point. A coding agent you control lets you change the model behind each step without changing the agent. That is what makes the cascade possible: the draft step, the check and the escalation are three calls you can point at three different models. If the agent can only call one vendor's model, you cannot run that test at all, and the step stays on the most expensive tier.

Where do sales and marketing teams use this first?

On the calls a person makes on every lead, reply and inbound message. Those are yes-or-no or pick-one questions asked thousands of times a week, which is the shape a decision model is priced for.

Four places we would test first:

  • Lead scoring against your ideal customer profile. One score per lead, with a probability your rules can threshold. We wrote up lead scoring with a decision model.
  • Reading a sales reply. Interested, not now, wrong person, or unsubscribe, so the reply reaches the right rep without someone sorting the inbox first.
  • Checking an outreach message fits the account before it sends, not after a complaint.
  • Tagging inbound at volume. Campaign replies, form fills and support mail, one category and its tags in one call.

The cost signal is OpenRouter's classification cookbook: about $0.026 per 1,000 items at roughly 610 input tokens each. A person still reviews the low-confidence cases, and the threshold is set from a labelled sample of your own data, as that cookbook does with 100 to 200 hand-labelled items. For a CRO or CMO the question is not which model is cheapest. It is which call your team makes most often, and whether a typed answer is enough for it.

What would we do first?

Price the steps before choosing the models. Shivanath has put 200+ agents into production, and the method we use for a first agent is Spec, shadow, gate: write the process down, run the agent in shadow next to the people doing it, then put a human gate on the write. The model choice comes out of the spec, step by step.

  1. 01List the steps in one agent and log tokens and cost per step for a week. Find the one or two steps that carry most of the bill.
  2. 02For each repeated step, test a cheaper workhorse model against your current one on 50 real cases.
  3. 03Put a decision model at each gate: routing, policy checks, and whether a person must see it. Measure its wrong approvals and needless escalations on a labelled set before it goes live.
  4. 04Run the cascade in shadow: cheap draft, check, escalate on fail. Compare wrong answers, cost and the gate's added latency (median and slowest 5%) with today.
  5. 05Set your data policy before any customer text moves: deny providers that store data, and keep free routes for public data.
Five steps: log cost per step, test a cheaper model, add a decision model at each gate, shadow the cascade, set data policy
Five steps before you change a model.

What to own, and what to rent

Rent the models and, if it helps, the router. Own four things: the list of steps with the cost of each, the labelled test sets, the gate thresholds, and the log of every gate decision. Those are what let you swap a model next quarter without starting again. If a router or a vendor holds them, you have traded model lock-in for router lock-in.

None of this needs OpenRouter. It is one way to reach every model in the stack through one bill. The choice that matters is the shape: a frontier model where judgment is hard, a workhorse where the step repeats, a decision model at the gates, and a person on the exceptions. See how we compare Agentforce and a custom stack, and what Jev does in an implementation.

Before you change a model, get three numbers for each step: tokens per run, cost per run, and wrong answers per hundred.

Filed under

OpenAI and other modelsopenrouter rankingsllm leaderboardai model rankingsdecision modelcheapest llm api
Useful to someone on your team?

FAQ

Questions teams ask next

Is OpenRouter market share the same as LLM market share?
No. OpenRouter ranks models by tokens and authors by requests on traffic routed through OpenRouter, and it says none of its charts measures users or spend. Enterprise contracts that go direct to a lab do not appear. Use it as a signal of what builders run, and use surveys or your own invoices for enterprise share.
What is a decision model on OpenRouter?
A model that reads the state you send and returns typed answers with probabilities instead of text: yes or no, one choice from a list, or a score on a scale. OpenRouter lists Jev 1.13, Perplexity's Decider V1 27B, Liquid AI's d1 and GPT-6 Luna Decisions in that category, priced on input tokens with output free.
Are free models on OpenRouter safe for company data?
Not by default. OpenRouter's notes on free routes say prompts and completions may be logged by the provider and used to improve the model. You can deny providers that store data, or require zero data retention, in account settings or per request. Keep customer data off free routes.
Do I need OpenRouter to use a decision model?
No. OpenRouter is one way to reach Jev and the other decision models, through its Decisions API and one bill. TypeSafe also offers Jev directly. The architecture choice, a decision model at each gate in front of your chat models, does not depend on the router you buy it through.

Sources

  1. 01DeepSeek API Docs: Models and Pricing (V4.1 Flash peak and off-peak rates)Accessed 07 Oct 2026
  2. 02OpenRouter: LLM Rankings (usage data through 5 October 2026)Accessed 07 Oct 2026
  3. 03OpenRouter: App and Agent RankingsAccessed 07 Oct 2026
  4. 04OpenRouter: Models in the Decisions categoryAccessed 07 Oct 2026
  5. 05OpenRouter: Jev 1.13 model pageAccessed 07 Oct 2026
  6. 06OpenRouter cookbook: Cut LLM Cost with a Jev-Verified CascadeAccessed 07 Oct 2026
  7. 07OpenRouter cookbook: Classify and Tag Text at Scale with JevAccessed 07 Oct 2026
  8. 08OpenRouter docs: Jev documentation hub (cookbooks, including gating agent tool calls)Accessed 07 Oct 2026
  9. 09OpenRouter: Free Models RouterAccessed 07 Oct 2026
  10. 10OpenRouter docs: API credit and rate limitsAccessed 07 Oct 2026
  11. 11OpenRouter docs: Provider logging and data retentionAccessed 07 Oct 2026
  12. 12OpenRouter: Space Bunny Alpha model pageAccessed 07 Oct 2026
  13. 13OpenRouter and a16z: State of AI, an empirical 100 trillion token study (December 2025)Accessed 07 Oct 2026
  14. 14Agency Enterprise, jev-recipes: Gate an agent tool call (measured accuracy, recorded 4 October 2026)Accessed 07 Oct 2026
  15. 15agent-autoguard v0.4.0 on PyPI (held-out evaluation)Accessed 07 Oct 2026
  16. 16AmpUp: A Faster Permission Gate for Claude Code (18 September 2026)Accessed 07 Oct 2026
  17. 17OpenRouter cookbook: Gate Agent Tool Calls with JevAccessed 07 Oct 2026
  18. 18OpenRouter: Claude Sonnet 5.5 vs DeepSeek V4.1 Flash (prices and p50 latency)Accessed 07 Oct 2026
  19. 19Menlo Ventures: 2025 Mid-Year LLM Market UpdateAccessed 07 Oct 2026
Shivanath Devinarayanan, founder of Mindcat

Written by

Shivanath Devinarayanan

Founder, Mindcat Consulting · Salesforce MVP Hall of Fame

Runs 200+ agents in production. Reads every brief that comes in and signs the work that goes out.

About Shivanath

First Agent in Production · the written assessment

We work in real estate and right now, the AI pilot never left the sandbox.

Industry
Problem

Brief: Real estate. Leads from portals and WhatsApp. Follow-up has no owner. Stalled AI or Salesforce pilot. Need a finish, cut, or rebuild call. From: /blog/llm-market-share-model-stack.