Speaking twice at Dreamforce · Sept 15-17 →

OpenAI and other models

US · Global

Agentic AI news, sorted for the team that runs the agents: route, approve, host or skip

By Shivanath DevinarayananPublished 14 min read
Written forCTO, CXO, or VP or director of engineering who owns the agent stack and the model bill“Which model should run which step, and what does each step cost?”

The short answer

What does this week's agentic AI news change for a team running agents in production?

Three decisions moved: how you route each step, who approves an agent's action, and where customer data is processed. OpenAI's Decisions API, Codex Auto-review, Mistral Large 4 and the Zapier SDK each touch one. The maths papers and the Photoshop clone are research. At Mindcat, we change nothing in production until a shadow test says so.

Harbour signal tower sorting boats into three channels: route each step, approve the write, choose where data runs

Takeaways

  1. 01OpenAI says its Decisions API answers in about a tenth of the Responses API's time on the same model, GPT-6 Luna.
  2. 02Red-teamers misled OpenAI's Auto-review, which OpenAI says approves about 99% of its own traffic, not yours.
  3. 03Mistral Large 4 weights arrive at the end of October. Self-hosting is the one route that puts the region in your hands.
  4. 04Zapier's SDK applies your policies to pre-built actions only. Direct .fetch() calls are not yet governed.
  5. 05PhotoCraft wires 100% of Photoshop's menus, and its own roadmap puts real parity well below 50%. Measure behaviour, not wiring.

What changed, sorted by decision?

Three decisions moved this week: how you route each step, who approves an agent's action, and where customer data is processed. Most of the rest is research. We started from a creator's daily news summary, then read every vendor's own page. Where the two disagreed, the vendor page wins, and the corrections are listed below the table.

Disclosure: Mindcat builds client agents on Claude and on OpenAI models, and we implement Jev, TypeSafe's decision model.

ReleaseDecision it touchesOur call
OpenAI Decisions API, public betaRouting each stepTest on one gate this month
Claude Code 2.1.292, effort per sub-agent callRouting each stepUse it if you build on Claude Code
OpenAI API tiers cut to Build, Launch, Grow (Grow at $500 paid, was $1,000)Rate limitsCheck your tier and monthly limit
Codex Auto-review, free on ChatGPT sign-inWho approves an actionFine for sandboxed dev work
Claude in Google Docs, Sheets and Slides, betaWho approves an actionPin the default if you allow it
Mistral Large 4, preview API, weights at month endWhere data runsWatch; test when weights ship
EmbeddingGemma 2, 740M, on-deviceWhere data runsTest for on-device search and routing
Anthropic and OpenAI cyber access tiersWho may use cyber capabilityOnly if your team does exploit or malware work
Zapier SDK, open betaAgent identity across appsBlock .fetch() until governed
OpenAI maths manuscripts; PhotoCraftNoneResearch. Skip
Four cards: route each step, approve the action, choose where data runs, and skip the research for now
This week, sorted by the decision each release touches.

Corrections to the week's coverage, from the vendors' own pages:

  • The product is the Decisions API, and OpenAI's guide puts its speed at about ten times the Responses API. It runs on GPT-6 Luna, so the comparison is Luna against Luna, through a narrower endpoint.
  • Mistral's live announcement says 52 billion active parameters. Early copies of the same text said 49 billion. The cyber figures in circulation ("18 of 19 challenges", a 12-minute malware run) are not on Mistral's page, so we left them out.
  • PhotoCraft is not "about 90 percent" of Photoshop. Its own roadmap says real parity "is still well below 50%".
If you already know which steps in your agent should go to which model, our page on routing LLM calls with a decision model shows the gate in practice.
A lamplit dispatcher's desk above a rail yard where one track splits three ways to a shed, a gated depot and a warehouse
One switch, three routes. Most of the work is deciding which.

Which model should run each step now?

A typed endpoint for the yes, no and pick-one steps, and a chat model only where the step has to write. OpenAI's Decisions API answers three kinds of question about text or images: the probability a condition is true, one choice from a list you define, or a score on ordered levels. OpenAI's own guide draws the boundary: use Structured Outputs or function calling when you need generated fields, written explanations or tool arguments.

The facts that matter for your bill and your contract:

  • gpt-6-luna is the only model, on POST /v1/decisions. Public beta; OpenAI expects general availability "in the coming weeks".
  • $0.10 per million input tokens, with no charge for output or caching. OpenAI's pricing page adds 10 percent to the price of regional processing (data residency) endpoints for models released since 5 March 2026.
  • Zero data retention and HIPAA use for eligible customers. Data residency and regional processing in the United States and Europe only.
  • Independent questions can share one request. Questions that depend on an earlier answer need separate requests.

What it costs, worked through at list price with no caching: a 600-token gate question costs $0.00006, or $0.06 per 1,000 checks. Decision models on OpenRouter list close to that: $0.042 per million input tokens for Jev 1.13 (about $0.025 per 1,000 checks) and $0.05 for Upstage's Solar Decide Flash. We have not compared their accuracy here, and Jev is the one we implement, so test any of them on your own labels before you pick.

At these prices the gate is not your bill. In the seeded simulation we published with our per-step cost breakdown, a CRM-update agent spent $0.0046 a write on models and $0.342 a write in all, because people reviewed about 11% of writes. Those are simulated numbers from our stated assumptions, not a client's. The escalation rate decides the bill. Pick a decision endpoint on latency, on accuracy against your labelled set, and on where it can process your data. OpenAI's guide gives the speed as a ratio, not in milliseconds, so measure your own median and slowest 5%.

This is the belief we hold hardest: not every hop is a model call. A routing step wants a typed answer you can threshold and log. OpenAI's guide says the same in its own words: "Use labeled examples from your application to set thresholds for routing, filtering, or review."

Lock-in has an exit here. Vercel's AI Gateway put an OpenAI-compatible /v1/decisions endpoint in front of several decision models, Jev and Liquid's d1 among them, on 7 October. If you write your gates to that request shape, you can swap the model behind a gate without rewriting it. A gateway is one more processor of your data, so read its retention terms as well.

Anthropic shipped a smaller routing control. Claude Code sub-agents could already carry an effort field in their definition. Version 2.1.292 adds an effort parameter per call, so Claude can run one delegated task at low effort and another at high. Levels run from low to max, depending on the model. If your team builds with Claude Code, set effort per sub-agent the same way you set a model per step: cheap where the task repeats, high where it is hard.

Who approves an agent's action?

A second agent for actions you can undo, and a person for writes that move money, change access or leave the company. A second agent is a good filter and a weak signature. The receipts come from the vendors themselves.

OpenAI made Codex Auto-review free for users signed in with a ChatGPT account on 6 October, and said it no longer draws on plan usage. Auto-review sends an approval request to a separate reviewer agent instead of a person. OpenAI's numbers from April: sessions stop for a person "roughly 200x less often", and the reviewer "approves around 99%" of the requests it sees. In one internal snapshot, 720 actions that would have interrupted a user were reviewed automatically. Seven were rejected. Four of those continued by a safer path, and three stopped for the user.

The same page states the limits. Red-teamers "identified cases where Auto-review could be misled into approving commands without user approval", and OpenAI says it "should not be treated as a guarantee of security". The docs add that it is "a reviewer swap, not a permission grant": it does not widen what the sandbox allows. OpenAI publishes no task mix or sample size beyond that snapshot, so treat 99% as OpenAI's traffic, not yours.

Anthropic's new Claude add-on for Google Docs, Sheets and Slides takes the opposite default. "Ask before edits" is on unless a user switches to "Accept all edits". Even then, it stops before it adds a link, inserts a web image or removes someone's access, with no "always allow" option. The mode switch sits in each user's sidebar, so decide your rule before you allow the install.

Two legal texts put this in writing. Article 14 of the EU AI Act requires high-risk AI systems to be built so that people can effectively oversee them, including the ability to override the output or stop the system. In the DIFC, the Commissioner's guidance on Regulation 10, which covers personal data processed by autonomous and semi-autonomous systems, holds the deployer accountable for how the system processes that data. Neither text tells you which approvals need a person. Both expect you to show who could step in.

Our own back-office agents follow the person-on-the-write rule. Agents that handle bank transfers, vendor payments, hiring requests and vendor onboarding ask a person before they act. Approval runs in steps: the submitter's manager, then the department head, then finance and the CEO when it touches budget. The approver sees the request, who submitted it, the reason and the cash impact. The request lives in a Slack List until the last approval, and only then writes back to Salesforce, which stays the system of record. Framing each process took days of discussion. Building each one took about a day. Over six months, one of our approval agents took 40 requests, and 12 of them reached a person. What we missed at first: contractors, external approvals, external reviews and auditing, all added after go-live. What changed because of it: we now write down the process, its exceptions and its scenarios before we automate any of it. This is back-office volume, not thousands of writes an hour, but the rule scales: a second agent can sort the queue, and a named person signs anything that moves cash. We tested what happens when an approval does not bind to the exact action in a $20 approval that ran as $2,000.

A second agent approves about 99% and can be misled; a named person signs anything that moves money
A second agent filters. A named person signs.

Does an open-weight model change where data runs?

Only when you host it. A hosted API runs where its vendor serves it. Open weights let you pick the region, at the price of running the model yourself.

Mistral Large 4 is in public preview on Mistral's API. Mistral describes it as 1 trillion parameters with 52 billion active, natively multimodal, with weights due by the end of October. Its docs list a 1.05 trillion total and a 1M-token context. Mistral says the model will run in multiple regions, "including a European deployment that Mistral operates end-to-end", and that security teams will be able to run it on private cloud or on-premise. The docs show preview prices of $0.68 per million input tokens and $2.09 per million output, marked as a sale from $1.36 and $4.18. Mistral chose the benchmarks on its page, some of them run by third parties, and the licence for the weights is not on the page yet.

At the small end, Google released EmbeddingGemma 2 under Apache 2.0: 740 million parameters, text, images, audio and video in one embedding space, and about 191MB of RAM for the text-only weights on a Pixel 11 Pro. Google's edge team also shows it matching inputs against labels for on-device routing with no training data. For search over files that must not leave a laptop or a phone, that is a model you can run today.

OptionWhere it runsWhat to check
OpenAI Decisions APIUS and Europe for data residencyZDR eligibility and the regional premium
Mistral Large 4 preview APIMistral's infrastructure; a European deployment it runs itselfRetention terms for the preview
Mistral Large 4 open weights (end of October)Wherever you host themLicence, GPU cost, who patches it
EmbeddingGemma 2On the device or your serversRecall on your own data before you trust it
OpenAI Decisions: US and Europe. Mistral API: its servers. Open weights: where you host. EmbeddingGemma 2: on device
Where each option can process your data, per the vendors' pages.

For data that has to stay in the UAE or India, none of the four pages above names those regions. Check whether your cloud provider serves the model you want in a region there, which is a separate contract we did not price here. Self-hosting open weights is the other route, and it is not small: 1.05 trillion parameters is about 1 TB of weights at 8 bits each, before any memory for context (our arithmetic). That is a multi-GPU server, not a spare box.

Security teams face the same split. Anthropic expanded its Cyber Verification Program to three tiers (Defense, Red Team and Specialized Access) covering Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1, with data retention required so Anthropic can monitor misuse, and phishing-resistant MFA with no API keys for Defense Access by 15 December 2026. OpenAI's equivalent is Daybreak, Blue on GPT-5.6 Sol and Red on GPT-5.6 Cyber. Mistral offers open weights under your own policies. With the verified tiers, the vendor watches the traffic. With open weights, you do.

What does one SDK into 9,000 apps do to identity?

It moves the question from "can the agent reach the app" to "whose account is it using, and who sees what it did". The Zapier SDK went to open beta, free during early access. Zapier's docs describe "any API call, on behalf of a user, with no OAuth setup required" across 9,000+ apps, with credentials kept away from the model and your code.

The governance line is in Zapier's own docs. App and action restrictions you set in Zapier apply to the SDK's pre-built actions. Direct API access through .fetch() "currently falls outside those policy controls", and Zapier lists that governance as in progress.

Our rules for any connector layer like this, Zapier or not:

  • Give each agent its own connection, not a person's. Otherwise the audit log names someone who never acted, and you cannot revoke one agent without cutting off a person.
  • Block or wrap the raw API path until your policies cover it: expose only the pre-built actions to the agent, or deny the route at your egress proxy.
  • Name the owner of the exceptions queue before the first write, not after the first bad one.
Three rules: one connection per agent, block raw API calls until governed, name the exceptions owner first
Three rules for any connector SDK.

The video we started from carried a paid Zapier segment. Everything in this section comes from Zapier's documentation instead.

Which headlines are research, not a purchase?

The maths papers and the Photoshop clone. Both are worth reading. Neither changes what you run on Monday.

OpenAI published a repository of mathematics manuscripts from an unreleased internal model: 722 at release and 719 in the README on 9 October, grouped into 372 families. The model was posed about 4,000 problems, and each result used on average about three hours of ChatGPT Pro thinking compute. The README says about 42% of top-line results are formalised in Lean, and that "some of the unformalized results could have issues". The model is not available to you, and three hours a result is the opposite of a cheap step.

PhotoCraft, an open-source Photoshop reimplementation in Rust, carries the lesson for anyone who runs evals. Its planning file measures remaining work in Claude Opus 5.5 agent-hours. Its own numbers, side by side:

MeasureResult
Menu items wired to a command100%
Weighted feature estimateAbout 82%, which the roadmap calls an upper bound
Feature surface a typical user touchesRoughly 60 to 70%
A professional could switch for daily workRoughly 25 to 35%
Real Photoshop parityWell below 50%

The roadmap names the gap itself: users found 214 shortcut failures and broken basics that "all counted as live" in the wiring report. Your agent evals have the same trap. A tool call that runs is not a task that succeeds, so score the outcome a person would accept.

Google's Nano Banana 2.1 image model ($0.0336 per 1K image on the Gemini API) belongs here for most agent teams too. The speed comparisons going around come from one creator's test.

What would we do this week?

Nothing in production until a shadow run says so. Shivanath has put 200+ agents into production, and the method we use is Spec, shadow, gate: write the step down, run the change in shadow next to today's path, then put a person on the write.

The pattern from our own back-office flows, without client data: the model was never the expensive part. Writing each process down took days of discussion, and building it took about a day. Every gap we found after go-live was a missing case in the process, such as a contractor, an outside reviewer or an audit trail, not a wrong model answer. So before you spend a week comparing models for a step, spend it writing down who approves, what they see and what happens when the request does not fit. That page is what any model swap gets tested against.

  1. 01List the yes, no and pick-one steps in one agent. Those are your candidates for a decision endpoint.
  2. 02Shadow one gate on the Decisions API or another decision model for a week. Compare accuracy on your labels, latency and cost.
  3. 03Write your approval rule: which actions a second agent may approve, and which need a named person.
  4. 04Map where each model processes customer data. Mark any step that needs a region no hosted option names.
  5. 05Block raw API paths in connector SDKs until your policies cover them, and give each agent its own identity.
Checklist: list yes or no steps, shadow one gate, write approval rules, map data, give each agent an identity
What we would do this week.

Before you change a model for any step, get three numbers: wrong answers per hundred, median latency, and the share of runs that reach a person.

The rule we suggest, so you can argue with it: move a step to a cheaper or faster model only if, in a shadow run on the same labelled requests, wrong answers per hundred are no higher than today's, the slowest 5% of answers still fits the time the step allows, and the share that reaches a person does not go up. If any one of the three gets worse, keep today's path. Run enough requests that one miss cannot swing the result: at 40 requests, the size of our approval agent's six months, a single wrong answer moves the rate by 2.5 points.

What to own and what to rent: rent the models, behind an endpoint you can swap. Own the labelled requests and the shadow test, because those are what let you switch.

Filed under

OpenAI and other modelsopen weight modelsmodel routingai agent costmistral large 4openai decisions api
Useful to someone on your team?

FAQ

Questions teams ask next

Is the OpenAI Decisions API the same as a decision model like Jev?
It is the same shape: you send text or images and a list of questions, and get back a probability, a pick from your options, or a score. OpenAI's version runs only gpt-6-luna, at $0.10 per million input tokens with no output charge, and is in public beta. Jev, Liquid's d1 and others answer the same question types through other providers, which gives you a second option if one model gets your labels wrong.
Should a second agent approve my agents' actions?
For actions you can undo inside a sandbox, a second agent cuts interruptions. OpenAI reports about 200 times fewer stops with Auto-review. For writes that move money, change access or leave the company, keep a person. OpenAI itself says Auto-review is not a guarantee of security, and its red-teamers found cases where it approved commands it should not have.
Does Mistral Large 4 solve data residency for UAE or India data?
Not on its own. Mistral's page names a European deployment it operates end to end and says the model will be available in multiple regions. The weights are due at the end of October, and self-hosting them is what lets you pick the region. Check the licence and the hosting cost when the weights ship.
Is the Zapier SDK safe for production agents?
It is in open beta and free during early access. Zapier keeps credentials away from the model and your code, and your Zapier app and action restrictions apply to pre-built actions. Direct API calls made with .fetch() are not yet covered by those policies, so block or wrap them until Zapier ships that governance.

Sources

  1. 01EU AI Act Service Desk: Article 14, Human oversightAccessed 09 Oct 2026
  2. 02DIFC Commissioner of Data Protection: guidance on Regulation 10 (autonomous and semi-autonomous systems)Accessed 09 Oct 2026
  3. 03OpenAI API docs: Decisions (public beta, pricing and data controls)Accessed 09 Oct 2026
  4. 04OpenAI Developer Community: 28 days of Shipping at OpenAI (day 2 tracker)Accessed 09 Oct 2026
  5. 05OpenAI API pricing (regional processing uplift)Accessed 09 Oct 2026
  6. 06OpenAI API docs: Rate limits and usage tiers (Build, Launch, Grow)Accessed 09 Oct 2026
  7. 07OpenAI Codex docs: Auto-reviewAccessed 09 Oct 2026
  8. 08OpenAI Alignment: Auto-review of agent actions without synchronous human oversight (30 April 2026)Accessed 09 Oct 2026
  9. 09Claude Help Center: Use Claude in Google Docs, Sheets, and SlidesAccessed 09 Oct 2026
  10. 10Claude Code docs: Subagents (effort field and per-call effort parameter)Accessed 09 Oct 2026
  11. 11Claude Code changelog (2.1.292)Accessed 09 Oct 2026
  12. 12Anthropic: Expanding the Cyber Verification Program (6 October 2026)Accessed 09 Oct 2026
  13. 13Claude Help Center: Cyber Verification ProgramAccessed 09 Oct 2026
  14. 14OpenAI Help Center: Daybreak, Trusted Access for Cyber overviewAccessed 09 Oct 2026
  15. 15Mistral AI: Introducing Mistral Large 4 (6 October 2026)Accessed 09 Oct 2026
  16. 16Mistral docs: Mistral Large 4 model card and pricesAccessed 09 Oct 2026
  17. 17Google: EmbeddingGemma 2 (6 October 2026)Accessed 09 Oct 2026
  18. 18Google Developers Blog: Bring multimodal semantic search to the edge with EmbeddingGemma 2Accessed 09 Oct 2026
  19. 19Google AI for Developers: Gemini Nano Banana 2.1Accessed 09 Oct 2026
  20. 20Gemini API pricing (Nano Banana 2.1 per-image price)Accessed 09 Oct 2026
  21. 21OpenRouter: Decisions category (Solar Decide Flash input price)Accessed 09 Oct 2026
  22. 22Mindcat: seeded gated CRM-agent simulation, in Price each agent step before you pick a modelAccessed 09 Oct 2026
  23. 23OpenRouter: Jev 1.13 model page (input price)Accessed 09 Oct 2026
  24. 24OpenAI Developer Community: First look at mathematics manuscripts from an internal frontier modelAccessed 09 Oct 2026
  25. 25PhotoCraft feature-parity estimate (remaining Opus 5.5 work)Accessed 09 Oct 2026
  26. 26Zapier SDK documentation (open beta)Accessed 09 Oct 2026
  27. 27Vercel changelog: OpenAI Decisions API on AI Gateway (7 October 2026)Accessed 09 Oct 2026
  28. 28openai/math repository READMEAccessed 09 Oct 2026
  29. 29PhotoCraft roadmap and honest parity assessment (updated 8 October 2026)Accessed 09 Oct 2026
Shivanath Devinarayanan, founder of Mindcat

Written by

Shivanath Devinarayanan

Founder, Mindcat Consulting · Salesforce MVP Hall of Fame

Runs 200+ agents in production. Reads every brief that comes in and signs the work that goes out.

About Shivanath

First Agent in Production · the written assessment

We work in real estate and right now, the AI pilot never left the sandbox.

Industry
Problem

Brief: Real estate. Leads from portals and WhatsApp. Follow-up has no owner. Stalled AI or Salesforce pilot. Need a finish, cut, or rebuild call. From: /blog/agentic-ai-news-agent-stack.