What changed, sorted by decision?
Three decisions moved this week: how you route each step, who approves an agent's action, and where customer data is processed. Most of the rest is research. We started from a creator's daily news summary, then read every vendor's own page. Where the two disagreed, the vendor page wins, and the corrections are listed below the table.
Disclosure: Mindcat builds client agents on Claude and on OpenAI models, and we implement Jev, TypeSafe's decision model.

Corrections to the week's coverage, from the vendors' own pages:
- The product is the Decisions API, and OpenAI's guide puts its speed at about ten times the Responses API. It runs on GPT-6 Luna, so the comparison is Luna against Luna, through a narrower endpoint.
- Mistral's live announcement says 52 billion active parameters. Early copies of the same text said 49 billion. The cyber figures in circulation ("18 of 19 challenges", a 12-minute malware run) are not on Mistral's page, so we left them out.
- PhotoCraft is not "about 90 percent" of Photoshop. Its own roadmap says real parity "is still well below 50%".

Which model should run each step now?
A typed endpoint for the yes, no and pick-one steps, and a chat model only where the step has to write. OpenAI's Decisions API answers three kinds of question about text or images: the probability a condition is true, one choice from a list you define, or a score on ordered levels. OpenAI's own guide draws the boundary: use Structured Outputs or function calling when you need generated fields, written explanations or tool arguments.
The facts that matter for your bill and your contract:
gpt-6-lunais the only model, onPOST /v1/decisions. Public beta; OpenAI expects general availability "in the coming weeks".- $0.10 per million input tokens, with no charge for output or caching. OpenAI's pricing page adds 10 percent to the price of regional processing (data residency) endpoints for models released since 5 March 2026.
- Zero data retention and HIPAA use for eligible customers. Data residency and regional processing in the United States and Europe only.
- Independent questions can share one request. Questions that depend on an earlier answer need separate requests.
What it costs, worked through at list price with no caching: a 600-token gate question costs $0.00006, or $0.06 per 1,000 checks. Decision models on OpenRouter list close to that: $0.042 per million input tokens for Jev 1.13 (about $0.025 per 1,000 checks) and $0.05 for Upstage's Solar Decide Flash. We have not compared their accuracy here, and Jev is the one we implement, so test any of them on your own labels before you pick.
At these prices the gate is not your bill. In the seeded simulation we published with our per-step cost breakdown, a CRM-update agent spent $0.0046 a write on models and $0.342 a write in all, because people reviewed about 11% of writes. Those are simulated numbers from our stated assumptions, not a client's. The escalation rate decides the bill. Pick a decision endpoint on latency, on accuracy against your labelled set, and on where it can process your data. OpenAI's guide gives the speed as a ratio, not in milliseconds, so measure your own median and slowest 5%.
This is the belief we hold hardest: not every hop is a model call. A routing step wants a typed answer you can threshold and log. OpenAI's guide says the same in its own words: "Use labeled examples from your application to set thresholds for routing, filtering, or review."
Lock-in has an exit here. Vercel's AI Gateway put an OpenAI-compatible /v1/decisions endpoint in front of several decision models, Jev and Liquid's d1 among them, on 7 October. If you write your gates to that request shape, you can swap the model behind a gate without rewriting it. A gateway is one more processor of your data, so read its retention terms as well.
Anthropic shipped a smaller routing control. Claude Code sub-agents could already carry an effort field in their definition. Version 2.1.292 adds an effort parameter per call, so Claude can run one delegated task at low effort and another at high. Levels run from low to max, depending on the model. If your team builds with Claude Code, set effort per sub-agent the same way you set a model per step: cheap where the task repeats, high where it is hard.
Who approves an agent's action?
A second agent for actions you can undo, and a person for writes that move money, change access or leave the company. A second agent is a good filter and a weak signature. The receipts come from the vendors themselves.
OpenAI made Codex Auto-review free for users signed in with a ChatGPT account on 6 October, and said it no longer draws on plan usage. Auto-review sends an approval request to a separate reviewer agent instead of a person. OpenAI's numbers from April: sessions stop for a person "roughly 200x less often", and the reviewer "approves around 99%" of the requests it sees. In one internal snapshot, 720 actions that would have interrupted a user were reviewed automatically. Seven were rejected. Four of those continued by a safer path, and three stopped for the user.
The same page states the limits. Red-teamers "identified cases where Auto-review could be misled into approving commands without user approval", and OpenAI says it "should not be treated as a guarantee of security". The docs add that it is "a reviewer swap, not a permission grant": it does not widen what the sandbox allows. OpenAI publishes no task mix or sample size beyond that snapshot, so treat 99% as OpenAI's traffic, not yours.
Anthropic's new Claude add-on for Google Docs, Sheets and Slides takes the opposite default. "Ask before edits" is on unless a user switches to "Accept all edits". Even then, it stops before it adds a link, inserts a web image or removes someone's access, with no "always allow" option. The mode switch sits in each user's sidebar, so decide your rule before you allow the install.
Two legal texts put this in writing. Article 14 of the EU AI Act requires high-risk AI systems to be built so that people can effectively oversee them, including the ability to override the output or stop the system. In the DIFC, the Commissioner's guidance on Regulation 10, which covers personal data processed by autonomous and semi-autonomous systems, holds the deployer accountable for how the system processes that data. Neither text tells you which approvals need a person. Both expect you to show who could step in.
Our own back-office agents follow the person-on-the-write rule. Agents that handle bank transfers, vendor payments, hiring requests and vendor onboarding ask a person before they act. Approval runs in steps: the submitter's manager, then the department head, then finance and the CEO when it touches budget. The approver sees the request, who submitted it, the reason and the cash impact. The request lives in a Slack List until the last approval, and only then writes back to Salesforce, which stays the system of record. Framing each process took days of discussion. Building each one took about a day. Over six months, one of our approval agents took 40 requests, and 12 of them reached a person. What we missed at first: contractors, external approvals, external reviews and auditing, all added after go-live. What changed because of it: we now write down the process, its exceptions and its scenarios before we automate any of it. This is back-office volume, not thousands of writes an hour, but the rule scales: a second agent can sort the queue, and a named person signs anything that moves cash. We tested what happens when an approval does not bind to the exact action in a $20 approval that ran as $2,000.

Does an open-weight model change where data runs?
Only when you host it. A hosted API runs where its vendor serves it. Open weights let you pick the region, at the price of running the model yourself.
Mistral Large 4 is in public preview on Mistral's API. Mistral describes it as 1 trillion parameters with 52 billion active, natively multimodal, with weights due by the end of October. Its docs list a 1.05 trillion total and a 1M-token context. Mistral says the model will run in multiple regions, "including a European deployment that Mistral operates end-to-end", and that security teams will be able to run it on private cloud or on-premise. The docs show preview prices of $0.68 per million input tokens and $2.09 per million output, marked as a sale from $1.36 and $4.18. Mistral chose the benchmarks on its page, some of them run by third parties, and the licence for the weights is not on the page yet.
At the small end, Google released EmbeddingGemma 2 under Apache 2.0: 740 million parameters, text, images, audio and video in one embedding space, and about 191MB of RAM for the text-only weights on a Pixel 11 Pro. Google's edge team also shows it matching inputs against labels for on-device routing with no training data. For search over files that must not leave a laptop or a phone, that is a model you can run today.

For data that has to stay in the UAE or India, none of the four pages above names those regions. Check whether your cloud provider serves the model you want in a region there, which is a separate contract we did not price here. Self-hosting open weights is the other route, and it is not small: 1.05 trillion parameters is about 1 TB of weights at 8 bits each, before any memory for context (our arithmetic). That is a multi-GPU server, not a spare box.
Security teams face the same split. Anthropic expanded its Cyber Verification Program to three tiers (Defense, Red Team and Specialized Access) covering Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1, with data retention required so Anthropic can monitor misuse, and phishing-resistant MFA with no API keys for Defense Access by 15 December 2026. OpenAI's equivalent is Daybreak, Blue on GPT-5.6 Sol and Red on GPT-5.6 Cyber. Mistral offers open weights under your own policies. With the verified tiers, the vendor watches the traffic. With open weights, you do.
What does one SDK into 9,000 apps do to identity?
It moves the question from "can the agent reach the app" to "whose account is it using, and who sees what it did". The Zapier SDK went to open beta, free during early access. Zapier's docs describe "any API call, on behalf of a user, with no OAuth setup required" across 9,000+ apps, with credentials kept away from the model and your code.
The governance line is in Zapier's own docs. App and action restrictions you set in Zapier apply to the SDK's pre-built actions. Direct API access through .fetch() "currently falls outside those policy controls", and Zapier lists that governance as in progress.
Our rules for any connector layer like this, Zapier or not:
- Give each agent its own connection, not a person's. Otherwise the audit log names someone who never acted, and you cannot revoke one agent without cutting off a person.
- Block or wrap the raw API path until your policies cover it: expose only the pre-built actions to the agent, or deny the route at your egress proxy.
- Name the owner of the exceptions queue before the first write, not after the first bad one.

The video we started from carried a paid Zapier segment. Everything in this section comes from Zapier's documentation instead.
Which headlines are research, not a purchase?
The maths papers and the Photoshop clone. Both are worth reading. Neither changes what you run on Monday.
OpenAI published a repository of mathematics manuscripts from an unreleased internal model: 722 at release and 719 in the README on 9 October, grouped into 372 families. The model was posed about 4,000 problems, and each result used on average about three hours of ChatGPT Pro thinking compute. The README says about 42% of top-line results are formalised in Lean, and that "some of the unformalized results could have issues". The model is not available to you, and three hours a result is the opposite of a cheap step.
PhotoCraft, an open-source Photoshop reimplementation in Rust, carries the lesson for anyone who runs evals. Its planning file measures remaining work in Claude Opus 5.5 agent-hours. Its own numbers, side by side:
The roadmap names the gap itself: users found 214 shortcut failures and broken basics that "all counted as live" in the wiring report. Your agent evals have the same trap. A tool call that runs is not a task that succeeds, so score the outcome a person would accept.
Google's Nano Banana 2.1 image model ($0.0336 per 1K image on the Gemini API) belongs here for most agent teams too. The speed comparisons going around come from one creator's test.
What would we do this week?
Nothing in production until a shadow run says so. Shivanath has put 200+ agents into production, and the method we use is Spec, shadow, gate: write the step down, run the change in shadow next to today's path, then put a person on the write.
The pattern from our own back-office flows, without client data: the model was never the expensive part. Writing each process down took days of discussion, and building it took about a day. Every gap we found after go-live was a missing case in the process, such as a contractor, an outside reviewer or an audit trail, not a wrong model answer. So before you spend a week comparing models for a step, spend it writing down who approves, what they see and what happens when the request does not fit. That page is what any model swap gets tested against.
- 01List the yes, no and pick-one steps in one agent. Those are your candidates for a decision endpoint.
- 02Shadow one gate on the Decisions API or another decision model for a week. Compare accuracy on your labels, latency and cost.
- 03Write your approval rule: which actions a second agent may approve, and which need a named person.
- 04Map where each model processes customer data. Mark any step that needs a region no hosted option names.
- 05Block raw API paths in connector SDKs until your policies cover them, and give each agent its own identity.

Before you change a model for any step, get three numbers: wrong answers per hundred, median latency, and the share of runs that reach a person.
The rule we suggest, so you can argue with it: move a step to a cheaper or faster model only if, in a shadow run on the same labelled requests, wrong answers per hundred are no higher than today's, the slowest 5% of answers still fits the time the step allows, and the share that reaches a person does not go up. If any one of the three gets worse, keep today's path. Run enough requests that one miss cannot swing the result: at 40 requests, the size of our approval agent's six months, a single wrong answer moves the rate by 2.5 points.
What to own and what to rent: rent the models, behind an endpoint you can swap. Own the labelled requests and the shadow test, because those are what let you switch.

