What happened when a $20 approval ran as $2,000?
An approver said yes to a transfer of 20 units, and the system sent 2,000 to a different account. In a controlled test published on 21 September, one researcher, Adithyan Arun Kumar, set up two roles on a LangGraph Agent Server that exposes an A2A route. The test used units on a mock ledger, and the dollar figures come from press coverage. A "maker" could create work and edit it while it waited. An "approver" could read the waiting request and approve it. The maker could not approve or run the transfer.
The agent proposed mock_wire_transfer(20, approved-vendor). The approver read it. Then the maker sent one more message to the same task, replacing the waiting call with mock_wire_transfer(2000, attacker-sink). The approver's yes, given for the first call, was then applied to the current one. The test's ledger recorded 2,000 units moved under the approver's authority, and no record of the 20.
Here is the record, from the field note:
- 01Approval view: id=transfer-call, mock_wire_transfer(20, approved-vendor). The approver reads this.
- 02Update from the maker, same task and same call ID: mock_wire_transfer(2000, attacker-sink).
- 03Approver sends: approve.
- 04The tool receives mock_wire_transfer(2000, attacker-sink).

The researcher calls this class of failure Loopjacking: a person approves one action, and the system uses that approval for a different one.

Why does a task ID not prove what a person approved?
Because a task ID says which work a message belongs to, and nothing about which action a person looked at. An A2A task is a continuing unit of work. A client can send more messages to it, and an agent can pause it and ask for authorization. The ID stays the same through all of that. If the approval is recorded against "this task", a later edit to the task's pending action rides along with the approval.
The A2A project saw the ambiguity. Issue 2080 asked for clarity, and pull request 2081 added section 7.6.4 to the specification on 30 July 2026. It says the authorization-required state is a coordination signal, not an authorization grant. The wording the AGNTCY team quotes from the section, as it stood in the working draft, is direct: agents must not treat that state transition, by itself, as authorization for any particular operation. What a yes covers, how long it lasts and how it is revoked is left to the implementation, the credential issuer or an extension.
The field note does not claim that A2A itself authorized the swap, or that the older wording caused the product result. It also says the latest tagged A2A release was v1.0.1 when it was written, so do not assume a tagged release carries the new wording.
The part that was missing is the one your team owns. What exactly did the person approve, and where is that checked before the write?

Which agent stacks has this been shown on?
On three of the four product lines tested, and not on the fourth. The paper reports results for a purposive set of released products, not a survey of the market.

The paper separates two variants. In the first, the dangerous action is already in the request but hidden or misrepresented when the human looks. In the second, the human sees the right action and a change to the waiting request swaps it afterwards. The A2A test is the second.
The negative control is the useful part. In the two OpenAI Agents SDK versions tested, continuing after an approval keeps the exact call that was approved, and a mutated call is rejected. MindPattern's summary of the paper puts the lesson this way: bind the approval to an immutable serialized operation, not to a mutable state handle. In the LangGraph test, the later approve rebuilt the call from current thread state, so the edit went through. The OpenAI SDK versions tested keep the call that was approved.
What does it mean to bind an approval to the action?
It means the yes is stored with the exact action it was given for, and the system compares that action with the one about to run. There are two routes, and the field note describes both.
The first is to bind the decision. The application builds the approval screen from a complete, canonical form of the action. It stores a protected record with the approved action, a hash of it and the approver's identity, and it checks the action again after the last step that can change it, whether a state update, a wrapper, a default or a retry. Any change to a field that matters means a new decision. Going from 20 to 2,000 units, or from approved-vendor to attacker-sink, is material even when the task ID, request ID and call ID are unchanged.
The second is to forbid the swap. In the tested setup, a supported policy that denied the maker's updates while an approval was pending stopped the attack. The approver still executed the 20-unit transfer. That fits a workflow where nobody should edit a request once it is waiting. The note is plain that it does not cover cases where authorized updates or retries must change a pending request. For those, bind the decision.
Two ways to close the gap

Neither route is a task ID check, and neither depends on the approver noticing a change. The test scripted the approver, so it measures the system's binding, not a person's attention.
Does this matter for the approvals you run in Salesforce?
It is the same question, asked of a different mechanism. A Salesforce approval process works on a record, not on an agent's tool call, so the attack in the test does not carry over as it is. But Salesforce's own documentation shows where to look.
Salesforce locks a record while it is in an approval process. Its Apex reference says admins can still edit a locked record, and that depending on the process's configuration, an assigned approver can too. That is a reasonable design for a person-run process. It is also the thing to check when an agent proposes the change. If your approvals happen as a button in Slack, the same goes for the message: what did it show, and what does the write compare it with? After someone approves, who or what can still change the record or the action, and does the final write compare against what the approver saw?
For an agent that writes to Salesforce on someone's behalf, we put the check where the write happens, not where the approval was clicked. That is our practice, in line with the belief we start from: the write is the risk.
How do you test your own agent for it?
Run the swap test in a sandbox before the agent writes to a live system. It is our test, built from the four control paths in the field note.
- 01Pick one approval your agent asks for. Write down the action in full: who, how much, to where, which record, which tool.
- 02Have the agent propose it, and have an approver open the approval screen. Record exactly what the screen shows.
- 03Before the approver clicks, change one material field through every route that can touch the waiting request: another message to the same task, an API update, a retry, a second agent.
- 04Approve. Then read what reached the system, not the status. Check the arguments the tool received.
- 05Pass: the changed action is blocked, or the system asks again. Fail: it runs, or the screen never showed the changed field.

Do one extra run with nothing changed. The unchanged action must still run, or you have only proved the approval is broken in a different way. The field note ran the same control.
Five fields to treat as material: who the action is for, the amount, the destination, the record it changes and the tool it calls. A change to any of them means ask again. Anything you leave off that list, write down why.

What do these results not tell you?
Less than the headlines. This is one researcher's work. The A2A test used an in-memory server, synthetic identities, a deterministic local model and a mock ledger. The note says it did not test the production Postgres image, which stopped at a license-key requirement before the client test ran, and it says independent reproduction is not claimed. The setup also depended on a policy that let the maker update the waiting request. The note did not measure how common that policy is, and it does not show the same result under the default policy.
The version evidence is conditional. The LangGraph path was executed on 12 sampled versions from 0.7.5 to 0.14.0, and found in source across all 130 releases from 0.7.5 to 0.13.4, which is not 130 live runs. The note identifies no vendor-fixed LangGraph release. The paper says its results do not estimate how common this is across the ecosystem.
And our own piece is a method, not a finding. We have not run the swap test on your stack, and we do not claim any product is safe or unsafe beyond what the sources tested. Check the version you run.
What would we do first?
Start with one agent action a person approves today, and run the swap test on it. At Mindcat, we start any work by reading what already exists, then writing down what is worth fixing. For an agent that means the approval screen, the code path that releases the write, and every route that can edit a waiting request. You get the result as a written note, with a named person on the exceptions, not a dashboard.
We work across Salesforce, MuleSoft and custom agent stacks, so we can read the approval wherever the agent runs. Our first agent in production path puts a human gate on the write after the agent has run in shadow mode. See also how we think about stopping an agent and switching on read first.
Before the next agent writes to a real system, get three answers in writing: what the approver was shown, what the system compares at the write, and which routes can change a request while it waits.

