Speaking twice at Dreamforce · Sept 15-17 →

Frameworks and protocols

US · Global

A $20 approval ran as $2,000: make human approval for AI agents cover what runs

By Shivanath DevinarayananPublished 10 min read
Written forCIO, CISO, or compliance lead who approves agents on real systems“Who acted, and can we stop it?”

The short answer

Does a human approval for an AI agent cover the action that actually runs?

Not always. In a published A2A test, an approver said yes to a $20 transfer, and the system ran a $2,000 one under that yes, because someone edited the waiting request and the task ID stayed the same. At Mindcat, we make the approval name the exact action and check it again at the write.

A harbour desk at dusk with a stamped slip on brass scales, under the title: a $20 approval ran as $2,000

Takeaways

  1. 01In a LangGraph Agent Server test, a yes for a 20-unit transfer released 2,000 units to another account. The task ID did not change.
  2. 02A2A's own specification now says the authorization-required state is a signal, not a grant. The implementation decides what a yes covers.
  3. 03Bind the decision to the exact action and compare it at the write, or forbid edits to a request that is waiting.
  4. 04Before an agent writes to a real system, run a swap test: change one field after the approver looks, approve, read what arrived.

What happened when a $20 approval ran as $2,000?

An approver said yes to a transfer of 20 units, and the system sent 2,000 to a different account. In a controlled test published on 21 September, one researcher, Adithyan Arun Kumar, set up two roles on a LangGraph Agent Server that exposes an A2A route. The test used units on a mock ledger, and the dollar figures come from press coverage. A "maker" could create work and edit it while it waited. An "approver" could read the waiting request and approve it. The maker could not approve or run the transfer.

The agent proposed mock_wire_transfer(20, approved-vendor). The approver read it. Then the maker sent one more message to the same task, replacing the waiting call with mock_wire_transfer(2000, attacker-sink). The approver's yes, given for the first call, was then applied to the current one. The test's ledger recorded 2,000 units moved under the approver's authority, and no record of the 20.

Here is the record, from the field note:

  1. 01Approval view: id=transfer-call, mock_wire_transfer(20, approved-vendor). The approver reads this.
  2. 02Update from the maker, same task and same call ID: mock_wire_transfer(2000, attacker-sink).
  3. 03Approver sends: approve.
  4. 04The tool receives mock_wire_transfer(2000, attacker-sink).
Five steps: agent asks to send 20, approver says yes, request edited to 2,000, the yes is reused, 2,000 leaves
How a $20 yes became $2,000, from the field note.

The researcher calls this class of failure Loopjacking: a person approves one action, and the system uses that approval for a different one.

If an approver is the last control between an agent and a real system, our control plane page shows where the approval, the identity and the log sit.
A lamp-lit harbour desk at night with a stamped slip of paper and a second sheet sliding out from beneath it
The yes was stamped on one sheet. A different one was underneath when the write ran.

Why does a task ID not prove what a person approved?

Because a task ID says which work a message belongs to, and nothing about which action a person looked at. An A2A task is a continuing unit of work. A client can send more messages to it, and an agent can pause it and ask for authorization. The ID stays the same through all of that. If the approval is recorded against "this task", a later edit to the task's pending action rides along with the approval.

The A2A project saw the ambiguity. Issue 2080 asked for clarity, and pull request 2081 added section 7.6.4 to the specification on 30 July 2026. It says the authorization-required state is a coordination signal, not an authorization grant. The wording the AGNTCY team quotes from the section, as it stood in the working draft, is direct: agents must not treat that state transition, by itself, as authorization for any particular operation. What a yes covers, how long it lasts and how it is revoked is left to the implementation, the credential issuer or an extension.

The field note does not claim that A2A itself authorized the swap, or that the older wording caused the product result. It also says the latest tagged A2A release was v1.0.1 when it was written, so do not assume a tagged release carries the new wording.

The part that was missing is the one your team owns. What exactly did the person approve, and where is that checked before the write?

A task ID tells you which work a message belongs to, not which exact action a person approved
A task ID answers a different question from the one the approver answered.

Which agent stacks has this been shown on?

On three of the four product lines tested, and not on the fourth. The paper reports results for a purposive set of released products, not a survey of the market.

ProductResultVersions
Agno AgentOSPost-approval substitution reproduced7 tested releases, ending at 3.0.9
LangGraph Agent Server (conditional, in-memory setup)Post-approval substitution reproduced12 tested versions, ending at 0.14.0
OpenClawMismatch between what was shown and what ran reproduced, then rejectedReproduced in 2026.2.23, rejected in 2026.2.24
OpenAI Agents SDKNegative control: the changed call was rejected0.22.0 and 0.22.2
The swap ran on 7 Agno AgentOS releases and 12 LangGraph versions, and was rejected by 2 OpenAI SDK versions
What the paper tested. One researcher, a mock ledger, no independent reproduction.

The paper separates two variants. In the first, the dangerous action is already in the request but hidden or misrepresented when the human looks. In the second, the human sees the right action and a change to the waiting request swaps it afterwards. The A2A test is the second.

The negative control is the useful part. In the two OpenAI Agents SDK versions tested, continuing after an approval keeps the exact call that was approved, and a mutated call is rejected. MindPattern's summary of the paper puts the lesson this way: bind the approval to an immutable serialized operation, not to a mutable state handle. In the LangGraph test, the later approve rebuilt the call from current thread state, so the edit went through. The OpenAI SDK versions tested keep the call that was approved.

What does it mean to bind an approval to the action?

It means the yes is stored with the exact action it was given for, and the system compares that action with the one about to run. There are two routes, and the field note describes both.

The first is to bind the decision. The application builds the approval screen from a complete, canonical form of the action. It stores a protected record with the approved action, a hash of it and the approver's identity, and it checks the action again after the last step that can change it, whether a state update, a wrapper, a default or a retry. Any change to a field that matters means a new decision. Going from 20 to 2,000 units, or from approved-vendor to attacker-sink, is material even when the task ID, request ID and call ID are unchanged.

The second is to forbid the swap. In the tested setup, a supported policy that denied the maker's updates while an approval was pending stopped the attack. The approver still executed the 20-unit transfer. That fits a workflow where nobody should edit a request once it is waiting. The note is plain that it does not cover cases where authorized updates or retries must change a pending request. For those, bind the decision.

Two ways to close the gap

CriterionHow it worksFits when
Bind the decision to the actionStore the exact action with the yes, then compare it with what is about to run. A changed material field needs a new yes.Requests may change while waiting, for example a retry or a corrected amount
Forbid edits while waitingDeny changes to a request once it is waiting for approvalNobody should edit a request after the approver has seen it
Both are described in the field note. Which one you choose depends on whether a waiting request may legitimately change.
Two fixes: bind the yes to the exact action and compare it before it runs, or lock the request while it waits
Two ways to close the gap. Choose binding when edits can be legitimate.

Neither route is a task ID check, and neither depends on the approver noticing a change. The test scripted the approver, so it measures the system's binding, not a person's attention.

Diagram: one task ID across four steps, with a check at the write comparing the call to the approved one
The check belongs at the write, because the task ID does not change.

Does this matter for the approvals you run in Salesforce?

It is the same question, asked of a different mechanism. A Salesforce approval process works on a record, not on an agent's tool call, so the attack in the test does not carry over as it is. But Salesforce's own documentation shows where to look.

Salesforce locks a record while it is in an approval process. Its Apex reference says admins can still edit a locked record, and that depending on the process's configuration, an assigned approver can too. That is a reasonable design for a person-run process. It is also the thing to check when an agent proposes the change. If your approvals happen as a button in Slack, the same goes for the message: what did it show, and what does the write compare it with? After someone approves, who or what can still change the record or the action, and does the final write compare against what the approver saw?

For an agent that writes to Salesforce on someone's behalf, we put the check where the write happens, not where the approval was clicked. That is our practice, in line with the belief we start from: the write is the risk.

How do you test your own agent for it?

Run the swap test in a sandbox before the agent writes to a live system. It is our test, built from the four control paths in the field note.

  1. 01Pick one approval your agent asks for. Write down the action in full: who, how much, to where, which record, which tool.
  2. 02Have the agent propose it, and have an approver open the approval screen. Record exactly what the screen shows.
  3. 03Before the approver clicks, change one material field through every route that can touch the waiting request: another message to the same task, an API update, a retry, a second agent.
  4. 04Approve. Then read what reached the system, not the status. Check the arguments the tool received.
  5. 05Pass: the changed action is blocked, or the system asks again. Fail: it runs, or the screen never showed the changed field.
The swap test: write the action, change one field, approve, read what the system received, expect a block
The swap test in five steps.

Do one extra run with nothing changed. The unchanged action must still run, or you have only proved the approval is broken in a different way. The field note ran the same control.

Five fields to treat as material: who the action is for, the amount, the destination, the record it changes and the tool it calls. A change to any of them means ask again. Anything you leave off that list, write down why.

Five fields that need a new approval when they change: who it is for, how much, where it goes, which record, which tool
Five material fields. Change any of them and the approver is asked again.

What do these results not tell you?

Less than the headlines. This is one researcher's work. The A2A test used an in-memory server, synthetic identities, a deterministic local model and a mock ledger. The note says it did not test the production Postgres image, which stopped at a license-key requirement before the client test ran, and it says independent reproduction is not claimed. The setup also depended on a policy that let the maker update the waiting request. The note did not measure how common that policy is, and it does not show the same result under the default policy.

The version evidence is conditional. The LangGraph path was executed on 12 sampled versions from 0.7.5 to 0.14.0, and found in source across all 130 releases from 0.7.5 to 0.13.4, which is not 130 live runs. The note identifies no vendor-fixed LangGraph release. The paper says its results do not estimate how common this is across the ecosystem.

And our own piece is a method, not a finding. We have not run the swap test on your stack, and we do not claim any product is safe or unsafe beyond what the sources tested. Check the version you run.

What would we do first?

Start with one agent action a person approves today, and run the swap test on it. At Mindcat, we start any work by reading what already exists, then writing down what is worth fixing. For an agent that means the approval screen, the code path that releases the write, and every route that can edit a waiting request. You get the result as a written note, with a named person on the exceptions, not a dashboard.

We work across Salesforce, MuleSoft and custom agent stacks, so we can read the approval wherever the agent runs. Our first agent in production path puts a human gate on the write after the agent has run in shadow mode. See also how we think about stopping an agent and switching on read first.

Before the next agent writes to a real system, get three answers in writing: what the approver was shown, what the system compares at the write, and which routes can change a request while it waits.

Filed under

Frameworks and protocolslanggraph human in the loopa2a securityhuman in the loop ai agentsai agent security risks
Useful to someone on your team?

FAQ

Questions teams ask next

What is Loopjacking?
It is the name Adithyan Arun Kumar gives to a failure where a person approves operation A and the system uses that approval to run a materially different operation B. In the post-approval variant the approver sees the right action, and a later change to the waiting request swaps it.
Is A2A itself vulnerable to this?
The field note does not say so. It shows a failure in one tested LangGraph Agent Server setup and says the older A2A wording left the question of approval scope open. Section 7.6.4 of the specification now says the authorization-required state is not a grant and that the implementation defines scope. A server that asks again after a change behaves safely.
How do you stop an approved action being swapped?
Two ways. Bind the decision to the exact action, with a record the approver cannot be talked out of, and compare the action about to run with the one approved. Or forbid changes to a request while it waits for approval. If legitimate edits must stay possible, use the first.
Do Salesforce approval processes have this problem?
Not in the same form. Salesforce locks a record while it waits for approval, but admins can still edit it, and the assigned approver can too, depending on the process setting. So ask who can change the record after someone said yes, and what the approver was shown.

Sources

  1. 01Adithyan Arun Kumar: Loopjacking in A2A Implementations: Hijacking Human-in-the-Loop Approvals (21 September 2026)Accessed 01 Oct 2026
  2. 02arXiv 2609.21081: Loopjacking: Hijacking Human-in-the-Loop Approval (17 September 2026)Accessed 01 Oct 2026
  3. 03StartupHub.ai: A $20 Approval Ran as $2,000 in A2A Test (26 September 2026)Accessed 01 Oct 2026
  4. 04A2A project: pull request 2081, clarify in-task authorization scope semantics (merged 30 July 2026)Accessed 01 Oct 2026
  5. 05A2A project: issue 2080, clarify in-task authorization scope in the A2A specAccessed 01 Oct 2026
  6. 06AGNTCY: Agent Identity and Task-Based Access Control, how A2A and AGNTCY Identity work together (25 August 2026)Accessed 01 Oct 2026
  7. 07MindPattern: Loopjacking, human approval boundaries reproduced as broken in shipped Agno, LangGraph and OpenClaw releasesAccessed 01 Oct 2026
  8. 08Salesforce Developers: Apex Reference Guide, Approval ClassAccessed 01 Oct 2026
Shivanath Devinarayanan, founder of Mindcat

Written by

Shivanath Devinarayanan

Founder, Mindcat Consulting · Salesforce MVP Hall of Fame

Runs 200+ agents in production. Reads every brief that comes in and signs the work that goes out.

About Shivanath

First Agent in Production · the written assessment

We work in real estate and right now, the AI pilot never left the sandbox.

Industry
Problem

Brief: Real estate. Leads from portals and WhatsApp. Follow-up has no owner. Stalled AI or Salesforce pilot. Need a finish, cut, or rebuild call. From: /blog/human-approval-ai-agents.