Why Signed AI Agent Payments Can Still Be Hijacked
Cryptographic mandates can prove that a transaction was authorized. They cannot automatically prove that manipulated information did not shape the decision before the user or agent signed it.
The weakest point in an AI payment may occur before money moves and before anything is signed.
A product description, poisoned tool result, compromised merchant agent or deceptive Agent-to-Agent message can influence what an AI shopper believes it should buy. If that manipulation succeeds, the final cart and payment can be internally consistent, correctly signed and still inconsistent with what the user originally wanted.
The practical answer is not to discard cryptographic mandates. It is to surround them with verifiable intent, deterministic pre-action policy, session-bound credentials, execution limits and signed evidence across the complete transaction lifecycle.
Core evidence: The APort Vault benchmark replayed 4,371 human-authored attacks across 14 models from eight labs, five policy configurations and 225,964 completed evaluations. At policy Levels 2 to 4, its authors reported 140 transfers to recipients not permitted by policy among 76,842 model-only evaluations, compared with zero among 69,297 evaluations protected by the deterministic pre-action layer.
Separate AP2 evidence: A systematic analysis of AP2 v0.2 catalogued 48 threats across five attack families and found eight reaching its High band in at least one deployment architecture. A second study reported three pre-signing attack classes with experimental success rates of 90%, 56% and 73.3% in its stated setup.
Limitations: These are separate research environments. They do not establish a universal real-world loss rate, prove that every AP2 implementation is vulnerable, or justify combining their percentages into one statistic.
A signature can authenticate the wrong decision
Digital signatures are excellent at answering a narrow question: did the holder of a particular credential approve this exact payload, and was that payload altered after signing?
Agentic commerce adds an earlier question: how was the payload formed?
An AI shopping agent may read catalog descriptions, query MCP tools, communicate with merchant agents over A2A, retrieve account information and assemble a cart. Those inputs can shape the decision before a trusted surface presents the mandate for approval. Signing preserves the integrity of the resulting decision. It does not retroactively establish the integrity of every fact or message that caused it.
What AP2 actually protects
AP2 v0.2 is designed to secure agent-performed transactions using deterministic verification. It defines Checkout Mandates for proving authority to purchase an assembled checkout and linked Payment Mandates for authorizing payment. Receipts provide evidence after acceptance or rejection.
The specification supports human-present and human-not-present flows. Open mandates can express constraints before a specific cart exists. Closed mandates bind approval to a defined checkout. AP2 also states that validation or processing assigned to protocol roles must occur in deterministic code, regardless of whether a role is agentic.
These are meaningful controls. The important boundary is scope: AP2 describes its commerce-protocol details, including catalog APIs and the mechanism by which the shopping agent determines the user's task, as outside the payment specification. The specification says the Checkout and Payment Mandate contents are assembled after the shopping agent has determined what the user wants.
That boundary is where pre-authorization manipulation becomes consequential.
The five-stage agent payment lifecycle
| Stage | What must remain true | Example failure | Required control |
|---|---|---|---|
| 1. Intent formation | The agent accurately captures the user's goal, limits and exclusions. | Ambiguous language is converted into broader authority than intended. | Structured intent, trusted surface and explicit constraints. |
| 2. Context and selection | Catalogs, tools and agents provide authentic, relevant information. | A product description or MCP result steers the agent toward an unauthorized item. | Provenance, content binding, cross-checks and untrusted-input isolation. |
| 3. Mandate authorization | The mandate matches both user intent and the selected checkout. | A valid signature approves a cart formed from poisoned context. | Intent-to-cart consistency checks and human confirmation for uncertainty. |
| 4. Execution and settlement | Recipient, amount, asset, network and timing remain within policy. | The model calls a payment tool with a recipient outside its permitted set. | Deterministic pre-action authorization and transaction limits. |
| 5. Fulfilment and recovery | The buyer receives what was authorized and can prove failure. | Payment settles but the wrong service is delivered or no usable evidence survives. | Signed receipts, delivery binding, revocation and dispute process. |
DN Intent-to-Settlement Integrity Grader
How safe is an agent payment design?
Score the architecture, not the model's promises. Choose the implementation status for each control.
The system should not autonomously move value.
This self-assessment is educational and does not replace architecture review, penetration testing, legal advice or regulated payment controls.
What the latest research changes
1. Authorization belongs outside the model
APort Vault is important because it distinguishes an agent requesting a transfer from a system allowing that transfer. The model can interpret language and propose an action. A separate deterministic layer decides whether the tool call complies with a signed policy.
In the reported higher-policy configurations, the layer blocked unauthorized recipients without simply disabling payments: 25,370 payments still executed behind the policy layer, while 187 of 25,640 evaluated transfer calls were denied.
The lesson is architectural. A prompt telling an agent to “never send money to an unauthorized recipient” is guidance. A pre-action verifier that rejects a recipient not present in the signed policy is enforcement.
2. The cart needs provenance, not only a signature
The AP2 Whisper Attacks research focuses on manipulation before signing. It describes attacks in which ordinary product-description text changes credential selection, cart contents or spending decisions while the eventual cart remains cryptographically valid.
Its proposed A-VIP defense treats signed intent as a capability grant, binds credential lookup to the requesting session and connects each cart line to the listing the user or agent saw. When manipulation cannot be resolved structurally, it escalates unauthorized spending for confirmation.
This suggests a broader standard: every economically relevant claim should carry sufficient provenance for the authorizing system to determine what was seen, who supplied it and whether it changed.
3. Cross-protocol composition creates cross-stage risk
Agent commerce will not run on one protocol. A2A may carry messages between agents, MCP may expose payment or catalog tools, AP2 may carry authorization evidence, and x402 may settle a machine-to-machine payment.
Each protocol can work as designed while the combined system still loses an important binding. A payment may be valid under x402, authorized under a mandate and attached to an agent interaction, yet fail to prove that the delivered service corresponds to the original request.
A formal analysis of AP2, x402, MPP and ACP argues that delegated authority must remain consistent with the resulting economic and service effects across actors, states and protocol stages. This is the emerging security frontier: compositional integrity.
Eight controls every payment agent needs
Structured intent
Translate natural language into explicit products, maximum amounts, recipients, timing, recurrence and prohibited actions.
Context provenance
Record the origin and version of catalog claims, tool results, quotes and agent messages used to make the decision.
Cross-stage binding
Bind user intent to the displayed offer, cart, payment terms, credential and expected fulfilment.
Deterministic policy
Intercept every value-moving tool call and verify it against machine-enforceable limits outside the language model.
Least privilege
Use short-lived, narrowly scoped credentials and session-specific authority rather than a wallet with unrestricted access.
Velocity controls
Cap cumulative spend, frequency, price variance and new-recipient exposure, not merely the size of one payment.
Verifiable receipts
Preserve signed authorization, policy decision, transaction, delivery and exception evidence for audit and disputes.
Revocation and recovery
Give users and operators a fast kill switch, credential revocation, fallback approval path and defined dispute process.
What buyers, developers and merchants should do
For consumers and businesses
- Begin with low transaction and daily limits.
- Require confirmation for new recipients, subscriptions, high-value purchases and unusual price changes.
- Do not grant a general-purpose agent unrestricted wallet or bank access.
- Demand a readable record showing the request, selected offer, final cart, policy decision and receipt.
For developers
- Keep authorization code separate from probabilistic model reasoning.
- Treat merchant content, MCP output and A2A messages as untrusted input.
- Test prompt injection, tool-result poisoning, replay, credential confusion and cross-session attacks.
- Fail closed when mandate, checkout, recipient or session bindings cannot be verified.
For merchants and payment providers
- Sign offers and checkout data in a form that can be bound to mandates.
- Return machine-verifiable fulfilment evidence, not only payment confirmation.
- Expose transparent cancellation, refund and dispute states to agents.
- Avoid product content that can be interpreted as instructions to the buyer's agent.
The DN Alpha Thesis
The valuable security layer in agentic commerce will sit between intelligence and money. Models will propose. Protocols will communicate. Payment rails will settle. A deterministic authorization boundary will decide what is actually allowed.
That layer could become a new control point comparable to identity providers, card authorization networks or cloud security gateways. The winning products may issue portable agent credentials, evaluate policy before every consequential action, bind evidence across protocols and provide an independent audit trail.
This is also the strategic opening for AgentNotary: not another wallet and not another general agent framework, but an authorization-provenance layer that can prove which identity, intent, context, policy and approval produced an economic action.
DN research opportunity
The Intent-to-Settlement Integrity Score can become a recurring benchmark for agent wallets, payment protocols, commerce agents and autonomous financial tools. Vendors may submit documentation or test environments, but sponsorship must never purchase a score.
Submit a system for reviewExplore AgenticFi researchMethodology
This article compares protocol specifications with recent independent security research. It distinguishes documented protocol guarantees from experimental findings and proposed mitigations. The DN grader weights deterministic authorization most heavily because it directly controls whether a value-moving action can execute. Intent, context and cryptographic binding receive the next-largest weights because they determine whether the authorized action corresponds to the user's purpose.
No production protocol or vendor receives a numerical ranking in this edition. A future index should test concrete implementations against a disclosed threat suite and score the deployed architecture, not the protocol name alone.
Primary sources
- Google Agentic Commerce: AP2 v0.2 specification
- AP2 security and privacy considerations
- APort Vault: Benchmarking AI Agent Payment Authorization
- APort Vault benchmark dataset
- Beyond the Mandate: A Systematic Security Analysis of AP2
- Signing the Transaction but Not the Decision
- A Formal Analysis of Agent Payment Protocols
- Open Agent Passport specification
- x402 official documentation
Frequently asked questions
Can a correctly signed AI agent payment still be unauthorized?
Yes. A signature can prove approval of a payload while manipulated context caused the agent or user to approve the wrong payload. Systems need controls before, during and after mandate signing.
Does this mean AP2 is insecure?
No. AP2 provides important mandate, binding, deterministic-verification and receipt mechanisms. The research highlights deployment and pre-authorization risks that must be addressed around the protocol.
Why is a prompt not enough to control agent spending?
Prompts guide probabilistic model behavior. Deterministic authorization code can enforce recipient, amount, timing and credential constraints before a payment tool executes.
What is the safest way to give an AI agent payment authority?
Use narrowly scoped, short-lived authority with explicit limits, deterministic pre-action checks, new-recipient confirmation, signed receipts and immediate revocation.
Are x402 and AP2 competitors?
They can be complementary. AP2 focuses on authorization and evidence for agent commerce, while x402 provides an HTTP-native mechanism for machine payments. Secure composition still requires cross-stage bindings.
What does the DN Intent-to-Settlement Integrity Score measure?
It assesses whether user intent, decision context, mandates, credentials, execution policy, settlement evidence and recovery controls remain consistently bound across the payment lifecycle.
Disclosure
This content is independent research and general information. It is not financial, legal, compliance, cybersecurity or procurement advice. Research results are reported within their stated experimental conditions and should not be interpreted as universal loss rates. Protocols and implementations can change after publication. Decentralised News may earn revenue from clearly disclosed commercial relationships, but payment does not purchase inclusion, methodology changes or a favorable assessment.
Related reading:
The Best AI Agent Runtimes of 2027: OpenAI vs Google vs AWS vs Microsoft
How to Verify an AI Trading Bot Before Risking Any Money
AI Context Efficiency 2027: Useful Outcomes per Token
AI Model Routing 2027: Does It Save Money Without Losing Quality?
The Agentic Treasury Benchmark: Payments, Stablecoins and Financial Control






