Decentralised News · Agentic infrastructure
A2A Interoperability Test 2027: Can Your AI Agents Actually Work Together?
The handoff is where agent automation becomes an operational system. Test the outcome, the recovery path and the authority boundary.
By Heath Muchena · Published 30 September 2026 · DN-A2A v1.0 · 2027 planning edition
What Matters
A2A compatibility is a starting point, not proof that two agents can finish a useful task together. DN proposes testing discovery, delegation, task state, artifact integrity, recovery and authorization as one workflow. This edition provides an evidence register and interactive scorecard, not a vendor leaderboard. Use it to expose weak handoffs before connecting autonomous agents to sensitive systems.
DN Evidence Block
Last verified: 30 September 2026. Research period: documentation review on that date. Live endpoint sample: zero. Protocol reference: A2A v1.0.0. Research status: proposed DN methodology; independent implementation testing pending.
Decisive distinction: the A2A specification describes interoperable interfaces, while DN’s proposed test adds workflow-specific acceptance criteria. The dimensions, weights and release gates below are editorial proposals, not official A2A certification requirements.
Primary evidence: versioned specification, task lifecycle guide and enterprise features guide. No independent reviewer is recorded for this edition.
The real test is a handoff that survives interruption
Two agents can exchange valid messages while failing the business task. A research agent may return a convincing answer with the wrong reporting period. A procurement agent may restart a request after a timeout and create a second order. An assistant may deliver the right document to the wrong account. Those are distinct failure modes, even if the network traffic looks normal.
DN’s central thesis is that interoperability has three layers: interface agreement, workflow continuity and outcome acceptance. A protocol check covers the first layer. An operational evaluation must cover all three. The most valuable result is often a discovered boundary: a particular workflow that should remain supervised until recovery and authorization are demonstrated.
This article extends the reliability question from individual tools to independent agents. It does not repeat the broader MCP-versus-A2A comparison. The question here is practical: can one agent hand work to another without losing the task, corrupting the result or expanding its authority?
What the protocol supplies, and what the buyer still needs
A2A describes Agent Cards, messages, tasks and artifacts, with operations for sending work and managing tasks. Its v1.0.0 specification defines JSON-RPC, gRPC and HTTP+JSON/REST bindings. A buyer should record which interface and version were actually evaluated rather than treating “supports A2A” as a complete compatibility statement. [1]
The task guide distinguishes ongoing interaction from terminal task outcomes. DN’s added acceptance rule asks whether the delivered artifact is usable for the stated purpose, not merely whether the agent reports completion. [2]
The enterprise guide discusses authentication and authorization. DN’s proposed negative tests check those boundaries in the deployed workflow, including requests made with insufficient permission. A declared security scheme is evidence about configuration; it is not evidence that isolation worked in practice. [3]
The DN Handoff Integrity Score
For each dimension, divide passed applicable fixtures by attempted applicable fixtures, then multiply by its weight. Add the weighted rates and renormalize only for dimensions excluded in the predeclared workflow profile. Report attempted counts and failures beside the score. A high percentage based on a tiny sample is not strong evidence.
| Dimension | Weight | Proposed fixture | Acceptance evidence |
|---|---|---|---|
| Discovery and interface agreement | 15% | Read the card, select the supported interface and pin protocol/SDK versions. | A mutually supported interface; clear rejection of an incompatible profile. |
| Delegation and context | 20% | Send a bounded task, then a clarification under the correct task/context identifiers. | No unrelated context enters the task; the clarification reaches the intended work. |
| Lifecycle and artifact integrity | 20% | Observe progress and validate the final artifact against a fixture. | Permitted transitions and an artifact that passes the independent acceptance rule. |
| Updates and recovery | 15% | Interrupt delivery, reconnect or retrieve task state through the declared mechanism. | One coherent final result without lost essential output or duplicated effects. |
| Cancellation and side effects | 15% | Cancel during a controlled, reversible test and inspect downstream state. | Documented cancellation outcome; no unintended additional action. |
| Authorization and isolation | 15% | Repeat with expired credentials, wrong tenant and insufficient privileges. | Access denied without disclosing another tenant’s data or performing a forbidden action. |
Formula: score = Σ(weight × passed ÷ attempted) ÷ Σ(applicable weights) × 100. Fixtures are individual tests, not production task counts. These weights are DN’s initial policy choice and are not empirically calibrated.
Release gates: a confirmed unauthorized action, cross-tenant disclosure, duplicated consequential write or falsely accepted final result blocks DN’s proposed release recommendation regardless of score. Unknown gate evidence leaves the recommendation pending. The score is a diagnostic summary, not a security certificate.
DN A2A Handoff Scorecard
Enter your fixture results. Defaults are fictional examples. This tool does not contact agents or establish protocol conformance. An excluded dimension must be outside your declared requirements, not an inconvenient failure.
Discovery and interface agreement · 15%
Delegation and context · 20%
Lifecycle and artifact integrity · 20%
Updates and recovery · 15%
Cancellation and side effects · 15%
Authorization and isolation · 15%
Illustrative only. No release recommendation.
54 / 60 applicable fixtures passed.
The export records your inputs, exclusions and gate selection. It contains no independently verified endpoint data. Inputs stay in this page and are not transmitted by this tool.
A reproducible test pack, not a polished demo
Freeze the task contract before running the test: the requested work, permissions, deadline, allowed side effects, expected output fields and acceptance rule. Keep the evaluator separate from the agent being assessed. Otherwise an agent can effectively grade its own response.
A useful first fixture is a read-only reconciliation handoff. The caller supplies synthetic records with a known total and asks a remote agent to produce a structured discrepancy report. The evaluator checks identifiers, totals and the reporting period. A valid-looking response with the wrong total fails the outcome test.
Next add an input-required case. Omit a necessary field, provide it after the agent requests clarification, and verify the final artifact uses that field exactly once. This tests continuity without requiring a live financial action. Run the same contract in both directions when both systems act as callers and recipients; compatibility can be asymmetric.
For recovery, interrupt delivery after the work has been accepted but before the caller sees the final result. Inspect the task’s eventual state using the supported retrieval or update mechanism. Count the original request once. Record every retry and any duplicate external effect separately. A retry that produces another order is a critical failure, not a successful recovery.
For isolation, use controlled test identities and synthetic tenant data. Submit an otherwise valid request with the wrong tenant or reduced privileges. The pass criterion is both refusal of the prohibited action and absence of leaked data. A refusal message containing another tenant’s confidential artifact still fails.
Evidence register to retain for every run
| Record | Why it matters | Publication treatment |
|---|---|---|
| Caller/recipient build; protocol, SDK and binding versions | Makes compatibility claims reproducible | Publish exact versions; do not generalize to untested builds |
| Card snapshot, required profile and fixture ID | Separates advertised support from observed behavior | Publish sanitized manifests and predeclared exclusions |
| Task/context identifiers, timestamps and attempt count | Reconstructs continuity and retries | Redact credentials and sensitive identifiers |
| Expected artifact, actual artifact and evaluator verdict | Exposes business failures hidden by completion states | Provide synthetic fixtures or safe reproducible summaries |
| Final downstream state and gate evidence | Detects duplicated or unauthorized actions | Report unresolved state as unknown, never as a pass |
This table is a proposed field register, not an observed performance dataset. Retain raw evidence privately where publishing it would expose credentials, personal data or customer records.
Choose the operating model before choosing a vendor
| Operating model | Best for | Avoid if | Costs and access | Control and risks |
|---|---|---|---|---|
| Read-only sandbox | Discovery, synthetic analysis and early interoperability tests | The task needs real execution | Compute/API usage and test setup; least-privilege test access | No funds entrusted by the fixture; data leakage still matters |
| Supervised handoff | Useful production work with approval before consequential action | Approval queues exceed the workflow deadline | Provider usage plus operator review time | Human approval can bound authority; review fatigue can weaken it |
| Bounded autonomous handoff | Repeated, narrow tasks with independently checked outcomes | Recovery, isolation or final state remains unknown | Monitoring, recovery and incident handling as well as model usage | Authority depends on credentials and downstream controls; a protocol does not define custody |
No provider is ranked or promoted in this edition. Operational status, geographic access, contractual terms and any custody arrangements must be checked for the actual service before deployment.
Measure cost per accepted handoff
A cheap request can become expensive when it needs retries, manual correction or incident handling. DN proposes tracking total run cost divided by independently accepted workflows. Include caller and recipient inference, API charges, retry work and review time. Report both the unit cost and the acceptance rate; neither explains the other.
For example, a fictional batch costs $12 to run and requires $18 of review. If 20 handoffs pass the acceptance rule, the cost is $1.50 per accepted handoff. This is arithmetic illustration only, not a vendor estimate. Failed work still consumes the budget and belongs in the numerator.
For a business decision, compare that cost with the same bounded task performed by the current process. Faster message exchange is not a commercial win if the accepted output rate falls or approvals become a bottleneck.
Methodology, limits and maintenance
Method: review official protocol documentation, separate interface features from business outcomes, define six test dimensions and attach independent release gates. This edition has no live measurements, no vendor scores and no sample-based confidence intervals. The calculator evaluates user inputs, not implementations.
Limitations: fixture pass rates depend on coverage, difficulty and environment. Correlated tests can inflate apparent evidence. A small clean sample cannot establish the absence of rare failures. Do not compare two scores unless their contracts, required profiles, observation windows and acceptance rules are comparable.
Maintenance: review documentation quarterly and rerun implementation fixtures after material changes. Keep measured reports versioned. A 2027 title is the planning horizon; this edition’s evidence date remains September 2026.
Change log: 30 September 2026, v1.0: initial proposed test protocol and local scorecard. Corrections: use the DN contact page with the disputed claim, version and primary evidence.
Your next step
Select one low-risk handoff. Write the acceptance rule and required capability profile, run the synthetic fixtures, then use the scorecard to identify the weakest dimension. Publish the evidence alongside any future score. Keep consequential actions supervised until the critical gates have been checked.
Explore the MCP Server Reliability Index for tool-level evidence, the Global AI Agent Identity Map for identity questions, and DN Pathfinder for broader platform selection.
Sources
- A2A v1.0.0 specification: interface, binding and task definitions.
- A2A Life of a Task: task workflow concepts. The latest URL may change after this review.
- A2A Enterprise Features: enterprise security context. The latest URL may change after this review.
Frequently asked questions
What does A2A interoperability mean?
Independent agent systems can exchange requests and task updates using compatible interfaces. DN additionally evaluates whether the recipient delivers an accepted result within the agreed limits.
Is this a tested vendor ranking?
No. This edition provides a proposed test protocol, fixture register and scorecard. No live vendor endpoints were tested.
Does an Agent Card prove that an agent is trustworthy?
No. It describes an interface and advertised capabilities. Verify the operator, authentication, authorization and actual behavior separately.
Can the calculator connect to an agent?
No. It calculates a weighted score from manually entered test outcomes and exports a local JSON record. It does not send requests or verify evidence.
Is cancellation the same as undoing a transaction?
No. Stopping further work does not necessarily reverse an external action. Check downstream state and use a separately authorized compensation process where appropriate.
How should optional capabilities be scored?
Declare the required capability profile before testing. Mark genuinely irrelevant features not applicable. A feature required by the workflow cannot be excluded merely because it fails.
What should a completed task prove?
A protocol terminal state alone does not prove business correctness. Apply an independent acceptance rule to the final artifact and check downstream state.
When should the test be rerun?
Rerun after changes to the protocol version, SDK, advertised interface, authentication, workflow policy or output contract. Retain the previous report for comparison.






