Guide / Reliability

AI agent idempotency test plan

A safe retry is a proven system property. It is not an instruction asking the agent to remember what it already did.

September 7, 2026

AI Agent Idempotency Test Plan: Prove Tool Retries Cannot Duplicate Side Effects

An AI agent can call the right tool with the right arguments and still cause the wrong business outcome. A timeout arrives after the provider completed a refund. The orchestrator retries. The agent replans. A second refund, ticket, email, or shipment is created.

Prompting the agent to avoid duplicates does not solve this. A safe retry is a property of the tool boundary, durable state, provider contract, and recovery procedure. It must survive process crashes, delayed responses, concurrent workers, and a fresh agent session with no memory of the first attempt.

This test plan gives production teams an executable way to prove that property before an agent receives write access.

What the current evidence supports

Measured evidence: The current Runner observation produced no usable analytics opportunity for this topic. The article uses a controlled current-web fallback and makes no measured demand claim. The current FDE repository contains production-readiness, delivery, and incident guidance that mentions idempotency, but it does not contain a dedicated test plan.

External evidence: Stripe documents idempotency keys for safely retrying create and update requests, including parameter comparison on key reuse and returning the previously stored result. AWS Powertools documents a persistence record with an idempotency key, payload hash, in-progress state, expiry, and response data, plus an in-progress lock intended to block concurrent duplicates. These are vendor behaviors, not guarantees about an agent's entire workflow.

Engineering inference: A production team should generate stable operation identity outside the model, persist intent before the effect, serialize concurrent duplicates, store the provider reference, and reconcile any crash-after-effect ambiguity. The operation-ledger schema and failure matrix below are an original implementation pattern built from those constraints.

Define the exact side effect

Start with one tool and one business occurrence. Examples:

  • issue one refund for return authorization RMA-2048;
  • create one support ticket for incident INC-731;
  • send one customer notification for order event ORD-84:delayed;
  • provision one workspace for contract CTR-901;
  • schedule one shipment pickup for parcel PKG-55.

The word one is the invariant. A payload hash alone is often insufficient because two legitimate occurrences can have identical arguments. The idempotency key should represent the intended business occurrence, not simply whatever JSON the model produced.

A practical key can derive from:

tenant + workflow + business object + action + intended occurrence

For example:

acme:return:RMA-2048:refund:approved-v1

The orchestrator or application creates this key. Do not ask the language model to invent a fresh key on every attempt.

Put a durable operation ledger in front of the tool

Use one record per intended occurrence:

FieldPurpose
operation_idInternal stable identifier for the business occurrence
idempotency_keyKey sent to a provider or enforced locally
normalized_payload_hashDetects a different request reusing the same identity
statenew, in_progress, outcome_unknown, succeeded, failed_terminal, or reconciling
external_referenceProvider refund, ticket, message, or job ID
attempt_countDistinguishes execution attempts from business effects
lease_expiryReleases a crashed worker without allowing blind duplication
result_hashProves later retries returned the same logical result
reconciliation_statusRecords how an ambiguous outcome was resolved
test_evidenceTrace, fixture, fault, assertion, and timestamps

Enforce a unique constraint on the business occurrence or idempotency key. A read-then-insert sequence without an atomic constraint is a concurrency bug disguised as duplicate checking.

Define the state machine

Use explicit transitions:

new -> in_progress -> succeeded
          |              |
          |              +-> return stored result on duplicate
          |
          +-> failed_terminal
          |
          +-> outcome_unknown -> reconciling -> succeeded
                                           |-> safe_retry -> in_progress
                                           |-> operator_review

An outcome_unknown state is essential. A timeout does not prove failure. If the external effect may have completed, retrying without reconciliation is unsafe.

Write invariants before tests

The test suite should assert business outcomes, not only HTTP status codes.

  1. One intended occurrence produces at most one external side effect.
  2. Identical retries return the same logical result or a stable in-progress response.
  3. A different normalized payload cannot reuse an existing key silently.
  4. Concurrent duplicates cannot both enter the provider effect.
  5. A crash after the provider effect does not cause a blind second effect.
  6. A stale in-progress lease is reconciled before retry.
  7. A terminal validation error cannot be transformed into success by retry noise.
  8. Operators can find the operation, effect, attempts, and reconciliation evidence without reading model conversation history.

Make the external system observable in tests. Use a provider sandbox, a faithful fake with an effect ledger, or a test account where the team can query the resulting object.

Run the failure-injection matrix

ScenarioInjected faultRequired assertionEvidence
BaselineNo faultOne effect, one success resultLedger row and provider reference
Duplicate after successRepeat same request and keyNo second effect; stored result returnedSame external reference and result hash
Concurrent duplicateRelease two workers togetherOne worker owns the lease; one waits or returns in progressUnique-key outcome and provider effect count
Payload driftSame key, changed amount or recipientReject before effectHash mismatch event and zero new effects
Lost acknowledgementProvider succeeds, response is droppedOperation becomes unknown; reconciliation finds the effectProvider reference recovered; one effect total
Crash before effectProcess exits after ledger claimExpired lease permits safe executionOne effect after controlled retry
Crash after effectProcess exits after provider success but before local commitReconcile before retryProvider query or webhook proves existing effect
Provider 5xx before executionProvider confirms no effect or offers safe key semanticsRetry under the same keyOne effect or none, never two
Provider timeoutCompletion is ambiguousDo not blind retry; enter reconciliationUnknown-state trace and query result
Stale in-progress recordLease expires without resultReconcile ownership and provider stateLease transfer and resolution record
Duplicate from replanFresh agent session requests same occurrenceApplication maps it to existing operationSame operation ID despite new agent run
Delayed webhookSuccess event arrives after retry beginsEvent converges on the existing operationOne state transition and one effect

Each test should record the fault location precisely. "Timeout test passed" is not useful if the team cannot say whether the timeout occurred before submission, during provider execution, after provider commit, or while returning the acknowledgement.

Build a deterministic test harness

Place the agent behind a tool adapter that the test controls:

agent or workflow
      |
      v
tool contract -> operation service -> durable ledger -> provider adapter
                                      |                |
                                      +-> fault switch +-> effect ledger

The harness needs five capabilities:

  1. Freeze the business occurrence and expected payload.
  2. Trigger one tool request through the same boundary used in production.
  3. Inject a fault at a named checkpoint.
  4. restart the worker or whole orchestration where required.
  5. Query both the local operation ledger and the provider-side effect.

Count effects from the external reference or provider test ledger. Counting tool-call attempts is not enough. Five attempts may correctly produce one effect, while one recorded attempt may hide two effects if the provider retried internally.

Test concurrency with a real barrier

A loop that invokes the tool twice in sequence does not test a race. Use a barrier so two workers reach the operation claim at the same time.

The test should prove:

  • the unique constraint or conditional write selects one owner;
  • only the owner can enter the provider call;
  • the other worker returns a stable in-progress or completed result;
  • an expired lease cannot be stolen while the current owner is healthy;
  • lease renewal and clock behavior are observable;
  • both callers eventually receive a consistent business result.

Repeat the race enough times to expose timing bugs, but define a deterministic assertion: external effect count equals one.

Treat payload drift as a release blocker

Suppose the first refund request is for USD 49 and a retry arrives with the same operation key but USD 94. Returning the original result without warning hides a serious mismatch. Executing the new payload creates a second or incorrect effect.

Normalize the fields that define the effect, hash them, and compare that hash on every reuse. Reject drift with an explicit error that sends the workflow to review. Do not normalize away fields that change business meaning.

Provider behavior varies. Stripe, for example, documents that its idempotency layer compares parameters when a key is reused. Your application should still enforce its own business-occurrence contract before calling any provider.

Reconcile crash-after-effect ambiguity

The hardest window is:

provider committed effect -> process crashed -> local success not saved

The operation ledger now says in_progress or outcome_unknown, while the provider may already contain the effect.

Use one of these recovery paths, in order:

  1. Retrieve the original result under the same provider idempotency key.
  2. Query the provider by your operation reference or metadata.
  3. Consume a provider webhook that includes the same occurrence identity.
  4. Search a bounded provider time window with additional matching fields.
  5. Route to an operator if the outcome cannot be proven.

Do not convert uncertainty into failure just because a lease expired. The release test must demonstrate the chosen reconciliation path.

Worked example: refund, ticket, and email

An agent handles an approved product return. The intended workflow has three side effects:

  1. issue one refund for RMA-2048;
  2. create one support ticket linked to the refund;
  3. send one customer confirmation after both records exist.

Treat them as three operation records, not one giant transaction:

OperationStable occurrenceDependencyDuplicate risk
RefundRMA-2048:refund:approved-v1Human approvalDuplicate money movement
TicketRMA-2048:ticket:refund-followup-v1Refund referenceDuplicate queues and conflicting ownership
EmailRMA-2048:email:refund-confirmed-v1Refund and ticket referencesRepeated customer message

The refund provider commits, but the response is dropped. The local refund operation moves to outcome_unknown. The workflow must not create the ticket or email from an assumed success, and it must not send another refund request with a new key.

The reconciler queries the original key, recovers refund rf_782, stores it, and moves the refund operation to succeeded. The ticket tool then executes once with its own key and stores T-991. The email tool sends once only after both references exist.

Now inject a complete orchestrator restart before the customer email. A fresh agent may plan the email again, but the application maps the business occurrence to the existing operation. If the first email already succeeded, the duplicate returns the stored message result. If its outcome is unknown, reconciliation runs before any resend.

This example separates workflow replay from side-effect replay. Replaying the workflow is acceptable. Duplicating a committed effect is not.

Add negative tests

The system should also refuse unsafe uses:

  • missing business object or intended occurrence;
  • caller-supplied random key where a stable occurrence is required;
  • reuse after the provider's documented retention window without local evidence;
  • key reuse across tenants;
  • changed amount, recipient, workspace, or permission scope;
  • manual provider action that was not reconciled locally;
  • an operation marked succeeded without an external reference where one should exist;
  • retry after a terminal policy or validation denial;
  • unbounded automatic reconciliation attempts.

Tenant identity must participate in the uniqueness boundary. A key collision between customers is not deduplication; it is cross-tenant corruption.

Set production release gates

Do not approve the tool for consequential writes until all applicable gates pass:

GatePass condition
IdentityStable business occurrence maps to one operation ID
AtomicityConcurrent claims produce one execution owner
DriftSame key plus changed effect fields is rejected
Provider contractRetention, retry, response, and query semantics are documented
AmbiguityCrash-after-effect is reconciled without blind retry
EvidenceOperation, attempts, provider reference, and test fault are traceable
OperationsAlerts exist for unknown, stuck, repeated, and failed reconciliation states
RecoveryAn operator can resolve ambiguity and resume safely
LoadLease, persistence, and provider behavior pass at expected concurrency
OwnershipPlatform, tool, provider, and business owners are named

Production monitoring should include effect duplicates, payload-drift rejections, operations stuck in progress, unknown outcomes, reconciliation age, retry count, and provider-contract errors. An idempotency layer that silently absorbs every mismatch is not healthy.

Use this test-case record

Test case ID:
Tool and side effect:
Business occurrence:
Operation ID and idempotency key:
Normalized payload hash:
Initial ledger state:
Fault checkpoint:
Concurrent callers or restart condition:
Expected ledger transitions:
Expected provider effect count:
Expected external reference:
Expected caller results:
Actual trace and timestamps:
Reconciliation result:
Pass, fail, owner, and follow-up:

Prove the boundary, not the prompt

The production question is not whether the agent usually remembers an earlier action. It is whether the system can prove that retries, restarts, concurrency, and uncertain provider outcomes cannot create a second business effect.

Use the AI pilot-to-production readiness checklist to place idempotency inside the broader release review and the production AI incident runbook to define response when a duplicate or unknown outcome reaches production.

A forward deployed engineer should run these tests against the actual repository, persistence layer, provider sandbox, queues, and operator workflow. The deliverable is not a claim of exactly-once behavior. It is versioned evidence showing where duplicates are prevented, how uncertainty is reconciled, and who owns the exceptions.

Review status

Technical reliability review: Dhruv Khatri, pending. Provider-specific behavior must be retested against the integration and current provider documentation before release.

Sources

Need to harden an agent before production?

FDE embeds a senior engineer in your repository to test one bounded AI workflow against real integrations, retries, failures, and operating controls.

Book a 15-minute scoping call