AI agent idempotency test plan
A safe retry is a proven system property. It is not an instruction asking the agent to remember what it already did.
September 7, 2026
AI Agent Idempotency Test Plan: Prove Tool Retries Cannot Duplicate Side Effects
An AI agent can call the right tool with the right arguments and still cause the wrong business outcome. A timeout arrives after the provider completed a refund. The orchestrator retries. The agent replans. A second refund, ticket, email, or shipment is created.
Prompting the agent to avoid duplicates does not solve this. A safe retry is a property of the tool boundary, durable state, provider contract, and recovery procedure. It must survive process crashes, delayed responses, concurrent workers, and a fresh agent session with no memory of the first attempt.
This test plan gives production teams an executable way to prove that property before an agent receives write access.
What the current evidence supports
Measured evidence: The current Runner observation produced no usable analytics opportunity for this topic. The article uses a controlled current-web fallback and makes no measured demand claim. The current FDE repository contains production-readiness, delivery, and incident guidance that mentions idempotency, but it does not contain a dedicated test plan.
External evidence: Stripe documents idempotency keys for safely retrying create and update requests, including parameter comparison on key reuse and returning the previously stored result. AWS Powertools documents a persistence record with an idempotency key, payload hash, in-progress state, expiry, and response data, plus an in-progress lock intended to block concurrent duplicates. These are vendor behaviors, not guarantees about an agent's entire workflow.
Engineering inference: A production team should generate stable operation identity outside the model, persist intent before the effect, serialize concurrent duplicates, store the provider reference, and reconcile any crash-after-effect ambiguity. The operation-ledger schema and failure matrix below are an original implementation pattern built from those constraints.
Define the exact side effect
Start with one tool and one business occurrence. Examples:
- issue one refund for return authorization
RMA-2048; - create one support ticket for incident
INC-731; - send one customer notification for order event
ORD-84:delayed; - provision one workspace for contract
CTR-901; - schedule one shipment pickup for parcel
PKG-55.
The word one is the invariant. A payload hash alone is often insufficient because two legitimate occurrences can have identical arguments. The idempotency key should represent the intended business occurrence, not simply whatever JSON the model produced.
A practical key can derive from:
tenant + workflow + business object + action + intended occurrence
For example:
acme:return:RMA-2048:refund:approved-v1
The orchestrator or application creates this key. Do not ask the language model to invent a fresh key on every attempt.
Put a durable operation ledger in front of the tool
Use one record per intended occurrence:
| Field | Purpose |
|---|---|
| operation_id | Internal stable identifier for the business occurrence |
| idempotency_key | Key sent to a provider or enforced locally |
| normalized_payload_hash | Detects a different request reusing the same identity |
| state | new, in_progress, outcome_unknown, succeeded, failed_terminal, or reconciling |
| external_reference | Provider refund, ticket, message, or job ID |
| attempt_count | Distinguishes execution attempts from business effects |
| lease_expiry | Releases a crashed worker without allowing blind duplication |
| result_hash | Proves later retries returned the same logical result |
| reconciliation_status | Records how an ambiguous outcome was resolved |
| test_evidence | Trace, fixture, fault, assertion, and timestamps |
Enforce a unique constraint on the business occurrence or idempotency key. A read-then-insert sequence without an atomic constraint is a concurrency bug disguised as duplicate checking.
Define the state machine
Use explicit transitions:
new -> in_progress -> succeeded
| |
| +-> return stored result on duplicate
|
+-> failed_terminal
|
+-> outcome_unknown -> reconciling -> succeeded
|-> safe_retry -> in_progress
|-> operator_review
An outcome_unknown state is essential. A timeout does not prove failure. If the external effect may have completed, retrying without reconciliation is unsafe.
Write invariants before tests
The test suite should assert business outcomes, not only HTTP status codes.
- One intended occurrence produces at most one external side effect.
- Identical retries return the same logical result or a stable in-progress response.
- A different normalized payload cannot reuse an existing key silently.
- Concurrent duplicates cannot both enter the provider effect.
- A crash after the provider effect does not cause a blind second effect.
- A stale in-progress lease is reconciled before retry.
- A terminal validation error cannot be transformed into success by retry noise.
- Operators can find the operation, effect, attempts, and reconciliation evidence without reading model conversation history.
Make the external system observable in tests. Use a provider sandbox, a faithful fake with an effect ledger, or a test account where the team can query the resulting object.
Run the failure-injection matrix
| Scenario | Injected fault | Required assertion | Evidence |
|---|---|---|---|
| Baseline | No fault | One effect, one success result | Ledger row and provider reference |
| Duplicate after success | Repeat same request and key | No second effect; stored result returned | Same external reference and result hash |
| Concurrent duplicate | Release two workers together | One worker owns the lease; one waits or returns in progress | Unique-key outcome and provider effect count |
| Payload drift | Same key, changed amount or recipient | Reject before effect | Hash mismatch event and zero new effects |
| Lost acknowledgement | Provider succeeds, response is dropped | Operation becomes unknown; reconciliation finds the effect | Provider reference recovered; one effect total |
| Crash before effect | Process exits after ledger claim | Expired lease permits safe execution | One effect after controlled retry |
| Crash after effect | Process exits after provider success but before local commit | Reconcile before retry | Provider query or webhook proves existing effect |
| Provider 5xx before execution | Provider confirms no effect or offers safe key semantics | Retry under the same key | One effect or none, never two |
| Provider timeout | Completion is ambiguous | Do not blind retry; enter reconciliation | Unknown-state trace and query result |
| Stale in-progress record | Lease expires without result | Reconcile ownership and provider state | Lease transfer and resolution record |
| Duplicate from replan | Fresh agent session requests same occurrence | Application maps it to existing operation | Same operation ID despite new agent run |
| Delayed webhook | Success event arrives after retry begins | Event converges on the existing operation | One state transition and one effect |
Each test should record the fault location precisely. "Timeout test passed" is not useful if the team cannot say whether the timeout occurred before submission, during provider execution, after provider commit, or while returning the acknowledgement.
Build a deterministic test harness
Place the agent behind a tool adapter that the test controls:
agent or workflow
|
v
tool contract -> operation service -> durable ledger -> provider adapter
| |
+-> fault switch +-> effect ledger
The harness needs five capabilities:
- Freeze the business occurrence and expected payload.
- Trigger one tool request through the same boundary used in production.
- Inject a fault at a named checkpoint.
- restart the worker or whole orchestration where required.
- Query both the local operation ledger and the provider-side effect.
Count effects from the external reference or provider test ledger. Counting tool-call attempts is not enough. Five attempts may correctly produce one effect, while one recorded attempt may hide two effects if the provider retried internally.
Test concurrency with a real barrier
A loop that invokes the tool twice in sequence does not test a race. Use a barrier so two workers reach the operation claim at the same time.
The test should prove:
- the unique constraint or conditional write selects one owner;
- only the owner can enter the provider call;
- the other worker returns a stable in-progress or completed result;
- an expired lease cannot be stolen while the current owner is healthy;
- lease renewal and clock behavior are observable;
- both callers eventually receive a consistent business result.
Repeat the race enough times to expose timing bugs, but define a deterministic assertion: external effect count equals one.
Treat payload drift as a release blocker
Suppose the first refund request is for USD 49 and a retry arrives with the same operation key but USD 94. Returning the original result without warning hides a serious mismatch. Executing the new payload creates a second or incorrect effect.
Normalize the fields that define the effect, hash them, and compare that hash on every reuse. Reject drift with an explicit error that sends the workflow to review. Do not normalize away fields that change business meaning.
Provider behavior varies. Stripe, for example, documents that its idempotency layer compares parameters when a key is reused. Your application should still enforce its own business-occurrence contract before calling any provider.
Reconcile crash-after-effect ambiguity
The hardest window is:
provider committed effect -> process crashed -> local success not saved
The operation ledger now says in_progress or outcome_unknown, while the provider may already contain the effect.
Use one of these recovery paths, in order:
- Retrieve the original result under the same provider idempotency key.
- Query the provider by your operation reference or metadata.
- Consume a provider webhook that includes the same occurrence identity.
- Search a bounded provider time window with additional matching fields.
- Route to an operator if the outcome cannot be proven.
Do not convert uncertainty into failure just because a lease expired. The release test must demonstrate the chosen reconciliation path.
Worked example: refund, ticket, and email
An agent handles an approved product return. The intended workflow has three side effects:
- issue one refund for
RMA-2048; - create one support ticket linked to the refund;
- send one customer confirmation after both records exist.
Treat them as three operation records, not one giant transaction:
| Operation | Stable occurrence | Dependency | Duplicate risk |
|---|---|---|---|
| Refund | RMA-2048:refund:approved-v1 | Human approval | Duplicate money movement |
| Ticket | RMA-2048:ticket:refund-followup-v1 | Refund reference | Duplicate queues and conflicting ownership |
RMA-2048:email:refund-confirmed-v1 | Refund and ticket references | Repeated customer message |
The refund provider commits, but the response is dropped. The local refund operation moves to outcome_unknown. The workflow must not create the ticket or email from an assumed success, and it must not send another refund request with a new key.
The reconciler queries the original key, recovers refund rf_782, stores it, and moves the refund operation to succeeded. The ticket tool then executes once with its own key and stores T-991. The email tool sends once only after both references exist.
Now inject a complete orchestrator restart before the customer email. A fresh agent may plan the email again, but the application maps the business occurrence to the existing operation. If the first email already succeeded, the duplicate returns the stored message result. If its outcome is unknown, reconciliation runs before any resend.
This example separates workflow replay from side-effect replay. Replaying the workflow is acceptable. Duplicating a committed effect is not.
Add negative tests
The system should also refuse unsafe uses:
- missing business object or intended occurrence;
- caller-supplied random key where a stable occurrence is required;
- reuse after the provider's documented retention window without local evidence;
- key reuse across tenants;
- changed amount, recipient, workspace, or permission scope;
- manual provider action that was not reconciled locally;
- an operation marked succeeded without an external reference where one should exist;
- retry after a terminal policy or validation denial;
- unbounded automatic reconciliation attempts.
Tenant identity must participate in the uniqueness boundary. A key collision between customers is not deduplication; it is cross-tenant corruption.
Set production release gates
Do not approve the tool for consequential writes until all applicable gates pass:
| Gate | Pass condition |
|---|---|
| Identity | Stable business occurrence maps to one operation ID |
| Atomicity | Concurrent claims produce one execution owner |
| Drift | Same key plus changed effect fields is rejected |
| Provider contract | Retention, retry, response, and query semantics are documented |
| Ambiguity | Crash-after-effect is reconciled without blind retry |
| Evidence | Operation, attempts, provider reference, and test fault are traceable |
| Operations | Alerts exist for unknown, stuck, repeated, and failed reconciliation states |
| Recovery | An operator can resolve ambiguity and resume safely |
| Load | Lease, persistence, and provider behavior pass at expected concurrency |
| Ownership | Platform, tool, provider, and business owners are named |
Production monitoring should include effect duplicates, payload-drift rejections, operations stuck in progress, unknown outcomes, reconciliation age, retry count, and provider-contract errors. An idempotency layer that silently absorbs every mismatch is not healthy.
Use this test-case record
Test case ID:
Tool and side effect:
Business occurrence:
Operation ID and idempotency key:
Normalized payload hash:
Initial ledger state:
Fault checkpoint:
Concurrent callers or restart condition:
Expected ledger transitions:
Expected provider effect count:
Expected external reference:
Expected caller results:
Actual trace and timestamps:
Reconciliation result:
Pass, fail, owner, and follow-up:
Prove the boundary, not the prompt
The production question is not whether the agent usually remembers an earlier action. It is whether the system can prove that retries, restarts, concurrency, and uncertain provider outcomes cannot create a second business effect.
Use the AI pilot-to-production readiness checklist to place idempotency inside the broader release review and the production AI incident runbook to define response when a duplicate or unknown outcome reaches production.
A forward deployed engineer should run these tests against the actual repository, persistence layer, provider sandbox, queues, and operator workflow. The deliverable is not a claim of exactly-once behavior. It is versioned evidence showing where duplicates are prevented, how uncertainty is reconciled, and who owns the exceptions.
Review status
Technical reliability review: Dhruv Khatri, pending. Provider-specific behavior must be retested against the integration and current provider documentation before release.
Sources
Need to harden an agent before production?
FDE embeds a senior engineer in your repository to test one bounded AI workflow against real integrations, retries, failures, and operating controls.
Book a 15-minute scoping call