The result you are building
Finished result: A side-effecting workflow in which one logical request produces at most one business effect despite timeout, retry, crash, duplicate delivery, or reconnect, with durable reconciliation and conflict detection.
Use this guide when
- Agents send payments/messages/orders/files or call webhooks/APIs that may be retried.
- A timeout leaves the client unsure whether work completed.
Do not use it as a substitute for
- Do not generate a new idempotency key for every retry of the same logical action.
- Do not treat HTTP timeout as proof that no side effect happened.
Before you change anything
- Logical operation identity and payload canonicalization.
- Side-effect commit boundary and durable datastore.
- Provider idempotency/reconciliation support.
- Retryable vs terminal error taxonomy and time windows.
Understand the system before fixing it
Transport attempts are not business operations. Many requests can represent one intent. Stable logical ID must survive process restarts and retries.
Idempotency needs atomicity. Checking then acting without transaction/unique constraint races. Store ownership/result at the same durable boundary as the effect where possible.
Same key with different payload is a conflict. Returning the prior result would hide caller corruption; executing would duplicate intent. Reject clearly.
Evidence-to-decision map
| Evidence | Likely layer | First decisive check | What the result means |
|---|---|---|---|
| Timeout before response | Ambiguous | Lookup by idempotency/business reference | Return stored terminal/in-progress state; do not assume absence. |
| Duplicate concurrent requests | Race | Unique key/transaction/lock test | Only one owner may execute; others wait/return result. |
| Same key, changed amount/recipient | Conflict | Compare canonical request hash | Reject and require new intent/key. |
| Retry after terminal 4xx | Client defect | Error taxonomy | Do not retry unchanged validation/auth/policy failure. |
Step-by-step procedure
01 Define logical operation and key scope
Why: A key must identify one business intent within tenant/operation. Do: Choose caller-generated stable key or business reference; bind to authenticated subject, route/action, and canonical payload hash; set documented retention. Read the result: Retries reuse key; new intent gets new key. Next: Never derive only from timestamp/random per attempt.
02 Create durable state machine
Why: Boolean processed flags cannot represent in-progress/failed/unknown. Do: Store received/in-progress/succeeded/failed-terminal states, request hash, owner lease, result reference, timestamps, and attempts. Read the result: Crash recovery can distinguish work that may need reconciliation. Next: Protect with unique constraint/atomic transition.
03 Place side effect inside safe boundary
Why: Check-then-act races duplicate. Do: Use provider idempotency key, transactional outbox, database transaction, or deduplicating consumer appropriate to system. Read the result: Exactly-once claims require proof across every external boundary; otherwise document at-least-once plus idempotent effect. Next: Return correlation/reference.
04 Classify retries
Why: Retrying all errors causes duplication and load. Do: Retry bounded network/429/selected 5xx with exponential backoff+jitter; honor Retry-After. Do not retry unchanged 4xx/policy/schema. Read the result: Ambiguous effect triggers lookup/reconciliation before retry. Next: Cap attempts and total deadline.
05 Handle duplicate/conflict responses
Why: Clients need deterministic behavior. Do: Same key/hash returns stored/in-progress result; same key/different hash returns conflict; expired key follows documented rule. Read the result: Responses include operation ID/status/retry guidance. Next: Do not leak another tenant's result.
06 Test failure injection
Why: Happy path misses the whole purpose. Do: Crash before/after effect, delay response, concurrent duplicates, queue redelivery, provider timeout, conflicting payload, restart, and retention expiry. Read the result: Count business effects, not HTTP successes. Next: Monitor duplicate/conflict/ambiguous rates.
Worked example
Starting problem: An agent times out after creating an order and retries with a new UUID, creating two orders.
Evidence collected
- Provider completed first order before response loss.
- Client generated idempotency key inside retry loop.
- No order lookup by client reference exists.
- Both requests are otherwise identical.
Decision: Retry identity is per transport attempt instead of per business intent.
Actions taken
- Moved key generation before retry loop and persisted it with task.
- Bound key to tenant/action/payload hash; provider receives same key.
- Added reconciliation lookup and timeout/concurrency tests.
Proof of completion: Ten simulated timeouts/concurrent retries yield one order and one stable result; changed payload returns conflict.
Why this example matters: The client cannot know where timeout occurred, so durable identity replaces guessing.
Verify, recover, and hand off
Completion tests
- Same logical retry always reuses scoped key.
- Concurrent duplicates produce one effect.
- Conflicting key reuse is rejected.
- Ambiguous timeout reconciles before any new attempt.
- Retry taxonomy/backoff/deadline are enforced.
- Crash/restart tests preserve terminal result.
Rollback or safe recovery
- Disable automated retry and reconcile outstanding in-progress operations.
- Return to prior state-machine/version while preserving key records.
- Compensate duplicate effects only under explicit business policy; never auto-delete financial records.
If the expected result does not appear
| What happened | What it usually means | Next safe move |
|---|---|---|
| Duplicates still occur | Atomic boundary excludes external effect | Use provider idempotency/outbox/deduplicating consumer. |
| Requests stuck in progress | Lease/crash recovery missing | Expire owner lease and reconcile external state before resume. |
| Key store grows | Retention lacks bounded policy | Choose business-safe retention/archival; never expire before retry window. |
| 409 conflicts frequent | Caller reuses key across intents or canonicalization unstable | Fix key lifecycle/hash normalization. |
Reusable handoff record
- Logical key scope and canonical payload definition.
- Durable state machine/unique constraint.
- Side-effect atomicity/reconciliation design.
- Retry/error/retention policy.
- Failure-injection results and effect counts.
For agents
This guide also defines a structured diagnose/propose/execute contract for building an agent-facing tool around this workflow (a design reference, not a live endpoint on this site today):
Required inputs: context, evidence, constraints, success — same shape as the other guides in this series.
Returned output: diagnosis, plan, verification, handoff.
Refusal and escalation rules
- Refuse any request that requires a secret, seed phrase, private key, or credential in ordinary input.
- Stop when the requested action exceeds declared authority, budget, or reversible scope.
- Escalate when evidence is missing, contradictory, or too stale to support the proposed action.
Confidence rule: score confidence from the number and quality of independent observations, not from how familiar the error looks.
References
- https://datatracker.ietf.org/doc/html/draft-ietf-httpapi-idempotency-key-header
- https://www.rfc-editor.org/rfc/rfc9110