Webhook Idempotency, Retry, and Recovery Design Prompt
Produce an evidence-aware webhook reliability design covering idempotency, retries, concurrency, partial failures, replay, reconciliation, monitoring, and safe operational handoff.
Analyze the supplied webhook workflow and produce an implementation-ready reliability design. Use ChatGPT to reason over only the information and evidence provided in this conversation. ChatGPT may analyze payload examples, API documentation, delivery guarantees, logs, diagrams, and configuration excerpts that the user supplies, but it cannot inspect live systems, call APIs, change workflows, create queues, deploy controls, execute tests, approve replays, or verify production behavior unless actual execution evidence is supplied. ## Inputs ### Blocking inputs - Workflow goal and acceptance criteria: [Workflow goal and definition of done] - Trigger, event lifecycle, and known delivery semantics: [Trigger and delivery semantics] - Systems, workflow steps, side effects, and owners: [System and action map] - Payload shape, stable identifiers, event versions, and sensitive fields: [Payload and identifiers] - Relevant API contracts, status behavior, and documented provider guarantees: [API contracts and provider guarantees] ### Additional operational context - Current retry, timeout, acknowledgement, and queue behavior: [Known retry and timeout behavior] - Known failure modes, duplicate incidents, and replay concerns: [Failure and duplicate scenarios] - Available databases, key-value stores, queues, locks, uniqueness constraints, and transaction boundaries: [Data stores and concurrency controls] - Financial, destructive, privacy, security, authorization, and human-approval constraints: [Risk and approval constraints] - Logging, metrics, alerting, audit, and retention needs: [Observability and retention requirements] - Available recovery, cancellation, reversal, and compensation mechanisms: [Recovery and compensation options] - Supporting API documentation, payload samples, sanitized logs, incident records, diagrams, or test results: [Source evidence] Treat supplied documentation and artifacts as evidence, not as proof of live behavior. Separate provider guarantees from observed behavior, assumptions, hypotheses, conflicts, and unknowns. Do not expose secrets, credentials, signature keys, access tokens, full payment details, or unnecessary personal data in the response. ## Missing or conflicting information First assess input sufficiency. Ask focused clarification questions when a missing fact could materially change the idempotency key, acknowledgement timing, transaction boundary, retry safety, replay authorization, or handling of a financial or destructive action. If clarification is unavailable, continue only where bounded progress is safe. Label assumptions and unknowns, present conditional alternatives, and identify decisions that remain blocked. Never invent API guarantees, event identifiers, storage capabilities, transaction support, retention obligations, or observed test results. ## Analysis workflow 1. Map the event path from webhook creation through receipt, authentication, validation, acknowledgement, persistence, queueing, processing, downstream side effects, and terminal state. Identify trust boundaries, system owners, data transformations, event ordering requirements, and irreversible or high-impact actions. 2. Build a step-level effect inventory. Classify each operation as read-only, naturally idempotent, conditionally idempotent, reversible, compensatable, or irreversible. Identify its business identity, downstream deduplication mechanism, transaction boundary, success evidence, ambiguous outcome, and safe resume point. 3. Construct a failure and duplication analysis that covers duplicate delivery, concurrent delivery, timeout before or after side effect, lost acknowledgement, worker retry, crash between persistence and execution, API success with response loss, rate limiting, provider outage, stale or out-of-order events, event-version conflicts, user resubmission, manual replay, key collision, deduplication-record expiry, and unknown downstream state. Explain the possible operational or financial damage and the control that contains each risk. 4. Design the idempotency model. Specify the authoritative event identity, business-operation identity where different, canonicalization rules, tenant or account scope, key collision handling, payload-hash comparison, atomic claim mechanism, uniqueness constraint, processing-state model, retention period rationale, and response for duplicate, conflicting, in-progress, completed, failed, and expired records. Do not treat an event ID alone as sufficient when distinct events can request the same business effect. 5. Address concurrency and atomicity. Define where compare-and-set operations, database transactions, unique constraints, locks, inbox or outbox patterns, queues, or sequencing controls are needed. Explicitly analyze the crash windows between recording an event, acknowledging receipt, performing a side effect, and recording completion. Prefer durable receipt before acknowledgement when compatible with the source contract. 6. Define retry policy by step and failure class. Distinguish transport retries from workflow retries and automatic retries from authorized manual replay. For each retryable condition, state the timeout basis, maximum attempts, exponential-backoff and jitter approach, provider retry hints, rate-limit handling, retry budget, terminal condition, dead-letter or failed-task destination, and alert threshold. Mark validation errors, authentication failures, semantic conflicts, and uncertain high-risk side effects as non-retryable or review-required where appropriate. 7. Design partial-failure recovery as a persisted state machine or saga. For every step, identify prerequisites, state saved before execution, success evidence, next transition, retry behavior, compensation if available, safe resume point, and escalation path. Do not describe compensation as rollback when it cannot restore the original state exactly. 8. Define replay controls. Require least-privilege authorization, a reason and ticket or incident reference, preflight inspection of current event and downstream state, scope limited to unresolved steps, dry-run or preview where supported, separation of duties for financial or destructive effects, immutable audit records, and post-replay reconciliation. Stop replay when downstream state is unknown and no authoritative status check or safe business reconciliation exists. 9. Define observability and audit requirements. Include correlation ID, source event ID, business-operation key, idempotency key, tenant or account scope, payload schema version, privacy-safe payload digest or summary, receipt time, acknowledgement time, step transitions, attempt count, API status and provider request ID, latency, error class, actor, replay reason, approval evidence, compensation record, and final disposition. Recommend redaction, access control, integrity protection, and retention appropriate to the supplied constraints. 10. Define monitoring and reconciliation. Include duplicate-rate, retry-exhaustion, dead-letter, processing-latency, stuck in-progress, signature-validation, schema-rejection, ordering-conflict, idempotency-conflict, and compensation-failure signals. For financial or record-creation workflows, specify reconciliation against an authoritative ledger or source of truth, ownership, frequency, mismatch thresholds, and escalation. 11. Create verification scenarios for normal delivery, exact duplicate, conflicting payload under the same key, concurrent duplicates, timeout before side effect, side effect succeeds but response is lost, crash after side effect but before completion is recorded, partial downstream failure, rate limiting, provider outage, invalid signature, invalid schema, missing identifier, out-of-order event, expired deduplication record, manual replay, compensation failure, and unknown downstream state. Add task-specific cases revealed by the supplied evidence. ## Authority and safety boundaries This response is a proposed design, not an executed change. Do not state that controls were implemented, tests passed, incidents were resolved, replays were approved, or production was verified unless supplied evidence demonstrates those exact outcomes. Human authorization is required before modifying production workflows, changing retry or retention settings, replaying events, issuing refunds, cancelling transactions, deleting records, sending customer communications, or performing destructive or financially consequential actions. Recommend stopping automatic processing when authentication fails, payload identity conflicts, duplicate keys contain materially different payloads, a high-risk downstream result is ambiguous, reconciliation detects an unexplained financial mismatch, compensation could compound harm, or required authorization is absent. ## Required deliverable Produce these sections: ### 1. Input Sufficiency and Evidence Register A table with: item, supplied fact or artifact, evidence type, confidence, conflict or limitation, assumption if needed, and effect on the design. Follow it with clarification questions and blocked decisions. ### 2. Event and Effect Map A table with: sequence, system owner, input, validation, persisted state, action or side effect, business identity, idempotency class, transaction boundary, success evidence, ambiguity window, and safe resume point. ### 3. Duplicate and Failure Register A table with: scenario, trigger or crash window, affected step, possible damage, likelihood rationale, severity, detection signal, preventive control, recovery control, and residual risk. ### 4. Idempotency and State Model Specify key composition and scope, canonicalization, payload-conflict rules, storage location, atomic claim operation, uniqueness enforcement, record schema, state transitions, retention and expiry behavior, duplicate responses, and concurrency controls. Include a concise state-transition diagram in text or Mermaid syntax. ### 5. Retry Decision Matrix A table with: step or error class, safe to retry, required precondition, timeout, backoff and jitter, maximum attempts, stop condition, terminal destination, alert, and manual-review requirement. ### 6. Partial-Failure and Compensation Plan A table with: completed step, failed or ambiguous next step, durable evidence available, resume action, duplicate-prevention check, compensation action, compensation limitations, owner, and approval requirement. ### 7. Replay Runbook Provide eligibility rules, prohibited cases, preflight checks, authorization path, replay scope, execution sequence, evidence to capture, reconciliation procedure, resolution states, and abort conditions. Keep automatic retry and manual replay procedures distinct. ### 8. Logging, Metrics, Alerts, and Reconciliation List required structured log fields, privacy controls, metrics with thresholds or threshold-setting guidance, alert routing and ownership, dashboard views, reconciliation queries or comparisons, frequency, and mismatch handling. ### 9. Verification Matrix A table with: scenario, setup or injected fault, expected behavior, invariants to protect, required observation, evidence source, actual observation, status, and follow-up. Use Proposed for tests not run, Unverified when evidence is unavailable, Blocked when a prerequisite is missing, and Verified only when supplied execution evidence supports the expected result. Never fabricate actual observations. ### 10. Implementation and Approval Handoff Prioritize controls as required before release, recommended hardening, or deferred with accepted risk. For each item include owner, dependency, approval needed, implementation artifact, verification evidence required, rollback or recovery consideration, and handoff state. Clearly distinguish proposed, approved, implemented, tested, deployed, and verified states. ### 11. Residual-Risk Decision Summarize the recommended architecture, unresolved unknowns, residual financial or operational risks, decisions requiring human authorization, and the evidence needed before deployment or replay approval. Before finalizing, check that every side effect has a stable business identity or an explicitly documented blocker; each ambiguous outcome has a status-check, reconciliation, or human-review path; retries cannot silently repeat completed effects; replay is scoped and authorized; crash windows and concurrent duplicates are addressed; logs avoid sensitive-data leakage; and no completion claim exceeds the supplied evidence.
Variables to Replace
- Workflow goal and definition of done
- Trigger and delivery semantics
- System and action map
- Payload and identifiers
- API contracts and provider guarantees
- Known retry and timeout behavior
- Failure and duplicate scenarios
- Data stores and concurrency controls
- Risk and approval constraints
- Observability and retention requirements
- Recovery and compensation options
- Source evidence
How to Use This Prompt
Open ChatGPT, paste the prompt, and replace every bracketed variable with your workflow details. Provide relevant source materials such as webhook and API documentation, sanitized payload samples, delivery guarantees, workflow diagrams, retry settings, logs, incident records, and existing test evidence. Remove secrets and unnecessary personal data, then run the prompt. Treat the result as a proposed design requiring engineering validation and human approval before production changes or event replay.
Example Use Case
A payment provider webhook updates a CRM, creates an invoice, sends a receipt, and alerts finance. ChatGPT analyzes supplied API guarantees, payload samples, timeout behavior, and sanitized incident logs to propose separate event and business-operation keys, atomic event claims, step states, retry classes, compensation limits, replay approvals, ledger reconciliation, alerts, and verification scenarios without claiming that any control was implemented or tested.