Reusable AI capability
Recover Durable Agent Execution State
Reconcile checkpoints, durable state, side effects, approvals, and idempotency after interruption to select and gate a safe resume, replay, compensate, or abort path.
This Skill packages a reusable way to use the linked Prompt or Workflow; Amo.ng does not run it for you.
# Recover Durable Agent Execution State Skill ID: AMO-S-000020 Skill URL: https://amo.ng/skills/recover-durable-agent-execution-state Purpose: Give agent platform and service owners a reusable recovery-state method for interrupted long-running executions across workflows, tools, and external systems without claiming that recovery actions occurred. Required inputs: - Interrupted run identity, intended outcome, workflow and configuration versions - Durable state, checkpoints, messages, tool records, approvals, and external side effects - Idempotency keys, transaction or compensation semantics, retry history, and known-good state - Incident timeline, containment status, recovery options, constraints, and accountable owners How to use: When to use: - An agent run stopped after partial execution and its safe continuation state is uncertain. - Duplicate, omitted, or conflicting side effects must be reconciled before restart. When not to use: - Stateless request retries with no durable state or external side effect. - Authorizing production replay, compensation, or rollback. - Treating the latest timestamp as proof of the latest valid checkpoint. Reusable method: 1. Freeze the run, workflow and configuration versions, intended outcome, interruption window, and authority boundary. 2. Build a checkpoint and side-effect ledger from state, messages, tools, approvals, and external-system evidence. 3. Classify each action as committed, pending, failed, duplicated, compensated, ambiguous, or not evidenced. 4. Establish the latest trustworthy recovery boundary and identify stale, corrupt, or conflicting state. 5. Compare Resume, Replay, Compensate, Abort, and Continue investigation paths for idempotency, dependency, data, customer, and reversibility risk. 6. Define prerequisites, smallest safe action, owner authorization, stop condition, verification, rollback, and monitoring for the chosen path. Expected output: A recovery-state record with checkpoint lineage, side-effect ledger, ambiguity and evidence gaps, path comparison, selected disposition, prerequisites, compensation plan, verification, and owner gates. Boundaries: Do not invent state, side effects, commands, approvals, or recovery. Service and incident owners choose the operating path; data, financial, security, and external-system owners authorize actions in their systems; the release owner controls restart. Source: AMO-P-000297. Applicable Workflow: Recover and Requalify an Interrupted AI Agent Runtime. Powered by Prompt: Long-Running Agent State Recovery Decision Source ID: AMO-P-000297 https://amo.ng/prompts/long-running-agent-state-recovery-decision Completion criteria: Complete when every material side effect and checkpoint is classified; the latest trustworthy boundary and uncertainty are explicit; one recovery disposition has prerequisites, owner, stop and rollback conditions; and observed recovery is separated from the proposed plan. Use this Amo.ng Skill with your preferred AI tool. Supply the required inputs and follow the usage instructions. # Recover Durable Agent Execution State Skill ID: AMO-S-000020 Skill URL: https://amo.ng/skills/recover-durable-agent-execution-state Purpose: Give agent platform and service owners a reusable recovery-state method for interrupted long-running executions across workflows, tools, and external systems without claiming that recovery actions occurred. Required inputs: - Interrupted run identity, intended outcome, workflow and configuration versions - Durable state, checkpoints, messages, tool records, approvals, and external side effects - Idempotency keys, transaction or compensation semantics, retry history, and known-good state - Incident timeline, containment status, recovery options, constraints, and accountable owners How to use: When to use: - An agent run stopped after partial execution and its safe continuation state is uncertain. - Duplicate, omitted, or conflicting side effects must be reconciled before restart. When not to use: - Stateless request retries with no durable state or external side effect. - Authorizing production replay, compensation, or rollback. - Treating the latest timestamp as proof of the latest valid checkpoint. Reusable method: 1. Freeze the run, workflow and configuration versions, intended outcome, interruption window, and authority boundary. 2. Build a checkpoint and side-effect ledger from state, messages, tools, approvals, and external-system evidence. 3. Classify each action as committed, pending, failed, duplicated, compensated, ambiguous, or not evidenced. 4. Establish the latest trustworthy recovery boundary and identify stale, corrupt, or conflicting state. 5. Compare Resume, Replay, Compensate, Abort, and Continue investigation paths for idempotency, dependency, data, customer, and reversibility risk. 6. Define prerequisites, smallest safe action, owner authorization, stop condition, verification, rollback, and monitoring for the chosen path. Expected output: A recovery-state record with checkpoint lineage, side-effect ledger, ambiguity and evidence gaps, path comparison, selected disposition, prerequisites, compensation plan, verification, and owner gates. Boundaries: Do not invent state, side effects, commands, approvals, or recovery. Service and incident owners choose the operating path; data, financial, security, and external-system owners authorize actions in their systems; the release owner controls restart. Source: AMO-P-000297. Applicable Workflow: Recover and Requalify an Interrupted AI Agent Runtime. Powered by Prompt: Long-Running Agent State Recovery Decision Source ID: AMO-P-000297 https://amo.ng/prompts/long-running-agent-state-recovery-decision Completion criteria: Complete when every material side effect and checkpoint is classified; the latest trustworthy boundary and uncertainty are explicit; one recovery disposition has prerequisites, owner, stop and rollback conditions; and observed recovery is separated from the proposed plan.Copy skill copies the Skill details. Use with AI adds a short instruction for your preferred AI tool; neither action runs the Skill.
Purpose
Give agent platform and service owners a reusable recovery-state method for interrupted long-running executions across workflows, tools, and external systems without claiming that recovery actions occurred.
Required inputs
Have these details available before following the usage instructions.
- Interrupted run identity, intended outcome, workflow and configuration versions
- Durable state, checkpoints, messages, tool records, approvals, and external side effects
- Idempotency keys, transaction or compensation semantics, retry history, and known-good state
- Incident timeline, containment status, recovery options, constraints, and accountable owners
How to use this Skill
When to use:
- An agent run stopped after partial execution and its safe continuation state is uncertain.
- Duplicate, omitted, or conflicting side effects must be reconciled before restart.
When not to use:
- Stateless request retries with no durable state or external side effect.
- Authorizing production replay, compensation, or rollback.
- Treating the latest timestamp as proof of the latest valid checkpoint.
Reusable method:
1. Freeze the run, workflow and configuration versions, intended outcome, interruption window, and authority boundary.
2. Build a checkpoint and side-effect ledger from state, messages, tools, approvals, and external-system evidence.
3. Classify each action as committed, pending, failed, duplicated, compensated, ambiguous, or not evidenced.
4. Establish the latest trustworthy recovery boundary and identify stale, corrupt, or conflicting state.
5. Compare Resume, Replay, Compensate, Abort, and Continue investigation paths for idempotency, dependency, data, customer, and reversibility risk.
6. Define prerequisites, smallest safe action, owner authorization, stop condition, verification, rollback, and monitoring for the chosen path.
Expected output:
A recovery-state record with checkpoint lineage, side-effect ledger, ambiguity and evidence gaps, path comparison, selected disposition, prerequisites, compensation plan, verification, and owner gates.
Boundaries:
Do not invent state, side effects, commands, approvals, or recovery. Service and incident owners choose the operating path; data, financial, security, and external-system owners authorize actions in their systems; the release owner controls restart. Source: AMO-P-000297. Applicable Workflow: Recover and Requalify an Interrupted AI Agent Runtime.
Powered by an Amo.ng Prompt
Long-Running Agent State Recovery Decision
Open the linked prompt to use the instructions that power this Skill.
Completion criteria
Complete when every material side effect and checkpoint is classified; the latest trustworthy boundary and uncertainty are explicit; one recovery disposition has prerequisites, owner, stop and rollback conditions; and observed recovery is separated from the proposed plan.
Explore related Workflows
Browse WorkflowsRecover and Requalify an Interrupted AI Agent Runtime
Reconcile interrupted durable state and side effects, conditionally recertify tool access and fallback behavior, then calibrate escalation controls before resuming an AI agent runtime.
Related Prompts
Browse PromptsAgent Escalation Threshold Calibration
Tune agent escalation triggers using incident severity, uncertainty, false-positive and false-negative evidence, queue capacity, delay, and owner authority.
Calibrate escalation thresholds for an AI agent using evidence about severity, uncertainty, control failures, reviewer capacity, delay, and actual outcomes. Distinguish thresholds that pause work from those that route routine review or trigger urgent containment. Provide: - Agent tasks, user groups, decisions/actions, risk classes, protected outcomes, and unacceptable failures: [Agent decisions and risk classes] - Rules, scores, confidence/uncertainty signals, triggers, routes, priorities, context package, and fallback behavior: [Current escalation logic and routes] - Incidents, near misses, reviewer decisions, overrides, missed escalations, unnecessary escalations, user requests, and downstream outcomes: [Incident feedback and outcome evidence] - Score distributions, calibration, labels, disagreement, slices, drift, missingness, and historical threshold changes: [Score uncertainty and threshold evidence] - Arrival volume, service targets, reviewer skill, coverage hours, queue age, abandonment, and surge capacity: [Queue capacity and service constraints] - Non-delegable approvals, service owner, product owner, review owner, security/privacy reviewer, and emergency authority: [Authority boundaries and accountable owners] Do not invent score distributions, error rates, or queue behavior. Distinguish observed outcome evidence from inference. Do not optimize for fewer escalations without pricing missed harm, or for maximum sensitivity without considering delay and reviewer overload. Separate uncertain model scores from independently observed risk triggers. Preserve user-requested escalation and mandatory policy triggers even if statistical tuning suggests otherwise. Calibration: 1. Define escalation classes. Specify Continue, Clarify, Abstain, Routine review, Priority review, Immediate pause/containment, and Emergency response as applicable. State purpose, owner, service target, and allowed agent behavior while waiting. 2. Build the outcome ledger. Reconcile current trigger decisions with reviewer outcomes, incidents, corrections, and downstream consequences. Identify true positive, false positive, true negative, false negative, and not-assessable cases only where labels support them. 3. Audit signals. Review definition, calibration, stability, missing values, manipulability, protected slices, and independence. Identify hard rules that must override probabilistic thresholds. 4. Model threshold trade-offs. Compare candidate thresholds using supplied confusion, severity, workload, and delay evidence. Show how volume and queueing change, and identify where performance is not estimable. 5. Protect critical slices and events. Define zero- or low-tolerance triggers for irreversible action, sensitive data, safety, legal/compliance, identity/authorization, user distress, or explicit review requests where applicable. Assign qualified routes. 6. Test operational feasibility. Check reviewer capacity, skills, hours, context quality, queue priority, fallback, and outage behavior. A threshold is not deployable if the designated route cannot meet the required response. 7. Recommend thresholds and rollout. Specify thresholds, hard triggers, hysteresis or cooldown if needed, context package, staged rollout, shadow comparison, monitoring, drift and queue alerts, stop conditions, and recalibration cadence. Use ChatGPT to analyze the supplied incidents, thresholds, and reviewer-capacity evidence; do not claim simulation, deployment, or threshold validation occurred unless results are supplied. Keep proposed threshold changes distinct from approved configuration, and identify the accountable operations owner and release owner for authorization. Required deliverable: # Agent Escalation Threshold Calibration ## Escalation Classes and Authority | Class | Trigger purpose | Agent behavior while pending | Route/owner | Service target | Mandatory rule | |---|---|---|---|---|---| ## Outcome and Error Ledger | Slice/trigger | Escalations | Confirmed needed | Unnecessary | Missed | Consequence | Evidence quality | |---|---:|---:|---:|---:|---|---| ## Signal Fitness | Signal | Meaning | Calibration/stability | Missing/manipulation risk | Fit for threshold? | |---|---|---|---|---| ## Candidate Thresholds | Slice/class | Threshold/rule | Expected miss/over-route effect | Queue impact | Risk | Evidence limit | |---|---|---|---|---|---| ## Recommended Policy | Slice/event | Continue threshold | Review threshold | Pause/containment trigger | Route | Owner | |---|---|---|---|---|---| ## Rollout and Recalibration - Shadow/canary plan: - Queue and incident stop conditions: - Protected hard rules: - Drift/recalibration triggers: - Evidence still needed: Completion requires outcome-linked thresholds, feasible review routes, preserved mandatory authority boundaries, and monitoring that detects both missed harm and overload.Model Fallback Failure Analysis
Reconstruct a failed model fallback decision, test contract compatibility across routes, and determine whether to repair, restrict, or disable fallback behavior.
Analyze a production failure in which traffic was routed from a primary model or configuration to a fallback and the fallback did not preserve the required service contract. Evidence to provide: - Routing rules, triggers, priorities, circuit breakers, and fallback chain: [Routing and fallback policy] - Required and observed capabilities, context limits, schemas, tool behavior, safety controls, and response contracts: [Primary and fallback model contracts] - Timestamped routing decisions, requests, responses, errors, retries, and downstream effects: [Incident traces and outputs] - Relevant offline evaluations, canary results, service metrics, user feedback, and accepted thresholds: [Evaluation and operational evidence] - Data residency, privacy, security, policy, tool-access, and contractual constraints: [Data tool and compliance constraints] - Repair, restriction, disablement, and communication options with accountable owners: [Recovery options and owners] Do not assume that a fallback is a drop-in substitute because it accepts the same request shape. Separate routing evidence, model behavior, integration behavior, and downstream handling. Do not claim a live reproduction or test unless results are supplied. Label missing route decisions, hidden provider behavior, and unobserved failures as unknown. Analysis: 1. Restate the service contract. Define the minimum quality, safety, structured-output, tool-use, latency, availability, data-boundary, and observability requirements that every fallback route must preserve. Distinguish mandatory invariants from degradable features. 2. Reconstruct the fallback path. Identify the trigger, routing decision, model/config selected, request transformation, context truncation, tool or schema adaptation, response validation, retries, and final downstream action. Mark the first point where evidence diverges from expected behavior. 3. Classify incompatibilities. Examine capability gaps, prompt/config differences, unsupported tools, schema drift, context loss, safety-policy differences, modality gaps, tokenizer or stop behavior, latency budget exhaustion, data-location conflicts, and error-normalization problems. Distinguish confirmed causes from plausible contributors. 4. Assess blast radius. Determine affected traffic slices, time window, user groups, task types, regions, and downstream systems using supplied evidence. Identify silent failures, incorrect success signals, unsafe outputs, repeated side effects, and records needing revalidation. 5. Evaluate controls. Review pre-route eligibility, health checks, compatibility tests, response validation, confidence or abstention rules, circuit breaking, observability, and disablement controls. Explain which control should have detected or contained the failure and why it did not. 6. Compare dispositions. Assess Repair and re-enable, Restrict fallback to compatible slices, Add an intermediate degraded mode, Route to manual handling, or Disable fallback. For each, specify evidence prerequisites, residual risk, user impact, and rollback. 7. Define proof before re-enable. Create task- and slice-specific compatibility tests, negative tests, schema/tool checks, canary gates, monitoring thresholds, and an owner decision. Avoid claiming equivalence beyond tested slices. Treat this as model-routing incident analysis and fallback-design evidence for release readiness, not as permission to alter routing. The service owner and release owner must approve any re-enablement or production change. Require regression-testing acceptance evidence that names each fallback contract check, its expected observation, the actual observation when supplied, and any unreconciled reliability gap. Proposed repairs and tests remain not executed unless their results are provided. Required deliverable: # Model Fallback Failure Analysis ## Service Contract | Requirement | Mandatory/degradable | Primary behavior | Required fallback behavior | Evidence | |---|---|---|---|---| ## Fallback Reconstruction | Sequence | Routing or transformation event | Expected | Observed | Evidence | Confidence | |---:|---|---|---|---|---| ## First Divergence and Cause Tree - First evidenced divergence: - Confirmed causes: - Contributing conditions: - Unresolved hypotheses: ## Compatibility and Control Gaps | Gap | Affected slice | Consequence | Existing control | Control failure | Remediation | |---|---|---|---|---|---| ## Blast Radius | Slice/window | Exposure evidence | Failure mode | Downstream action | Revalidation needed | |---|---|---|---|---| ## Fallback Disposition - Decision: Re-enable / Restrict / Degraded mode / Manual route / Disable - Allowed scope: - Preconditions: - Accountable service and release owners: - Rollback trigger: - Residual uncertainty: ## Re-Enablement Gate | Test or control | Slice | Acceptance threshold | Evidence owner | Status | |---|---|---|---|---| Completion requires a trace-supported first divergence, bounded blast radius, an explicit fallback disposition, and compatibility evidence for every slice proposed for re-enablement.Tool Permission Drift Investigation
Compare approved and effective agent tool permissions over time, reconstruct permission drift, contain excess access, and define evidence-based recertification actions.
Investigate suspected permission drift across an AI agent or automation toolchain. Compare what was approved with what identities, tokens, connectors, roles, and tools could actually do at each material point in time. Inputs: - Approved roles, scopes, tools, actions, resources, environments, and expiration conditions: [Approved permission baseline] - Effective grants, token scopes, policy evaluations, connector capabilities, and resource permissions: [Current effective permission evidence] - Agent identities, service accounts, role assumptions, delegated grants, and ownership: [Identity role and delegation records] - Configuration, deployment, policy, connector, group, and credential changes: [Change and deployment history] - Timestamped tool calls, denied requests, access logs, and material side effects: [Tool usage and access logs] - Containment limits, operational dependencies, recertification cadence, and accountable owners: [Containment constraints and owners] Use supplied evidence only. Do not claim live access to identity providers, cloud consoles, MCP servers, tools, or logs. Treat capability descriptions and configured scopes as claims until corroborated by authoritative effective-access evidence. Distinguish the ability to invoke a tool from the ability to affect a particular resource. Mark inherited, conditional, time-bound, and environment-specific access separately. Investigation method: 1. Define the expected boundary. - Translate the approval baseline into testable subject-action-resource-environment conditions. - Identify the accountable tool owner, identity owner, security reviewer, and service owner from supplied evidence. - Record missing approvals, owners, expiry dates, or purpose limitations. 2. Build an effective-access timeline. - Map identities, credentials, groups, roles, connectors, and downstream policies at each relevant change point. - Explain how effective access was derived, including inheritance, wildcard scopes, role chaining, default permissions, cached tokens, and stale sessions. - Separate observed grants from inferred reachability. 3. Create a permission-drift register. Classify each difference as Intended approved change, Unapproved expansion, Stale retained access, Incorrect reduction, Ambiguous baseline, Compensating control dependency, or Evidence gap. Record when it began, likely cause, affected resources, exercised use, and exposure window. 4. Determine actual use and consequence. - Link effective permissions to observed tool calls without assuming unused access caused an incident. - Identify sensitive read, write, delete, execute, impersonate, delegation, and secret-access capabilities. - Distinguish latent exposure from confirmed use and confirmed side effect. 5. Recommend the smallest safe containment. Prioritize expiring or reducing the specific grant, token, role, or connector path responsible for the drift. Account for availability dependencies and emergency access. Do not recommend broad credential revocation when a narrower verified control would contain the exposure. 6. Define restoration and recertification. Specify the intended least-privilege state, evidence required to restore any removed capability, owner approvals, token/session invalidation checks, negative permission tests, and recurring drift detection. If the approved baseline, effective-access evidence, or change history is missing or contradictory, request the blocking evidence once and do not close the investigation. Continue only with bounded analysis, label unresolved permissions Unknown, and state how each missing item limits containment or recertification. Required deliverable: # Tool Permission Drift Investigation ## Approved Boundary | Subject or role | Allowed action | Resource | Environment | Purpose/condition | Expiry | Approval evidence | |---|---|---|---|---|---|---| ## Effective-Access Timeline | Time/change | Identity/credential | Effective capability | Derivation evidence | Confidence | Exposure window | |---|---|---|---|---|---| ## Drift Register | Drift | Classification | Baseline | Effective state | Cause evidence | Exercised? | Risk | Owner | |---|---|---|---|---|---|---|---| ## Containment Decision | Priority | Smallest safe action | Capability affected | Dependency risk | Authorization | Verification | |---|---|---|---|---|---| ## Recertification Plan - Target permission state: - Required negative tests: - Token/session invalidation checks: - Owners and approvals: - Monitoring and review cadence: - Unresolved evidence: The investigation is complete only when every material effective permission is reconciled to an approved purpose or listed as unresolved, excess access has a bounded containment owner, and recertification can prove both required access and denied unauthorized access.Agent Identity and Delegated Authorization Failure Review
Investigate wrong-principal agent actions and delegated authority failures across agent, tool, and downstream-system boundaries.
Review the supplied evidence to determine which principal acted, under which delegated authority, at each agent, tool, and downstream-system boundary, and where that principal-authority binding failed. Focus on agent identity spoofing, confused-deputy behavior, stale delegation, wrong-principal actions, and authorization-context loss. Do not perform a broad agent security audit. Do not infer inspection, execution, approval, containment, or remediation unless the supplied evidence supports it. Context to provide: - Incident or workflow name: [Incident or workflow name] - Time window: [Time window] - Agent and tool boundary description: [Agent and tool boundary description] - Evidence bundle: [Evidence bundle] - Delegation policy sources: [Delegation policy sources] - Accountable owners: [Accountable owners] Evidence discipline: - Separate observed facts from inference, hypothesis, and missing information. - Treat logs, traces, request IDs, token claims, policy excerpts, configuration snapshots, audit events, tickets, and owner statements as evidence only when supplied. - Preserve uncertainty when timestamps conflict, identities are aliased, policies are incomplete, or downstream enforcement behavior is not evidenced. - Do not assume that the initiating user, agent runtime identity, tool credential, service account, and downstream effective principal are the same. - If a claim depends on unavailable evidence, state what evidence would be needed and which owner should provide or verify it. Authority boundaries: - Do not declare the issue fixed, contained, approved, or closed without evidence tied to a security reviewer, system owner, data owner, release owner, or service owner named in the supplied context. - Do not recommend bypassing authorization checks, expanding privileges for convenience, or relying on logging as a substitute for enforcement. - Prefer the smallest containment and repair steps that restore correct principal binding while preserving legitimate workflow behavior. Produce the following deliverable. 1. Review scope and evidence inventory - State the workflow, systems, principals, tools, downstream services, and time window actually covered by the evidence. - List supplied evidence by type and relevance. - List important evidence not provided, including why it matters. - State confidence level for the review and the main reasons for that confidence. 2. Principal and delegation chain Create a step-by-step chain showing how authority was carried or transformed across boundaries. Use a table with these columns: - Step - Timestamp or sequence marker - Boundary crossed - Request, trace, job, or session identifier - Initiating principal - Agent runtime principal - Tool or connector principal - Downstream effective principal - Delegated authority claimed - Delegation source or policy reference - Token, credential, session, or capability involved - Observed action - Evidence citation from supplied material - Confidence - Missing evidence or ambiguity Call out every point where identity was asserted, translated, delegated, cached, proxied, impersonated, or lost. 3. Authorization failure chronology Build a concise chronology of the failure path. For each event, identify: - What happened - Which principal appeared to act - Which principal should have been authoritative, if determinable - What delegated authority was present, stale, missing, overbroad, or misapplied - Whether the event indicates spoofing, confused deputy behavior, stale delegation, wrong-principal action, authorization-context loss, or another clearly named failure mode - What evidence supports the classification - What remains uncertain 4. Action-to-authority matrix Create a matrix mapping each consequential action to the authority required and the authority actually observed. Use these columns: - Action or operation - Resource or data scope - Required principal or role - Required delegation condition, scope, audience, tenant, time limit, or consent - Observed principal - Observed delegated authority - Enforcement point expected - Enforcement point evidenced - Decision observed, such as allowed, denied, skipped, inherited, cached, or unknown - Mismatch type - Potential impact - Evidence citation Highlight actions where the system accepted an authority context that was missing, stale, meant for another audience, meant for another tenant, too broad, inherited from the wrong actor, or not revalidated downstream. 5. Affected scope Separate confirmed, likely, possible, and not evidenced impact. Cover: - Users, tenants, accounts, workspaces, repositories, environments, datasets, secrets, transactions, or external systems affected - Time interval of exposure - Data or operations reachable under the mistaken authority - Whether unauthorized read, write, execute, approve, delete, publish, spend, or disclose actions are evidenced - Whether lateral movement, replay, token reuse, cached delegation, or downstream propagation is evidenced or merely plausible - Constraints that limit impact 6. Containment and identity-control repair gates Define repair gates that accountable owners can use to decide whether the workflow may resume or continue operating. Include immediate containment and durable control gates. For each gate, provide: - Gate name - Failure it addresses - Required control or change - Owner accountable for verification - Evidence required to pass - Pass criteria - Fail or block condition - Residual risk if accepted Include gates where relevant for: - Principal provenance at agent start and tool-call time - Delegation freshness and revocation handling - Audience, tenant, resource, and scope binding - Prevention of confused-deputy delegation reuse - Downstream reauthorization rather than blind trust in upstream context - Service account and connector credential separation - Audit correlation across user, agent, tool, and downstream identifiers - Token, session, capability, or cached grant invalidation - Least-privilege restoration without breaking legitimate workflow paths 7. Owner decision record Prepare a decision-ready summary for the named accountable owners. Include: - Most likely root cause, stated with confidence and evidence basis - Highest-risk unresolved uncertainty - Minimum containment required before further use - Minimum identity-control repair required before normal operation - Verification tasks by owner role - Open questions that block a defensible decision - Items that can be safely deferred, with rationale Completion checks: - Confirm that the principal and delegation chain covers each supplied boundary or states why a boundary could not be assessed. - Confirm that every consequential action is mapped to required and observed authority, or marked unknown with missing evidence. - Confirm that affected scope distinguishes observed impact from plausible but unproven exposure. - Confirm that each repair gate has an accountable owner, evidence requirement, and pass/fail criterion. - Do not present remediation as complete unless evidence supplied in the prompt demonstrates completion.Agent Memory Integrity and Poisoning Investigation
Investigate persistent agent memory for poisoning, misattribution, over-retention, or unauthorized alteration and produce a defensible containment and recovery decision.
Investigate whether persistent agent memory was altered, misattributed, over-retained, or poisoned, and produce an evidence-backed containment and recovery decision. Context to provide: - Agent or system name: [Agent or system name] - Investigation window: [Investigation window] - Memory stores and schemas: [Memory stores and schemas] - Available evidence: [Available evidence] - Known suspicious symptoms: [Known suspicious symptoms] - Decision owner: [Decision owner] - Recovery authority boundaries: [Recovery authority boundaries] Evidence rules: - Use only the evidence provided in [Available evidence]. Do not claim that logs, traces, memory stores, tickets, approvals, tests, commands, or source files were inspected unless they are included or quoted. - Separate observations from inference. Label each inferred conclusion as low, medium, or high confidence. - Preserve uncertainty. If a required fact is missing, state exactly what evidence is needed and why it matters. - Before attributing compromise or selecting containment, separate blocking gaps from non-blocking gaps. Request all blocking evidence in one consolidated clarification and stop the affected conclusion until it is supplied. Continue past non-blocking gaps only when they are recorded as Unknown with their effect on confidence, scope, containment, and recovery. - Do not broaden this into a general agent security audit or a generic trace taxonomy. Stay focused on persistent memory integrity, provenance, poisoning, retention, attribution, and recovery. - Treat containment and recovery as accountable operational decisions. If an action exceeds [Recovery authority boundaries], identify the responsible owner or function that must decide. Investigation method: 1. Establish scope - Define which memory stores, records, embeddings, summaries, preference stores, tool-state caches, user profile memories, system memories, and derived memories are in scope based on [Memory stores and schemas]. - Identify what is out of scope and any assumptions required because evidence is missing. 2. Build a memory provenance ledger Create a ledger table with these columns: - Memory item or record ID - Current stored claim or value - Memory type or store - First observed timestamp, if known - Last modified timestamp, if known - Claimed source or author - Evidence supporting source attribution - Write or update path - Retention basis or deletion expectation - Integrity concern: none, altered, misattributed, over-retained, suspicious insertion, suspicious deletion, unverifiable - Confidence level - Evidence references - Required follow-up evidence 3. Reconstruct write paths For each memory store or memory type, reconstruct: - Authorized writers and expected write triggers - Observed or reported writes during [Investigation window] - Transformation steps from raw input to persisted memory - Summarization, embedding, deduplication, merge, overwrite, deletion, or compaction behavior - Identity and attribution mechanisms used at write time - Retention and deletion controls - Plausible bypasses, race conditions, stale-cache paths, ingestion errors, or cross-user contamination paths - Evidence gaps that prevent confirmation 4. Develop a poisoning hypothesis matrix Create a matrix with these columns: - Hypothesis - Poisoning or integrity failure mechanism - Memory items affected - Supporting observations - Contradicting observations - Missing evidence needed - Likelihood: low, medium, high, or indeterminate - Potential blast radius - Immediate containment implication Include at least these hypothesis categories when relevant to the evidence: - Malicious prompt or tool output caused durable memory insertion - Benign user input was misattributed to another user, tenant, workflow, or authority - Summarization or consolidation distorted source meaning - Outdated memory was over-retained after deletion, revocation, policy change, or user correction - Memory merge or deduplication combined incompatible identities, accounts, or contexts - Retrieval surfaced poisoned or stale memory into later decisions - Administrative, migration, sync, or backfill process altered memory incorrectly - Evidence does not support a memory poisoning finding, but confidence is limited 5. Map affected decisions Create an affected-decision map with these columns: - Decision, action, response, workflow, or automation potentially influenced - Memory items used or likely retrieved - Evidence of use versus plausible exposure - Business, safety, security, privacy, financial, or user impact - Reversibility - Owner accountable for review - Required validation before relying on the decision again Do not assume a decision was affected solely because a memory item exists. Distinguish confirmed use, likely use, possible use, and no evidence of use. 6. Decide containment posture Recommend one of these containment states for each affected memory store or memory class: - No containment needed based on current evidence - Monitor only - Read-only quarantine - Retrieval suppression - Write suspension - Record-level deletion or correction pending owner approval - Full memory store isolation - Agent workflow suspension For each recommendation, provide: - Evidence basis - Risk reduced - Operational cost or user impact - Decision owner required - Time sensitivity - Reversal condition 7. Define recovery gate Produce a quarantine and recovery gate that [Decision owner] can use. Include: - Minimum evidence required before restoring memory retrieval or writes - Records that must be corrected, deleted, re-attributed, or re-generated - Regression checks or replay checks needed, if evidence supports them - Owner approvals required under [Recovery authority boundaries] - Monitoring signals after restoration - Conditions that require escalation rather than recovery 8. Final decision artifact End with a concise decision record: - Investigation conclusion: confirmed compromise, likely compromise, possible compromise, no compromise found, or indeterminate - Confidence and primary uncertainty - Memory stores or records requiring action - Affected decisions requiring review - Immediate containment decision - Recovery gate status: not ready, conditionally ready, or ready - Named accountable owner for the next decision - Evidence not available that would materially change the conclusion Completion checks before finalizing: - Every material conclusion cites evidence from [Available evidence] or is explicitly labeled as inference. - The memory provenance ledger, write-path reconstruction, poisoning hypothesis matrix, affected-decision map, and quarantine/recovery gate are all included. - Missing evidence is stated plainly rather than filled with assumptions. - No source, system, command, test, approval, or inspection is claimed without corresponding evidence. - The recommended action stays within [Recovery authority boundaries] or identifies the accountable owner who must authorize it.Multi-Agent Coordination Failure Reconstruction
Reconstructs a multi-agent failure to find the first coordination divergence and produce a corrected handoff contract with testable verification.
Reconstruct the coordination failure across multiple agents and identify the first coordination divergence that can be corrected and tested. Inputs to provide: - Agent roles and authority: [Agent roles and authority] - Coordination evidence: [Coordination evidence] - Expected handoff contract: [Expected handoff contract] - Shared-state context: [Shared-state context] - Observed failure and impact: [Observed failure and impact] - Constraints and authorized changes: [Constraints and authorized changes] Evidence discipline: - Separate observed evidence from inference at every major step. - Do not claim that a log, trace, source file, system, test, approval, or action was inspected unless it is present in the provided material. - If timestamps, message ordering, state versions, ownership rules, or agent responsibilities are missing, name the gap and explain how it affects confidence. - Preserve uncertainty where the evidence supports more than one explanation. - Do not redesign the entire multi-agent workflow. Focus on the smallest correction to the failed coordination point. - Do not produce a broad trace taxonomy. Use classification only where it helps isolate the first divergence. Reconstruction method: 1. Build an evidence inventory listing each provided artifact, what it can prove, and what it cannot prove. 2. Create a coordination chronology from the earliest relevant trigger through the failure manifestation. 3. Map responsibility, authority, and shared-state access for each involved agent. 4. Compare expected coordination behavior against observed behavior. 5. Identify the first point where delegation, handoff, ownership, message ordering, or shared state diverged from the expected contract. 6. Decide the most likely failure mechanism and explain why competing mechanisms are less supported. 7. Define discriminating tests that could confirm or falsify the proposed mechanism. 8. Draft a corrected handoff contract that is narrow, testable, and preserves intended existing behavior. Deliverable format: 1. Evidence inventory Provide a table with columns: - Artifact - Evidence observed - Reliability or limitation - Coordination question it helps answer - Missing information, if any 2. Coordination chronology Provide an ordered table with columns: - Step or timestamp - Agent - Message, action, or state operation - Intended recipient or owner - Shared state read or written - Expected behavior - Observed behavior - Evidence reference - Observation vs inference - Confidence: high, medium, or low 3. Responsibility and state map For each relevant agent, list: - Delegated responsibility - Decision authority - Required inputs - State it may read - State it may write - Handoff obligations - Acknowledgment or completion signal - Actual behavior seen in evidence - Ownership ambiguity or conflict, if any Then list shared-state objects or records with: - State object - Expected owner - Writers - Readers - Versioning or ordering assumption - Observed mutation or stale-read risk - Evidence supporting the risk 4. First coordination divergence State the earliest supported divergence in one sentence. Then provide: - Divergence type: delegation gap, handoff ambiguity, ownership conflict, message-ordering violation, stale shared state, unauthorized state mutation, missing acknowledgment, retry/idempotency failure, or other specified mechanism - Exact expected contract at that point - Exact observed deviation - Why this is earlier than downstream symptoms - Evidence supporting the decision - Confidence level - What evidence would change the conclusion 5. Failure mechanism decision Compare the leading mechanism against plausible alternatives in a table: - Candidate mechanism - Supporting evidence - Contradicting or missing evidence - Predicted observable signal - Decision: selected, possible, unlikely, or not assessable 6. Discriminating tests Propose 3 to 6 tests or checks that the accountable engineering, automation, or operations owner can run. For each, include: - Hypothesis tested - Setup or evidence required - Expected signal if the mechanism is correct - Expected signal if the mechanism is wrong - Owner best placed to verify - Risk of false positive or false negative Prefer tests that distinguish between message-ordering, ownership, handoff, and shared-state causes rather than merely reproducing the incident. 7. Corrected handoff contract Draft the smallest safe correction. Include: - Trigger condition - Sending agent - Receiving agent - Required payload or state reference - Preconditions - State version or ordering requirement - Acknowledgment rule - Completion signal - Timeout or retry behavior - Idempotency requirement - Conflict handling rule - Audit fields or log events needed for future reconstruction - Backward-compatibility or behavior-preservation note 8. Implementation boundary State what should not be changed yet because the evidence does not justify it. Call out any broad workflow redesign, new orchestration pattern, or policy change that would be premature. 9. Completion checks List observable criteria for considering this reconstruction complete, including: - The first divergence is identified or the blocking evidence gap is explicit. - The selected mechanism has at least one discriminating test. - The corrected handoff contract is narrow enough for the responsible owner to implement or reject. - The release owner, automation owner, or service owner can verify the proposed correction against logs, replay, tests, or production telemetry before adoption.Was this useful?