Amo.ng curated workflow

Recover and Requalify an Interrupted AI Agent Runtime

Reconcile interrupted durable state and side effects, conditionally recertify tool access and fallback behavior, then calibrate escalation controls before resuming an AI agent runtime.

Workflow ID
AMO-W-000022
Steps
4
Published
Download Markdown

Copy workflow includes every step and the full linked Prompt content. Use with AI copies a shorter guide with Prompt links; neither action runs the Workflow.

Outcome

A controlled recovery package containing a state and side-effect ledger, resume or compensation decision, applicable permission and fallback dispositions, recalibrated escalation rules, verification evidence, and accountable restart conditions.

Before you begin

Have all or some of the following available before you start. The more relevant context you can provide, the stronger the workflow output will be.

  • Interrupted agent run, intended outcome, workflow version, checkpoints, durable state, messages, side effects, and idempotency evidence
  • Incident timeline, traces, logs, tool calls, approvals, external-system records, and current containment state
  • Approved and effective tool permissions, identity and delegation evidence when access drift may be material
  • Primary and fallback model contracts, routing evidence, outputs, errors, and compatibility criteria when fallback is implicated
  • Escalation policy, incident severity, uncertainty, queue capacity, owner availability, restart authority, and operating constraints

Ordered sequence

Workflow steps

Complete the steps in order. For each step, provide the listed context, carry its result into the next step, and pause wherever a review note is shown.

  1. Step 1 Reconcile durable state and choose a recovery path

    Reconstruct the latest trustworthy checkpoint, committed and uncertain side effects, pending work, approvals, and idempotency boundary. Choose Resume, Replay, Compensate, Abort, or Continue investigation.

    Prompt: Long-Running Agent State Recovery Decision

    Input for this step

    Provide run identifiers, workflow and configuration versions, state stores, checkpoints, traces, tool and message logs, external side effects, approval records, incident timing, and recovery constraints.

    Carry forward

    Carry the state ledger, side-effect reconciliation, affected scope, chosen path, prerequisites, compensation needs, evidence gaps, and verification gates into conditional access review.

    Review note

    The incident and service owners approve the recovery path; owners of affected external systems authorize compensation, replay, or rollback actions.

    Open prompt
  2. Step 2 Recertify effective tool access when applicable

    Run this step when tool permission drift could have caused, amplified, or outlived the interruption. Otherwise record Not applicable with evidence. Reconcile approved and effective access across identities, tools, environments, and time.

    Prompt: Tool Permission Drift Investigation

    Input for this step

    Use the state findings with approved access, runtime policy, role and group data, credential or token metadata, tool schemas, audit logs, changes, exceptions, and current containment evidence.

    Carry forward

    Carry the permission-drift register, excess or missing access, affected actions, containment and recertification needs, owners, and regression gates into fallback review.

    Review note

    Security, identity, tool, and service owners authorize access changes and recertification; absence of logs is not evidence that drift did not occur.

    Open prompt
  3. Step 3 Requalify model fallback when applicable

    Run this step when fallback selection or contract incompatibility contributed to the incident or could affect recovery. Otherwise record Not applicable and why. Compare route contracts, capabilities, outputs, safety behavior, latency, failure semantics, and recovery conditions.

    Prompt: Model Fallback Failure Analysis

    Input for this step

    Provide routing policy, primary and fallback model versions and contracts, request and response traces, tool or schema requirements, errors, evaluation evidence, constraints, and prior findings.

    Carry forward

    Carry the fallback failure mechanism, compatibility matrix, restricted or disabled routes, corrective tests, re-enable gate, and unresolved evidence into escalation calibration.

    Review note

    Model, product, security, and release owners approve fallback restrictions and any re-enablement decision.

    Open prompt
  4. Step 4 Calibrate escalation and restart conditions

    Tune escalation triggers using incident severity, uncertainty, false-positive and false-negative evidence, queue capacity, response delay, and owner authority. Reconcile the thresholds with the recovery, access, and fallback findings.

    Prompt: Agent Escalation Threshold Calibration

    Input for this step

    Provide current escalation rules, incident and near-miss data, severity and uncertainty definitions, queues and response times, owner schedules, containment options, and all prior outputs.

    Carry forward

    Produce the final recovery and requalification package with restart or containment decision, thresholds, routing, owners, verification, monitoring window, rollback, and re-review triggers.

    Review note

    The service and incident owners approve operating thresholds; the release owner authorizes restart after security and model owners close applicable gates.

    Open prompt

Completion criteria

The workflow is complete when:

  • Durable state, committed side effects, pending work, approvals, and uncertainty are reconciled for the interrupted run.
  • The path is explicitly Resume, Replay, Compensate, Abort, or Continue investigation with prerequisites and idempotency controls.
  • Permission and fallback reviews are completed when applicable or marked Not applicable with evidence.
  • Escalation triggers account for severity, uncertainty, response delay, queue capacity, and decision authority.
  • Restart or continued containment has observable verification, stop, rollback, and owner conditions; no recovery action is claimed without evidence.
Browse Workflows
AMO-W-000012 5 steps

Investigate an AI Agent Security Incident

Reconstruct an AI agent incident, trace delegated authority and sensitive context, conditionally investigate memory or RAG authorization, and prepare evidence-based containment and recovery gates.

Was this useful?