Reusable AI capability

Recover Failed Kubernetes Rollouts Safely

Apply a repeatable evidence and decision method to isolate a failed Kubernetes rollout, choose the smallest safe recovery path, and verify restored service before closing the incident.

This Skill packages a reusable way to use the linked Prompt or Workflow; Amo.ng does not run it for you.

Skill ID
AMO-S-000017
Powered by
Prompt
Published

Copy skill copies the Skill details. Use with AI adds a short instruction for your preferred AI tool; neither action runs the Skill.

Purpose

Give platform and service teams a reusable rollout-recovery capability that separates diagnosis, recovery authorization, execution evidence, and service verification across different Kubernetes workloads and release mechanisms.

Required inputs

Have these details available before following the usage instructions.

  • Affected environment, namespace, workloads, services, and incident window
  • Release, manifest, image, configuration, secret, ingress, service, policy, or infrastructure changes
  • Available rollout status, events, logs, probes, endpoints, routing, metrics, alerts, and application evidence
  • User impact, SLO constraints, recovery window, rollback and forward-fix options
  • Platform, service, security, incident, and release owners plus authorization boundaries

How to use this Skill

When to use:
- A Kubernetes rollout is failed, stalled, or degraded and the team needs an evidence-based recovery decision.
- The same recovery discipline must be applied consistently across Deployments, StatefulSets, Services, ingress, probes, configuration, or dependencies.

When not to use:
- Designing a new cluster architecture or generic Kubernetes hardening.
- Executing production changes without the relevant platform and release authority.
- Declaring recovery when workload and application evidence are unavailable.

Reusable method:
1. Freeze the incident boundary, latest known-good release, impact, recent changes, and inspectable evidence.
2. Build a timeline and classify observations, inferences, assumptions, missing evidence, and actions already taken.
3. Localize failure across artifact, scheduling, configuration, secret, lifecycle, probe, service discovery, ingress, network, dependency, resource, and autoscaling layers.
4. Rank competing causes and define the next discriminating observation for unresolved hypotheses.
5. Compare rollback, configuration reversal, image correction, probe correction, scaling, traffic diversion, dependency recovery, and continue-investigation paths.
6. Select the smallest safe path with prerequisites, expected effect, blast radius, stop condition, rollback or compensation option, and authorized owner.
7. Verify workload availability, pod stability, probes, endpoints, routing, application transactions, metrics, logs, and customer-impact indicators over a defined monitoring window.

Expected output:
A recovery decision record containing evidence timeline, affected resources, layered hypothesis register, chosen and rejected recovery paths, authorization gates, execution checklist, service-verification matrix, monitoring window, unresolved evidence, and owner handoff.

Boundaries:
Do not invent cluster state, command output, changes, approvals, or recovery. Platform and service owners authorize configuration and workload actions; the security reviewer dispositions material security impact; the incident and release owners control production recovery and closure. AMO-P-000235 is the grounding source, not a prerequisite to invoke the Skill.

Powered by an Amo.ng Prompt

Kubernetes Deployment Recovery Runbook

Open the linked prompt to use the instructions that power this Skill.

Open prompt

Completion criteria

Complete when:
- Every material cause finding links to supplied evidence and unresolved causes have a discriminating check.
- The chosen path has prerequisites, owner, expected effect, risk, stop condition, and fallback.
- Production execution and observed results are distinguished from proposed commands.
- Workload and application verification have observable results or remain explicitly pending.
- The status is Ready for authorized recovery, Conditionally ready, Blocked, or Recovered with evidence.

Browse Prompts
Codex & Coding Expert Codex

FastAPI Production Readiness Gate

Assess a FastAPI service for secure deployment, validated API boundaries, resilient workers, dependency safety, observability, operational ownership, rollback readiness, and evidence-based release approval.

Updated Aug 6, 2026

View prompt Verified ✓ 211 views · 18 copies
Codex & Coding Expert Codex

PostgreSQL Slow Query Evidence Pack

Investigate a PostgreSQL query using plans, runtime statistics, locks, indexes, data shape, cache conditions, and controlled experiments before recommending a safe optimization.

Updated Aug 6, 2026

View prompt Verified ✓ 248 views · 22 copies

Was this useful?