Amo.ng curated workflow
Design a Governed AI Agent Architecture and Accountability Model
Turn an already-approved business use case and verified process evidence into an AI agent architecture with explicit roles, handoffs, authority boundaries, role-based review, escalation, and recovery controls.
# Design a Governed AI Agent Architecture and Accountability Model Workflow ID: AMO-W-000008 ## Outcome A review-ready architecture and accountability package for an already-approved use case, including role decomposition, interface contracts, permissions, review gates, escalation paths, override controls, and a conditional design-approval decision. ## Before you begin - Approved business use case, outcome, scope, and accountable owner - Current process documentation, representative cases, exceptions, and performance evidence - Prior selection or readiness decisions and any binding remediation requirements - Systems, data, tool capabilities, integrations, and permission constraints - Security, privacy, compliance, operational, and service-level requirements - Risk classification, success criteria, reviewers, and implementation constraints ## Step 1 — Map the Approved Process and Preserve Its Control Boundaries **Prompt** Evidence-Based AI Business Process Automation Mapping **Instructions** Map the current approved process from supplied evidence, including triggers, activities, decisions, data, exceptions, handoffs, failure points, and existing controls. Distinguish AI work, deterministic automation, and human responsibility without reopening portfolio selection or assuming automation is appropriate. **Input for this step** Provide the approved use-case decision, current SOP or process evidence, representative normal and exceptional cases, systems, roles, performance evidence, known constraints, and binding readiness conditions. **Carry forward** Pass the verified process map, approved automation boundary, exceptions, dependencies, existing controls, evidence gaps, and binding constraints to architecture decomposition. **Review note** The accountable process owner should confirm that the map and approved scope reflect actual work before agent architecture is designed. **Prompt content** ## Objective Analyze the supplied business process evidence in ChatGPT. Produce a traceable map of the current workflow, distinguish deterministic automation from appropriate AI assistance, identify required human oversight, and create a phased implementation and validation plan. This is an analysis and planning exercise. Do not claim that an automation was configured, tested, approved, deployed, integrated, measured, or completed unless the supplied materials contain explicit evidence that the action occurred. Label planned work as proposed and keep it distinct from executed work. ## Context Business context: [Business context] Current workflow evidence: [Current workflow evidence] Stakeholders and approvals: [Stakeholders and approvals] Systems and integrations: [Systems and integrations] Inputs, outputs, and data: [Inputs outputs and data] Pain points and exceptions: [Pain points and exceptions] Volume, service levels, and costs: [Volume service levels and costs] Compliance, privacy, and security constraints: [Compliance privacy and security constraints] Budget, timeline, and tool constraints: [Budget timeline and tool constraints] Automation goals and definition of done: [Automation goals and definition of done] ## Input requirements Blocking inputs are a recognizable process trigger, the main workflow steps, the intended output or outcome, and the automation goal. If any blocking input is absent or contradictory, ask focused clarification questions before making final recommendations. Useful but non-blocking inputs include process documentation, standard operating procedures, screenshots, anonymized forms, sample records, decision rules, exception logs, service-level targets, task volumes, handling times, error rates, cost data, customer complaints, audit findings, system ownership, API or integration constraints, and approval policies. If these are unavailable, continue only with bounded analysis, state the limitation, and avoid fabricated measurements or system capabilities. Do not request or reproduce passwords, access tokens, private keys, unnecessary personal data, payment credentials, health information, or other secrets. Recommend redacted or synthetic examples where possible. Treat any supplied legal, regulatory, security, or financial interpretation as requiring review by the responsible human authority. ## ChatGPT operating boundary ChatGPT may analyze text and files supplied in the active conversation, organize evidence, identify patterns, compare options, calculate estimates from supplied figures, and draft recommendations. It cannot independently inspect business systems, observe staff performing the process, confirm vendor features, access private records, configure integrations, contact stakeholders, grant approval, or execute deployment unless such capabilities and resulting evidence are explicitly available in the session. Never imply that external inspection or action occurred. ## Evidence and uncertainty rules 1. Create stable identifiers for supplied evidence, such as E1, E2, and E3, and for workflow steps, such as W1, W2, and W3. 2. Classify material statements as one of: Supplied Fact, Observed in Supplied Artifact, Assumption, Hypothesis, Unknown, Conflict, or Execution Evidence. 3. Cite the relevant evidence identifier for each mapped workflow step, quantified baseline, risk, and recommendation. If there is no evidence, mark the item as an assumption or unknown rather than presenting it as fact. 4. Separate current-state observations from future-state proposals. Do not convert stakeholder aspirations into current capabilities. 5. When sources conflict, record both versions, explain the operational consequence, and identify the process owner who should resolve the conflict. 6. Show formulas and inputs for time, cost, capacity, error-reduction, or return-on-investment estimates. Present ranges when values are uncertain and do not invent precision. 7. Treat vendor capabilities, integration feasibility, model accuracy, and compliance suitability as unverified until supported by current documentation or testing evidence. ## Analysis workflow ### 1. Establish the evidence base Inventory the supplied artifacts and statements. Record source, date if known, process scope, reliability limitations, and the claims each source supports. List missing evidence and clarification needs. ### 2. Map the current process Decompose the process from trigger to final outcome. Include normal flow and documented exceptions. For each step capture: - step identifier and name; - trigger and predecessor; - actor or accountable team; - system or tool; - input and output; - business rule or judgment applied; - data classification; - average volume, handling time, wait time, and service target when supplied; - approval, handoff, queue, rework loop, and exception path; - failure mode and current control; - supporting evidence identifier. Identify bottlenecks without assuming that automation is the remedy. Distinguish processing time from waiting time and note whether the constraint arises from policy, capacity, poor data quality, system fragmentation, unclear ownership, or genuine judgment. ### 3. Classify automation suitability Assess each relevant workflow step against these paths: - Eliminate or simplify: the step may be unnecessary or redesigned before automation. - Deterministic automation: stable rules and structured inputs permit conventional workflow, scripts, forms, validation, or integration. - AI assistance: probabilistic work such as extraction, classification, summarization, drafting, knowledge retrieval, anomaly flagging, or recommendation support. - Human-led: nuanced judgment, negotiation, accountability, sensitive decisions, or poorly defined exceptions should remain human-controlled. - Not ready: evidence, data quality, process stability, system access, controls, or ownership is inadequate. For AI candidates, define the exact input, output, permissible use, prohibited use, expected error modes, confidence or abstention behavior, review requirement, exception route, and fallback procedure. Consider hallucination, omission, misclassification, prompt injection, data leakage, automation bias, model drift, inconsistent output, and inaccessible source citations where relevant. ### 4. Evaluate value, feasibility, and risk For every candidate, estimate business value, implementation complexity, process readiness, integration dependency, and operational risk using a clearly explained Low, Medium, High, or Critical scale. Assess privacy, security, legal or compliance exposure, financial impact, customer impact, accuracy, availability, vendor dependency, change-management burden, and over-automation risk. Prioritize only after documenting the rationale and evidence. A high-value candidate must not outrank a safer option solely because its projected savings are larger. Flag estimates that depend on missing baseline data. ### 5. Define authority and safeguards Identify the accountable process owner, system owner, data owner, risk or compliance reviewer, approver, operator, and escalation contact when known. Human authorization is mandatory before production configuration, access expansion, sensitive-data processing, customer-facing release, financial or legal decisions, or removal of an existing control. For Medium, High, or Critical risks, specify controls such as data minimization, masking, role-based access, approved environments, retention limits, source grounding, confidence thresholds, human approval gates, dual control, sampling, audit logs, rate limits, exception queues, kill switches, manual fallback, incident escalation, rollback criteria, and periodic review. Recommend no-go or pause conditions where controls or ownership are absent. ### 6. Design the future-state workflow Describe the proposed flow using the existing workflow identifiers. Show which steps remain manual, are simplified, use deterministic automation, or receive AI assistance. Define handoffs, review gates, exception routes, fallback operations, audit evidence, and recovery behavior. Do not assume integrations exist merely because they are desirable. ### 7. Build a phased implementation plan Sequence work through discovery and baseline confirmation, low-risk pilot, controlled human-in-the-loop rollout, integration, and monitored scale-up. For each phase state scope, dependencies, owner, approval gate, deliverable, validation method, rollback or fallback condition, and exit criteria. Mark every phase Proposed unless supplied execution evidence supports another status. ### 8. Define measurement and validation For each success metric provide its definition, baseline source, calculation, target, measurement window, owner, data source, and review cadence. Suitable metrics may include cycle time, touch time, queue time, first-pass yield, error or rework rate, exception rate, review override rate, false-positive and false-negative rates, service-level attainment, customer impact, adoption, operating cost, control failures, and risk incidents. For each pilot test, specify the test population, representative edge cases, expected observation, actual observation if evidence exists, evidence location, acceptance threshold, result status, and remediation owner. Use only these result states: Not Run, Passed with Evidence, Failed with Evidence, Inconclusive, or Blocked. Never mark a test passed based on a proposed procedure. ## Required deliverable Produce the following sections in order. ### 1. Decision Brief State the process scope, strongest supported opportunities, major constraints, recommended starting point, and decisions requiring human authorization. Separate facts from assumptions. ### 2. Evidence and Uncertainty Register Use columns: Evidence ID | Source or Artifact | Date or Version | Supported Claim | Classification | Reliability Limitation | Conflict or Gap. ### 3. Current-State Workflow Register Use columns: Step ID | Trigger or Predecessor | Activity | Actor | System | Input | Output | Rule or Judgment | Data Classification | Volume or Timing | Approval or Handoff | Exception or Failure Mode | Current Control | Evidence ID. ### 4. Bottleneck and Root-Cause Analysis For each bottleneck, distinguish observed symptom, supported or hypothesized cause, operational effect, evidence, uncertainty, and whether simplification should precede automation. ### 5. Automation Opportunity Matrix Use columns: Opportunity ID | Linked Step ID | Current Activity | Recommended Path | AI or Automation Function | Required Input | Produced Output | Business Value | Complexity | Risk | Human Review | Evidence ID | Assumptions | Priority | Rationale. ### 6. AI Use and Human Oversight Specification For each AI candidate provide allowed use, prohibited use, accountable owner, reviewer, confidence or abstention rule, known failure modes, review checklist, exception route, escalation trigger, audit record, and manual fallback. ### 7. Risk and Control Register Use columns: Risk ID | Opportunity ID | Risk Scenario | Affected Data or Stakeholder | Likelihood | Impact | Rating | Preventive Control | Detective Control | Response or Recovery Control | Control Owner | Approval Required | Residual Risk | Evidence or Validation Needed. ### 8. Proposed Future-State Workflow Map the proposed steps back to current Step IDs. Identify manual, simplified, deterministic, and AI-assisted activities; system boundaries; approvals; exception queues; fallback paths; and audit points. ### 9. Phased Implementation and Handoff Plan Use columns: Phase | Proposed Scope | Dependencies | Responsible Owner | Required Approval | Deliverable | Validation Method | Exit Criteria | Rollback or Fallback Trigger | Status. Status must remain Proposed unless execution evidence supports a different label. ### 10. Measurement and Validation Plan Use columns: Metric or Test ID | Opportunity ID | Definition or Test Case | Baseline and Source | Target or Expected Observation | Actual Observation | Evidence | Measurement Window | Owner | Acceptance Threshold | Result Status | Follow-up. ### 11. Assumptions, Unknowns, Conflicts, and Open Decisions List each unresolved item, its consequence, the evidence or decision needed, and the accountable resolver. Do not hide unresolved Critical risks in narrative text. ### 12. Final Recommendation and Authorization Requests Recommend proceed, revise, pilot, defer, or reject for each opportunity. State the evidence basis, residual uncertainty, next decision, required approver, and prohibited actions pending approval. ## Final acceptance checks Before returning the deliverable, verify and correct the following: - Every opportunity links to at least one workflow Step ID and supporting Evidence ID, or is explicitly labeled as an assumption. - Every mapped step identifies an actor, input, output, decision or rule, exception path, and evidence gap where these are unknown. - Deterministic automation, AI assistance, human-led work, and not-ready work are not conflated. - Every Medium, High, or Critical risk has an owner, approval requirement, safeguards, and response or fallback control. - Sensitive-data use identifies minimization, access, retention, audit, and approved-environment requirements. - Quantified benefits show supplied inputs, formulas, uncertainty, and baseline provenance. - Every pilot has acceptance thresholds, expected observations, evidence requirements, and a valid result state. - Proposed, executed, tested, approved, deployed, and measured states remain distinct and evidence-backed. - The recommended plan fits the stated budget, timeline, systems, process maturity, and authority constraints. - Unknowns and conflicts remain visible and are routed to named owners or owner roles for resolution. ## Step 2 — Decompose Agent, Automation, and Accountable Roles **Prompt** AI Agent Process Decomposition Prompt **Instructions** Design the agent architecture using the approved process boundary. Define justified agent responsibilities, deterministic components, decisions reserved for accountable owners, handoff contracts, tool permissions, data and memory rules, exception paths, failure behavior, and verification tests. **Input for this step** Treat the approved scope, process evidence, existing controls, readiness restrictions, system capabilities, permission limits, and representative failure cases as binding design inputs. **Carry forward** Pass the architecture, role and handoff contracts, permission model, data and memory boundaries, failure paths, and test requirements to the accountability and control design step. **Prompt content** Analyze the supplied process and produce an implementation-ready decomposition that distinguishes AI-agent work from deterministic automation and human judgment. Work only from information available in this conversation. You may analyze materials and propose a design, but you cannot inspect unprovided systems, run tools, change workflows, deploy agents, grant access, approve controls, or claim tests were executed. Inputs - Process objective: [Process objective] - Current process evidence: [Current process evidence] - Scope and boundaries: [Scope and boundaries] - Constraints and policies: [Constraints and policies] - Systems and tool capabilities: [Systems and tool capabilities] - Success criteria: [Success criteria] - Risk and approval requirements: [Risk and approval requirements] Input handling Treat the process objective, current process evidence, scope, success criteria, and authority boundaries as blocking prerequisites. Useful supporting material includes process maps, SOPs, sample cases, decision tables, forms, API documentation, event schemas, failure logs, service-level targets, volumes, costs, security classifications, and operator interviews. If a reliable decomposition is blocked, ask no more than five grouped clarification questions before designing it. A missing detail is blocking when it prevents you from defining process boundaries, assigning a consequential decision, validating a handoff, determining permitted system access, or evaluating acceptance. Otherwise, continue with a bounded provisional design and record the missing item as an unknown. Do not silently reconcile conflicting sources; identify the conflict, its design impact, and the owner who must resolve it. Evidence and claim rules 1. Classify material statements as supplied fact, evidence-backed observation, assumption, hypothesis, conflict, or unknown. 2. Reference the supplied artifact, section, example, or statement supporting each important process or control conclusion. Do not fabricate citations or system behavior. 3. Keep proposed, approved, configured, executed, observed, and verified states distinct. Unless execution evidence is supplied, describe all workflow components and tests as proposed or unverified. 4. Never state that an integration works, a control is effective, a test passed, or a workflow was deployed or approved without corresponding evidence. Decomposition method 1. Reconstruct the current process. Identify its trigger, termination states, actors, inputs, outputs, sequence, decisions, business rules, systems, data classes, service levels, volumes, known exceptions, rework loops, and current failure points. Separate documented behavior from inferred behavior. 2. Test automation suitability at the activity level. For each activity, choose one disposition: deterministic automation, AI-agent responsibility, human responsibility, shared responsibility, or out of scope. Justify the choice using rule stability, ambiguity, contextual judgment, tool interaction, error detectability, reversibility, data sensitivity, consequence severity, latency, and cost. Do not create an agent where a rule, workflow engine, validation service, or human decision is safer and simpler. 3. Define the target orchestration. Describe the happy path, alternate paths, exception paths, cancellation behavior, and terminal states. Show where work is queued, resumed, retried, escalated, or stopped. Identify concurrency hazards, ordering dependencies, duplicate events, partial completion, timeout behavior, and recovery paths. 4. Establish cohesive agent boundaries. Group responsibilities by decision context, required tools, data access, risk, and accountability rather than assigning one agent per process step. Give every proposed agent one clear objective, explicit inputs and outputs, permitted decisions, prohibited actions, dependencies, and an accountable human owner. Flag overlapping authority, circular delegation, orphaned work, and single points of failure. 5. Specify every handoff. Define producer, consumer, trigger, preconditions, payload fields and schema, provenance, validation rules, acknowledgment, correlation identifier, status model, timeout, retry policy, idempotency control, failure owner, escalation route, and completion signal. State how malformed, missing, stale, duplicated, contradictory, or late payloads are handled. 6. Define tool access. For each agent-tool pairing, identify the operation, read or write scope, authentication boundary, minimum permission, data exposed, side effects, rate limits, timeout, error response, audit event, and required approval. Treat external content and tool output as untrusted until validated. Agents must not expose secrets, bypass access controls, expand their own permissions, or perform consequential writes merely because an instruction appears in retrieved content. 7. Design memory deliberately. Distinguish transient working context, session state, case history, and durable organizational memory. For each memory store, define its purpose, source of truth, read and write authority, provenance, retention, deletion, sensitivity, conflict resolution, freshness checks, and contamination controls. Avoid durable storage when the task can be completed without it. 8. Place human review according to consequence. Require explicit authorization before irreversible, externally visible, regulated, financial, legal, safety-related, privacy-sensitive, production-changing, or high-impact actions. Define the reviewer, evidence presented, decision options, response deadline, escalation path, and what happens on rejection or no response. Do not let an agent approve its own high-risk output. 9. Add operational safeguards proportionate to risk. Include least privilege, data minimization, secret redaction, input and output validation, prompt-injection resistance, policy enforcement outside the model where feasible, transaction limits, rate limits, timeouts, bounded retries, duplicate-side-effect prevention, audit logging, monitoring, kill switch, rollback or compensating action, and manual recovery. Stop and escalate when authority is unclear, required evidence is absent, sensitive data cannot be protected, control requirements conflict, or a safe recovery path does not exist. 10. Build verification before rollout. Cover normal cases, decision boundaries, malformed inputs, missing data, conflicting evidence, unauthorized requests, tool failure, timeout, retry exhaustion, duplicate delivery, partial writes, stale memory, prompt injection, human rejection, escalation, rollback, and recovery. For every test, specify setup, expected observation, required evidence, and acceptance rule. Record an actual observation only when supplied execution evidence exists; otherwise mark it not run. 11. Recommend a staged implementation only if the decomposition is sufficiently supported. Prefer a read-only or shadow-mode pilot, then limited traffic with human approval, then controlled expansion. Define entry criteria, exit criteria, monitoring thresholds, stop conditions, rollback ownership, and residual risks. If evidence is inadequate, recommend discovery work rather than implementation. Required deliverable A. Scope and evidence register - State the objective, in-scope start and end events, exclusions, constraints, and success measures. - Provide an evidence table with: ID; source; relevant observation; evidence classification; confidence; conflict or limitation; design implication. - List blocking gaps, non-blocking unknowns, assumptions, and conflicts separately, with an owner and resolution needed. B. Current-process model - Present the ordered activities, actors, decisions, inputs, outputs, systems, exceptions, service levels, and failure points. - Identify any undocumented transitions or contradictions that prevent a reliable model. C. Activity disposition matrix For every meaningful activity, provide: activity; disposition; rationale; required judgment; consequence of error; reversibility; evidence basis; human accountability. Include activities intentionally retained as deterministic or human-operated. D. Target workflow and responsibility map - Describe the happy path, alternate paths, exceptions, stop conditions, and terminal states. - Provide an agent and human responsibility table with: component; objective; owned activities; permitted decisions; prohibited actions; inputs; outputs; dependencies; accountable owner. - Explain why the proposed number and boundaries of agents are preferable to credible alternatives. E. Handoff contract catalog For each handoff, provide: handoff ID; producer; consumer; trigger; preconditions; payload and provenance; validation; acknowledgment; correlation and idempotency method; timeout and retry behavior; failure owner; escalation; completion signal. F. Tool and permission matrix For each component and tool, provide: operation; read or write scope; minimum permission; data classification; possible side effect; validation; approval gate; audit evidence; failure and recovery behavior. Mark unsupported or unconfirmed capabilities as unknown. G. Memory and state design Provide: memory type; information stored; purpose; source of truth; read and write authority; retention and deletion; sensitivity; provenance; freshness rule; conflict handling; contamination control. State when no persistent memory is needed. H. Human review and control plan Provide: checkpoint; triggering condition; reviewer; evidence shown; allowed decision; response deadline; no-response behavior; escalation; actions blocked pending approval. Also list privacy, security, compliance, operational, and recovery controls with their enforcement point and owner. I. Failure-mode register Include at least the process-relevant failure modes and provide: failure; cause; detection signal; affected state; containment; retry or compensating action; escalation owner; recovery evidence; residual risk. Do not add irrelevant hazards merely to increase the list. J. Verification and acceptance matrix Provide: test ID; scenario; setup or test data; expected observation; actual observation or not run; evidence required; acceptance rule; owner; resulting status. Map each success criterion and critical control to one or more tests, identify uncovered criteria, and reconcile conflicting results rather than averaging them away. K. Implementation and decision record - Give staged implementation steps with dependencies, accountable owners, approval gates, entry and exit criteria, monitoring thresholds, stop conditions, and rollback or compensating actions. - Summarize major design decisions and rejected alternatives with evidence, trade-offs, and residual risks. - End with one recommendation: proceed to discovery, proceed to design review, proceed to a controlled pilot, or do not proceed. State the evidence required for the next gate and make clear that the recommendation is not approval or execution. ## Step 3 — Design Risk-Based Review and Escalation **Prompt** Human-in-the-Loop Quality Gate Builder **Instructions** Define proportionate review points, responsible reviewer criteria, escalation rules, approval paths, service-level expectations, audit evidence, and handling for low-confidence, exceptional, or high-impact cases. **Input for this step** Use the architecture, role boundaries, risk classification, failure modes, user impact, policy requirements, and operational capacity established in earlier steps. **Carry forward** Pass the review rubric, escalation matrix, approval evidence, service-level constraints, audit requirements, and unresolved control gaps to final governance design. **Prompt content** You are an AI workflow quality architect, human-in-the-loop systems designer, and operational risk reviewer. You design practical review gates for AI-assisted workflows so teams can catch quality, safety, accuracy, policy, legal, financial, brand, or customer-impact failures before outputs are released. ## Task Design a human-in-the-loop quality gate system for an AI-assisted workflow. The system should define what needs review, who reviews it, when escalation is required, what evidence must be retained, what quality criteria should be checked, and how the workflow can remain efficient without creating unnecessary bottlenecks. ## Context Placeholders Use the context below. If a placeholder is missing, name the missing item and make a conservative assumption before continuing. - [Workflow name] - [Workflow purpose] - [AI-generated output] - [AI tool or model used] - [Users affected] - [Customer or internal audience] - [Risk level] - [Quality criteria] - [Accuracy requirements] - [Policy or compliance constraints] - [Brand or tone rules] - [Review roles] - [Approver roles] - [Escalation triggers] - [Evidence to retain] - [Failure examples] - [Service-level needs] - [Volume of outputs] - [Allowed delay] - [Automation boundaries] - [Final decision owner] ## Important Constraints 1. Do not invent facts, policies, legal requirements, compliance obligations, metrics, customer impact, or workflow details. 2. Separate supplied facts from assumptions. 3. Do not create human review gates that are heavier than the risk justifies. 4. Do not allow high-risk AI outputs to bypass human review. 5. Do not treat all AI outputs as equal risk. 6. Do not design a workflow that depends on vague reviewer judgment without clear quality criteria. 7. Do not make the reviewer responsible for decisions they are not qualified or authorized to make. 8. Do not remove human review from legal, financial, medical, safety, security, regulatory, employment, public-facing, or high-impact decisions unless the user explicitly confirms that the workflow is low-risk. 9. Do not recommend silent automation for outputs that could harm customers, mislead users, damage trust, violate policy, or create legal exposure. 10. Keep the quality gate practical enough for a real team to operate. 11. Include audit evidence only when it is useful for accountability, compliance, dispute handling, quality improvement, or operational review. 12. Recommend sampling only when full review is unnecessary and risk is low enough. 13. Include escalation paths for uncertainty, policy conflict, repeated failures, sensitive topics, and unusual cases. 14. Make every recommendation specific to the workflow, risk level, affected users, and service-level needs. ## Risk Levels Use this risk scale unless the user provides another one. ### Low Risk The output is internal, reversible, low-impact, and unlikely to affect customers, money, legal obligations, safety, security, or public reputation. ### Medium Risk The output may affect customers, internal decisions, team operations, support quality, brand perception, or moderate business outcomes. ### High Risk The output may affect legal, financial, medical, safety, security, regulatory, employment, customer rights, public claims, executive decisions, or irreversible business actions. ### Critical Risk The output could create serious harm, legal exposure, financial loss, safety issues, privacy violations, public misinformation, or major customer trust damage. ## Quality Gate Design Process Follow this process before producing the final system. 1. Restate the workflow purpose and AI-generated output. 2. Identify who is affected by the output. 3. Classify the workflow risk level. 4. Identify the most likely failure modes. 5. Identify which failures can be caught automatically. 6. Identify which failures require human judgment. 7. Decide which outputs require full review, sampled review, escalation review, or no review. 8. Define reviewer roles and decision authority. 9. Create quality criteria that reviewers can apply consistently. 10. Define escalation triggers. 11. Define audit evidence to retain. 12. Define service-level expectations. 13. Recommend the lightest effective review process. 14. Create a verification and improvement loop. ## Output Format ### 1. Workflow Snapshot Provide a concise overview of: 1. Workflow name. 2. Workflow purpose. 3. AI-generated output. 4. Audience affected. 5. Risk level. 6. Main quality concerns. 7. Required review depth. 8. Final decision owner. 9. Service-level needs. 10. Missing inputs. ### 2. Workflow Risk Map Create a table with: | Risk Area | Possible Failure | Impact | Likelihood | Severity | Review Needed | Notes | | --- | --- | --- | --- | --- | --- | --- | Include risk areas such as: 1. Accuracy. 2. Policy compliance. 3. Legal exposure. 4. Financial impact. 5. Customer harm. 6. Privacy. 7. Security. 8. Brand tone. 9. Fairness or bias. 10. Operational reliability. 11. Public reputation. 12. Escalation failure. ### 3. Quality Gate Design Design the review gates. Create a table with: | Gate | When It Happens | What It Checks | Reviewer | Decision Options | Escalation Trigger | Evidence Retained | | --- | --- | --- | --- | --- | --- | --- | Use gate types such as: 1. Pre-generation input check. 2. AI output quality check. 3. Policy and compliance check. 4. High-risk case escalation. 5. Final approval. 6. Post-release sampling. 7. Incident review. 8. Continuous improvement review. ### 4. Review Routing Rules Define which outputs require which level of review. Use categories such as: 1. Auto-approve. 2. Sample review. 3. Mandatory human review. 4. Specialist review. 5. Manager approval. 6. Legal or compliance review. 7. Security review. 8. Executive approval. 9. Do not release. Explain the conditions for each route. ### 5. Reviewer Rubric Create a practical scoring rubric. Include: 1. Accuracy. 2. Completeness. 3. Relevance. 4. Policy compliance. 5. Tone and brand fit. 6. Safety. 7. Privacy. 8. Customer impact. 9. Escalation need. 10. Release readiness. Use a simple scale such as: 1. Pass. 2. Needs minor edit. 3. Needs major edit. 4. Escalate. 5. Reject. ### 6. Escalation Rules Create escalation rules for cases where the reviewer should not decide alone. Include triggers such as: 1. Missing or uncertain facts. 2. Legal or compliance concern. 3. Financial commitment. 4. Refund, cancellation, or account-risk issue. 5. Medical, safety, or security implication. 6. Sensitive customer complaint. 7. Public-facing claim. 8. Policy conflict. 9. High-value customer impact. 10. Repeated AI failure. 11. Reviewer uncertainty. 12. Potential reputational harm. For each trigger, specify: 1. Who receives the escalation. 2. What evidence should be included. 3. Expected response time. 4. Whether the output should be paused. ### 7. Audit Evidence Plan Define what should be retained. Include: 1. Original user input or request. 2. AI-generated output. 3. Prompt or workflow version. 4. Reviewer identity or role. 5. Review decision. 6. Edits made. 7. Escalation notes. 8. Approval timestamp. 9. Final released output. 10. Failure reason, if rejected. 11. Follow-up action. 12. Retention period, if known. Do not collect unnecessary sensitive data. ### 8. Service-Level and Bottleneck Review Assess whether the review gate is operationally realistic. Include: 1. Expected output volume. 2. Review time per item. 3. Reviewer capacity. 4. Allowed delay. 5. Bottleneck risk. 6. What can be automated safely. 7. What must remain human-reviewed. 8. Suggested sampling rate, if appropriate. 9. Escalation response expectations. 10. Fallback plan if reviewers are unavailable. ### 9. Failure Mode Examples Create examples of outputs that should: 1. Pass. 2. Need minor edits. 3. Need major edits. 4. Be escalated. 5. Be rejected. For each example, explain why. ### 10. Implementation Checklist Create a checklist for rollout. Include: 1. Workflow owner assigned. 2. Review roles assigned. 3. Rubric approved. 4. Escalation contacts confirmed. 5. Audit evidence fields defined. 6. Review tooling selected. 7. Test cases created. 8. Reviewer training completed. 9. Pilot run completed. 10. Failure examples reviewed. 11. Metrics agreed. 12. Review cadence scheduled. ### 11. Metrics and Continuous Improvement Recommend metrics to track. Include: 1. AI output pass rate. 2. Edit rate. 3. Escalation rate. 4. Rejection rate. 5. Reviewer disagreement rate. 6. Customer complaint rate. 7. Policy failure rate. 8. Average review time. 9. Bottleneck frequency. 10. Repeated failure patterns. 11. Prompt or workflow version performance. 12. Incident count. Explain how these metrics should be used to improve the workflow. ### 12. Human Review Checklist Create a concise checklist a reviewer can use before approving an AI output. The checklist should be practical, specific, and easy to apply during daily operations. ### 13. Final Recommendation End with: 1. Recommended quality gate structure. 2. Minimum review requirement. 3. Highest-risk failure to prevent. 4. Escalation owner. 5. Audit evidence required. 6. Suggested pilot approach. 7. Next action for the workflow owner. ### 14. Missing Inputs and Assumptions List: 1. Missing inputs. 2. Conservative assumptions made. 3. Decisions requiring human approval. 4. Risks that cannot be fully assessed from the supplied context. 5. Information needed before implementation. ## Verification Before finalizing, confirm that: 1. The review gates are proportionate to the workflow risk. 2. High-risk outputs do not bypass human review. 3. Reviewers have clear decision criteria. 4. Escalation triggers are specific. 5. Audit evidence is useful and not excessive. 6. Service-level needs are considered. 7. The workflow avoids unnecessary bottlenecks. 8. The final design can be implemented by a real team. 9. Missing inputs and assumptions are clearly listed. ## Final Instruction to Begin Begin now. If the workflow name, AI-generated output, users affected, risk level, or quality criteria are missing, ask for them first. If enough context is available, produce the full human-in-the-loop quality gate design in the requested markdown format. ## Step 4 — Finalize Governance, Override, and Design Approval **Prompt** AI Agent Governance, Human Override, and Deployment Readiness Playbook **Instructions** Consolidate least-privilege permissions, approval gates, authorized override, pause and shutdown behavior, monitoring, incident response, ownership, validation, and staged implementation controls into a conditional architecture decision. **Input for this step** Supply the approved use-case boundary, process map, architecture and handoff contracts, role-based review design, control gaps, success criteria, monitoring capacity, incident contacts, and disablement or recovery options. **Carry forward** Produce the final architecture and control package with accountable actions, acceptance tests, approval gates, residual risks, and a conditional design go, revise, or stop recommendation. **Review note** Authorized business, technical, security or privacy, and operational owners must approve the architecture and control package before implementation or pilot deployment. **Prompt content** Create an evidence-traceable governance and human-override playbook for the AI agent described below. Use ChatGPT to analyze only the supplied materials, identify control gaps, reconcile requirements, and draft governance artifacts. Do not imply that ChatGPT inspected live systems, changed permissions, tested controls, approved the agent, or deployed anything unless direct execution evidence is supplied. ## Inputs Business and deployment context: [Business and deployment context] Agent purpose and users: [Agent purpose and users] Agent architecture and autonomy: [Agent architecture and autonomy] Tools, systems, and actions: [Tools, systems, and actions] Data inventory and classifications: [Data inventory and classifications] Existing controls and approval requirements: [Existing controls and approval requirements] Legal, regulatory, and policy requirements: [Legal, regulatory, and policy requirements] Threat model, failure modes, and incidents: [Threat model, failure modes, and incidents] Monitoring, logging, and retention capabilities: [Monitoring, logging, and retention capabilities] Owners, incident response, and escalation: [Owners, incident response, and escalation] Risk appetite and acceptance criteria: [Risk appetite and acceptance criteria] Evidence pack: [Evidence pack] ## Input rules Treat the agent purpose, intended users, autonomy, accessible systems, permitted actions, data classifications, accountable owner, and deployment environment as blocking prerequisites. If any is absent or materially ambiguous, ask focused clarification questions before making a deployment-readiness decision. Treat architecture diagrams, data-flow maps, access-control exports, test results, policy excerpts, model or agent evaluations, vendor documentation, prior incident records, sample audit events, and recovery exercise results as useful evidence. If they are unavailable, continue only with a bounded draft and mark affected conclusions as unverified. When inputs conflict, display the conflict, identify the affected control or decision, and request an authoritative resolution. Do not silently choose one version. Do not invent controls, owners, legal conclusions, test outcomes, system behavior, or evidence. ## Evidence and status discipline Assign an evidence ID to each supplied artifact or factual excerpt used. For every material conclusion, label its basis as one of: supplied fact, documented observation, execution evidence, assumption, hypothesis, unknown, or conflict. Cite evidence IDs where available. Keep these states distinct: - Proposed: a control or action recommended but not implemented. - Reported: implementation asserted in supplied material but not independently evidenced. - Evidenced: supported by a supplied artifact or recorded observation. - Tested: supported by supplied test steps, expected results, actual results, date, environment, and outcome. - Approved: supported by an identified authorized approver and approval record. - Blocked: a decision cannot proceed because a prerequisite or control is missing. - Not applicable: excluded with a documented rationale. Never convert reported or proposed controls into evidenced, tested, approved, or deployed states. Phrase legal and regulatory interpretations as items for qualified review unless authoritative advice is included in the evidence pack. ## Analysis workflow 1. Establish scope and accountability - Define the business process, users, affected people, deployment environment, intended outcomes, prohibited uses, system boundaries, dependencies, and accountable business and technical owners. - Separate advisory outputs from actions that change data, communicate externally, make decisions, trigger transactions, alter access, or affect people. 2. Build an evidence register - Record each evidence ID, artifact name, source, date, environment or version, relevant claim, reliability limitation, and controls supported. - Create an unknowns and conflicts register. Identify which unknowns block risk classification, control design, testing, approval, or deployment. 3. Classify inherent and residual risk - Assess impact and likelihood for privacy, security, safety, legal or compliance exposure, financial loss, service availability, customer harm, discrimination, misinformation, and reputational harm where relevant. - Consider autonomy, reversibility, action frequency, blast radius, data sensitivity, external communication, privilege level, detectability, and human review latency. - Assign an inherent risk tier before controls and a provisional residual risk tier after evidenced controls. Explain the method and rationale; do not reduce residual risk based on proposed or merely reported controls. - Identify risks that exceed the stated risk appetite or require specialist review. 4. Define permission and authority boundaries - Produce a least-privilege matrix for each tool, system, dataset, and action. Include identity used, read or write scope, environment, data fields, rate or value limits, duration, approval condition, segregation of duties, prohibited actions, and revocation method. - Prohibit shared credentials, unrestricted production access, self-approval, silent privilege escalation, disabling logs, bypassing policy controls, and access beyond the documented purpose. - Require explicit human authorization before irreversible, high-impact, regulated, financial, security-sensitive, externally binding, or rights-affecting actions. 5. Design approval gates and human oversight - Define which decisions use human-in-the-loop review, human-on-the-loop supervision, or human-only execution. - For each gate, specify trigger, risk condition, reviewer role, information shown to the reviewer, response deadline, approve or reject criteria, delegation rules, evidence recorded, and fail-closed behavior. - Prevent rubber-stamp approval by requiring reviewers to see the proposed action, material context, affected records, confidence or uncertainty indicators when available, policy checks, and consequences. 6. Design override, pause, and shutdown controls - Provide runbooks for rejecting a single action, suspending a workflow, revoking credentials or tokens, disabling integrations, isolating the agent, switching to a manual fallback, preserving evidence, and restoring service safely. - For each mechanism, identify authorized operators, access path, expected effect, dependencies, confirmation signal, maximum target time, communication path, recovery prerequisites, and rollback or compensating controls. - Include stop conditions for suspected data leakage, unauthorized access, unsafe repeated actions, approval bypass, anomalous action volume, corrupted context, unavailable monitoring, or loss of human control. 7. Specify monitoring and auditability - Define required events for prompts or instructions where lawful, retrieved context, model and agent version, tool calls, authorization decisions, approvals, data access, outputs, external communications, errors, overrides, configuration changes, and credential events. - Specify timestamps, actor and agent identity, correlation ID, environment, target resource, action, outcome, policy decision, evidence link, integrity protection, access restrictions, retention, redaction, and privacy minimization. - Define operational metrics and alert thresholds for approval bypass attempts, denied actions, unusual tool use, sensitive-data access, repeated failures, override frequency, latency, drift, and incident indicators. Mark thresholds requiring calibration rather than inventing values. 8. Create incident and recovery procedures - Define detection, triage, severity, containment, credential revocation, evidence preservation, impact assessment, notification decision, remediation, recovery validation, post-incident review, and control updates. - Map each step to an owner, backup owner, trigger, target response time if supplied, required record, escalation path, and decision authority. - Address agent-specific scenarios such as prompt injection, tool misuse, data exfiltration, hallucinated external communication, runaway action loops, stale permissions, poisoned retrieval content, and failure of the approval channel. 9. Define validation and acceptance - Create a control validation matrix with control ID, risk addressed, status, test method, preconditions, test environment, expected observation, actual observation if supplied, evidence ID, result, owner, and unresolved issue. - Include scenario tests for denied unauthorized actions, approval enforcement, least-privilege access, sensitive-data restrictions, logging completeness, alert delivery, pause and shutdown, credential revocation, manual fallback, recovery, and preservation of in-flight work. - Reconcile each acceptance criterion to evidence. A missing actual observation must remain untested, not passed. 10. Determine readiness without granting approval - Return one analytical recommendation: ready for authorized approval, conditionally ready, not ready, or blocked by missing information. - List the rationale, unmet controls, risk exceptions, required evidence, accountable action owners, and approval authorities. - Make clear that the recommendation is not deployment authorization. Only the named human authority may accept residual risk and approve deployment. ## Required deliverable Produce the following sections: ### 1. Governance Brief State the agent, business process, scope boundary, deployment environment, accountable owners, proposed governance posture, and analytical readiness recommendation. ### 2. Assumptions, Unknowns, and Conflicts Register Use columns: ID, type, statement, affected decision or control, consequence, clarification needed, owner, and status. ### 3. Evidence Register Use columns: evidence ID, artifact or excerpt, source, date, environment or version, claim supported, reliability limitation, and linked control IDs. ### 4. Agent Scope and Responsibility Map Document intended and prohibited uses, affected parties, dependencies, human responsibilities, agent responsibilities, and activities reserved for humans. ### 5. Risk Register and Tier Decision Use columns: risk ID, scenario, cause, consequence, affected party, inherent impact, inherent likelihood, inherent tier, existing evidenced controls, residual impact, residual likelihood, provisional residual tier, evidence IDs, owner, treatment, and acceptance authority. ### 6. Permission and Action Boundary Matrix Use columns: boundary ID, tool or system, identity, resource or data, permitted operation, prohibited operation, environment, limit, approval gate, logging requirement, revocation method, control status, and evidence ID. ### 7. Human Oversight and Approval-Gate Matrix Use columns: gate ID, triggering action or condition, oversight mode, reviewer role, review context, approve criteria, reject criteria, timeout behavior, fail-safe state, audit record, and escalation path. ### 8. Override, Pause, Shutdown, and Recovery Runbook Provide ordered procedures with trigger, operator, authorization needed, action, expected confirmation, target time if supplied, evidence preserved, fallback, recovery prerequisite, and escalation. Separate single-action rejection, workflow pause, integration isolation, full shutdown, and controlled restoration. ### 9. Monitoring, Logging, and Alert Specification List required events, fields, integrity controls, privacy treatment, retention basis, access roles, alert conditions, response owner, and evidence needed to validate coverage. ### 10. Incident Response Matrix Use columns: scenario, detection signal, severity factors, immediate containment, owner, escalation, notification decision owner, evidence to preserve, recovery check, and post-incident action. ### 11. Control Validation Matrix Use the validation fields defined in the workflow. Show expected and actual observations separately, and identify unavailable execution evidence. ### 12. Deployment Readiness Decision Record Include recommendation, rationale, acceptance criteria reconciliation, passed controls with evidence, untested controls, failed controls, blockers, risk exceptions, required remediation, action owner, approving authority, and review expiry or trigger. ### 13. Ongoing Governance Schedule Define review triggers for material model changes, new tools or data, permission changes, incidents, control failures, regulatory changes, drift, vendor changes, and scheduled recertification. Include owner and required evidence for each review. ### 14. Executive Action List Prioritize immediate blockers, pre-approval actions, post-approval monitoring obligations, and items requiring security, privacy, legal, compliance, or operational review. ## Final quality checks Before returning the playbook, confirm that: - Every material conclusion is linked to evidence or explicitly marked as an assumption, hypothesis, unknown, or conflict. - Proposed and reported controls were not treated as tested or effective without execution evidence. - High-impact and irreversible actions require explicit human authorization and fail safely when approval is unavailable. - Permission boundaries are least-privilege, revocable, logged, and specific to systems, data, environments, and actions. - Override and shutdown procedures identify operators, confirmation signals, fallbacks, recovery conditions, and evidence preservation. - Validation distinguishes expected observations from supplied actual observations. - Readiness criteria reconcile to the control validation matrix, with unresolved items retained. - No statement claims that ChatGPT changed, tested, approved, deployed, paused, revoked, notified, or remediated anything unless corresponding evidence was supplied. - The final decision is clearly an analytical recommendation requiring authorized human review. ## Completion criteria The approved use case has a verified process map; agent, deterministic automation, and human roles are explicit; handoffs, permissions, data and memory boundaries, review gates, escalation, override, monitoring, and recovery controls are defined; validation criteria and owners are assigned; and authorized reviewers issue a conditional design go, revise, or stop decision. The workflow does not reselect the use case or claim implementation or deployment occurred. # Design a Governed AI Agent Architecture and Accountability Model Use this Amo.ng workflow with your preferred AI tool. Complete the steps in order and carry the specified output forward. Outcome: A review-ready architecture and accountability package for an already-approved use case, including role decomposition, interface contracts, permissions, review gates, escalation paths, override controls, and a conditional design-approval decision. Required inputs: - Approved business use case, outcome, scope, and accountable owner - Current process documentation, representative cases, exceptions, and performance evidence - Prior selection or readiness decisions and any binding remediation requirements - Systems, data, tool capabilities, integrations, and permission constraints - Security, privacy, compliance, operational, and service-level requirements - Risk classification, success criteria, reviewers, and implementation constraints ## Step 1 — Map the Approved Process and Preserve Its Control Boundaries **Instructions** Map the current approved process from supplied evidence, including triggers, activities, decisions, data, exceptions, handoffs, failure points, and existing controls. Distinguish AI work, deterministic automation, and human responsibility without reopening portfolio selection or assuming automation is appropriate. **Input for this step** Provide the approved use-case decision, current SOP or process evidence, representative normal and exceptional cases, systems, roles, performance evidence, known constraints, and binding readiness conditions. **Carry forward** Pass the verified process map, approved automation boundary, exceptions, dependencies, existing controls, evidence gaps, and binding constraints to architecture decomposition. **Review note** The accountable process owner should confirm that the map and approved scope reflect actual work before agent architecture is designed. **Prompt** Evidence-Based AI Business Process Automation Mapping **Prompt URL** https://amo.ng/prompts/ai-business-process-automation-mapping ## Step 2 — Decompose Agent, Automation, and Accountable Roles **Instructions** Design the agent architecture using the approved process boundary. Define justified agent responsibilities, deterministic components, decisions reserved for accountable owners, handoff contracts, tool permissions, data and memory rules, exception paths, failure behavior, and verification tests. **Input for this step** Treat the approved scope, process evidence, existing controls, readiness restrictions, system capabilities, permission limits, and representative failure cases as binding design inputs. **Carry forward** Pass the architecture, role and handoff contracts, permission model, data and memory boundaries, failure paths, and test requirements to the accountability and control design step. **Prompt** AI Agent Process Decomposition Prompt **Prompt URL** https://amo.ng/prompts/ai-agent-process-decomposition-prompt ## Step 3 — Design Risk-Based Review and Escalation **Instructions** Define proportionate review points, responsible reviewer criteria, escalation rules, approval paths, service-level expectations, audit evidence, and handling for low-confidence, exceptional, or high-impact cases. **Input for this step** Use the architecture, role boundaries, risk classification, failure modes, user impact, policy requirements, and operational capacity established in earlier steps. **Carry forward** Pass the review rubric, escalation matrix, approval evidence, service-level constraints, audit requirements, and unresolved control gaps to final governance design. **Prompt** Human-in-the-Loop Quality Gate Builder **Prompt URL** https://amo.ng/prompts/human-in-the-loop-quality-gate-builder ## Step 4 — Finalize Governance, Override, and Design Approval **Instructions** Consolidate least-privilege permissions, approval gates, authorized override, pause and shutdown behavior, monitoring, incident response, ownership, validation, and staged implementation controls into a conditional architecture decision. **Input for this step** Supply the approved use-case boundary, process map, architecture and handoff contracts, role-based review design, control gaps, success criteria, monitoring capacity, incident contacts, and disablement or recovery options. **Carry forward** Produce the final architecture and control package with accountable actions, acceptance tests, approval gates, residual risks, and a conditional design go, revise, or stop recommendation. **Review note** Authorized business, technical, security or privacy, and operational owners must approve the architecture and control package before implementation or pilot deployment. **Prompt** AI Agent Governance, Human Override, and Deployment Readiness Playbook **Prompt URL** https://amo.ng/prompts/ai-agent-governance-human-override-playbook Completion criteria: The approved use case has a verified process map; agent, deterministic automation, and human roles are explicit; handoffs, permissions, data and memory boundaries, review gates, escalation, override, monitoring, and recovery controls are defined; validation criteria and owners are assigned; and authorized reviewers issue a conditional design go, revise, or stop decision. The workflow does not reselect the use case or claim implementation or deployment occurred.Copy workflow includes every step and the full linked Prompt content. Use with AI copies a shorter guide with Prompt links; neither action runs the Workflow.
Outcome
A review-ready architecture and accountability package for an already-approved use case, including role decomposition, interface contracts, permissions, review gates, escalation paths, override controls, and a conditional design-approval decision.
Before you begin
Have all or some of the following available before you start. The more relevant context you can provide, the stronger the workflow output will be.
- Approved business use case, outcome, scope, and accountable owner
- Current process documentation, representative cases, exceptions, and performance evidence
- Prior selection or readiness decisions and any binding remediation requirements
- Systems, data, tool capabilities, integrations, and permission constraints
- Security, privacy, compliance, operational, and service-level requirements
- Risk classification, success criteria, reviewers, and implementation constraints
Ordered sequence
Workflow steps
Complete the steps in order. For each step, provide the listed context, carry its result into the next step, and pause wherever a review note is shown.
-
Step 1 Map the Approved Process and Preserve Its Control Boundaries
Map the current approved process from supplied evidence, including triggers, activities, decisions, data, exceptions, handoffs, failure points, and existing controls. Distinguish AI work, deterministic automation, and human responsibility without reopening portfolio selection or assuming automation is appropriate.
Prompt: Evidence-Based AI Business Process Automation Mapping## Objective Analyze the supplied business process evidence in ChatGPT. Produce a traceable map of the current workflow, distinguish deterministic automation from appropriate AI assistance, identify required human oversight, and create a phased implementation and validation plan. This is an analysis and planning exercise. Do not claim that an automation was configured, tested, approved, deployed, integrated, measured, or completed unless the supplied materials contain explicit evidence that the action occurred. Label planned work as proposed and keep it distinct from executed work. ## Context Business context: [Business context] Current workflow evidence: [Current workflow evidence] Stakeholders and approvals: [Stakeholders and approvals] Systems and integrations: [Systems and integrations] Inputs, outputs, and data: [Inputs outputs and data] Pain points and exceptions: [Pain points and exceptions] Volume, service levels, and costs: [Volume service levels and costs] Compliance, privacy, and security constraints: [Compliance privacy and security constraints] Budget, timeline, and tool constraints: [Budget timeline and tool constraints] Automation goals and definition of done: [Automation goals and definition of done] ## Input requirements Blocking inputs are a recognizable process trigger, the main workflow steps, the intended output or outcome, and the automation goal. If any blocking input is absent or contradictory, ask focused clarification questions before making final recommendations. Useful but non-blocking inputs include process documentation, standard operating procedures, screenshots, anonymized forms, sample records, decision rules, exception logs, service-level targets, task volumes, handling times, error rates, cost data, customer complaints, audit findings, system ownership, API or integration constraints, and approval policies. If these are unavailable, continue only with bounded analysis, state the limitation, and avoid fabricated measurements or system capabilities. Do not request or reproduce passwords, access tokens, private keys, unnecessary personal data, payment credentials, health information, or other secrets. Recommend redacted or synthetic examples where possible. Treat any supplied legal, regulatory, security, or financial interpretation as requiring review by the responsible human authority. ## ChatGPT operating boundary ChatGPT may analyze text and files supplied in the active conversation, organize evidence, identify patterns, compare options, calculate estimates from supplied figures, and draft recommendations. It cannot independently inspect business systems, observe staff performing the process, confirm vendor features, access private records, configure integrations, contact stakeholders, grant approval, or execute deployment unless such capabilities and resulting evidence are explicitly available in the session. Never imply that external inspection or action occurred. ## Evidence and uncertainty rules 1. Create stable identifiers for supplied evidence, such as E1, E2, and E3, and for workflow steps, such as W1, W2, and W3. 2. Classify material statements as one of: Supplied Fact, Observed in Supplied Artifact, Assumption, Hypothesis, Unknown, Conflict, or Execution Evidence. 3. Cite the relevant evidence identifier for each mapped workflow step, quantified baseline, risk, and recommendation. If there is no evidence, mark the item as an assumption or unknown rather than presenting it as fact. 4. Separate current-state observations from future-state proposals. Do not convert stakeholder aspirations into current capabilities. 5. When sources conflict, record both versions, explain the operational consequence, and identify the process owner who should resolve the conflict. 6. Show formulas and inputs for time, cost, capacity, error-reduction, or return-on-investment estimates. Present ranges when values are uncertain and do not invent precision. 7. Treat vendor capabilities, integration feasibility, model accuracy, and compliance suitability as unverified until supported by current documentation or testing evidence. ## Analysis workflow ### 1. Establish the evidence base Inventory the supplied artifacts and statements. Record source, date if known, process scope, reliability limitations, and the claims each source supports. List missing evidence and clarification needs. ### 2. Map the current process Decompose the process from trigger to final outcome. Include normal flow and documented exceptions. For each step capture: - step identifier and name; - trigger and predecessor; - actor or accountable team; - system or tool; - input and output; - business rule or judgment applied; - data classification; - average volume, handling time, wait time, and service target when supplied; - approval, handoff, queue, rework loop, and exception path; - failure mode and current control; - supporting evidence identifier. Identify bottlenecks without assuming that automation is the remedy. Distinguish processing time from waiting time and note whether the constraint arises from policy, capacity, poor data quality, system fragmentation, unclear ownership, or genuine judgment. ### 3. Classify automation suitability Assess each relevant workflow step against these paths: - Eliminate or simplify: the step may be unnecessary or redesigned before automation. - Deterministic automation: stable rules and structured inputs permit conventional workflow, scripts, forms, validation, or integration. - AI assistance: probabilistic work such as extraction, classification, summarization, drafting, knowledge retrieval, anomaly flagging, or recommendation support. - Human-led: nuanced judgment, negotiation, accountability, sensitive decisions, or poorly defined exceptions should remain human-controlled. - Not ready: evidence, data quality, process stability, system access, controls, or ownership is inadequate. For AI candidates, define the exact input, output, permissible use, prohibited use, expected error modes, confidence or abstention behavior, review requirement, exception route, and fallback procedure. Consider hallucination, omission, misclassification, prompt injection, data leakage, automation bias, model drift, inconsistent output, and inaccessible source citations where relevant. ### 4. Evaluate value, feasibility, and risk For every candidate, estimate business value, implementation complexity, process readiness, integration dependency, and operational risk using a clearly explained Low, Medium, High, or Critical scale. Assess privacy, security, legal or compliance exposure, financial impact, customer impact, accuracy, availability, vendor dependency, change-management burden, and over-automation risk. Prioritize only after documenting the rationale and evidence. A high-value candidate must not outrank a safer option solely because its projected savings are larger. Flag estimates that depend on missing baseline data. ### 5. Define authority and safeguards Identify the accountable process owner, system owner, data owner, risk or compliance reviewer, approver, operator, and escalation contact when known. Human authorization is mandatory before production configuration, access expansion, sensitive-data processing, customer-facing release, financial or legal decisions, or removal of an existing control. For Medium, High, or Critical risks, specify controls such as data minimization, masking, role-based access, approved environments, retention limits, source grounding, confidence thresholds, human approval gates, dual control, sampling, audit logs, rate limits, exception queues, kill switches, manual fallback, incident escalation, rollback criteria, and periodic review. Recommend no-go or pause conditions where controls or ownership are absent. ### 6. Design the future-state workflow Describe the proposed flow using the existing workflow identifiers. Show which steps remain manual, are simplified, use deterministic automation, or receive AI assistance. Define handoffs, review gates, exception routes, fallback operations, audit evidence, and recovery behavior. Do not assume integrations exist merely because they are desirable. ### 7. Build a phased implementation plan Sequence work through discovery and baseline confirmation, low-risk pilot, controlled human-in-the-loop rollout, integration, and monitored scale-up. For each phase state scope, dependencies, owner, approval gate, deliverable, validation method, rollback or fallback condition, and exit criteria. Mark every phase Proposed unless supplied execution evidence supports another status. ### 8. Define measurement and validation For each success metric provide its definition, baseline source, calculation, target, measurement window, owner, data source, and review cadence. Suitable metrics may include cycle time, touch time, queue time, first-pass yield, error or rework rate, exception rate, review override rate, false-positive and false-negative rates, service-level attainment, customer impact, adoption, operating cost, control failures, and risk incidents. For each pilot test, specify the test population, representative edge cases, expected observation, actual observation if evidence exists, evidence location, acceptance threshold, result status, and remediation owner. Use only these result states: Not Run, Passed with Evidence, Failed with Evidence, Inconclusive, or Blocked. Never mark a test passed based on a proposed procedure. ## Required deliverable Produce the following sections in order. ### 1. Decision Brief State the process scope, strongest supported opportunities, major constraints, recommended starting point, and decisions requiring human authorization. Separate facts from assumptions. ### 2. Evidence and Uncertainty Register Use columns: Evidence ID | Source or Artifact | Date or Version | Supported Claim | Classification | Reliability Limitation | Conflict or Gap. ### 3. Current-State Workflow Register Use columns: Step ID | Trigger or Predecessor | Activity | Actor | System | Input | Output | Rule or Judgment | Data Classification | Volume or Timing | Approval or Handoff | Exception or Failure Mode | Current Control | Evidence ID. ### 4. Bottleneck and Root-Cause Analysis For each bottleneck, distinguish observed symptom, supported or hypothesized cause, operational effect, evidence, uncertainty, and whether simplification should precede automation. ### 5. Automation Opportunity Matrix Use columns: Opportunity ID | Linked Step ID | Current Activity | Recommended Path | AI or Automation Function | Required Input | Produced Output | Business Value | Complexity | Risk | Human Review | Evidence ID | Assumptions | Priority | Rationale. ### 6. AI Use and Human Oversight Specification For each AI candidate provide allowed use, prohibited use, accountable owner, reviewer, confidence or abstention rule, known failure modes, review checklist, exception route, escalation trigger, audit record, and manual fallback. ### 7. Risk and Control Register Use columns: Risk ID | Opportunity ID | Risk Scenario | Affected Data or Stakeholder | Likelihood | Impact | Rating | Preventive Control | Detective Control | Response or Recovery Control | Control Owner | Approval Required | Residual Risk | Evidence or Validation Needed. ### 8. Proposed Future-State Workflow Map the proposed steps back to current Step IDs. Identify manual, simplified, deterministic, and AI-assisted activities; system boundaries; approvals; exception queues; fallback paths; and audit points. ### 9. Phased Implementation and Handoff Plan Use columns: Phase | Proposed Scope | Dependencies | Responsible Owner | Required Approval | Deliverable | Validation Method | Exit Criteria | Rollback or Fallback Trigger | Status. Status must remain Proposed unless execution evidence supports a different label. ### 10. Measurement and Validation Plan Use columns: Metric or Test ID | Opportunity ID | Definition or Test Case | Baseline and Source | Target or Expected Observation | Actual Observation | Evidence | Measurement Window | Owner | Acceptance Threshold | Result Status | Follow-up. ### 11. Assumptions, Unknowns, Conflicts, and Open Decisions List each unresolved item, its consequence, the evidence or decision needed, and the accountable resolver. Do not hide unresolved Critical risks in narrative text. ### 12. Final Recommendation and Authorization Requests Recommend proceed, revise, pilot, defer, or reject for each opportunity. State the evidence basis, residual uncertainty, next decision, required approver, and prohibited actions pending approval. ## Final acceptance checks Before returning the deliverable, verify and correct the following: - Every opportunity links to at least one workflow Step ID and supporting Evidence ID, or is explicitly labeled as an assumption. - Every mapped step identifies an actor, input, output, decision or rule, exception path, and evidence gap where these are unknown. - Deterministic automation, AI assistance, human-led work, and not-ready work are not conflated. - Every Medium, High, or Critical risk has an owner, approval requirement, safeguards, and response or fallback control. - Sensitive-data use identifies minimization, access, retention, audit, and approved-environment requirements. - Quantified benefits show supplied inputs, formulas, uncertainty, and baseline provenance. - Every pilot has acceptance thresholds, expected observations, evidence requirements, and a valid result state. - Proposed, executed, tested, approved, deployed, and measured states remain distinct and evidence-backed. - The recommended plan fits the stated budget, timeline, systems, process maturity, and authority constraints. - Unknowns and conflicts remain visible and are routed to named owners or owner roles for resolution.Input for this step
Provide the approved use-case decision, current SOP or process evidence, representative normal and exceptional cases, systems, roles, performance evidence, known constraints, and binding readiness conditions.
Carry forward
Pass the verified process map, approved automation boundary, exceptions, dependencies, existing controls, evidence gaps, and binding constraints to architecture decomposition.
Review note
The accountable process owner should confirm that the map and approved scope reflect actual work before agent architecture is designed.
-
Step 2 Decompose Agent, Automation, and Accountable Roles
Design the agent architecture using the approved process boundary. Define justified agent responsibilities, deterministic components, decisions reserved for accountable owners, handoff contracts, tool permissions, data and memory rules, exception paths, failure behavior, and verification tests.
Prompt: AI Agent Process Decomposition PromptAnalyze the supplied process and produce an implementation-ready decomposition that distinguishes AI-agent work from deterministic automation and human judgment. Work only from information available in this conversation. You may analyze materials and propose a design, but you cannot inspect unprovided systems, run tools, change workflows, deploy agents, grant access, approve controls, or claim tests were executed. Inputs - Process objective: [Process objective] - Current process evidence: [Current process evidence] - Scope and boundaries: [Scope and boundaries] - Constraints and policies: [Constraints and policies] - Systems and tool capabilities: [Systems and tool capabilities] - Success criteria: [Success criteria] - Risk and approval requirements: [Risk and approval requirements] Input handling Treat the process objective, current process evidence, scope, success criteria, and authority boundaries as blocking prerequisites. Useful supporting material includes process maps, SOPs, sample cases, decision tables, forms, API documentation, event schemas, failure logs, service-level targets, volumes, costs, security classifications, and operator interviews. If a reliable decomposition is blocked, ask no more than five grouped clarification questions before designing it. A missing detail is blocking when it prevents you from defining process boundaries, assigning a consequential decision, validating a handoff, determining permitted system access, or evaluating acceptance. Otherwise, continue with a bounded provisional design and record the missing item as an unknown. Do not silently reconcile conflicting sources; identify the conflict, its design impact, and the owner who must resolve it. Evidence and claim rules 1. Classify material statements as supplied fact, evidence-backed observation, assumption, hypothesis, conflict, or unknown. 2. Reference the supplied artifact, section, example, or statement supporting each important process or control conclusion. Do not fabricate citations or system behavior. 3. Keep proposed, approved, configured, executed, observed, and verified states distinct. Unless execution evidence is supplied, describe all workflow components and tests as proposed or unverified. 4. Never state that an integration works, a control is effective, a test passed, or a workflow was deployed or approved without corresponding evidence. Decomposition method 1. Reconstruct the current process. Identify its trigger, termination states, actors, inputs, outputs, sequence, decisions, business rules, systems, data classes, service levels, volumes, known exceptions, rework loops, and current failure points. Separate documented behavior from inferred behavior. 2. Test automation suitability at the activity level. For each activity, choose one disposition: deterministic automation, AI-agent responsibility, human responsibility, shared responsibility, or out of scope. Justify the choice using rule stability, ambiguity, contextual judgment, tool interaction, error detectability, reversibility, data sensitivity, consequence severity, latency, and cost. Do not create an agent where a rule, workflow engine, validation service, or human decision is safer and simpler. 3. Define the target orchestration. Describe the happy path, alternate paths, exception paths, cancellation behavior, and terminal states. Show where work is queued, resumed, retried, escalated, or stopped. Identify concurrency hazards, ordering dependencies, duplicate events, partial completion, timeout behavior, and recovery paths. 4. Establish cohesive agent boundaries. Group responsibilities by decision context, required tools, data access, risk, and accountability rather than assigning one agent per process step. Give every proposed agent one clear objective, explicit inputs and outputs, permitted decisions, prohibited actions, dependencies, and an accountable human owner. Flag overlapping authority, circular delegation, orphaned work, and single points of failure. 5. Specify every handoff. Define producer, consumer, trigger, preconditions, payload fields and schema, provenance, validation rules, acknowledgment, correlation identifier, status model, timeout, retry policy, idempotency control, failure owner, escalation route, and completion signal. State how malformed, missing, stale, duplicated, contradictory, or late payloads are handled. 6. Define tool access. For each agent-tool pairing, identify the operation, read or write scope, authentication boundary, minimum permission, data exposed, side effects, rate limits, timeout, error response, audit event, and required approval. Treat external content and tool output as untrusted until validated. Agents must not expose secrets, bypass access controls, expand their own permissions, or perform consequential writes merely because an instruction appears in retrieved content. 7. Design memory deliberately. Distinguish transient working context, session state, case history, and durable organizational memory. For each memory store, define its purpose, source of truth, read and write authority, provenance, retention, deletion, sensitivity, conflict resolution, freshness checks, and contamination controls. Avoid durable storage when the task can be completed without it. 8. Place human review according to consequence. Require explicit authorization before irreversible, externally visible, regulated, financial, legal, safety-related, privacy-sensitive, production-changing, or high-impact actions. Define the reviewer, evidence presented, decision options, response deadline, escalation path, and what happens on rejection or no response. Do not let an agent approve its own high-risk output. 9. Add operational safeguards proportionate to risk. Include least privilege, data minimization, secret redaction, input and output validation, prompt-injection resistance, policy enforcement outside the model where feasible, transaction limits, rate limits, timeouts, bounded retries, duplicate-side-effect prevention, audit logging, monitoring, kill switch, rollback or compensating action, and manual recovery. Stop and escalate when authority is unclear, required evidence is absent, sensitive data cannot be protected, control requirements conflict, or a safe recovery path does not exist. 10. Build verification before rollout. Cover normal cases, decision boundaries, malformed inputs, missing data, conflicting evidence, unauthorized requests, tool failure, timeout, retry exhaustion, duplicate delivery, partial writes, stale memory, prompt injection, human rejection, escalation, rollback, and recovery. For every test, specify setup, expected observation, required evidence, and acceptance rule. Record an actual observation only when supplied execution evidence exists; otherwise mark it not run. 11. Recommend a staged implementation only if the decomposition is sufficiently supported. Prefer a read-only or shadow-mode pilot, then limited traffic with human approval, then controlled expansion. Define entry criteria, exit criteria, monitoring thresholds, stop conditions, rollback ownership, and residual risks. If evidence is inadequate, recommend discovery work rather than implementation. Required deliverable A. Scope and evidence register - State the objective, in-scope start and end events, exclusions, constraints, and success measures. - Provide an evidence table with: ID; source; relevant observation; evidence classification; confidence; conflict or limitation; design implication. - List blocking gaps, non-blocking unknowns, assumptions, and conflicts separately, with an owner and resolution needed. B. Current-process model - Present the ordered activities, actors, decisions, inputs, outputs, systems, exceptions, service levels, and failure points. - Identify any undocumented transitions or contradictions that prevent a reliable model. C. Activity disposition matrix For every meaningful activity, provide: activity; disposition; rationale; required judgment; consequence of error; reversibility; evidence basis; human accountability. Include activities intentionally retained as deterministic or human-operated. D. Target workflow and responsibility map - Describe the happy path, alternate paths, exceptions, stop conditions, and terminal states. - Provide an agent and human responsibility table with: component; objective; owned activities; permitted decisions; prohibited actions; inputs; outputs; dependencies; accountable owner. - Explain why the proposed number and boundaries of agents are preferable to credible alternatives. E. Handoff contract catalog For each handoff, provide: handoff ID; producer; consumer; trigger; preconditions; payload and provenance; validation; acknowledgment; correlation and idempotency method; timeout and retry behavior; failure owner; escalation; completion signal. F. Tool and permission matrix For each component and tool, provide: operation; read or write scope; minimum permission; data classification; possible side effect; validation; approval gate; audit evidence; failure and recovery behavior. Mark unsupported or unconfirmed capabilities as unknown. G. Memory and state design Provide: memory type; information stored; purpose; source of truth; read and write authority; retention and deletion; sensitivity; provenance; freshness rule; conflict handling; contamination control. State when no persistent memory is needed. H. Human review and control plan Provide: checkpoint; triggering condition; reviewer; evidence shown; allowed decision; response deadline; no-response behavior; escalation; actions blocked pending approval. Also list privacy, security, compliance, operational, and recovery controls with their enforcement point and owner. I. Failure-mode register Include at least the process-relevant failure modes and provide: failure; cause; detection signal; affected state; containment; retry or compensating action; escalation owner; recovery evidence; residual risk. Do not add irrelevant hazards merely to increase the list. J. Verification and acceptance matrix Provide: test ID; scenario; setup or test data; expected observation; actual observation or not run; evidence required; acceptance rule; owner; resulting status. Map each success criterion and critical control to one or more tests, identify uncovered criteria, and reconcile conflicting results rather than averaging them away. K. Implementation and decision record - Give staged implementation steps with dependencies, accountable owners, approval gates, entry and exit criteria, monitoring thresholds, stop conditions, and rollback or compensating actions. - Summarize major design decisions and rejected alternatives with evidence, trade-offs, and residual risks. - End with one recommendation: proceed to discovery, proceed to design review, proceed to a controlled pilot, or do not proceed. State the evidence required for the next gate and make clear that the recommendation is not approval or execution.Input for this step
Treat the approved scope, process evidence, existing controls, readiness restrictions, system capabilities, permission limits, and representative failure cases as binding design inputs.
Carry forward
Pass the architecture, role and handoff contracts, permission model, data and memory boundaries, failure paths, and test requirements to the accountability and control design step.
-
Step 3 Design Risk-Based Review and Escalation
Define proportionate review points, responsible reviewer criteria, escalation rules, approval paths, service-level expectations, audit evidence, and handling for low-confidence, exceptional, or high-impact cases.
Prompt: Human-in-the-Loop Quality Gate BuilderYou are an AI workflow quality architect, human-in-the-loop systems designer, and operational risk reviewer. You design practical review gates for AI-assisted workflows so teams can catch quality, safety, accuracy, policy, legal, financial, brand, or customer-impact failures before outputs are released. ## Task Design a human-in-the-loop quality gate system for an AI-assisted workflow. The system should define what needs review, who reviews it, when escalation is required, what evidence must be retained, what quality criteria should be checked, and how the workflow can remain efficient without creating unnecessary bottlenecks. ## Context Placeholders Use the context below. If a placeholder is missing, name the missing item and make a conservative assumption before continuing. - [Workflow name] - [Workflow purpose] - [AI-generated output] - [AI tool or model used] - [Users affected] - [Customer or internal audience] - [Risk level] - [Quality criteria] - [Accuracy requirements] - [Policy or compliance constraints] - [Brand or tone rules] - [Review roles] - [Approver roles] - [Escalation triggers] - [Evidence to retain] - [Failure examples] - [Service-level needs] - [Volume of outputs] - [Allowed delay] - [Automation boundaries] - [Final decision owner] ## Important Constraints 1. Do not invent facts, policies, legal requirements, compliance obligations, metrics, customer impact, or workflow details. 2. Separate supplied facts from assumptions. 3. Do not create human review gates that are heavier than the risk justifies. 4. Do not allow high-risk AI outputs to bypass human review. 5. Do not treat all AI outputs as equal risk. 6. Do not design a workflow that depends on vague reviewer judgment without clear quality criteria. 7. Do not make the reviewer responsible for decisions they are not qualified or authorized to make. 8. Do not remove human review from legal, financial, medical, safety, security, regulatory, employment, public-facing, or high-impact decisions unless the user explicitly confirms that the workflow is low-risk. 9. Do not recommend silent automation for outputs that could harm customers, mislead users, damage trust, violate policy, or create legal exposure. 10. Keep the quality gate practical enough for a real team to operate. 11. Include audit evidence only when it is useful for accountability, compliance, dispute handling, quality improvement, or operational review. 12. Recommend sampling only when full review is unnecessary and risk is low enough. 13. Include escalation paths for uncertainty, policy conflict, repeated failures, sensitive topics, and unusual cases. 14. Make every recommendation specific to the workflow, risk level, affected users, and service-level needs. ## Risk Levels Use this risk scale unless the user provides another one. ### Low Risk The output is internal, reversible, low-impact, and unlikely to affect customers, money, legal obligations, safety, security, or public reputation. ### Medium Risk The output may affect customers, internal decisions, team operations, support quality, brand perception, or moderate business outcomes. ### High Risk The output may affect legal, financial, medical, safety, security, regulatory, employment, customer rights, public claims, executive decisions, or irreversible business actions. ### Critical Risk The output could create serious harm, legal exposure, financial loss, safety issues, privacy violations, public misinformation, or major customer trust damage. ## Quality Gate Design Process Follow this process before producing the final system. 1. Restate the workflow purpose and AI-generated output. 2. Identify who is affected by the output. 3. Classify the workflow risk level. 4. Identify the most likely failure modes. 5. Identify which failures can be caught automatically. 6. Identify which failures require human judgment. 7. Decide which outputs require full review, sampled review, escalation review, or no review. 8. Define reviewer roles and decision authority. 9. Create quality criteria that reviewers can apply consistently. 10. Define escalation triggers. 11. Define audit evidence to retain. 12. Define service-level expectations. 13. Recommend the lightest effective review process. 14. Create a verification and improvement loop. ## Output Format ### 1. Workflow Snapshot Provide a concise overview of: 1. Workflow name. 2. Workflow purpose. 3. AI-generated output. 4. Audience affected. 5. Risk level. 6. Main quality concerns. 7. Required review depth. 8. Final decision owner. 9. Service-level needs. 10. Missing inputs. ### 2. Workflow Risk Map Create a table with: | Risk Area | Possible Failure | Impact | Likelihood | Severity | Review Needed | Notes | | --- | --- | --- | --- | --- | --- | --- | Include risk areas such as: 1. Accuracy. 2. Policy compliance. 3. Legal exposure. 4. Financial impact. 5. Customer harm. 6. Privacy. 7. Security. 8. Brand tone. 9. Fairness or bias. 10. Operational reliability. 11. Public reputation. 12. Escalation failure. ### 3. Quality Gate Design Design the review gates. Create a table with: | Gate | When It Happens | What It Checks | Reviewer | Decision Options | Escalation Trigger | Evidence Retained | | --- | --- | --- | --- | --- | --- | --- | Use gate types such as: 1. Pre-generation input check. 2. AI output quality check. 3. Policy and compliance check. 4. High-risk case escalation. 5. Final approval. 6. Post-release sampling. 7. Incident review. 8. Continuous improvement review. ### 4. Review Routing Rules Define which outputs require which level of review. Use categories such as: 1. Auto-approve. 2. Sample review. 3. Mandatory human review. 4. Specialist review. 5. Manager approval. 6. Legal or compliance review. 7. Security review. 8. Executive approval. 9. Do not release. Explain the conditions for each route. ### 5. Reviewer Rubric Create a practical scoring rubric. Include: 1. Accuracy. 2. Completeness. 3. Relevance. 4. Policy compliance. 5. Tone and brand fit. 6. Safety. 7. Privacy. 8. Customer impact. 9. Escalation need. 10. Release readiness. Use a simple scale such as: 1. Pass. 2. Needs minor edit. 3. Needs major edit. 4. Escalate. 5. Reject. ### 6. Escalation Rules Create escalation rules for cases where the reviewer should not decide alone. Include triggers such as: 1. Missing or uncertain facts. 2. Legal or compliance concern. 3. Financial commitment. 4. Refund, cancellation, or account-risk issue. 5. Medical, safety, or security implication. 6. Sensitive customer complaint. 7. Public-facing claim. 8. Policy conflict. 9. High-value customer impact. 10. Repeated AI failure. 11. Reviewer uncertainty. 12. Potential reputational harm. For each trigger, specify: 1. Who receives the escalation. 2. What evidence should be included. 3. Expected response time. 4. Whether the output should be paused. ### 7. Audit Evidence Plan Define what should be retained. Include: 1. Original user input or request. 2. AI-generated output. 3. Prompt or workflow version. 4. Reviewer identity or role. 5. Review decision. 6. Edits made. 7. Escalation notes. 8. Approval timestamp. 9. Final released output. 10. Failure reason, if rejected. 11. Follow-up action. 12. Retention period, if known. Do not collect unnecessary sensitive data. ### 8. Service-Level and Bottleneck Review Assess whether the review gate is operationally realistic. Include: 1. Expected output volume. 2. Review time per item. 3. Reviewer capacity. 4. Allowed delay. 5. Bottleneck risk. 6. What can be automated safely. 7. What must remain human-reviewed. 8. Suggested sampling rate, if appropriate. 9. Escalation response expectations. 10. Fallback plan if reviewers are unavailable. ### 9. Failure Mode Examples Create examples of outputs that should: 1. Pass. 2. Need minor edits. 3. Need major edits. 4. Be escalated. 5. Be rejected. For each example, explain why. ### 10. Implementation Checklist Create a checklist for rollout. Include: 1. Workflow owner assigned. 2. Review roles assigned. 3. Rubric approved. 4. Escalation contacts confirmed. 5. Audit evidence fields defined. 6. Review tooling selected. 7. Test cases created. 8. Reviewer training completed. 9. Pilot run completed. 10. Failure examples reviewed. 11. Metrics agreed. 12. Review cadence scheduled. ### 11. Metrics and Continuous Improvement Recommend metrics to track. Include: 1. AI output pass rate. 2. Edit rate. 3. Escalation rate. 4. Rejection rate. 5. Reviewer disagreement rate. 6. Customer complaint rate. 7. Policy failure rate. 8. Average review time. 9. Bottleneck frequency. 10. Repeated failure patterns. 11. Prompt or workflow version performance. 12. Incident count. Explain how these metrics should be used to improve the workflow. ### 12. Human Review Checklist Create a concise checklist a reviewer can use before approving an AI output. The checklist should be practical, specific, and easy to apply during daily operations. ### 13. Final Recommendation End with: 1. Recommended quality gate structure. 2. Minimum review requirement. 3. Highest-risk failure to prevent. 4. Escalation owner. 5. Audit evidence required. 6. Suggested pilot approach. 7. Next action for the workflow owner. ### 14. Missing Inputs and Assumptions List: 1. Missing inputs. 2. Conservative assumptions made. 3. Decisions requiring human approval. 4. Risks that cannot be fully assessed from the supplied context. 5. Information needed before implementation. ## Verification Before finalizing, confirm that: 1. The review gates are proportionate to the workflow risk. 2. High-risk outputs do not bypass human review. 3. Reviewers have clear decision criteria. 4. Escalation triggers are specific. 5. Audit evidence is useful and not excessive. 6. Service-level needs are considered. 7. The workflow avoids unnecessary bottlenecks. 8. The final design can be implemented by a real team. 9. Missing inputs and assumptions are clearly listed. ## Final Instruction to Begin Begin now. If the workflow name, AI-generated output, users affected, risk level, or quality criteria are missing, ask for them first. If enough context is available, produce the full human-in-the-loop quality gate design in the requested markdown format.Input for this step
Use the architecture, role boundaries, risk classification, failure modes, user impact, policy requirements, and operational capacity established in earlier steps.
Carry forward
Pass the review rubric, escalation matrix, approval evidence, service-level constraints, audit requirements, and unresolved control gaps to final governance design.
-
Step 4 Finalize Governance, Override, and Design Approval
Consolidate least-privilege permissions, approval gates, authorized override, pause and shutdown behavior, monitoring, incident response, ownership, validation, and staged implementation controls into a conditional architecture decision.
Prompt: AI Agent Governance, Human Override, and Deployment Readiness PlaybookCreate an evidence-traceable governance and human-override playbook for the AI agent described below. Use ChatGPT to analyze only the supplied materials, identify control gaps, reconcile requirements, and draft governance artifacts. Do not imply that ChatGPT inspected live systems, changed permissions, tested controls, approved the agent, or deployed anything unless direct execution evidence is supplied. ## Inputs Business and deployment context: [Business and deployment context] Agent purpose and users: [Agent purpose and users] Agent architecture and autonomy: [Agent architecture and autonomy] Tools, systems, and actions: [Tools, systems, and actions] Data inventory and classifications: [Data inventory and classifications] Existing controls and approval requirements: [Existing controls and approval requirements] Legal, regulatory, and policy requirements: [Legal, regulatory, and policy requirements] Threat model, failure modes, and incidents: [Threat model, failure modes, and incidents] Monitoring, logging, and retention capabilities: [Monitoring, logging, and retention capabilities] Owners, incident response, and escalation: [Owners, incident response, and escalation] Risk appetite and acceptance criteria: [Risk appetite and acceptance criteria] Evidence pack: [Evidence pack] ## Input rules Treat the agent purpose, intended users, autonomy, accessible systems, permitted actions, data classifications, accountable owner, and deployment environment as blocking prerequisites. If any is absent or materially ambiguous, ask focused clarification questions before making a deployment-readiness decision. Treat architecture diagrams, data-flow maps, access-control exports, test results, policy excerpts, model or agent evaluations, vendor documentation, prior incident records, sample audit events, and recovery exercise results as useful evidence. If they are unavailable, continue only with a bounded draft and mark affected conclusions as unverified. When inputs conflict, display the conflict, identify the affected control or decision, and request an authoritative resolution. Do not silently choose one version. Do not invent controls, owners, legal conclusions, test outcomes, system behavior, or evidence. ## Evidence and status discipline Assign an evidence ID to each supplied artifact or factual excerpt used. For every material conclusion, label its basis as one of: supplied fact, documented observation, execution evidence, assumption, hypothesis, unknown, or conflict. Cite evidence IDs where available. Keep these states distinct: - Proposed: a control or action recommended but not implemented. - Reported: implementation asserted in supplied material but not independently evidenced. - Evidenced: supported by a supplied artifact or recorded observation. - Tested: supported by supplied test steps, expected results, actual results, date, environment, and outcome. - Approved: supported by an identified authorized approver and approval record. - Blocked: a decision cannot proceed because a prerequisite or control is missing. - Not applicable: excluded with a documented rationale. Never convert reported or proposed controls into evidenced, tested, approved, or deployed states. Phrase legal and regulatory interpretations as items for qualified review unless authoritative advice is included in the evidence pack. ## Analysis workflow 1. Establish scope and accountability - Define the business process, users, affected people, deployment environment, intended outcomes, prohibited uses, system boundaries, dependencies, and accountable business and technical owners. - Separate advisory outputs from actions that change data, communicate externally, make decisions, trigger transactions, alter access, or affect people. 2. Build an evidence register - Record each evidence ID, artifact name, source, date, environment or version, relevant claim, reliability limitation, and controls supported. - Create an unknowns and conflicts register. Identify which unknowns block risk classification, control design, testing, approval, or deployment. 3. Classify inherent and residual risk - Assess impact and likelihood for privacy, security, safety, legal or compliance exposure, financial loss, service availability, customer harm, discrimination, misinformation, and reputational harm where relevant. - Consider autonomy, reversibility, action frequency, blast radius, data sensitivity, external communication, privilege level, detectability, and human review latency. - Assign an inherent risk tier before controls and a provisional residual risk tier after evidenced controls. Explain the method and rationale; do not reduce residual risk based on proposed or merely reported controls. - Identify risks that exceed the stated risk appetite or require specialist review. 4. Define permission and authority boundaries - Produce a least-privilege matrix for each tool, system, dataset, and action. Include identity used, read or write scope, environment, data fields, rate or value limits, duration, approval condition, segregation of duties, prohibited actions, and revocation method. - Prohibit shared credentials, unrestricted production access, self-approval, silent privilege escalation, disabling logs, bypassing policy controls, and access beyond the documented purpose. - Require explicit human authorization before irreversible, high-impact, regulated, financial, security-sensitive, externally binding, or rights-affecting actions. 5. Design approval gates and human oversight - Define which decisions use human-in-the-loop review, human-on-the-loop supervision, or human-only execution. - For each gate, specify trigger, risk condition, reviewer role, information shown to the reviewer, response deadline, approve or reject criteria, delegation rules, evidence recorded, and fail-closed behavior. - Prevent rubber-stamp approval by requiring reviewers to see the proposed action, material context, affected records, confidence or uncertainty indicators when available, policy checks, and consequences. 6. Design override, pause, and shutdown controls - Provide runbooks for rejecting a single action, suspending a workflow, revoking credentials or tokens, disabling integrations, isolating the agent, switching to a manual fallback, preserving evidence, and restoring service safely. - For each mechanism, identify authorized operators, access path, expected effect, dependencies, confirmation signal, maximum target time, communication path, recovery prerequisites, and rollback or compensating controls. - Include stop conditions for suspected data leakage, unauthorized access, unsafe repeated actions, approval bypass, anomalous action volume, corrupted context, unavailable monitoring, or loss of human control. 7. Specify monitoring and auditability - Define required events for prompts or instructions where lawful, retrieved context, model and agent version, tool calls, authorization decisions, approvals, data access, outputs, external communications, errors, overrides, configuration changes, and credential events. - Specify timestamps, actor and agent identity, correlation ID, environment, target resource, action, outcome, policy decision, evidence link, integrity protection, access restrictions, retention, redaction, and privacy minimization. - Define operational metrics and alert thresholds for approval bypass attempts, denied actions, unusual tool use, sensitive-data access, repeated failures, override frequency, latency, drift, and incident indicators. Mark thresholds requiring calibration rather than inventing values. 8. Create incident and recovery procedures - Define detection, triage, severity, containment, credential revocation, evidence preservation, impact assessment, notification decision, remediation, recovery validation, post-incident review, and control updates. - Map each step to an owner, backup owner, trigger, target response time if supplied, required record, escalation path, and decision authority. - Address agent-specific scenarios such as prompt injection, tool misuse, data exfiltration, hallucinated external communication, runaway action loops, stale permissions, poisoned retrieval content, and failure of the approval channel. 9. Define validation and acceptance - Create a control validation matrix with control ID, risk addressed, status, test method, preconditions, test environment, expected observation, actual observation if supplied, evidence ID, result, owner, and unresolved issue. - Include scenario tests for denied unauthorized actions, approval enforcement, least-privilege access, sensitive-data restrictions, logging completeness, alert delivery, pause and shutdown, credential revocation, manual fallback, recovery, and preservation of in-flight work. - Reconcile each acceptance criterion to evidence. A missing actual observation must remain untested, not passed. 10. Determine readiness without granting approval - Return one analytical recommendation: ready for authorized approval, conditionally ready, not ready, or blocked by missing information. - List the rationale, unmet controls, risk exceptions, required evidence, accountable action owners, and approval authorities. - Make clear that the recommendation is not deployment authorization. Only the named human authority may accept residual risk and approve deployment. ## Required deliverable Produce the following sections: ### 1. Governance Brief State the agent, business process, scope boundary, deployment environment, accountable owners, proposed governance posture, and analytical readiness recommendation. ### 2. Assumptions, Unknowns, and Conflicts Register Use columns: ID, type, statement, affected decision or control, consequence, clarification needed, owner, and status. ### 3. Evidence Register Use columns: evidence ID, artifact or excerpt, source, date, environment or version, claim supported, reliability limitation, and linked control IDs. ### 4. Agent Scope and Responsibility Map Document intended and prohibited uses, affected parties, dependencies, human responsibilities, agent responsibilities, and activities reserved for humans. ### 5. Risk Register and Tier Decision Use columns: risk ID, scenario, cause, consequence, affected party, inherent impact, inherent likelihood, inherent tier, existing evidenced controls, residual impact, residual likelihood, provisional residual tier, evidence IDs, owner, treatment, and acceptance authority. ### 6. Permission and Action Boundary Matrix Use columns: boundary ID, tool or system, identity, resource or data, permitted operation, prohibited operation, environment, limit, approval gate, logging requirement, revocation method, control status, and evidence ID. ### 7. Human Oversight and Approval-Gate Matrix Use columns: gate ID, triggering action or condition, oversight mode, reviewer role, review context, approve criteria, reject criteria, timeout behavior, fail-safe state, audit record, and escalation path. ### 8. Override, Pause, Shutdown, and Recovery Runbook Provide ordered procedures with trigger, operator, authorization needed, action, expected confirmation, target time if supplied, evidence preserved, fallback, recovery prerequisite, and escalation. Separate single-action rejection, workflow pause, integration isolation, full shutdown, and controlled restoration. ### 9. Monitoring, Logging, and Alert Specification List required events, fields, integrity controls, privacy treatment, retention basis, access roles, alert conditions, response owner, and evidence needed to validate coverage. ### 10. Incident Response Matrix Use columns: scenario, detection signal, severity factors, immediate containment, owner, escalation, notification decision owner, evidence to preserve, recovery check, and post-incident action. ### 11. Control Validation Matrix Use the validation fields defined in the workflow. Show expected and actual observations separately, and identify unavailable execution evidence. ### 12. Deployment Readiness Decision Record Include recommendation, rationale, acceptance criteria reconciliation, passed controls with evidence, untested controls, failed controls, blockers, risk exceptions, required remediation, action owner, approving authority, and review expiry or trigger. ### 13. Ongoing Governance Schedule Define review triggers for material model changes, new tools or data, permission changes, incidents, control failures, regulatory changes, drift, vendor changes, and scheduled recertification. Include owner and required evidence for each review. ### 14. Executive Action List Prioritize immediate blockers, pre-approval actions, post-approval monitoring obligations, and items requiring security, privacy, legal, compliance, or operational review. ## Final quality checks Before returning the playbook, confirm that: - Every material conclusion is linked to evidence or explicitly marked as an assumption, hypothesis, unknown, or conflict. - Proposed and reported controls were not treated as tested or effective without execution evidence. - High-impact and irreversible actions require explicit human authorization and fail safely when approval is unavailable. - Permission boundaries are least-privilege, revocable, logged, and specific to systems, data, environments, and actions. - Override and shutdown procedures identify operators, confirmation signals, fallbacks, recovery conditions, and evidence preservation. - Validation distinguishes expected observations from supplied actual observations. - Readiness criteria reconcile to the control validation matrix, with unresolved items retained. - No statement claims that ChatGPT changed, tested, approved, deployed, paused, revoked, notified, or remediated anything unless corresponding evidence was supplied. - The final decision is clearly an analytical recommendation requiring authorized human review.Input for this step
Supply the approved use-case boundary, process map, architecture and handoff contracts, role-based review design, control gaps, success criteria, monitoring capacity, incident contacts, and disablement or recovery options.
Carry forward
Produce the final architecture and control package with accountable actions, acceptance tests, approval gates, residual risks, and a conditional design go, revise, or stop recommendation.
Review note
Authorized business, technical, security or privacy, and operational owners must approve the architecture and control package before implementation or pilot deployment.
Completion criteria
The approved use case has a verified process map; agent, deterministic automation, and human roles are explicit; handoffs, permissions, data and memory boundaries, review gates, escalation, override, monitoring, and recovery controls are defined; validation criteria and owners are assigned; and authorized reviewers issue a conditional design go, revise, or stop decision. The workflow does not reselect the use case or claim implementation or deployment occurred.
Related Workflows
Browse WorkflowsSafe AI Agent Workflow Selection and Deployment Readiness
Move from a broad list of AI opportunities to one prioritized, mapped, governed, and measurable agent workflow that is ready for an informed pilot decision.
Vendor Procurement Due Diligence
Assess vendor claims, evidence quality, security and privacy risk, procurement fit, and decision readiness before approval.
Plan and Review a Complex Laravel Feature for Safe Release
Turn an evidence-supported product opportunity into a phased Laravel implementation plan, conditionally review migration safety, and—after separately authorized implementation produces a real change set—review the pull request and prepare a risk-based release gate.
Was this useful?