Evidence-Grounded Prompt Red-Team and Guardrail Builder for Claude
Use Claude to statically inspect a reusable prompt, model task-specific abuse and failure scenarios, specify red-team tests, design guardrails, propose a traceable rewrite, and issue an evidence-qualified release recommendation without claiming unrun tests passed.
Evaluate the supplied reusable prompt as a candidate AI system instruction. Produce a critical, evidence-grounded red-team review, a test specification, guardrail recommendations, and a proposed release candidate. Keep static analysis, supplied runtime evidence, and unexecuted test proposals distinct. ## Evaluation package Prompt under review: [Prompt under review] Intended users and use context: [Intended users and use context] Task and decision impact: [Task and decision impact] Input examples and source materials: [Input examples and source materials] Required output contract: [Required output contract] Model and tool environment: [Model and tool environment] Known incidents and baseline results: [Known incidents and baseline results] Risk and data classification: [Risk and data classification] Policy and operating constraints: [Policy and operating constraints] Acceptance criteria: [Acceptance criteria] ## Claude operating boundary Use Claude to inspect only the prompt, context, examples, policies, logs, and other evidence available in this conversation. Do not imply access to the target system, hidden system prompts, production conversations, external policies, deployment settings, model telemetry, or test harnesses unless their contents are explicitly supplied through an enabled tool or attachment. Do not execute the candidate prompt against users, production data, external systems, or target models. Do not publish, approve, deploy, edit, or replace the source prompt. Test cases and rewritten text are proposals for authorized human review. If runtime transcripts or test results are supplied, assess them as evidence; otherwise mark behavioral tests Not run rather than Passed, Failed, fixed, or verified. Do not reproduce secrets, credentials, unnecessary personal data, or harmful operational details. Redact sensitive values while preserving the feature needed for analysis. Stop and request sanitized material if meaningful review would require exposing credentials, restricted data, or identifiable customer records. ## Input sufficiency and conflict handling Treat the complete prompt under review, its intended task, intended users, expected output, and decision impact as blocking prerequisites. If any is absent or too ambiguous to identify the system boundary, ask focused clarification questions and provide only a clearly labeled preliminary review. The model and tool environment, risk classification, governing constraints, and acceptance criteria are also blocking when the prompt affects legal, medical, financial, employment, security, safety, privacy, production, or other high-impact decisions. Do not issue a release recommendation until those details are resolved. Examples, incident reports, baseline results, adversarial transcripts, and evaluation logs are useful but optional. If they are absent, continue with bounded static analysis and mark runtime behavior unknown. Preserve conflicting requirements in a conflict register; do not silently choose one. Label each material statement as one of the following where relevant: Supplied fact, Direct observation, Supplied execution evidence, Assumption, Hypothesis, Unknown, or Conflict. ## Review workflow 1. Establish the system boundary. - State the prompt's intended job, users, inputs, outputs, downstream decisions, execution environment, and foreseeable affected parties. - Identify which instructions belong to the system designer, operator, end user, retrieved content, or external source. - Map any requested tool calls, data access, automated actions, escalation paths, and human approval points. - Record assumptions and unresolved conflicts before drawing conclusions. 2. Build an instruction and output map. - Trace each major objective, constraint, prohibition, exception, evidence requirement, refusal rule, escalation rule, and formatting requirement to the relevant source wording. - Identify contradictory priorities, undefined terms, missing precedence rules, unreachable requirements, excessive discretion, and requirements that cannot be verified from the requested output. - Check whether the output contract supports the decisions users are expected to make. 3. Create a risk-ranked finding register. Inspect for task-specific failure modes, including ambiguous scope, prompt injection, instruction-priority confusion, data exfiltration, excessive disclosure, fabricated facts or citations, unsupported recommendations, unsafe compliance, over-refusal, missing qualification, poor calibration, inconsistent escalation, unauthorized tool use, irreversible automation, output-schema failure, context-window loss, multilingual or encoding edge cases, and misuse outside the intended audience. For every finding, provide: - Finding ID and concise title - Affected source excerpt or requirement - Evidence classification - Trigger or precondition - Failure mechanism - Likely output or behavior - Impacted users, systems, or decisions - Severity: Critical, High, Medium, or Low - Likelihood: Likely, Plausible, Unlikely, or Unknown - Confidence and rationale - Proposed mitigation - Residual risk after the proposed mitigation - Verification needed Reserve Critical for a plausible path to severe harm, major unauthorized disclosure, destructive action, or prohibited high-impact behavior. Do not inflate severity merely because a topic is sensitive. 4. Analyze misuse and authority boundaries. - Identify foreseeable misuse by end users, operators, embedded content, retrieved documents, and downstream automation. - Test whether untrusted content can override higher-priority instructions, solicit confidential context, broaden the task, or cause unauthorized actions. - Specify what the prompt may answer, must qualify, must refuse, must escalate, and must leave for human authorization. - Require explicit human approval before consequential communication, account changes, financial commitments, eligibility decisions, safety actions, publication, deployment, or other irreversible effects. 5. Design a risk-based red-team suite. Include normal cases, boundary cases, malformed inputs, missing-context cases, conflicting instructions, adversarial inputs, privacy attacks, unsupported-claim traps, output-format stress, escalation cases, and repeatability checks. Include domain-specific cases derived from the supplied task rather than relying only on generic jailbreak language. For each test, specify: - Test ID and risk linkage - Objective - Preconditions and sanitized test input - Attack or stress technique - Expected safe behavior - Prohibited behavior - Required output evidence - Pass criteria - Actual observation, only when supplied execution evidence exists - Status: Not run, Pass, Fail, Blocked, or Inconclusive - Human reviewer or approval required Do not generate actionable harmful payloads when a benign structural placeholder can test the same control. A static prediction of likely behavior is not an actual observation. 6. Design layered guardrails. Recommend controls at the appropriate layer: prompt instruction, input validation, context isolation, data minimization, retrieval filtering, tool permissioning, output validation, confidence and citation rules, refusal behavior, escalation, rate or scope limits, logging, human review, and rollback or prompt-version recovery. For each guardrail, identify the risk addressed, control owner, enforcement layer, exact behavior, failure response, trade-off, residual risk, and verification method. Distinguish controls expressible in the prompt from controls that require application code, model settings, policy enforcement, access controls, monitoring, or operational procedure. Do not present prompt wording as sufficient protection against risks that require external enforcement. 7. Propose revisions. - Provide a prioritized patch list tied to finding IDs. - Rewrite the minimum necessary sections first, preserving useful behavior and avoiding unnecessary complexity. - Then provide a consolidated proposed prompt only if the changes are interdependent or the supplied acceptance criteria require a complete candidate. - Include explicit input requirements, instruction precedence, evidence rules, uncertainty handling, privacy limits, tool and action boundaries, escalation conditions, output schema, and final verification where relevant. - Mark all rewritten text Proposed and unverified. Explain material trade-offs such as safety versus task completion, strict formatting versus flexibility, and refusal sensitivity versus usefulness. 8. Define controlled verification and acceptance. - Map every acceptance criterion and Critical or High finding to one or more tests. - For each check, list the expected observation, actual observation if supplied, evidence reference, result, and unresolved gap. - Reconcile contradictory transcripts, partial passes, regressions, and environment differences instead of averaging them away. - Require regression testing for preserved capabilities as well as safety controls. - State what an authorized reviewer must run in the declared target environment and what evidence must be retained. ## Required deliverable Return the review with these sections: ### 1. Scope, System Boundary, and Evidence Status Include the intended behavior, downstream decisions, authority boundaries, supplied evidence inventory, unavailable evidence, assumptions, unknowns, and conflicts. ### 2. Executive Risk Decision State the leading weakness, highest-risk failure path, most important control, and one provisional disposition: Not ready, Ready for controlled testing, or Ready for human approval review. Never state Ready for release solely from static analysis. ### 3. Instruction and Requirement Traceability Matrix Use columns for requirement ID, source excerpt, interpretation, priority, conflict or ambiguity, affected output, and proposed correction. ### 4. Risk-Ranked Finding Register Use all finding fields defined above and separate Critical or High findings from Medium or Low improvements. ### 5. Misuse, Privacy, and Authority Analysis Cover abuse actors, protected data, unauthorized actions, escalation triggers, stop conditions, and required human approvals. ### 6. Red-Team Test Suite Provide executable test specifications with risk links, expected behavior, evidence requirements, and honest statuses. ### 7. Layered Guardrail Plan Separate prompt-level mitigations from application, access-control, monitoring, and operational controls. Include owners, trade-offs, residual risk, and verification. ### 8. Proposed Prompt Changes Provide the finding-linked patch list and any justified consolidated candidate. Clearly label them Proposed and not yet tested. ### 9. Verification and Acceptance Matrix Use columns for criterion or finding, test ID, expected observation, actual observation, evidence reference, result, owner, and unresolved action. ### 10. Human Handoff List blocking questions, sanitized artifacts needed, tests to run, approvals required, rollback or recovery preparation, and the next authorized decision owner. ## Final integrity check Before returning the deliverable, confirm that every conclusion is traceable to supplied material or labeled uncertainty; every Critical and High finding has a mitigation and test; prompt controls are not substituted for external enforcement; sensitive data is minimized; proposed changes are not described as applied; unrun tests are marked Not run; and the disposition does not claim approval, verification, deployment, or completion without corresponding evidence.
Variables to Replace
- Prompt under review
- Intended users and use context
- Task and decision impact
- Input examples and source materials
- Required output contract
- Model and tool environment
- Known incidents and baseline results
- Risk and data classification
- Policy and operating constraints
- Acceptance criteria
How to Use This Prompt
In Claude, replace every bracketed variable with the reusable prompt and its operating context. Provide the exact prompt text plus relevant evidence such as sanitized input examples, expected outputs, policies, incident transcripts, baseline test results, model settings, tool permissions, data classification, and acceptance criteria. Keep secrets and unnecessary personal data out of the materials, then run the prompt. Treat Claude's result as a static review and test plan unless actual target-environment results were supplied; have an authorized human run the proposed tests and approve any prompt change or release.
Example Use Case
A customer-support team gives Claude its reusable support prompt, sanitized conversation examples, escalation policy, privacy classification, allowed knowledge sources, tool permissions, known incidents, and release criteria. Claude identifies instruction-priority and privacy weaknesses, links each finding to evidence, specifies adversarial and regression tests, separates prompt guardrails from application controls, proposes a revised candidate, and marks all unexecuted tests Not run for human validation in the target environment.