Published version comparison

Evidence-Grounded Prompt Red-Team and Guardrail Builder for Claude

1.0.0 → 2.0.0

Source version 1.0.0

Published

Initial: Initial published snapshot.

Destination version 2.0.0

Published

Major: Replace the legacy Prompt System Red-Team and Guardrail Builder template with a domain-specific input, evidence, authority, safety, workflow, output, and verification contract.

Public field comparison

Title Changed

1.0.0
Prompt System Red-Team and Guardrail Builder
2.0.0
Evidence-Grounded Prompt Red-Team and Guardrail Builder for Claude

Summary Changed

1.0.0
Red-team a reusable prompt system, identify failure modes, unsafe outputs, ambiguity, missing constraints, and create guardrails, tests, and improvement rules.
2.0.0
Use Claude to statically inspect a reusable prompt, model task-specific abuse and failure scenarios, specify red-team tests, design guardrails, propose a traceable rewrite, and issue an evidence-qualified release recommendation without claiming unrun tests passed.

Share-purpose line Changed

1.0.0
Red-teaming reusable prompt systems to identify failure modes, unsafe outputs, ambiguity, missing constraints, and practical guardrail improvements.
2.0.0
Evaluate reusable prompts for ambiguity, instruction conflicts, unsafe behavior, privacy exposure, unsupported claims, weak escalation rules, and other task-specific failure modes before controlled testing or release.

Best use cases Changed

1.0.0
Prompt System Red-Teaming
Guardrail Design
Failure Mode Testing
Unsafe Output Review
Reusable Prompt QA
Prompt Evaluation Checklist
2.0.0
Pre-release static review of reusable prompt systems
Risk-based red-team test design for prompts
Prompt-injection and instruction-conflict analysis
Privacy, tool-permission, and authority-boundary review
Evidence-linked guardrail and prompt revision planning
Human-governed prompt release-readiness assessment

Variables Changed

1.0.0
Prompt to evaluate
Intended users
Intended task
Expected output
Tools or models used
Known failure cases
Sensitive risks
Business or user context
Constraints
Definition of done
2.0.0
Prompt under review
Intended users and use context
Task and decision impact
Input examples and source materials
Required output contract
Model and tool environment
Known incidents and baseline results
Risk and data classification
Policy and operating constraints
Acceptance criteria

How to Use Changed

1.0.0
Paste the reusable prompt you want to test, along with its intended users, task, expected output, known failure cases, risks, constraints, and definition of done. Use the output to strengthen the prompt before publishing, reusing, or deploying it in a real workflow.
2.0.0
In Claude, replace every bracketed variable with the reusable prompt and its operating context. Provide the exact prompt text plus relevant evidence such as sanitized input examples, expected outputs, policies, incident transcripts, baseline test results, model settings, tool permissions, data classification, and acceptance criteria. Keep secrets and unnecessary personal data out of the materials, then run the prompt. Treat Claude's result as a static review and test plan unless actual target-environment results were supplied; have an authorized human run the proposed tests and approve any prompt change or release.

Example use case Changed

1.0.0
A team has a reusable AI customer support prompt and wants to prevent unsafe replies, vague answers, privacy leaks, unsupported claims, and inconsistent escalation behavior before using it with real customer conversations.
2.0.0
A customer-support team gives Claude its reusable support prompt, sanitized conversation examples, escalation policy, privacy classification, allowed knowledge sources, tool permissions, known incidents, and release criteria. Claude identifies instruction-priority and privacy weaknesses, links each finding to evidence, specifies adversarial and regression tests, separates prompt guardrails from application controls, proposes a revised candidate, and marks all unexecuted tests Not run for human validation in the target environment.

Difficulty Unchanged

1.0.0
Expert
2.0.0
Expert

Tool Unchanged

1.0.0
Claude
2.0.0
Claude

Prompt type Unchanged

1.0.0
analysis
2.0.0
analysis

Tags Changed

1.0.0
prompt-engineering
prompt-red-team
guardrail-design
failure-modes
prompt-testing
unsafe-outputs
reusable-prompts
prompt-quality
evaluation
ai-safety
2.0.0
prompt-engineering
claude
prompt-red-team
guardrail-design
prompt-injection
failure-mode-analysis
ai-safety
prompt-evaluation
privacy
release readiness

SEO title Changed

1.0.0
Prompt System Red-Team and Guardrail Builder Prompt
2.0.0
Claude Prompt Red-Team and Guardrail Builder

SEO description Changed

1.0.0
Red-team reusable prompts, identify failure modes, unsafe outputs, ambiguity, missing constraints, and create guardrails and tests.
2.0.0
Use Claude to audit reusable prompts, model failure paths, design guardrails and tests, and produce an evidence-qualified release recommendation.

Prompt-body line comparison

Removed Added Unchanged context

You are an expert prompt engineer and AI system evaluator specializing in prompt red-teaming, guardrail design, failure-mode analysis, unsafe output detection, ambiguity review, prompt evaluation, and reusable AI workflow quality.

Your task is to evaluate a reusable prompt system and improve its safety, clarity, reliability, usefulness, and output quality before it is published, reused, or deployed in a real workflow.

Context:
Prompt to evaluate: [Prompt to evaluate]
Intended users: [Intended users]
Intended task: [Intended task]
Expected output: [Expected output]
Tools or models used: [Tools or models used]
Known failure cases: [Known failure cases]
Sensitive risks: [Sensitive risks]
Business or user context: [Business or user context]
Constraints: [Constraints]
Definition of done: [Definition of done]

Important constraints:

* Do not only praise the prompt.
* Do not assume the prompt is safe, complete, or clear.
* Look for ambiguity, missing context, weak instructions, unsafe assumptions, overbroad requests, and poor verification.
* Identify where the prompt may produce vague, misleading, low-quality, harmful, privacy-risky, or unsupported outputs.
* Do not add unnecessary complexity.
* Improve the prompt while keeping it practical and reusable.
* Include realistic red-team test cases.
* Include guardrails that are specific to the prompt’s intended task.
* Separate critical issues from minor improvements.
* If information is missing, state the assumption clearly before giving recommendations.

Task:

1. Summarize the prompt’s intended job.
   Explain:
Evaluate the supplied reusable prompt as a candidate AI system instruction. Produce a critical, evidence-grounded red-team review, a test specification, guardrail recommendations, and a proposed release candidate. Keep static analysis, supplied runtime evidence, and unexecuted test proposals distinct.

* What the prompt is trying to help the user do
* Who it is designed for
* What output it should produce
* What decisions or actions may depend on the output
* Where quality, safety, accuracy, or reliability matters most
## Evaluation package

2. Review ambiguity and unclear instructions.
   Identify:
Prompt under review: [Prompt under review]
Intended users and use context: [Intended users and use context]
Task and decision impact: [Task and decision impact]
Input examples and source materials: [Input examples and source materials]
Required output contract: [Required output contract]
Model and tool environment: [Model and tool environment]
Known incidents and baseline results: [Known incidents and baseline results]
Risk and data classification: [Risk and data classification]
Policy and operating constraints: [Policy and operating constraints]
Acceptance criteria: [Acceptance criteria]

* Vague wording
* Missing definitions
* Unclear success criteria
* Confusing role instructions
* Weak task boundaries
* Unclear output expectations
* Missing examples or constraints
* Instructions that could be interpreted in multiple ways
## Claude operating boundary

3. Identify missing context.
   List the missing information that would improve the prompt, such as:
Use Claude to inspect only the prompt, context, examples, policies, logs, and other evidence available in this conversation. Do not imply access to the target system, hidden system prompts, production conversations, external policies, deployment settings, model telemetry, or test harnesses unless their contents are explicitly supplied through an enabled tool or attachment.

* User goal
* Audience
* Input format
* Source material
* Risk level
* Tool or model constraints
* Legal, financial, medical, safety, privacy, or business constraints
* Output format
* Verification requirements
* Human review requirements
Do not execute the candidate prompt against users, production data, external systems, or target models. Do not publish, approve, deploy, edit, or replace the source prompt. Test cases and rewritten text are proposals for authorized human review. If runtime transcripts or test results are supplied, assess them as evidence; otherwise mark behavioral tests Not run rather than Passed, Failed, fixed, or verified.

4. Identify failure modes.
   Analyze how the prompt could fail, including:
Do not reproduce secrets, credentials, unnecessary personal data, or harmful operational details. Redact sensitive values while preserving the feature needed for analysis. Stop and request sanitized material if meaningful review would require exposing credentials, restricted data, or identifiable customer records.

* Generic output
* Hallucinated claims
* Unsupported recommendations
* Overconfident answers
* Missing edge cases
* Weak reasoning
* Poor formatting
* Unsafe instructions
* Privacy leaks
* Misuse by the user
* Inconsistent output across runs
* Failure to ask for missing information
## Input sufficiency and conflict handling

For each failure mode, explain the likely cause and the potential impact.
Treat the complete prompt under review, its intended task, intended users, expected output, and decision impact as blocking prerequisites. If any is absent or too ambiguous to identify the system boundary, ask focused clarification questions and provide only a clearly labeled preliminary review.

5. Identify misuse and sensitive-risk scenarios.
   Review whether the prompt could be misused or produce risky outputs in areas such as:
The model and tool environment, risk classification, governing constraints, and acceptance criteria are also blocking when the prompt affects legal, medical, financial, employment, security, safety, privacy, production, or other high-impact decisions. Do not issue a release recommendation until those details are resolved.

* Personal data
* Financial decisions
* Legal decisions
* Health or safety
* Security
* Customer communication
* Public claims
* Hiring or career decisions
* Business-critical operations
* Automated actions without human review
Examples, incident reports, baseline results, adversarial transcripts, and evaluation logs are useful but optional. If they are absent, continue with bounded static analysis and mark runtime behavior unknown. Preserve conflicting requirements in a conflict register; do not silently choose one. Label each material statement as one of the following where relevant: Supplied fact, Direct observation, Supplied execution evidence, Assumption, Hypothesis, Unknown, or Conflict.

6. Create red-team test cases.
   Create practical test cases that challenge the prompt.
## Review workflow

For each test case, include:
1. Establish the system boundary.
   - State the prompt's intended job, users, inputs, outputs, downstream decisions, execution environment, and foreseeable affected parties.
   - Identify which instructions belong to the system designer, operator, end user, retrieved content, or external source.
   - Map any requested tool calls, data access, automated actions, escalation paths, and human approval points.
   - Record assumptions and unresolved conflicts before drawing conclusions.

* Test name
* Test input
* What could go wrong
* Expected safe behavior
* What the prompt should refuse, question, qualify, or verify
* How to judge whether the prompt passed the test
2. Build an instruction and output map.
   - Trace each major objective, constraint, prohibition, exception, evidence requirement, refusal rule, escalation rule, and formatting requirement to the relevant source wording.
   - Identify contradictory priorities, undefined terms, missing precedence rules, unreachable requirements, excessive discretion, and requirements that cannot be verified from the requested output.
   - Check whether the output contract supports the decisions users are expected to make.

7. Recommend guardrails.
   Create specific guardrails for:
3. Create a risk-ranked finding register.
   Inspect for task-specific failure modes, including ambiguous scope, prompt injection, instruction-priority confusion, data exfiltration, excessive disclosure, fabricated facts or citations, unsupported recommendations, unsafe compliance, over-refusal, missing qualification, poor calibration, inconsistent escalation, unauthorized tool use, irreversible automation, output-schema failure, context-window loss, multilingual or encoding edge cases, and misuse outside the intended audience.

* Missing information
* Unsupported claims
* Sensitive topics
* Privacy and confidential data
* High-impact decisions
* Human review
* Output quality
* Evidence and citations, where relevant
* Formatting and structure
* Verification before final output
   For every finding, provide:
   - Finding ID and concise title
   - Affected source excerpt or requirement
   - Evidence classification
   - Trigger or precondition
   - Failure mechanism
   - Likely output or behavior
   - Impacted users, systems, or decisions
   - Severity: Critical, High, Medium, or Low
   - Likelihood: Likely, Plausible, Unlikely, or Unknown
   - Confidence and rationale
   - Proposed mitigation
   - Residual risk after the proposed mitigation
   - Verification needed

8. Rewrite weak prompt sections.
   Rewrite the parts of the prompt that need improvement.
   Reserve Critical for a plausible path to severe harm, major unauthorized disclosure, destructive action, or prohibited high-impact behavior. Do not inflate severity merely because a topic is sensitive.

Include:
4. Analyze misuse and authority boundaries.
   - Identify foreseeable misuse by end users, operators, embedded content, retrieved documents, and downstream automation.
   - Test whether untrusted content can override higher-priority instructions, solicit confidential context, broaden the task, or cause unauthorized actions.
   - Specify what the prompt may answer, must qualify, must refuse, must escalate, and must leave for human authorization.
   - Require explicit human approval before consequential communication, account changes, financial commitments, eligibility decisions, safety actions, publication, deployment, or other irreversible effects.

* Improved role instruction
* Improved task instruction
* Improved context placeholders
* Improved constraints
* Improved output format
* Improved verification section
* Improved final instruction
5. Design a risk-based red-team suite.
   Include normal cases, boundary cases, malformed inputs, missing-context cases, conflicting instructions, adversarial inputs, privacy attacks, unsupported-claim traps, output-format stress, escalation cases, and repeatability checks. Include domain-specific cases derived from the supplied task rather than relying only on generic jailbreak language.

Do not rewrite the entire prompt unless the whole prompt is weak. Focus on the sections that will create the highest improvement.
   For each test, specify:
   - Test ID and risk linkage
   - Objective
   - Preconditions and sanitized test input
   - Attack or stress technique
   - Expected safe behavior
   - Prohibited behavior
   - Required output evidence
   - Pass criteria
   - Actual observation, only when supplied execution evidence exists
   - Status: Not run, Pass, Fail, Blocked, or Inconclusive
   - Human reviewer or approval required

9. Create an output verification checklist.
   Build a checklist the user can apply after the AI produces an answer.
   Do not generate actionable harmful payloads when a benign structural placeholder can test the same control. A static prediction of likely behavior is not an actual observation.

The checklist should confirm:
6. Design layered guardrails.
   Recommend controls at the appropriate layer: prompt instruction, input validation, context isolation, data minimization, retrieval filtering, tool permissioning, output validation, confidence and citation rules, refusal behavior, escalation, rate or scope limits, logging, human review, and rollback or prompt-version recovery.

* The output follows the requested structure
* The output uses the provided context
* Assumptions are clearly labeled
* Missing information is identified
* Sensitive risks are handled carefully
* Claims are not invented
* Recommendations are practical
* Human review is included where needed
* The output is safe to use for the intended purpose
   For each guardrail, identify the risk addressed, control owner, enforcement layer, exact behavior, failure response, trade-off, residual risk, and verification method. Distinguish controls expressible in the prompt from controls that require application code, model settings, policy enforcement, access controls, monitoring, or operational procedure. Do not present prompt wording as sufficient protection against risks that require external enforcement.

10. Provide final prompt improvement recommendations.
    Summarize:
7. Propose revisions.
   - Provide a prioritized patch list tied to finding IDs.
   - Rewrite the minimum necessary sections first, preserving useful behavior and avoiding unnecessary complexity.
   - Then provide a consolidated proposed prompt only if the changes are interdependent or the supplied acceptance criteria require a complete candidate.
   - Include explicit input requirements, instruction precedence, evidence rules, uncertainty handling, privacy limits, tool and action boundaries, escalation conditions, output schema, and final verification where relevant.
   - Mark all rewritten text Proposed and unverified. Explain material trade-offs such as safety versus task completion, strict formatting versus flexibility, and refusal sensitivity versus usefulness.

* The most important weakness
* The highest-risk failure mode
* The most important guardrail to add
* The strongest rewrite recommendation
* Whether the prompt is ready to use, needs revision, or should not be used yet
8. Define controlled verification and acceptance.
   - Map every acceptance criterion and Critical or High finding to one or more tests.
   - For each check, list the expected observation, actual observation if supplied, evidence reference, result, and unresolved gap.
   - Reconcile contradictory transcripts, partial passes, regressions, and environment differences instead of averaging them away.
   - Require regression testing for preserved capabilities as well as safety controls.
   - State what an authorized reviewer must run in the declared target environment and what evidence must be retained.

Output format:
## Required deliverable

## Prompt Purpose Summary
Return the review with these sections:

## Ambiguity Review
### 1. Scope, System Boundary, and Evidence Status
Include the intended behavior, downstream decisions, authority boundaries, supplied evidence inventory, unavailable evidence, assumptions, unknowns, and conflicts.

## Missing Context
### 2. Executive Risk Decision
State the leading weakness, highest-risk failure path, most important control, and one provisional disposition: Not ready, Ready for controlled testing, or Ready for human approval review. Never state Ready for release solely from static analysis.

## Failure Modes
### 3. Instruction and Requirement Traceability Matrix
Use columns for requirement ID, source excerpt, interpretation, priority, conflict or ambiguity, affected output, and proposed correction.

## Misuse and Sensitive-Risk Scenarios
### 4. Risk-Ranked Finding Register
Use all finding fields defined above and separate Critical or High findings from Medium or Low improvements.

## Red-Team Test Cases
### 5. Misuse, Privacy, and Authority Analysis
Cover abuse actors, protected data, unauthorized actions, escalation triggers, stop conditions, and required human approvals.

## Guardrail Recommendations
### 6. Red-Team Test Suite
Provide executable test specifications with risk links, expected behavior, evidence requirements, and honest statuses.

## Rewritten Prompt Sections
### 7. Layered Guardrail Plan
Separate prompt-level mitigations from application, access-control, monitoring, and operational controls. Include owners, trade-offs, residual risk, and verification.

## Output Verification Checklist
### 8. Proposed Prompt Changes
Provide the finding-linked patch list and any justified consolidated candidate. Clearly label them Proposed and not yet tested.

## Final Recommendations
### 9. Verification and Acceptance Matrix
Use columns for criterion or finding, test ID, expected observation, actual observation, evidence reference, result, owner, and unresolved action.

Verification:
Before finalizing, check that:
### 10. Human Handoff
List blocking questions, sanitized artifacts needed, tests to run, approvals required, rollback or recovery preparation, and the next authorized decision owner.

* The review is critical, not only complimentary.
* Failure modes are specific to the prompt being evaluated.
* Red-team test cases are realistic.
* Guardrails are practical and not generic.
* Rewritten sections improve clarity and safety.
* Missing information is clearly identified.
* High-risk outputs include human review.
* The final recommendation clearly states whether the prompt is ready to use, needs revision, or should not be used yet.
## Final integrity check

Begin the prompt system red-team and guardrail review now.
Before returning the deliverable, confirm that every conclusion is traceable to supplied material or labeled uncertainty; every Critical and High finding has a mitigation and test; prompt controls are not substituted for external enforcement; sensitive data is minimized; proposed changes are not described as applied; unrun tests are marked Not run; and the disposition does not claim approval, verification, deployment, or completion without corresponding evidence.