# Prompt Evaluation Harness Prompt

Public URL: https://amo.ng/prompts/prompt-evaluation-harness-prompt

Summary: Design prompt tests with sample inputs, expected behaviors, failure cases, scoring criteria, and regression checks.

Use this for: Testing, scoring, comparing, and improving reusable AI prompts before using them in production.

Category: Prompt Engineering
Tool: ChatGPT
Difficulty: Expert
Prompt type: evaluation

## Best Use Cases

1. Prompt Evaluation
2. Scoring Rubric Design
3. Evaluation Harness Design
4. Output Quality Review
5. Regression Testing
6. Prompt Iteration Planning

## Prompt Body

Act as a senior Prompt Engineering specialist using ChatGPT. Your task is: [Goal or task].

Context:
- Current situation: [Current context]
- Constraints: [Constraints]
- Available materials: [Files, data, examples, URLs, logs, notes]
- Success criteria: [Definition of done]

Workflow:
1. Restate the objective in operational terms and identify any missing information that would block a reliable answer.
2. Make reasonable assumptions only when they are low risk, and label them clearly.
3. Produce the main deliverable for "Prompt Evaluation Harness Prompt" with enough detail that a skilled operator can execute it immediately.
4. Include edge cases, failure modes, dependencies, and tradeoffs that a junior prompt would usually miss.
5. Add a verification checklist with concrete tests, review questions, metrics, or acceptance criteria.
6. End with the smallest safe next action.

Output format:
- Executive summary
- Detailed plan or implementation
- Risks and mitigations
- Verification checklist
- Next action

Do not give generic advice. Optimize for a production-quality evaluation outcome.

## Variables to Replace

1. Goal or task
2. Current context
3. Constraints
4. Files, data, or examples
5. Definition of done

## How to Use

Replace every bracketed placeholder before running. Give the model enough context to inspect assumptions, ask only blocking questions, and produce a concrete deliverable. For code prompts, include relevant files, errors, logs, and test commands.

## Example Use Case

Use this when you need a production-ready evaluation result in Prompt Engineering, not a generic brainstorm. The expected output should include findings, implementation steps, risks, and verification checks.

## Tags

1. chatgpt
2. evals
3. prompt-testing
4. prompt-engineering

## Dates

Published: 2026-06-04
Updated: 2026-06-11
