# Production AI Evaluation Drift Detection Plan

Amo ID: AMO-P-000275
Version: 1.0.0
Public URL: https://amo.ng/prompts/production-ai-evaluation-drift-detection-plan

Summary: Design a control plan that detects when a once-approved AI evaluation is no longer reliable for production decisions.

Use this for: Use this to define drift controls that keep production AI evaluations decision-reliable as traffic, labels, graders, prompts, models, retrieval, and outcomes change.

Category: Prompt Engineering
Tool: General AI
Difficulty: Expert
Prompt type: evaluation

## Best Use Cases

1. Production Evaluation Monitoring
2. AI Evaluation Drift Review
3. Prompt And Model Change Oversight
4. Retrieval Quality Surveillance
5. Outcome-Based Revalidation Planning
6. Evaluation Governance Evidence Pack

## Prompt Body

Design a production AI evaluation-drift detection plan for an already-approved evaluation. The goal is to detect when production conditions make the evaluation no longer decision-reliable.

Do not design a generic evaluation harness. Do not perform one-time model upgrade regression testing. Focus on ongoing production drift controls for an evaluation that already exists.

Context to provide:
- System or product under evaluation: [System or product under evaluation]
- Approved evaluation decision use: [Approved evaluation decision use]
- Current evaluation artifact summary: [Current evaluation artifact summary]
- Production telemetry and outcome evidence available: [Production telemetry and outcome evidence available]
- Known recent or planned changes: [Known recent or planned changes]
- Accountable owners and operating constraints: [Accountable owners and operating constraints]
- Risk tolerance or escalation policy: [Risk tolerance or escalation policy]

Evidence discipline:
- Separate observed evidence from inference.
- Do not claim logs, datasets, tests, graders, prompts, retrieval systems, production traffic, or approvals were inspected unless they are included in the provided context.
- Flag missing information that prevents firm threshold-setting or assignment of ownership.
- Preserve uncertainty where evidence is incomplete.
- If assumptions are necessary, label them as assumptions and explain how the evaluation owner should verify them.

Produce the following deliverable:

1. Evaluation reliability boundary
Define what decision the evaluation is approved to support, what production population it is intended to represent, what conditions must remain comparable, and what conditions would make the evaluation no longer decision-reliable.

2. Drift taxonomy
Create a task-specific taxonomy covering at minimum:
- Traffic or user-intent drift
- Input-format or language drift
- Label-policy or ground-truth drift
- Human reviewer or labeling-team drift
- LLM-as-grader, rubric, or judge-prompt drift
- Application prompt or system-instruction drift
- Model, provider, parameter, or routing drift
- Retrieval corpus, embedding, ranking, or freshness drift
- Tool, API, data dependency, or integration drift
- Outcome, complaint, incident, conversion, safety, or business-metric drift
For each drift type, state observable signals, likely false positives, decision impact, accountable owner, and evidence needed.

3. Sentinel and sampling plan
Specify sentinel checks and production sampling methods that can detect each meaningful drift type. Include:
- Always-on metrics versus periodic review samples
- Stratified slices that must be monitored
- Minimum viable sample size or sample logic when exact sizes cannot be justified
- Triggered sampling after incidents, launches, prompt changes, retrieval updates, model routing changes, policy changes, or unusual outcome shifts
- Treatment of low-volume but high-risk slices
- Owner responsible for sample collection, review, and documentation

4. Comparability ledger
Design a ledger that records whether the current production environment remains comparable to the approved evaluation baseline. Include ledger fields for:
- Evaluation version and approved decision use
- Baseline dataset or traffic window
- Production traffic window reviewed
- Prompt, model, retrieval, grader, rubric, label policy, and tool versions
- Known changes since approval
- Evidence source for each comparison
- Comparability status: comparable, degraded, not comparable, or unknown
- Owner attestation required from evaluation owner, product owner, data owner, security reviewer, or other named accountable role as appropriate
- Revalidation requirement and due date

5. Detection thresholds
Propose initial thresholds using the evidence provided. Where evidence is insufficient, provide threshold-setting rules instead of invented numbers. Cover:
- Statistical or distributional thresholds
- Operational thresholds
- Quality and safety thresholds
- Business or outcome thresholds
- Grader agreement or calibration thresholds
- Retrieval freshness or coverage thresholds
- Incident-based hard stops
For each threshold, state the metric, comparison baseline, trigger level, rationale, owner, action required, and expected review cadence.

6. Investigation triggers
Define when the team must investigate before continuing to rely on the evaluation. Include triggers for traffic anomalies, label disagreement, grader instability, prompt or model changes, retrieval degradation, unexplained outcome shifts, safety incidents, and stakeholder challenges to evaluation validity.

7. Revalidation triggers
Define when the approved evaluation must be refreshed, rerun, recalibrated, or retired. Include criteria for partial revalidation, full revalidation, temporary suspension of evaluation-based decisions, and formal owner signoff.

8. Runbook
Write a practical runbook with:
- Daily, weekly, monthly, and release-event checks where appropriate
- Required inputs and evidence sources
- Step-by-step investigation path
- Decision states: continue relying, rely with caveat, pause reliance, revalidate, or retire
- Communication path to evaluation owner, product owner, data owner, security reviewer, release owner, and incident owner where relevant
- Documentation artifacts to retain

9. Completion checks
End with observable completion criteria. The plan is complete only if it identifies monitored drift types, assigns accountable owners, defines sentinel and sampling mechanisms, records comparability evidence, states thresholds or threshold-setting rules, specifies investigation and revalidation triggers, and explains what evidence is still missing.

## Variables to Replace

1. System or product under evaluation
2. Approved evaluation decision use
3. Current evaluation artifact summary
4. Production telemetry and outcome evidence available
5. Known recent or planned changes
6. Accountable owners and operating constraints
7. Risk tolerance or escalation policy

## How to Use

Use this in a general AI assistant. Paste or upload the approved evaluation summary, baseline dataset notes, production telemetry definitions, outcome metrics, grader or rubric details, prompt/model/retrieval version history, and any relevant incident or change records. Replace every bracketed placeholder, then run the prompt. Afterward, have the evaluation owner, product owner, data owner, and any required security or release owner verify thresholds, owners, evidence sources, and revalidation triggers before adopting the plan.

## Example Use Case

A product team relies on an approved AI support-quality evaluation to decide whether automated responses can remain enabled. Traffic mix, retrieval content, judge prompts, and complaint rates have changed since approval. The team uses this prompt to create a drift taxonomy, sentinel checks, sampling plan, comparability ledger, thresholds, and revalidation triggers so the evaluation owner can decide whether the evaluation is still reliable.

## Tags

1. prompt-engineering
2. llm-evaluation
3. verification
4. evaluation-design
5. monitoring
6. quality-metrics

## Dates

Published: 2026-08-19
Updated: 2026-08-19
