# Review Production AI Evaluation Reliability

Workflow ID: AMO-W-000013
Workflow URL: https://amo.ng/workflows/production-ai-evaluation-reliability-review

## Outcome

A decision-ready evaluation reliability record containing dataset admissibility, judge permitted-use boundaries, production drift controls, conditional retrieval experiment evidence, unresolved limitations, and revalidation or release conditions.

## Before you begin

- Evaluation decision or release claim and accountable owners
- Dataset records, provenance, splits, labels, coverage targets, and contamination references
- Judge prompts, rubrics, calibration cases, repeated scores, and arbitration evidence
- Production samples, metrics, model or system changes, monitoring history, and comparability records
- Retrieval configuration, benchmark questions, corpus samples, and answer standards when retrieval is material

## Step 1 — Establish evaluation-dataset admissibility

**Prompt**

Evaluation Dataset Coverage and Contamination Audit

**Instructions**

Audit provenance, coverage, duplication, contamination, label quality, and decision risk against the exact release or operating claim the dataset is expected to support.

**Input for this step**

Supply the intended claim, dataset or representative records, schema, sources, collection dates, transformations, labels, adjudication evidence, target slices, and known training or benchmark overlap.

**Carry forward**

Carry the provenance ledger, coverage matrix, contamination register, label-quality findings, admissibility decision, remediation needs, and release-claim limits to judge review.

**Review note**

The evaluation owner, dataset owner, data or privacy owner, and release owner confirm whether the dataset is admissible for the stated decision.

**Prompt ID**

AMO-P-000273

**Prompt URL**

https://amo.ng/prompts/evaluation-dataset-coverage-and-contamination-audit

**Prompt content**

Audit the evaluation dataset for decision fitness before it is used to support release claims. Treat the dataset as inadmissible until the evidence below supports a narrower conclusion.

Context and inputs to provide:
- Dataset and intended release claim: [Dataset description and intended release claim]
- Dataset records, schema, samples, or files available for review: [Dataset records or sample with schema]
- Source lineage, collection method, licenses, collection dates, and transformations: [Provenance sources and collection dates]
- Label definitions, labeling instructions, adjudication process, annotator metadata, and quality checks: [Labeling guidelines and annotator metadata]
- Target tasks, risk categories, domains, locales, languages, user populations, and expected operating conditions: [Target tasks risks and user populations]
- Known model training corpora, exclusion lists, benchmark sources, public datasets, or other contamination references: [Known training data or exclusion sources]

Evidence discipline:
- Separate observed facts from inference. Mark each material claim as Observed, Inferred, Not provided, or Not assessable from supplied evidence.
- Do not claim that a file, source, test, command, repository, dataset split, or system was inspected unless it is present in the supplied material.
- Do not invent counts, percentages, inter-annotator agreement, contamination rates, licenses, or collection dates. If a metric cannot be calculated from the supplied evidence, state what is missing and whether a proxy assessment is possible.
- Preserve uncertainty. Use confidence levels only when tied to available evidence.
- Focus on this dataset audit, not on building an evaluation harness or comparing model versions.

Authority and action boundaries:
- This audit may assess evidence and recommend a disposition; it does not authorize release claims or modify the dataset.
- Do not quarantine, remove, relabel, deduplicate, rebalance, augment, disclose, or release dataset records, and do not approve or revise a release claim, unless the responsible owner has separately authorized that action.
- Dataset changes require authorization from the dataset owner and evaluation owner. Rights, consent, privacy, or license decisions require the data owner and legal or privacy reviewer. Final use of the dataset for a release claim remains with the release owner.

Audit procedure:
1. Define the admissibility question.
   - Restate the release claim the dataset is expected to support.
   - Identify the accountable evaluation owner, dataset owner, data owner, and release owner if named in the supplied material; otherwise list them as missing accountability assignments.
   - Define what the dataset must demonstrate to be fit for that release claim.

2. Build a dataset provenance ledger.
   For each source, split, subset, or major record group, capture:
   - Source name or origin
   - Collection date or time window
   - Collection method
   - Rights, license, consent, or usage restriction evidence
   - Transformation, filtering, augmentation, or generation steps
   - Label source and labeling workflow
   - Known exclusions or quarantine rules
   - Traceability gaps
   - Confidence in provenance

3. Build a coverage matrix.
   Map the dataset against the supplied target tasks, risk categories, domains, user populations, languages/locales, difficulty bands, failure modes, and operating conditions. Include:
   - Available counts or proportions where directly calculable
   - Coverage status: Adequate, Thin, Missing, Overrepresented, Not assessable
   - Evidence basis
   - Release-claim consequence of each gap
   - Minimum additional evidence or data needed to close the gap

4. Audit leakage, duplication, and contamination risk.
   Create a contamination register covering:
   - Exact duplicate records within the dataset
   - Near duplicates or paraphrase clusters, if detectable from supplied records
   - Train/eval split leakage, if split information is supplied
   - Prompt-answer leakage, rubric leakage, or label leakage
   - Overlap with known training data, public benchmarks, synthetic data sources, vendor examples, documentation, or previous evaluation sets
   - Temporal leakage relative to the intended release claim
   - Source reuse that could inflate performance claims
   For each item, state the evidence, detection method available from supplied material, severity, uncertainty, owner, and remediation.

5. Review label quality and decision reliability.
   Assess:
   - Label definition clarity and mutual exclusivity
   - Alignment between labels, rubric, and release claim
   - Ambiguous or underspecified cases
   - Annotator qualification evidence
   - Adjudication and dispute-resolution process
   - Inter-annotator agreement or audit sample results, only if provided
   - Gold-standard or expert-review evidence, only if provided
   - Label drift across sources, time periods, or task categories
   - Examples where the label appears inconsistent with the provided guideline

6. Identify decision risks.
   Distinguish risks that affect:
   - Statistical validity
   - External validity and representativeness
   - Safety or policy risk coverage
   - Bias across user populations or locales
   - Claim wording and overgeneralization
   - Reproducibility and auditability
   - Legal, license, privacy, or data-rights admissibility

7. Recommend remediation.
   Provide targeted actions only. Avoid broad rebuilds unless the evidence shows the dataset cannot be repaired. For each action include:
   - Remediation action
   - Specific defect addressed
   - Priority
   - Responsible owner: dataset owner, evaluation owner, data owner, security reviewer, policy reviewer, legal reviewer, or release owner as applicable
   - Acceptance check
   - Whether the dataset must be quarantined, relabeled, deduplicated, rebalanced, restricted to a narrower claim, or rejected

8. Make an admissibility decision.
   Choose one:
   - Admissible for the stated release claim
   - Conditionally admissible after named remediation
   - Admissible only for a narrower claim
   - Not admissible for release claims
   Explain the decision in terms of evidence sufficiency, unresolved uncertainty, contamination risk, coverage gaps, label quality, and provenance.

Required output format:

# Evaluation Dataset Coverage and Contamination Audit

## 1. Admissibility Question
- Intended release claim:
- Dataset use in the release decision:
- Accountable owners named in evidence:
- Missing accountability assignments:
- Fitness threshold for this audit:

## 2. Evidence Inventory
| Evidence item | Provided | Used for | Limitations | Missing information |
|---|---:|---|---|---|

## 3. Dataset Provenance Ledger
| Dataset segment/source | Origin | Collection window | Collection method | Rights/consent evidence | Transformations | Label source | Restrictions | Traceability gaps | Confidence |
|---|---|---|---|---|---|---|---|---|---|

## 4. Coverage Matrix
| Task/risk/user-population dimension | Expected coverage | Observed coverage | Status | Evidence basis | Release-claim consequence | Data or evidence needed |
|---|---|---|---|---|---|---|

## 5. Duplication, Leakage, and Contamination Register
| Issue | Type | Evidence observed | Detection possible from supplied material | Severity | Uncertainty | Owner | Remediation |
|---|---|---|---|---|---|---|---|

## 6. Label-Quality Findings
| Finding | Evidence status | Affected records or segment | Decision impact | Confidence | Remediation |
|---|---|---|---|---|---|

## 7. Decision Risks
List the material risks that remain after the audit. For each, state whether it affects statistical validity, external validity, safety coverage, fairness, reproducibility, rights/privacy, or claim wording.

## 8. Remediation Plan
| Priority | Action | Defect addressed | Responsible owner | Acceptance check | Release impact |
|---|---|---|---|---|---|

## 9. Admissibility Decision
- Decision:
- Claim supported, if any:
- Claim not supported:
- Required conditions before use:
- Residual uncertainty:
- Final verification checks for the release owner:

Completion checks before finalizing:
- The provenance ledger, coverage matrix, contamination register, label-quality findings, remediation plan, and admissibility decision are all present.
- Every major conclusion cites supplied evidence or is explicitly marked as inference or not assessable.
- No unavailable inspection, test, source comparison, or metric is claimed.
- The release owner has enough information to decide whether the dataset may support the stated claim, must be narrowed, or must be rejected.


## Step 2 — Calibrate the automated judge and permitted uses

**Prompt**

Judge-Model Reliability and Calibration Review

**Instructions**

Evaluate the judge contract, blinded calibration evidence, disagreement patterns, bias and instability, attack susceptibility, arbitration needs, and the decisions the judge may or may not support.

**Input for this step**

Use the dataset disposition plus judge instructions, rubric, calibration cases, reference labels, repeated runs, comparison scores, disagreements, model settings, and intended decision threshold.

**Carry forward**

Carry judge reliability findings, permitted and prohibited uses, uncertainty, arbitration rules, recalibration triggers, and evidence gaps into production drift design.

**Review note**

The evaluation owner and domain reviewer approve the judge’s permitted-use boundary; the release owner does not treat an aggregate score as sufficient evidence by itself.

**Prompt ID**

AMO-P-000274

**Prompt URL**

https://amo.ng/prompts/judge-model-reliability-and-calibration-review

**Prompt content**

Determine whether the proposed automated judge is sufficiently reliable and calibrated for the specified evaluation decision. Focus only on judge fitness for the stated scoring or classification decision; do not design a full evaluation platform or generic prompt evaluation harness.

Context to provide:
- Evaluation decision: [Evaluation decision]
- Judge prompt or rubric: [Judge prompt or rubric]
- Candidate outputs and source items: [Candidate outputs and source items]
- Human reference labels or adjudications: [Human reference labels or adjudications]
- Risk profile and protected attributes: [Risk profile and protected attributes]
- Acceptance thresholds: [Acceptance thresholds]

Evidence discipline:
- Separate observed evidence from inference.
- State when evidence is missing, weak, imbalanced, or not blinded.
- Do not claim that tests, files, systems, reviewers, or production behavior were inspected unless the provided material supports that claim.
- Preserve uncertainty where the sample is too small, labels are disputed, or reference judgments are not independent.
- Treat human reference labels as evidence to be assessed, not as automatically correct.

Deliverable required:

1. Evaluation decision and judge contract
Define the exact decision the judge is being asked to make:
- Decision type: classification, ordinal rating, pairwise preference, threshold pass/fail, ranking, or other.
- Intended users and accountable owner, such as evaluation owner, product owner, domain reviewer, security reviewer, or data owner.
- Inputs the judge is allowed to use.
- Inputs or knowledge the judge must not use.
- Score scale, labels, thresholds, and tie-breaking rules.
- What counts as a correct judgment versus an acceptable judgment.
- Known boundary cases where the judge’s authority should stop.
- Consequences of a false positive, false negative, over-score, and under-score.

2. Blinded calibration design
Propose a calibration design suitable for the provided decision and evidence:
- How examples should be blinded, randomized, deduplicated, and stratified.
- Minimum case coverage needed across easy cases, close calls, failures, adversarial examples, protected or sensitive attributes, and domain-specific edge cases.
- How many independent human adjudications are needed and when domain reviewer arbitration is required.
- Which metrics are appropriate and why, such as exact agreement, weighted agreement, Cohen’s kappa, Krippendorff’s alpha, rank correlation, threshold confusion matrix, false positive and false negative rates, calibration by score bucket, and self-consistency across repeated runs.
- How to avoid leakage from model identity, author identity, expected answer wording, ordering effects, or rubric hints.
- What must be held out for future regression checks.

3. Current evidence assessment
Using only the provided evidence, assess whether calibration can be judged now:
- Evidence available.
- Evidence missing.
- Sample quality concerns.
- Label quality concerns.
- Whether the provided material is enough to support a permitted-use decision.

4. Disagreement analysis
Analyze judge disagreement against reference labels or adjudications:
- Where the judge agrees reliably.
- Where disagreement clusters by score band, topic, task type, output length, language, ambiguity, or source quality.
- Whether disagreements are random, systematic, rubric-driven, or caused by unclear source material.
- Whether the judge is too lenient, too strict, overconfident, inconsistent near thresholds, or sensitive to irrelevant style features.
- Distinguish clear judge errors from cases where the human reference may be ambiguous or under-specified.

5. Bias, instability, and attack findings
Evaluate fitness risks that could invalidate the judge for the specified decision:
- Bias or disparate error patterns related to [Risk profile and protected attributes].
- Sensitivity to superficial wording, formatting, verbosity, fluency, dialect, language variety, or model identity.
- Instability across repeated judgments, order changes, paraphrases, or equivalent source presentations.
- Prompt injection exposure, including whether candidate text can influence the judge’s rubric, authority, scoring scale, or refusal behavior.
- Boundary-case behavior, including ambiguous answers, partially correct answers, missing citations, conflicting sources, unsafe but persuasive content, and cases near the pass/fail threshold.

6. Permitted-use and arbitration gate
Make a clear, bounded decision using [Acceptance thresholds]:
- Permitted use: where the judge may be used without routine human review.
- Conditional use: where the judge may assist but must be sampled, audited, or reviewed by an accountable owner.
- Prohibited use: where the judge should not be used for this decision.
- Arbitration triggers: exact conditions that require domain reviewer, security reviewer, data owner, or product owner review.
- Monitoring requirements: what should be logged, sampled, and periodically recalibrated.
- Regression triggers: what changes to the judge prompt, model, rubric, data distribution, or product policy require renewed calibration.

7. Completion check
End with a concise readiness statement:
- Fit for use, conditionally fit, or not fit for the specified evaluation decision.
- Main evidence supporting that conclusion.
- Main unresolved risks.
- Minimum additional evidence needed before expanding use.
- Accountable owner who should accept or reject the permitted-use gate.

Use precise professional language. Avoid generic AI governance commentary. Do not recommend broad rewrites of the judge unless a specific reliability failure requires a targeted change.


## Step 3 — Design production evaluation-drift controls

**Prompt**

Production AI Evaluation Drift Detection Plan

**Instructions**

Define drift taxonomy, sentinels, sampling, comparability records, thresholds, investigation triggers, and revalidation rules that detect when previously accepted evaluation evidence is no longer reliable.

**Input for this step**

Provide the dataset and judge dispositions, production traffic and sample evidence, model or prompt versions, system changes, historical metrics, known incidents, decision cadence, and monitoring constraints.

**Carry forward**

Carry the reliability boundary, sentinel plan, drift thresholds, runbook, revalidation conditions, owner assignments, and unresolved observability gaps into the conditional retrieval step or final decision.

**Review note**

The model or product owner, evaluation owner, and release owner approve monitoring coverage and the conditions that suspend automated release evidence.

**Prompt ID**

AMO-P-000275

**Prompt URL**

https://amo.ng/prompts/production-ai-evaluation-drift-detection-plan

**Prompt content**

Design a production AI evaluation-drift detection plan for an already-approved evaluation. The goal is to detect when production conditions make the evaluation no longer decision-reliable.

Do not design a generic evaluation harness. Do not perform one-time model upgrade regression testing. Focus on ongoing production drift controls for an evaluation that already exists.

Context to provide:
- System or product under evaluation: [System or product under evaluation]
- Approved evaluation decision use: [Approved evaluation decision use]
- Current evaluation artifact summary: [Current evaluation artifact summary]
- Production telemetry and outcome evidence available: [Production telemetry and outcome evidence available]
- Known recent or planned changes: [Known recent or planned changes]
- Accountable owners and operating constraints: [Accountable owners and operating constraints]
- Risk tolerance or escalation policy: [Risk tolerance or escalation policy]

Evidence discipline:
- Separate observed evidence from inference.
- Do not claim logs, datasets, tests, graders, prompts, retrieval systems, production traffic, or approvals were inspected unless they are included in the provided context.
- Flag missing information that prevents firm threshold-setting or assignment of ownership.
- Preserve uncertainty where evidence is incomplete.
- If assumptions are necessary, label them as assumptions and explain how the evaluation owner should verify them.

Produce the following deliverable:

1. Evaluation reliability boundary
Define what decision the evaluation is approved to support, what production population it is intended to represent, what conditions must remain comparable, and what conditions would make the evaluation no longer decision-reliable.

2. Drift taxonomy
Create a task-specific taxonomy covering at minimum:
- Traffic or user-intent drift
- Input-format or language drift
- Label-policy or ground-truth drift
- Human reviewer or labeling-team drift
- LLM-as-grader, rubric, or judge-prompt drift
- Application prompt or system-instruction drift
- Model, provider, parameter, or routing drift
- Retrieval corpus, embedding, ranking, or freshness drift
- Tool, API, data dependency, or integration drift
- Outcome, complaint, incident, conversion, safety, or business-metric drift
For each drift type, state observable signals, likely false positives, decision impact, accountable owner, and evidence needed.

3. Sentinel and sampling plan
Specify sentinel checks and production sampling methods that can detect each meaningful drift type. Include:
- Always-on metrics versus periodic review samples
- Stratified slices that must be monitored
- Minimum viable sample size or sample logic when exact sizes cannot be justified
- Triggered sampling after incidents, launches, prompt changes, retrieval updates, model routing changes, policy changes, or unusual outcome shifts
- Treatment of low-volume but high-risk slices
- Owner responsible for sample collection, review, and documentation

4. Comparability ledger
Design a ledger that records whether the current production environment remains comparable to the approved evaluation baseline. Include ledger fields for:
- Evaluation version and approved decision use
- Baseline dataset or traffic window
- Production traffic window reviewed
- Prompt, model, retrieval, grader, rubric, label policy, and tool versions
- Known changes since approval
- Evidence source for each comparison
- Comparability status: comparable, degraded, not comparable, or unknown
- Owner attestation required from evaluation owner, product owner, data owner, security reviewer, or other named accountable role as appropriate
- Revalidation requirement and due date

5. Detection thresholds
Propose initial thresholds using the evidence provided. Where evidence is insufficient, provide threshold-setting rules instead of invented numbers. Cover:
- Statistical or distributional thresholds
- Operational thresholds
- Quality and safety thresholds
- Business or outcome thresholds
- Grader agreement or calibration thresholds
- Retrieval freshness or coverage thresholds
- Incident-based hard stops
For each threshold, state the metric, comparison baseline, trigger level, rationale, owner, action required, and expected review cadence.

6. Investigation triggers
Define when the team must investigate before continuing to rely on the evaluation. Include triggers for traffic anomalies, label disagreement, grader instability, prompt or model changes, retrieval degradation, unexplained outcome shifts, safety incidents, and stakeholder challenges to evaluation validity.

7. Revalidation triggers
Define when the approved evaluation must be refreshed, rerun, recalibrated, or retired. Include criteria for partial revalidation, full revalidation, temporary suspension of evaluation-based decisions, and formal owner signoff.

8. Runbook
Write a practical runbook with:
- Daily, weekly, monthly, and release-event checks where appropriate
- Required inputs and evidence sources
- Step-by-step investigation path
- Decision states: continue relying, rely with caveat, pause reliance, revalidate, or retire
- Communication path to evaluation owner, product owner, data owner, security reviewer, release owner, and incident owner where relevant
- Documentation artifacts to retain

9. Completion checks
End with observable completion criteria. The plan is complete only if it identifies monitored drift types, assigns accountable owners, defines sentinel and sampling mechanisms, records comparability evidence, states thresholds or threshold-setting rules, specifies investigation and revalidation triggers, and explains what evidence is still missing.


## Step 4 — Design a retrieval experiment when configuration is material

**Prompt**

Retrieval Chunking and Metadata Experiment Design

**Instructions**

Use this step only when chunking or metadata configuration could materially affect evaluation reliability. Otherwise mark it Not applicable with the supporting system boundary. When applicable, define a controlled experiment across representative query slices without redesigning the entire RAG system.

**Input for this step**

Supply corpus samples, current retrieval configuration, candidate chunking and metadata options, query logs or benchmarks, answer standards, and the evaluation limitations from prior steps.

**Carry forward**

Produce the final evaluation reliability decision: admissible evidence, restricted uses, required experiment or monitoring work, acceptance thresholds, revalidation triggers, owners, and release or operating limitations.

**Review note**

The retrieval owner, evaluation owner, data owner, and release owner approve experiment scope and decide whether current evidence supports release, restricted use, revalidation, or hold.

**Prompt ID**

AMO-P-000284

**Prompt URL**

https://amo.ng/prompts/retrieval-chunking-and-metadata-experiment-design

**Prompt content**

Design a controlled retrieval experiment to isolate how chunking choices and metadata choices affect retrieval and answer quality. Keep the scope limited to corpus configuration decisions: do not redesign the full RAG architecture, diagnose the whole pipeline, change model selection, or introduce unrelated retrieval changes unless they are explicitly controlled constants.

Context to use:
- Corpus description and sample documents: [Corpus description and sample documents]
- Current retrieval setup: [Current retrieval setup]
- Candidate chunking options: [Candidate chunking options]
- Candidate metadata fields: [Candidate metadata fields]
- Query logs or benchmark questions: [Query logs or benchmark questions]
- Answer evaluation standard: [Answer evaluation standard]
- Decision owners and constraints: [Decision owners and constraints]

Evidence discipline:
- Separate observations supported by the provided inputs from inferences and assumptions.
- Do not claim that experiments, tests, retrieval runs, indexing jobs, or evaluations were completed unless results are provided.
- If required information is missing, state the missing evidence and design the smallest reasonable way to obtain it.
- Preserve uncertainty where the available evidence does not justify a firm conclusion.
- Before defining experiment arms, treat the baseline retrieval configuration, candidate changes, representative corpus and query evidence, and answer-evaluation standard as blocking when absent or materially conflicting. Request all blocking items in one consolidated clarification and stop the affected design decision until they are supplied. Record other gaps as Unknown with their effect on sizing, confounding, metrics, and acceptance criteria.
- Treat the retrieval owner as accountable for experiment execution, the data owner as accountable for corpus and metadata validity, and the product owner or domain owner as accountable for acceptance criteria tied to user impact.

Produce the following deliverable:

1. Experiment objective and decision boundary
- State the configuration decision this experiment will support.
- Define what is in scope: chunking strategy, overlap, boundary rules, parent-child or hierarchical chunking if relevant, metadata extraction, metadata normalization, and metadata use in filtering or ranking.
- Define what is out of scope and should be held constant: embedding model, reranker, generation prompt, answer model, UI behavior, permissions, and production rollout unless the inputs require otherwise.
- Identify the baseline configuration and the candidate configurations to compare.

2. Controlled variables and constants
Create a table with:
- Variable under test
- Candidate levels
- Why it matters
- Expected retrieval effect
- Required implementation evidence
- What must remain constant

Include at minimum:
- Chunk size or token range
- Chunk overlap
- Chunk boundary rule
- Structural handling of headings, tables, lists, or sections
- Metadata fields to attach
- Metadata normalization rules
- Metadata use pattern: display only, filter, boost, grouping, or attribution

3. Factorial experiment matrix
Build a practical factorial or fractional-factorial matrix that isolates chunking and metadata effects without creating unnecessary runs.
For each experiment arm include:
- Arm ID
- Chunking configuration
- Metadata configuration
- Retrieval settings held constant
- Indexing requirements
- Expected comparison value
- Risk of confounding
- Minimum evidence needed before execution

If the full factorial design is too large, propose a staged design:
- Screening stage to remove weak options
- Focused comparison stage for finalists
- Confirmation stage against the baseline

4. Dataset and query slices
Define the evaluation dataset and query slices needed to detect meaningful differences.
Include:
- Document sample selection rules
- Minimum corpus coverage requirements
- Query source: logs, expert-written questions, synthetic-but-reviewed questions, or known support cases
- Query slice taxonomy
- Required number of queries per slice if enough information is available; otherwise give a sizing rule

Use query slices such as:
- Exact lookup questions
- Multi-section synthesis questions
- Long-tail entity questions
- Recent or version-specific questions
- Table or list extraction questions
- Procedure or policy questions
- Ambiguous terminology questions
- Metadata-dependent questions
- Queries where the answer should not be in the corpus

For each slice, specify the retrieval failure mode it is intended to expose.

5. Relevance judgments and answer evaluation setup
Define how ground truth or reference judgments should be created.
Include:
- Who should judge relevance and answer correctness
- What evidence they need
- How to label relevant, partially relevant, irrelevant, stale, duplicate, and misleading chunks
- How to handle multiple valid source passages
- How to record uncertainty and disagreements
- How to prevent evaluators from seeing the tested configuration when practical

6. Retrieval metrics
Specify retrieval metrics that directly measure chunking and metadata effects.
Include:
- Recall@k
- Precision@k or context precision
- MRR or nDCG where graded relevance exists
- Source coverage for multi-source answers
- Duplicate or near-duplicate rate in top-k
- Metadata filter precision and filter fallout
- Freshness or version correctness when metadata includes time or version fields
- No-answer retrieval behavior for queries outside the corpus

For each metric, define:
- Formula or scoring method in plain language
- Required inputs
- Which query slices it applies to
- What kind of configuration failure it reveals

7. Answer-level metrics
Define answer metrics only as downstream checks of retrieval configuration, not as a full generation evaluation.
Include:
- Answer correctness against the evaluation standard
- Citation/source support
- Missing critical fact rate
- Unsupported claim rate
- Wrong-version or wrong-jurisdiction rate if relevant
- Refusal or no-answer correctness where the corpus lacks the answer

Explain how to attribute answer failures to retrieval versus generation when the top-k evidence is sufficient but the answer is wrong.

8. Error attribution rules
Create a concrete error taxonomy with decision rules. Include at minimum:
- Chunk boundary split error
- Chunk too broad or noisy
- Chunk too narrow or missing context
- Overlap redundancy error
- Metadata absent
- Metadata incorrect
- Metadata too coarse
- Metadata filter excludes relevant evidence
- Metadata boost overpromotes irrelevant evidence
- Duplicate chunk crowding
- Stale or wrong-version retrieval
- Relevant evidence not indexed
- Relevant evidence indexed but not retrieved
- Answer failure despite sufficient retrieved evidence

For each error type, define:
- Observable evidence
- How to distinguish it from similar errors
- Likely configuration implication
- Whether it should count against chunking, metadata, indexing, retrieval settings, or answer generation

9. Analysis plan by query slice
Define how results should be compared across arms and slices.
Include:
- Primary metric and secondary metrics
- Minimum practical improvement threshold
- Regression checks where a configuration improves one slice but harms another
- Treatment of ties
- Treatment of small sample sizes
- Required confidence or stability checks if repeated runs are possible
- How to summarize tradeoffs for decision owners without hiding slice-level failures

10. Configuration acceptance decision
Create a decision framework the retrieval owner, data owner, and product owner can use after the experiment is run.
Include:
- Acceptance criteria for adopting a new chunking and metadata configuration
- Rejection criteria
- Conditional acceptance criteria requiring remediation
- Required evidence package before approval
- Rollback or re-indexing considerations if adopted
- Open questions that must be resolved before production use

11. Completion checklist
End with a checklist that confirms whether the experiment design is ready to execute. Include checks for:
- Baseline defined
- Candidate arms defined
- Constants identified
- Query slices complete
- Relevance judgment process defined
- Metrics mapped to slices
- Error attribution rules defined
- Acceptance criteria agreed by named accountable owners
- Missing evidence listed
- No unsupported claims of completed execution


## Completion criteria

The workflow is complete when:

- Dataset admissibility is tied to the exact decision claim and material coverage or contamination limits are explicit.
- Automated-judge uses are classified as permitted, limited, or prohibited with arbitration rules.
- Production drift sentinels, thresholds, owners, investigation triggers, and revalidation gates are observable.
- A retrieval experiment is defined when retrieval configuration materially affects the evaluation, or is marked Not applicable with evidence.
- The final decision distinguishes measured results from assumptions, proxies, missing data, and work not run.
