Evaluation Dataset Coverage and Contamination Audit
Audit an evaluation dataset for provenance, coverage, leakage, contamination, duplication, label quality, and admissibility before it supports release claims.
Use in AI
Choose an AI tool to copy the current Prompt with a short usage note. Nothing is sent to that tool.
Audit the evaluation dataset for decision fitness before it is used to support release claims. Treat the dataset as inadmissible until the evidence below supports a narrower conclusion. Context and inputs to provide: - Dataset and intended release claim: [Dataset description and intended release claim] - Dataset records, schema, samples, or files available for review: [Dataset records or sample with schema] - Source lineage, collection method, licenses, collection dates, and transformations: [Provenance sources and collection dates] - Label definitions, labeling instructions, adjudication process, annotator metadata, and quality checks: [Labeling guidelines and annotator metadata] - Target tasks, risk categories, domains, locales, languages, user populations, and expected operating conditions: [Target tasks risks and user populations] - Known model training corpora, exclusion lists, benchmark sources, public datasets, or other contamination references: [Known training data or exclusion sources] Evidence discipline: - Separate observed facts from inference. Mark each material claim as Observed, Inferred, Not provided, or Not assessable from supplied evidence. - Do not claim that a file, source, test, command, repository, dataset split, or system was inspected unless it is present in the supplied material. - Do not invent counts, percentages, inter-annotator agreement, contamination rates, licenses, or collection dates. If a metric cannot be calculated from the supplied evidence, state what is missing and whether a proxy assessment is possible. - Preserve uncertainty. Use confidence levels only when tied to available evidence. - Focus on this dataset audit, not on building an evaluation harness or comparing model versions. Audit procedure: 1. Define the admissibility question. - Restate the release claim the dataset is expected to support. - Identify the accountable evaluation owner, dataset owner, data owner, and release owner if named in the supplied material; otherwise list them as missing accountability assignments. - Define what the dataset must demonstrate to be fit for that release claim. 2. Build a dataset provenance ledger. For each source, split, subset, or major record group, capture: - Source name or origin - Collection date or time window - Collection method - Rights, license, consent, or usage restriction evidence - Transformation, filtering, augmentation, or generation steps - Label source and labeling workflow - Known exclusions or quarantine rules - Traceability gaps - Confidence in provenance 3. Build a coverage matrix. Map the dataset against the supplied target tasks, risk categories, domains, user populations, languages/locales, difficulty bands, failure modes, and operating conditions. Include: - Available counts or proportions where directly calculable - Coverage status: Adequate, Thin, Missing, Overrepresented, Not assessable - Evidence basis - Release-claim consequence of each gap - Minimum additional evidence or data needed to close the gap 4. Audit leakage, duplication, and contamination risk. Create a contamination register covering: - Exact duplicate records within the dataset - Near duplicates or paraphrase clusters, if detectable from supplied records - Train/eval split leakage, if split information is supplied - Prompt-answer leakage, rubric leakage, or label leakage - Overlap with known training data, public benchmarks, synthetic data sources, vendor examples, documentation, or previous evaluation sets - Temporal leakage relative to the intended release claim - Source reuse that could inflate performance claims For each item, state the evidence, detection method available from supplied material, severity, uncertainty, owner, and remediation. 5. Review label quality and decision reliability. Assess: - Label definition clarity and mutual exclusivity - Alignment between labels, rubric, and release claim - Ambiguous or underspecified cases - Annotator qualification evidence - Adjudication and dispute-resolution process - Inter-annotator agreement or audit sample results, only if provided - Gold-standard or expert-review evidence, only if provided - Label drift across sources, time periods, or task categories - Examples where the label appears inconsistent with the provided guideline 6. Identify decision risks. Distinguish risks that affect: - Statistical validity - External validity and representativeness - Safety or policy risk coverage - Bias across user populations or locales - Claim wording and overgeneralization - Reproducibility and auditability - Legal, license, privacy, or data-rights admissibility 7. Recommend remediation. Provide targeted actions only. Avoid broad rebuilds unless the evidence shows the dataset cannot be repaired. For each action include: - Remediation action - Specific defect addressed - Priority - Responsible owner: dataset owner, evaluation owner, data owner, security reviewer, policy reviewer, legal reviewer, or release owner as applicable - Acceptance check - Whether the dataset must be quarantined, relabeled, deduplicated, rebalanced, restricted to a narrower claim, or rejected 8. Make an admissibility decision. Choose one: - Admissible for the stated release claim - Conditionally admissible after named remediation - Admissible only for a narrower claim - Not admissible for release claims Explain the decision in terms of evidence sufficiency, unresolved uncertainty, contamination risk, coverage gaps, label quality, and provenance. Required output format: # Evaluation Dataset Coverage and Contamination Audit ## 1. Admissibility Question - Intended release claim: - Dataset use in the release decision: - Accountable owners named in evidence: - Missing accountability assignments: - Fitness threshold for this audit: ## 2. Evidence Inventory | Evidence item | Provided | Used for | Limitations | Missing information | |---|---:|---|---|---| ## 3. Dataset Provenance Ledger | Dataset segment/source | Origin | Collection window | Collection method | Rights/consent evidence | Transformations | Label source | Restrictions | Traceability gaps | Confidence | |---|---|---|---|---|---|---|---|---|---| ## 4. Coverage Matrix | Task/risk/user-population dimension | Expected coverage | Observed coverage | Status | Evidence basis | Release-claim consequence | Data or evidence needed | |---|---|---|---|---|---|---| ## 5. Duplication, Leakage, and Contamination Register | Issue | Type | Evidence observed | Detection possible from supplied material | Severity | Uncertainty | Owner | Remediation | |---|---|---|---|---|---|---|---| ## 6. Label-Quality Findings | Finding | Evidence status | Affected records or segment | Decision impact | Confidence | Remediation | |---|---|---|---|---|---| ## 7. Decision Risks List the material risks that remain after the audit. For each, state whether it affects statistical validity, external validity, safety coverage, fairness, reproducibility, rights/privacy, or claim wording. ## 8. Remediation Plan | Priority | Action | Defect addressed | Responsible owner | Acceptance check | Release impact | |---|---|---|---|---|---| ## 9. Admissibility Decision - Decision: - Claim supported, if any: - Claim not supported: - Required conditions before use: - Residual uncertainty: - Final verification checks for the release owner: Completion checks before finalizing: - The provenance ledger, coverage matrix, contamination register, label-quality findings, remediation plan, and admissibility decision are all present. - Every major conclusion cites supplied evidence or is explicitly marked as inference or not assessable. - No unavailable inspection, test, source comparison, or metric is claimed. - The release owner has enough information to decide whether the dataset may support the stated claim, must be narrowed, or must be rejected.
Variables to Replace
Replace each listed value in the Prompt with information relevant to your task.
- Dataset description and intended release claim
- Dataset records or sample with schema
- Provenance sources and collection dates
- Labeling guidelines and annotator metadata
- Target tasks risks and user populations
- Known training data or exclusion sources
How to Use This Prompt
Use in Claude. Paste or upload the dataset records, schema, samples, provenance notes, labeling guidelines, and any known training-data or benchmark exclusion references. Replace every bracketed placeholder, then run the prompt. Afterward, have the evaluation owner verify technical findings, the data owner verify provenance and rights constraints, and the release owner confirm whether the admissibility decision supports the intended claim.
Example Use Case
A release owner wants to claim that a new assistant is safer on regulated financial advice tasks. Before using the evaluation dataset as evidence, the evaluation owner runs this audit to expose missing risk categories, duplicated examples, uncertain source rights, public benchmark overlap, and inconsistent labels, then decides whether the dataset is admissible or must be remediated.
Was this useful?