Reusable AI capability
Calibrate AI Regression Acceptance Thresholds
Set and maintain risk-based AI regression gates using measurement uncertainty, baseline variance, critical slices, practical significance, and explicit release trade-offs.
This Skill packages a reusable way to use the linked Prompt or Workflow; Amo.ng does not run it for you.
# Calibrate AI Regression Acceptance Thresholds Skill ID: AMO-S-000023 Skill URL: https://amo.ng/skills/calibrate-ai-regression-acceptance-thresholds Purpose: Give evaluation, domain, product, and release owners a repeatable threshold-calibration method that can be reused across versions without treating one aggregate score as the decision. Required inputs: - Exact release claim, affected behaviors, user or risk slices, and decision owner - Baseline and candidate measurements, sample sizes, uncertainty, repeatability, and historical variance - Severity, exposure, practical-significance, guardrail, and non-compensable failure criteria - False-pass and false-block consequences, release constraints, and revalidation cadence How to use: When to use: - AI regression results need a defensible release threshold or an existing gate requires recalibration. - Aggregate pass rates may hide critical slice failures or normal measurement noise. When not to use: - Choosing evaluation tasks or building a dataset from scratch. - Setting thresholds without repeated baseline evidence or a stated decision consequence. Reusable method: 1. State the exact release claim and segment behaviors into material decision slices. 2. Establish baseline distribution, measurement uncertainty, repeatability, and normal variance for each slice. 3. Define practical harm and exposure, not only statistical difference. 4. Set non-compensable gates, acceptable bands, warning bands, and investigation triggers by slice. 5. Test sensitivity to sample size, missing data, multiple comparisons, threshold placement, and false-pass versus false-block cost. 6. Record exceptions, expiry, monitoring, and recalibration triggers; apply the gate only to the version and evidence boundary reviewed. Expected output: A threshold register with baseline evidence, uncertainty, slice gates, non-compensable failures, sensitivity, decision trade-offs, exceptions, owner approvals, and revalidation triggers. Boundaries: Do not invent results, uncertainty, thresholds, approvals, or risk tolerance. Evaluation and domain owners approve measurement interpretation; product and risk owners approve trade-offs; the release owner controls use of the gate. Source: AMO-P-000304. Applicable Workflow: Assure an AI System Change for Production Release. Powered by Prompt: Regression Acceptance Threshold Calibration Source ID: AMO-P-000304 https://amo.ng/prompts/regression-acceptance-threshold-calibration Completion criteria: Complete when every material slice has baseline evidence, uncertainty, practical-significance rule, gate and rationale; non-compensable failures are explicit; false-pass and false-block trade-offs are bounded; and each threshold has owner, version scope, expiry or recalibration trigger. Use this Amo.ng Skill with your preferred AI tool. Supply the required inputs and follow the usage instructions. # Calibrate AI Regression Acceptance Thresholds Skill ID: AMO-S-000023 Skill URL: https://amo.ng/skills/calibrate-ai-regression-acceptance-thresholds Purpose: Give evaluation, domain, product, and release owners a repeatable threshold-calibration method that can be reused across versions without treating one aggregate score as the decision. Required inputs: - Exact release claim, affected behaviors, user or risk slices, and decision owner - Baseline and candidate measurements, sample sizes, uncertainty, repeatability, and historical variance - Severity, exposure, practical-significance, guardrail, and non-compensable failure criteria - False-pass and false-block consequences, release constraints, and revalidation cadence How to use: When to use: - AI regression results need a defensible release threshold or an existing gate requires recalibration. - Aggregate pass rates may hide critical slice failures or normal measurement noise. When not to use: - Choosing evaluation tasks or building a dataset from scratch. - Setting thresholds without repeated baseline evidence or a stated decision consequence. Reusable method: 1. State the exact release claim and segment behaviors into material decision slices. 2. Establish baseline distribution, measurement uncertainty, repeatability, and normal variance for each slice. 3. Define practical harm and exposure, not only statistical difference. 4. Set non-compensable gates, acceptable bands, warning bands, and investigation triggers by slice. 5. Test sensitivity to sample size, missing data, multiple comparisons, threshold placement, and false-pass versus false-block cost. 6. Record exceptions, expiry, monitoring, and recalibration triggers; apply the gate only to the version and evidence boundary reviewed. Expected output: A threshold register with baseline evidence, uncertainty, slice gates, non-compensable failures, sensitivity, decision trade-offs, exceptions, owner approvals, and revalidation triggers. Boundaries: Do not invent results, uncertainty, thresholds, approvals, or risk tolerance. Evaluation and domain owners approve measurement interpretation; product and risk owners approve trade-offs; the release owner controls use of the gate. Source: AMO-P-000304. Applicable Workflow: Assure an AI System Change for Production Release. Powered by Prompt: Regression Acceptance Threshold Calibration Source ID: AMO-P-000304 https://amo.ng/prompts/regression-acceptance-threshold-calibration Completion criteria: Complete when every material slice has baseline evidence, uncertainty, practical-significance rule, gate and rationale; non-compensable failures are explicit; false-pass and false-block trade-offs are bounded; and each threshold has owner, version scope, expiry or recalibration trigger.Copy skill copies the Skill details. Use with AI adds a short instruction for your preferred AI tool; neither action runs the Skill.
Purpose
Give evaluation, domain, product, and release owners a repeatable threshold-calibration method that can be reused across versions without treating one aggregate score as the decision.
Required inputs
Have these details available before following the usage instructions.
- Exact release claim, affected behaviors, user or risk slices, and decision owner
- Baseline and candidate measurements, sample sizes, uncertainty, repeatability, and historical variance
- Severity, exposure, practical-significance, guardrail, and non-compensable failure criteria
- False-pass and false-block consequences, release constraints, and revalidation cadence
How to use this Skill
When to use:
- AI regression results need a defensible release threshold or an existing gate requires recalibration.
- Aggregate pass rates may hide critical slice failures or normal measurement noise.
When not to use:
- Choosing evaluation tasks or building a dataset from scratch.
- Setting thresholds without repeated baseline evidence or a stated decision consequence.
Reusable method:
1. State the exact release claim and segment behaviors into material decision slices.
2. Establish baseline distribution, measurement uncertainty, repeatability, and normal variance for each slice.
3. Define practical harm and exposure, not only statistical difference.
4. Set non-compensable gates, acceptable bands, warning bands, and investigation triggers by slice.
5. Test sensitivity to sample size, missing data, multiple comparisons, threshold placement, and false-pass versus false-block cost.
6. Record exceptions, expiry, monitoring, and recalibration triggers; apply the gate only to the version and evidence boundary reviewed.
Expected output:
A threshold register with baseline evidence, uncertainty, slice gates, non-compensable failures, sensitivity, decision trade-offs, exceptions, owner approvals, and revalidation triggers.
Boundaries:
Do not invent results, uncertainty, thresholds, approvals, or risk tolerance. Evaluation and domain owners approve measurement interpretation; product and risk owners approve trade-offs; the release owner controls use of the gate. Source: AMO-P-000304. Applicable Workflow: Assure an AI System Change for Production Release.
Powered by an Amo.ng Prompt
Regression Acceptance Threshold Calibration
Open the linked prompt to use the instructions that power this Skill.
Completion criteria
Complete when every material slice has baseline evidence, uncertainty, practical-significance rule, gate and rationale; non-compensable failures are explicit; false-pass and false-block trade-offs are bounded; and each threshold has owner, version scope, expiry or recalibration trigger.
Explore related Workflows
Browse WorkflowsAssure an AI System Change for Production Release
Reconcile the deployed baseline, gate a proposed AI system change, calibrate acceptance thresholds, test adversarial and production coverage, and detect misleading evaluation proxies before release.
Related Prompts
Browse PromptsEvaluation Metric Gaming and Proxy Failure Audit
Determine whether optimization against an AI evaluation metric rewards undesirable behavior, hides target failure, or distorts release and operating decisions.
Audit whether an AI evaluation metric is being used as a trustworthy measure of the target outcome or has become a gameable proxy that rewards the wrong behavior. Evidence: - Real-world outcome, protected invariants, users, harms, and decision the metric informs: [Target outcome and decision context] - Metric definitions, labels, rubrics, graders, transformations, weights, aggregation, thresholds, and dashboards: [Metrics rubrics and aggregation rules] - Training, tuning, selection, prompting, vendor, team, and release incentives tied to the metric: [System incentives and optimization process] - Historical runs, production outcomes, slice results, reviewer evidence, downstream measures, and confidence information: [Evaluation production and outcome evidence] - Incidents, edge cases, suspicious improvements, regressions, complaints, and populations where the proxy may fail: [Known failure cases and affected slices] - Accepted error, non-compensable harms, evaluation owner, product owner, risk reviewer, and release owner: [Risk tolerance and accountable owners] Do not infer gaming solely from a score improvement. Do not assume correlation proves the proxy causes or predicts the target outcome. Do not invent raw data, incentives, or counterfactual results. Separate deliberate gaming, optimization pressure, specification error, measurement error, distribution shift, label leakage, and aggregation artifacts. Audit: 1. Define the target-proxy contract. State the real outcome or invariant, why the metric is expected to represent it, the decision it authorizes, and conditions under which that relationship may fail. 2. Trace the metric construction. Review data source, sampling, labels, rubric, judge, transformations, weights, exclusions, missing data, threshold, and aggregation. Identify degrees of freedom available to the optimized system or team. 3. Map incentives and attack surfaces. Identify who or what adapts to the metric, information exposed during optimization, repeated test use, feedback leakage, and ways to improve the score without improving the target. Keep hypothetical paths separate from observed behavior. 4. Test proxy validity from supplied evidence. Compare metric movement with direct outcome evidence, independent reviewer evidence, protected-slice results, incidents, and downstream measures. Look for divergence, saturation, nonlinearity, reversal, and improvements caused by case exclusion or aggregation. 5. Find compensating and hidden failures. Determine whether gains in common cases mask severe losses in rare slices; whether averages hide abstention, refusal, verbosity, cost, latency, or safety failures; and whether one metric rewards behavior another penalizes. 6. Classify findings. Use Valid within tested range, Weak proxy, Broken proxy, Gameable but controlled, Actively distorted, or Not assessable. State evidence and decision consequence for each metric/slice. 7. Design a safer measurement contract. Recommend direct measures, counter-metrics, holdouts, blind review, refreshed cases, protected-slice gates, independent outcome sampling, or reduced decision authority. Prefer the smallest change that restores decision integrity. 8. Make a use decision. Choose Retain, Retain with counter-metrics, Restrict to named scope, Recalibrate, Replace, or Suspend. Name the owner and evidence required before broader use. Required deliverable: # Evaluation Metric Gaming and Proxy Failure Audit ## Target-Proxy Contract | Target outcome/invariant | Proxy metric | Decision use | Assumed relationship | Known limit | |---|---|---|---|---| ## Construction and Incentive Map | Metric component | Evidence/source | Optimization exposure | Manipulation or distortion path | Observed/hypothetical | |---|---|---|---|---| ## Proxy Validity Findings | Metric/slice | Direct outcome evidence | Proxy movement | Divergence | Classification | Decision impact | |---|---|---|---|---|---| ## Hidden Trade-offs | Apparent gain | Compensating loss | Slice/condition | Evidence | Non-compensable? | |---|---|---|---|---| ## Measurement Repair | Priority | Change | Failure addressed | Acceptance evidence | Owner | |---|---|---|---|---| ## Metric Use Decision - Decision: - Authorized scope: - Required counter-metrics or gates: - Unsupported claims: - Reassessment trigger: Completion requires a documented target-proxy relationship, evidence on material divergence and incentives, and a decision boundary that prevents an aggregate score from authorizing outcomes it does not measure.Production Feedback-to-Evaluation Gap Analysis
Trace incidents, reviewer corrections, user feedback, and operational failures into evaluation coverage, exposing missing, stale, or misweighted regression evidence.
Analyze whether material production feedback is represented accurately in the evaluations used to approve and monitor an AI system. Convert observed failures and corrections into a prioritized, privacy-safe evaluation maintenance decision. Provide: - Current system, release claims, user populations, task slices, and decisions supported by evaluation: [System release and evaluation scope] - Sanitized incidents, complaints, reviewer corrections, escalations, overrides, support tickets, outcome signals, and trace summaries: [Production feedback and incident evidence] - Evaluation datasets, test cases, rubrics, graders, slices, thresholds, run history, and known exclusions: [Current evaluation assets and results] - Feedback sampling, severity, taxonomy, deduplication, triage, and follow-up methods: [Sampling taxonomy and triage rules] - Consent, privacy, retention, rights, security, review capacity, and evaluation/release owners: [Privacy constraints and accountable owners] Do not equate feedback frequency with prevalence unless collection supports that inference. Do not copy sensitive production content into tests without authorization and sanitization. Distinguish user preference, product defect, model failure, retrieval failure, policy issue, workflow issue, and measurement artifact. Do not claim that an evaluation would reproduce a failure unless execution evidence exists. Analysis: 1. Bound the feedback population. Describe collection channels, period, exposed traffic, sampling biases, missing populations, deduplication, and severity rules. Identify feedback that cannot support prevalence estimates. 2. Normalize production signals. Group evidence into stable failure or correction patterns while preserving important context such as task, user, language, model, prompt, tool, retrieval, policy, and downstream consequence. Keep ambiguous cases unresolved. 3. Map signals to evaluation assets. For each material pattern, find the closest test case, slice, rubric criterion, grader behavior, and threshold. Classify representation as Direct and current, Partial, Proxy, Stale, Misweighted, Mislabelled, or Missing. 4. Diagnose the gap mechanism. Determine whether the gap comes from feedback capture, triage, privacy constraints, taxonomy, sampling, dataset curation, rubric design, grader behavior, test execution, thresholding, or release decision use. 5. Assess decision impact. Explain which release or monitoring claims become weaker because of each gap. Identify failures that were present in evaluation but hidden by aggregation or accepted thresholds. 6. Design safe evaluation updates. Specify the smallest representative case, slice, rubric change, grader calibration sample, threshold change, or monitoring link needed. Include provenance, sanitization, consent/rights, label ownership, and expiry or refresh rules. 7. Prioritize and assign. Rank by harm, recurrence evidence, claim impact, coverage absence, and feasible remediation. Do not manufacture a numeric priority score when inputs do not support one. 8. Define closure. Require a trace from production evidence to approved sanitized case, executable result, release interpretation, and recurrence monitoring. The privacy or data owner must approve use of production-derived feedback, the evaluation owner approves benchmark changes, and the release owner retains authority over deployment gates. Do not copy sensitive records outside the authorized evidence boundary or treat this analysis as approval to change production evaluation policy. Required deliverable: # Production Feedback-to-Evaluation Gap Analysis ## Feedback Population and Limitations - Period and traffic represented: - Channels: - Sampling/deduplication method: - Biases and missing populations: ## Production Signal Taxonomy | Pattern | Evidence count/status | Context/slice | Consequence | Confidence | Prevalence claim allowed? | |---|---|---|---|---|---| ## Feedback-to-Evaluation Map | Production pattern | Existing asset | Representation | Gap mechanism | Release/monitoring impact | Owner | |---|---|---|---|---|---| ## Evaluation Maintenance Backlog | Priority | Update | Source evidence | Sanitization/rights check | Acceptance test | Decision affected | Owner | |---|---|---|---|---|---|---| ## Decision Summary - Evaluation claims still supported: - Claims narrowed or suspended: - Updates required before next release: - Signals retained for monitoring only: - Unresolved uncertainty: Completion requires every high-severity production pattern to be mapped or explicitly excluded with rationale, and every new evaluation asset to retain safe provenance back to the production learning it represents.Adversarial Evaluation Coverage Review
Map credible abuse and failure hypotheses to adversarial tests, exposed system surfaces, production controls, and release-blocking coverage gaps.
Review whether adversarial evaluation evidence covers the credible abuse cases and failure modes of an AI system's actual production boundary. Focus on threat-to-test traceability, not the volume or novelty of prompts. Evidence: - Intended users, capabilities, actions, data, deployment context, and release claims: [System scope and release claims] - Threat actors, assets, trust boundaries, misuse cases, failure hypotheses, and severity rationale: [Threat model and abuse hypotheses] - Test cases, mutations, datasets, harness settings, results, failures, and reviewer notes: [Evaluation cases and execution records] - Models, prompts, retrieval, memory, tools, identities, interfaces, filters, monitoring, and containment: [Architecture surfaces and controls] - Incidents, near misses, abuse reports, support feedback, red-team findings, and observed attack variants: [Production incidents and feedback] - Accepted risk, release threshold, policy constraints, security owner, evaluation owner, and release owner: [Risk tolerance and accountable owners] Do not treat a test title as proof of coverage. Do not invent an attack result or claim that a control was bypassed unless execution evidence is supplied. Separate test design coverage from executed coverage, successful defense, failure detection, and containment. Avoid producing novel exploit instructions beyond what is necessary to describe a bounded test objective. Coverage review: 1. Bound the evaluated system. Identify protected assets, allowed behavior, consequential actions, interfaces, user roles, data classes, and release claims. Flag architecture surfaces omitted from the supplied threat model. 2. Normalize threat hypotheses. Express each credible threat as actor/precondition, attack or failure path, target asset or invariant, expected harm, and observable success/failure condition. Consolidate duplicates without losing materially different preconditions. 3. Build a threat-to-test matrix. Map each hypothesis to test cases, variants, environments, system surfaces, evaluator/oracle, and execution evidence. Distinguish Direct, Partial, Proxy-only, Design-only, and Missing coverage. 4. Assess test strength. Review realism, adaptiveness, multi-turn behavior, encoded/indirect inputs, tool and retrieval paths, identity contexts, state persistence, rate effects, chained failures, and negative controls where relevant. Identify tests that exercise only a mock or obsolete configuration. 5. Evaluate oracle and control evidence. Determine whether the test can tell prevention, safe refusal, detection, containment, and recovery apart. Flag weak graders, ambiguous expected behavior, silent side effects, and controls whose operation is not observable. 6. Incorporate production learning. Map supplied incidents and feedback to existing hypotheses and tests. Identify recurring production failures absent from the evaluation or sanitized regression bank. 7. Prioritize gaps. Rank gaps by credible exposure, severity, control weakness, test feasibility, and release claim—not by arbitrary numeric scoring. Name the minimum test or evidence needed to close each gap. 8. Make a release-coverage decision. Choose Adequate for stated scope, Adequate with restrictions, Conditional on named tests, or Inadequate. State excluded claims and untested surfaces. Keep testing within the authorized evaluation environment and supplied data boundary; do not probe production systems or real users. The security reviewer approves threat coverage and the release owner approves residual-risk acceptance. For each material threat, require acceptance evidence showing the expected observation, actual observation when tested, coverage status, and unresolved gap. Proposed adversarial tests are not executed results. Required deliverable: # Adversarial Evaluation Coverage Review ## System and Threat Boundary | Asset/invariant | Surface | Actor/precondition | Consequence | Evidence source | |---|---|---|---|---| ## Threat-to-Test Traceability Matrix | Threat hypothesis | Test IDs | Variants/surfaces | Oracle | Execution evidence | Coverage status | Gap | |---|---|---|---|---|---|---| ## Test-Strength Findings | Test/group | Realism limitation | Obsolete/missing surface | Oracle weakness | Decision impact | |---|---|---|---|---| ## Production Feedback Coverage | Incident/feedback pattern | Existing test | Representation quality | Regression case needed | Owner | |---|---|---|---|---| ## Priority Coverage Backlog | Priority | Missing coverage | Minimum safe test | Required evidence | Owner | Release consequence | |---|---|---|---|---|---| ## Coverage Decision - Decision: - Release scope supported: - Restricted or unsupported scope: - Tests required before release: - Residual threat uncertainty: Completion requires traceability for every material threat hypothesis, explicit separation of designed and executed coverage, and a release scope no broader than the tested system surfaces and oracles support.Production Prompt and Configuration Drift Audit
Reconcile approved AI prompts and runtime configuration with deployed variants, trace unauthorized drift, and define rollback, adoption, or revalidation decisions.
Audit whether an AI system is running the exact prompt and configuration state that its owners approved. Include system and developer prompts, templates, policies, model and version, retrieval settings, tool definitions, decoding controls, feature flags, and environment-specific overrides. Provide: - Approved versions, hashes, manifests, release records, and intended environment matrix: [Approved prompt and configuration baseline] - Deployed artifacts, resolved configuration, environment values, gateway settings, and runtime version identifiers: [Deployed artifacts and environment inventory] - Pull requests, tickets, approvals, emergency changes, deployment records, and rollback history: [Change deployment and approval records] - Timestamped traces, rendered prompts, model routes, tool schemas, retrieval settings, outputs, and incidents: [Runtime traces and behavior evidence] - Evaluation results, regression thresholds, release claims, and accepted limitations: [Evaluation results and acceptance criteria] - Recovery options, production constraints, and prompt, service, security, and release owners: [Rollback constraints and owners] Do not reconstruct a secret or omitted prompt from behavior alone. Do not claim an environment or provider was inspected unless its artifacts are supplied. Treat version labels without content hashes or resolved configuration evidence as weak identifiers. Separate deployed drift from expected environment variation and from behavior drift with no proven configuration change. Audit procedure: 1. Establish the approved baseline. Create a component inventory with exact identifiers, content hashes where supplied, owning role, approved environment, dependency versions, and acceptance evidence. Flag mutable aliases, undocumented defaults, and missing baselines. 2. Resolve the effective production state. Map the artifacts and configuration actually used for representative runtime records. Include template rendering, injected policies, gateway transformations, model routing, tool schemas, retrieval settings, feature flags, and post-processing. 3. Compare baseline to effective state. Classify differences as Approved release, Expected environment variance, Emergency authorized change, Unauthorized drift, Stale deployment, Partial rollout, Mutable dependency change, or Not assessable. 4. Reconstruct material drift. For each material difference, identify first evidence, likely change path, affected environments and traffic, duration, accountable owner, observed behavior, and whether the drift invalidates prior evaluation evidence. 5. Assess decision impact. Determine whether the drift changes safety boundaries, tool authority, data handling, answer contract, retrieval behavior, quality claims, cost or latency assumptions, or regulatory/contractual commitments. Do not infer causation from temporal correlation without supporting evidence. 6. Select a disposition per drift item. Choose Roll back to approved baseline, Adopt through change control, Restrict affected traffic, Revalidate before decision, or Monitor as accepted variance. State smallest safe action, owner, approval, dependencies, and rollback trigger. 7. Design recurrence controls. Define immutable artifact identifiers, resolved-config capture, deployment attestation, runtime sampling, hash comparison, alert thresholds, emergency-change expiry, and evaluation invalidation rules proportionate to the risk. If an approved baseline, component hash, deployment record, or owner decision is missing or conflicts with another source, request the blocking evidence and keep the affected drift status Unknown. Do not infer the approved configuration from the currently running state, and do not close the audit while a material baseline conflict remains unresolved. Required deliverable: # Production Prompt and Configuration Drift Audit ## Approved Baseline Inventory | Component | Approved identifier/hash | Environment | Owner | Approval evidence | Evaluation evidence | |---|---|---|---|---|---| ## Effective Production Inventory | Component | Observed identifier/config | Evidence source | Traffic/window | Confidence | Missing proof | |---|---|---|---|---|---| ## Drift Register | Difference | Classification | First evidence | Scope | Behavior/claim impact | Prior evaluation still valid? | Owner | |---|---|---|---|---|---|---| ## Disposition Plan | Priority | Drift item | Decision | Smallest safe action | Approval | Verification | Rollback trigger | |---|---|---|---|---|---|---| ## Recurrence Controls | Control | Drift detected | Evidence retained | Threshold | Owner | Cadence | |---|---|---|---|---|---| ## Audit Conclusion - Production state matches approval: Yes / Partially / No / Not assessable - Release claims invalidated or narrowed: - Immediate restrictions: - Unresolved evidence: - Next owner decision: Complete the audit only when every material production component is reconciled to an approved state or a named disposition, and the release owner can tell which evaluation claims remain valid.Production AI Evaluation Drift Detection Plan
Design a control plan that detects when a once-approved AI evaluation is no longer reliable for production decisions.
Design a production AI evaluation-drift detection plan for an already-approved evaluation. The goal is to detect when production conditions make the evaluation no longer decision-reliable. Do not design a generic evaluation harness. Do not perform one-time model upgrade regression testing. Focus on ongoing production drift controls for an evaluation that already exists. Context to provide: - System or product under evaluation: [System or product under evaluation] - Approved evaluation decision use: [Approved evaluation decision use] - Current evaluation artifact summary: [Current evaluation artifact summary] - Production telemetry and outcome evidence available: [Production telemetry and outcome evidence available] - Known recent or planned changes: [Known recent or planned changes] - Accountable owners and operating constraints: [Accountable owners and operating constraints] - Risk tolerance or escalation policy: [Risk tolerance or escalation policy] Evidence discipline: - Separate observed evidence from inference. - Do not claim logs, datasets, tests, graders, prompts, retrieval systems, production traffic, or approvals were inspected unless they are included in the provided context. - Flag missing information that prevents firm threshold-setting or assignment of ownership. - Preserve uncertainty where evidence is incomplete. - If assumptions are necessary, label them as assumptions and explain how the evaluation owner should verify them. Produce the following deliverable: 1. Evaluation reliability boundary Define what decision the evaluation is approved to support, what production population it is intended to represent, what conditions must remain comparable, and what conditions would make the evaluation no longer decision-reliable. 2. Drift taxonomy Create a task-specific taxonomy covering at minimum: - Traffic or user-intent drift - Input-format or language drift - Label-policy or ground-truth drift - Human reviewer or labeling-team drift - LLM-as-grader, rubric, or judge-prompt drift - Application prompt or system-instruction drift - Model, provider, parameter, or routing drift - Retrieval corpus, embedding, ranking, or freshness drift - Tool, API, data dependency, or integration drift - Outcome, complaint, incident, conversion, safety, or business-metric drift For each drift type, state observable signals, likely false positives, decision impact, accountable owner, and evidence needed. 3. Sentinel and sampling plan Specify sentinel checks and production sampling methods that can detect each meaningful drift type. Include: - Always-on metrics versus periodic review samples - Stratified slices that must be monitored - Minimum viable sample size or sample logic when exact sizes cannot be justified - Triggered sampling after incidents, launches, prompt changes, retrieval updates, model routing changes, policy changes, or unusual outcome shifts - Treatment of low-volume but high-risk slices - Owner responsible for sample collection, review, and documentation 4. Comparability ledger Design a ledger that records whether the current production environment remains comparable to the approved evaluation baseline. Include ledger fields for: - Evaluation version and approved decision use - Baseline dataset or traffic window - Production traffic window reviewed - Prompt, model, retrieval, grader, rubric, label policy, and tool versions - Known changes since approval - Evidence source for each comparison - Comparability status: comparable, degraded, not comparable, or unknown - Owner attestation required from evaluation owner, product owner, data owner, security reviewer, or other named accountable role as appropriate - Revalidation requirement and due date 5. Detection thresholds Propose initial thresholds using the evidence provided. Where evidence is insufficient, provide threshold-setting rules instead of invented numbers. Cover: - Statistical or distributional thresholds - Operational thresholds - Quality and safety thresholds - Business or outcome thresholds - Grader agreement or calibration thresholds - Retrieval freshness or coverage thresholds - Incident-based hard stops For each threshold, state the metric, comparison baseline, trigger level, rationale, owner, action required, and expected review cadence. 6. Investigation triggers Define when the team must investigate before continuing to rely on the evaluation. Include triggers for traffic anomalies, label disagreement, grader instability, prompt or model changes, retrieval degradation, unexplained outcome shifts, safety incidents, and stakeholder challenges to evaluation validity. 7. Revalidation triggers Define when the approved evaluation must be refreshed, rerun, recalibrated, or retired. Include criteria for partial revalidation, full revalidation, temporary suspension of evaluation-based decisions, and formal owner signoff. 8. Runbook Write a practical runbook with: - Daily, weekly, monthly, and release-event checks where appropriate - Required inputs and evidence sources - Step-by-step investigation path - Decision states: continue relying, rely with caveat, pause reliance, revalidate, or retire - Communication path to evaluation owner, product owner, data owner, security reviewer, release owner, and incident owner where relevant - Documentation artifacts to retain 9. Completion checks End with observable completion criteria. The plan is complete only if it identifies monitored drift types, assigns accountable owners, defines sentinel and sampling mechanisms, records comparability evidence, states thresholds or threshold-setting rules, specifies investigation and revalidation triggers, and explains what evidence is still missing.Judge-Model Reliability and Calibration Review
Evaluate whether a proposed automated judge is calibrated enough for a specific scoring or classification decision.
Determine whether the proposed automated judge is sufficiently reliable and calibrated for the specified evaluation decision. Focus only on judge fitness for the stated scoring or classification decision; do not design a full evaluation platform or generic prompt evaluation harness. Context to provide: - Evaluation decision: [Evaluation decision] - Judge prompt or rubric: [Judge prompt or rubric] - Candidate outputs and source items: [Candidate outputs and source items] - Human reference labels or adjudications: [Human reference labels or adjudications] - Risk profile and protected attributes: [Risk profile and protected attributes] - Acceptance thresholds: [Acceptance thresholds] Evidence discipline: - Separate observed evidence from inference. - State when evidence is missing, weak, imbalanced, or not blinded. - Do not claim that tests, files, systems, reviewers, or production behavior were inspected unless the provided material supports that claim. - Preserve uncertainty where the sample is too small, labels are disputed, or reference judgments are not independent. - Treat human reference labels as evidence to be assessed, not as automatically correct. Deliverable required: 1. Evaluation decision and judge contract Define the exact decision the judge is being asked to make: - Decision type: classification, ordinal rating, pairwise preference, threshold pass/fail, ranking, or other. - Intended users and accountable owner, such as evaluation owner, product owner, domain reviewer, security reviewer, or data owner. - Inputs the judge is allowed to use. - Inputs or knowledge the judge must not use. - Score scale, labels, thresholds, and tie-breaking rules. - What counts as a correct judgment versus an acceptable judgment. - Known boundary cases where the judge’s authority should stop. - Consequences of a false positive, false negative, over-score, and under-score. 2. Blinded calibration design Propose a calibration design suitable for the provided decision and evidence: - How examples should be blinded, randomized, deduplicated, and stratified. - Minimum case coverage needed across easy cases, close calls, failures, adversarial examples, protected or sensitive attributes, and domain-specific edge cases. - How many independent human adjudications are needed and when domain reviewer arbitration is required. - Which metrics are appropriate and why, such as exact agreement, weighted agreement, Cohen’s kappa, Krippendorff’s alpha, rank correlation, threshold confusion matrix, false positive and false negative rates, calibration by score bucket, and self-consistency across repeated runs. - How to avoid leakage from model identity, author identity, expected answer wording, ordering effects, or rubric hints. - What must be held out for future regression checks. 3. Current evidence assessment Using only the provided evidence, assess whether calibration can be judged now: - Evidence available. - Evidence missing. - Sample quality concerns. - Label quality concerns. - Whether the provided material is enough to support a permitted-use decision. 4. Disagreement analysis Analyze judge disagreement against reference labels or adjudications: - Where the judge agrees reliably. - Where disagreement clusters by score band, topic, task type, output length, language, ambiguity, or source quality. - Whether disagreements are random, systematic, rubric-driven, or caused by unclear source material. - Whether the judge is too lenient, too strict, overconfident, inconsistent near thresholds, or sensitive to irrelevant style features. - Distinguish clear judge errors from cases where the human reference may be ambiguous or under-specified. 5. Bias, instability, and attack findings Evaluate fitness risks that could invalidate the judge for the specified decision: - Bias or disparate error patterns related to [Risk profile and protected attributes]. - Sensitivity to superficial wording, formatting, verbosity, fluency, dialect, language variety, or model identity. - Instability across repeated judgments, order changes, paraphrases, or equivalent source presentations. - Prompt injection exposure, including whether candidate text can influence the judge’s rubric, authority, scoring scale, or refusal behavior. - Boundary-case behavior, including ambiguous answers, partially correct answers, missing citations, conflicting sources, unsafe but persuasive content, and cases near the pass/fail threshold. 6. Permitted-use and arbitration gate Make a clear, bounded decision using [Acceptance thresholds]: - Permitted use: where the judge may be used without routine human review. - Conditional use: where the judge may assist but must be sampled, audited, or reviewed by an accountable owner. - Prohibited use: where the judge should not be used for this decision. - Arbitration triggers: exact conditions that require domain reviewer, security reviewer, data owner, or product owner review. - Monitoring requirements: what should be logged, sampled, and periodically recalibrated. - Regression triggers: what changes to the judge prompt, model, rubric, data distribution, or product policy require renewed calibration. 7. Completion check End with a concise readiness statement: - Fit for use, conditionally fit, or not fit for the specified evaluation decision. - Main evidence supporting that conclusion. - Main unresolved risks. - Minimum additional evidence needed before expanding use. - Accountable owner who should accept or reject the permitted-use gate. Use precise professional language. Avoid generic AI governance commentary. Do not recommend broad rewrites of the judge unless a specific reliability failure requires a targeted change.Was this useful?