Judge-Model Reliability and Calibration Review
Evaluate whether a proposed automated judge is calibrated enough for a specific scoring or classification decision.
Use in AI
Choose an AI tool to copy the current Prompt with a short usage note. Nothing is sent to that tool.
Determine whether the proposed automated judge is sufficiently reliable and calibrated for the specified evaluation decision. Focus only on judge fitness for the stated scoring or classification decision; do not design a full evaluation platform or generic prompt evaluation harness. Context to provide: - Evaluation decision: [Evaluation decision] - Judge prompt or rubric: [Judge prompt or rubric] - Candidate outputs and source items: [Candidate outputs and source items] - Human reference labels or adjudications: [Human reference labels or adjudications] - Risk profile and protected attributes: [Risk profile and protected attributes] - Acceptance thresholds: [Acceptance thresholds] Evidence discipline: - Separate observed evidence from inference. - State when evidence is missing, weak, imbalanced, or not blinded. - Do not claim that tests, files, systems, reviewers, or production behavior were inspected unless the provided material supports that claim. - Preserve uncertainty where the sample is too small, labels are disputed, or reference judgments are not independent. - Treat human reference labels as evidence to be assessed, not as automatically correct. Deliverable required: 1. Evaluation decision and judge contract Define the exact decision the judge is being asked to make: - Decision type: classification, ordinal rating, pairwise preference, threshold pass/fail, ranking, or other. - Intended users and accountable owner, such as evaluation owner, product owner, domain reviewer, security reviewer, or data owner. - Inputs the judge is allowed to use. - Inputs or knowledge the judge must not use. - Score scale, labels, thresholds, and tie-breaking rules. - What counts as a correct judgment versus an acceptable judgment. - Known boundary cases where the judge’s authority should stop. - Consequences of a false positive, false negative, over-score, and under-score. 2. Blinded calibration design Propose a calibration design suitable for the provided decision and evidence: - How examples should be blinded, randomized, deduplicated, and stratified. - Minimum case coverage needed across easy cases, close calls, failures, adversarial examples, protected or sensitive attributes, and domain-specific edge cases. - How many independent human adjudications are needed and when domain reviewer arbitration is required. - Which metrics are appropriate and why, such as exact agreement, weighted agreement, Cohen’s kappa, Krippendorff’s alpha, rank correlation, threshold confusion matrix, false positive and false negative rates, calibration by score bucket, and self-consistency across repeated runs. - How to avoid leakage from model identity, author identity, expected answer wording, ordering effects, or rubric hints. - What must be held out for future regression checks. 3. Current evidence assessment Using only the provided evidence, assess whether calibration can be judged now: - Evidence available. - Evidence missing. - Sample quality concerns. - Label quality concerns. - Whether the provided material is enough to support a permitted-use decision. 4. Disagreement analysis Analyze judge disagreement against reference labels or adjudications: - Where the judge agrees reliably. - Where disagreement clusters by score band, topic, task type, output length, language, ambiguity, or source quality. - Whether disagreements are random, systematic, rubric-driven, or caused by unclear source material. - Whether the judge is too lenient, too strict, overconfident, inconsistent near thresholds, or sensitive to irrelevant style features. - Distinguish clear judge errors from cases where the human reference may be ambiguous or under-specified. 5. Bias, instability, and attack findings Evaluate fitness risks that could invalidate the judge for the specified decision: - Bias or disparate error patterns related to [Risk profile and protected attributes]. - Sensitivity to superficial wording, formatting, verbosity, fluency, dialect, language variety, or model identity. - Instability across repeated judgments, order changes, paraphrases, or equivalent source presentations. - Prompt injection exposure, including whether candidate text can influence the judge’s rubric, authority, scoring scale, or refusal behavior. - Boundary-case behavior, including ambiguous answers, partially correct answers, missing citations, conflicting sources, unsafe but persuasive content, and cases near the pass/fail threshold. 6. Permitted-use and arbitration gate Make a clear, bounded decision using [Acceptance thresholds]: - Permitted use: where the judge may be used without routine human review. - Conditional use: where the judge may assist but must be sampled, audited, or reviewed by an accountable owner. - Prohibited use: where the judge should not be used for this decision. - Arbitration triggers: exact conditions that require domain reviewer, security reviewer, data owner, or product owner review. - Monitoring requirements: what should be logged, sampled, and periodically recalibrated. - Regression triggers: what changes to the judge prompt, model, rubric, data distribution, or product policy require renewed calibration. 7. Completion check End with a concise readiness statement: - Fit for use, conditionally fit, or not fit for the specified evaluation decision. - Main evidence supporting that conclusion. - Main unresolved risks. - Minimum additional evidence needed before expanding use. - Accountable owner who should accept or reject the permitted-use gate. Use precise professional language. Avoid generic AI governance commentary. Do not recommend broad rewrites of the judge unless a specific reliability failure requires a targeted change.
Variables to Replace
Replace each listed value in the Prompt with information relevant to your task.
- Evaluation decision
- Judge prompt or rubric
- Candidate outputs and source items
- Human reference labels or adjudications
- Risk profile and protected attributes
- Acceptance thresholds
How to Use This Prompt
Use in Claude. Paste the judge prompt or rubric, representative candidate outputs, source items, human reference labels or adjudications, known risk attributes, and proposed acceptance thresholds. Replace every bracketed placeholder, then run the prompt. Have the evaluation owner review the permitted-use gate, with domain reviewer, security reviewer, or data owner verification where the findings identify arbitration triggers.
Example Use Case
A product team wants to use an LLM judge to decide whether customer-support answers pass a policy-compliance rubric. They run this prompt with blinded examples, adjudicated human labels, edge cases, and pass/fail thresholds to determine whether the judge can auto-score routine cases or must remain advisory near policy boundaries.
Was this useful?