Rubric Design and Calibration Prompt
Design or revise an evidence-aligned assessment rubric with observable criteria, distinct performance levels, scoring rules, annotated anchors, fairness checks, and a practical calibration protocol.
Design or revise an assessment rubric and its calibration plan from the following inputs. Assessment purpose and context: [Assessment purpose and context] Learning objectives: [Learning objectives] Student task and submission format: [Student task and submission format] Learner population and accommodations: [Learner population and accommodations] Scoring model and performance levels: [Scoring model and performance levels] Constraints and policies: [Constraints and policies] Evidence, existing rubric, and exemplars: [Evidence and exemplars] Use Claude to analyze only the materials supplied in the conversation or available through explicitly enabled tool access. Do not claim to have opened a link, inspected a file, interviewed raters, scored submissions, or calculated agreement unless that action occurred and its evidence is available. Claude may draft and analyze the rubric, but it may not publish it, approve it, enter grades, modify an LMS, or make final decisions about students. Input requirements and uncertainty handling - Blocking inputs are the assessment purpose, assessable learning objectives, student task, expected evidence, and intended performance-level or scoring structure. If any is absent or materially contradictory, ask focused clarification questions before presenting a final rubric. - Useful non-blocking inputs include an existing rubric, standards, sample submissions, prior scoring data, accommodation rules, institutional policy, and the desired calibration threshold. - If clarification is unavailable, create only a clearly labeled provisional draft when bounded progress is educationally safe. Preserve unknowns in an unresolved-input register; do not invent standards, policies, student characteristics, exemplar quality, or scoring results. - Separate supplied facts, interpretations, assumptions, conflicts, and unknowns. Cite the relevant supplied objective, policy, exemplar, or source section for consequential design decisions when identifiers are available. Rubric design workflow 1. Parse the assessment context. Identify the decision the score will support, whether the rubric should be analytic or holistic, the intended users, stakes, submission modes, constraints, and approval owner. Explain the analytic-versus-holistic choice and its trade-offs. 2. Build an alignment map connecting each learning objective to task evidence and proposed criteria. Flag objectives that are unassessed, evidence that is irrelevant to an objective, and criteria that would double-count the same performance. 3. Define a concise set of criteria. Each criterion must measure a distinct construct, use observable evidence, state what is in and out of scope, and avoid grading effort, behavior, writing conventions, attendance, or presentation polish unless these are explicitly part of the objective. 4. Construct performance-level descriptors. Make levels ordered, parallel, observable, and meaningfully distinct. Describe evidence quality rather than relying on vague modifiers such as excellent, adequate, weak, or mostly. State boundary conditions between adjacent levels and address missing, incomplete, off-topic, inaccessible, or non-scorable work. 5. Define scoring mechanics. Specify criterion weights or points, score ranges, aggregation and rounding rules, treatment of not-applicable criteria, minimum-evidence rules, resubmissions, and any cut scores. Verify that weights total the intended maximum and that no rule produces impossible or contradictory scores. Identify policy choices that require human authorization. 6. Develop scoring guidance and anchors. For each criterion, explain positive evidence, common misconceptions, borderline cases, and evidence that must not influence scoring. Map supplied exemplars to levels only when their contents support the mapping. If no authentic exemplars are supplied, create short synthetic illustrations labeled as hypothetical and never present them as observed student work. 7. Design a calibration protocol proportionate to the stakes. Include rater orientation, blind independent scoring, a representative sample strategy, recording of criterion-level scores and rationales, agreement measures, discrepancy review, descriptor revision, a second scoring round, and escalation for unresolved disagreement. Use a supplied agreement threshold when authorized; otherwise recommend a threshold and label it as a proposal rather than policy. For ordinal levels, include exact and adjacent agreement and, when sample size and expertise permit, an appropriate statistic such as weighted kappa or intraclass correlation with its limitations. 8. Review validity, accessibility, and fairness. Check that criteria measure the intended construct, permitted accommodations do not create unintended penalties, descriptors work across allowed submission modes, and language does not encode cultural, linguistic, disability, socioeconomic, or demographic assumptions. Do not infer protected or sensitive characteristics. Recommend specialist or policy review where accessibility, legal, or high-stakes implications exceed the supplied authority. 9. Test the draft against representative cases: clear high and low performance, adjacent-level boundaries, mixed profiles, missing evidence, off-topic work, accommodation-dependent formats, and plausible rater disagreement. Record expected treatment, actual result from applying the written rules, evidence used, and any unresolved defect. If no submissions can actually be tested, provide test cases and expected outcomes while marking execution as unavailable. 10. Reconcile defects. Revise criteria, descriptors, weights, or scoring notes when tests reveal construct gaps, overlap, ambiguous boundaries, score inflation, score compression, or inconsistent handling. Keep a decision log showing what changed and why. Authority and safeguards - Treat the result as a draft for review unless an authorized owner explicitly approves it outside this interaction. - Do not assign or alter real grades, publish the rubric, communicate decisions to learners, or claim institutional compliance. - Remove or mask student names, identifiers, disability information, and other unnecessary personal data from examples. Do not reproduce sensitive student work beyond what is necessary for rubric analysis. - Stop and request human review if the rubric could determine high-stakes progression, discipline, access to services, or legal compliance; if policy conflicts remain unresolved; if the task asks for discriminatory criteria; or if the available evidence cannot support a defensible scoring model. - Preserve the original rubric and decision history when proposing revisions so an authorized reviewer can compare or reverse changes. Required deliverable 1. Rubric status and scope - Draft status, intended assessment, users, stakes, rubric type, scoring purpose, approval required, and unresolved blockers. 2. Input and evidence ledger - Item, classification as supplied fact, interpretation, assumption, conflict, or unknown, source reference, effect on the design, and resolution needed. 3. Objective-to-evidence alignment map - Learning objective, task evidence, proposed criterion, coverage status, overlap risk, and rationale. 4. Rubric table - Criterion name, construct measured, objective alignment, observable evidence, weight or points, one descriptor for every performance level, adjacent-level boundary notes, and non-scorable conditions. 5. Scoring specification - Formula, total possible score, weighting rationale, rounding, missing and not-applicable handling, cut-score status, resubmission rules, and worked score examples. Label unauthorized choices as recommendations. 6. Criterion scoring guide - Positive indicators, misconceptions, irrelevant evidence, edge cases, and scorer cautions for each criterion. 7. Anchor set - Exemplar identifier or hypothetical label, criterion-level placement, evidence-based rationale, ambiguity, and reviewer confirmation needed. 8. Calibration protocol - Participants, preparation, sample composition, scoring rounds, data to record, agreement measures, proposed or authorized thresholds, discrepancy rules, descriptor-revision trigger, and escalation owner. 9. Fairness, accessibility, and validity review - Check, evidence inspected, issue found, likely impact, mitigation, responsible reviewer, and unresolved status. 10. Verification and acceptance matrix - For each check, record expected result, actual observation if executed, evidence, and status as passed, failed, blocked, or not run. Include objective coverage, criterion independence, observable and monotonic descriptors, adjacent-level distinguishability, arithmetic totals, score-bound integrity, consistent treatment of identical evidence, accommodation compatibility, anchor traceability, and calibration-threshold attainment. 11. Decision and change log - Design decision or revision, alternatives considered, evidence, trade-off, authorization required, and resulting status. 12. Handoff - Separate completed drafting from unexecuted calibration, unavailable evidence, unresolved decisions, and approvals still required. End with the smallest safe action for the authorized educator or assessment owner. Do not describe the rubric as validated, calibrated, reliable, fair, approved, tested, or ready for grading unless the corresponding review or test was actually performed and documented. When only a plan or draft exists, use those exact states consistently.
Variables to Replace
- Assessment purpose and context
- Learning objectives
- Student task and submission format
- Learner population and accommodations
- Scoring model and performance levels
- Constraints and policies
- Evidence and exemplars
How to Use This Prompt
In Claude, replace every bracketed variable with the relevant assessment information. Provide the learning objectives, student task, scoring requirements, policies, existing rubric, anonymized exemplars, and any prior rater or calibration evidence. Then run the prompt. If essential evidence is missing, answer Claude’s focused clarification questions before treating the rubric as final.
Example Use Case
A university course team provides Claude with a capstone brief, four learning outcomes, an existing 100-point rubric, accommodation policy, and anonymized sample submissions. Claude maps each outcome to observable evidence, identifies overlapping criteria and ambiguous level boundaries, proposes a revised analytic rubric, annotates provisional anchors, and supplies a two-round calibration protocol. The faculty assessment lead reviews and authorizes the rubric before it is used for grading.