Blind AI Model Comparison Teaching Lab and Evaluation Pack
Design a reproducible learning lab that hides model identity, uses held-out tasks and calibrated scoring, and teaches students to interpret uncertainty and failure slices.
Design a blind comparison lab that teaches participants how to evaluate AI systems for a defined task without turning a small exercise into a universal model ranking.
## Lab context
Learners, learning outcomes and available session time:
{{audience_and_learning_outcomes}}
Task definition, candidate systems and permitted interfaces:
{{task_and_candidate_systems}}
Cases, reference evidence and scoring resources:
{{cases_and_reference_evidence}}
Privacy, access, budget and facilitation constraints:
{{lab_constraints}}
## Evidence and fairness rules
- Separate supplied facts, design choices, evaluator judgments, assumptions, conflicts, missing information and unresolved uncertainty.
- Do not claim that a model was run, blinded or scored unless corresponding records are supplied.
- Use only data and content that participants are authorized to submit. Remove personal, confidential, copyrighted or restricted material unless its use is licensed and approved.
- Hide model identity from scorers where practical; record order, interface, prompt, model/version, settings and date for the organizer's reproducibility record.
- Check whether masking is effective before the scored comparison. Record interface or output cues that reveal identity and either remove them consistently or document the residual unblinding risk.
- Hold testing conditions comparable for the target task, but do not claim that different interfaces, access tiers or tool capabilities are identical. Record every material difference and its effect on interpretation.
- Define the target task and population of cases before comparing systems. A result applies only to the sampled cases, rubric and conditions.
- Keep human relevance or quality judgments traceable. Measure scorer disagreement rather than presenting consensus that did not occur.
- Reserve procurement, deployment and teaching-policy decisions for the responsible owner.
## Design method
1. Convert the learning outcome into one bounded evaluation question and state what the lab cannot establish.
2. Define case strata, difficulty and failure slices. Separate development examples from held-out comparison cases and check for likely contamination or prior exposure.
3. Create a common task packet: identical inputs, allowed context, output format, refusal rules, time or cost limits and recording method.
4. Build an observable rubric with anchored examples. Include task success, evidence use, uncertainty, safety and usability only when relevant; avoid one opaque composite score.
5. Define blinding, randomized presentation order, independent scoring, disagreement handling and adjudication. Run a masking-effectiveness check on labels, formatting, interface cues and metadata before scoring, and document residual cues.
6. Specify repeated runs or matched cases where nondeterminism matters. Record failures, refusals, latency and cost without inventing unavailable measurements.
7. Plan analysis using per-criterion results, uncertainty intervals or ranges, scorer agreement and failure-slice comparisons. State when the sample is too small for stable conclusions.
8. Add a debrief that asks learners to distinguish observed results from explanations, identify validity threats and propose the next discriminating test.
## Output contract: Blind Comparison Lab Pack
Return:
1. **Learning and evaluation brief**: audience, learning outcomes, bounded question, scope and non-claims.
2. **Case manifest**: case ID, stratum, permission status, held-out status, reference evidence and contamination concern.
3. **Reproducible run protocol**: system labels, task packet, settings to record, order randomization, privacy controls and failure capture.
4. **Anchored scoring guide**: criterion, observable evidence, scale anchors, non-applicable rule and example.
5. **Scoring and adjudication sheets**: blinded result ID, independent judgments, rationale, disagreement and resolution status.
6. **Analysis plan**: comparisons, uncertainty treatment, scorer agreement, failure slices, cost/latency fields and claims that would be unsupported.
7. **Facilitator debrief**: questions that test interpretation, limitations and transfer rather than model preference.
8. **Decision boundary**: what this lab can support, what remains unknown and which owner may act on the result.
## Verification and completion
The design is complete only when the cases are authorized and task-relevant; held-out status is documented; all systems receive comparable inputs; scoring anchors are observable; identity masking and disagreement handling are feasible; analysis retains uncertainty; and the debrief tests learning.
For each completion condition, record the expected design evidence, the actual supplied evidence, `Met`, `Not met`, or `Blocked` status, and the unresolved organizer action.
If cases, permissions, candidate-system access or scoring capacity are missing, produce a provisional lab plan and a focused acquisition list. Reject requests to fabricate runs, scores, endorsements or a claim that one model is universally best.
Variables to Replace
- audience_and_learning_outcomes
- task_and_candidate_systems
- cases_and_reference_evidence
- lab_constraints
How to Use This Prompt
Run this prompt in any capable AI assistant while planning the lab. Paste the learning outcome and task, then attach sanitized candidate cases, reference evidence, available model/version information and constraints. Do not upload protected coursework or personal data. Have the facilitator test the blinding and rubric on synthetic examples before learners use the pack.
Example Use Case
A university AI club compares three assistants on extracting claims from openly licensed or otherwise authorized abstracts. The pack tests the masking, randomizes labelled outputs, uses held-out abstracts, records scorer disagreement and concludes only which system performed better under that specific protocol.