Blind AI Model Comparison Teaching Lab and Evaluation Pack
Design a reproducible learning lab that hides model identity, uses held-out tasks and calibrated scoring, and teaches students to interpret uncertainty and failure slices.
Use in AI
Choose an AI tool to copy the current Prompt with a short usage note. Nothing is sent to that tool.
Design a blind comparison lab that teaches participants how to evaluate AI systems for a defined task without turning a small exercise into a universal model ranking.
## Lab context
Learners, learning outcomes and available session time:
{{audience_and_learning_outcomes}}
Task definition, candidate systems and permitted interfaces:
{{task_and_candidate_systems}}
Cases, reference evidence and scoring resources:
{{cases_and_reference_evidence}}
Privacy, access, budget and facilitation constraints:
{{lab_constraints}}
## Evidence and fairness rules
- Separate supplied facts, design choices, evaluator judgments, assumptions, conflicts, missing information and unresolved uncertainty.
- Do not claim that a model was run, blinded or scored unless corresponding records are supplied.
- Use only data and content that participants are authorized to submit. Remove personal, confidential, copyrighted or restricted material unless its use is licensed and approved.
- Define the unit of comparison before blinding: a base model, configured endpoint or complete assistant product. Hide system identity from scorers where practical. Record the provider, product, model and version, interface, system instructions where available, enabled tools or retrieval, user prompt, settings, date and any unequal capabilities. Attribute findings only to the tested configuration.
- Check whether masking is effective before the scored comparison. Record interface or output cues that reveal identity and either remove them consistently or document the residual unblinding risk.
- Hold testing conditions comparable for the target task, but do not claim that different interfaces, access tiers or tool capabilities are identical. Record every material difference and its effect on interpretation.
- Define the target task and population of cases before comparing systems. A result applies only to the sampled cases, rubric and conditions.
- Keep human relevance or quality judgments traceable. Measure scorer disagreement rather than presenting consensus that did not occur.
- Reserve procurement, deployment and teaching-policy decisions for the responsible owner.
## Design method
1. Convert the learning outcome into one bounded evaluation question and state what the lab cannot establish.
2. Define case strata, difficulty and failure slices. Separate development examples from cases held out from lab design or tuning. Do not claim that a case was absent from a system's training data unless supported by evidence. Record known or suspected benchmark exposure, searchable answer leakage and contamination uncertainty.
3. Choose and document the comparison regime. For a controlled-capability comparison, standardize or disable tools, retrieval, context and interface assistance where possible. For a product-experience comparison, retain native capabilities, document them and attribute results to the complete configured product. Then create a common task packet covering inputs, allowed context, output format, refusal rules, time or cost limits and recording method.
4. Build an observable rubric with anchored examples. Include task success, evidence use, uncertainty, safety and usability only when relevant; avoid one opaque composite score.
5. Define blinding, randomized presentation order, independent scoring, disagreement handling and adjudication. Run a masking-effectiveness check on labels, formatting, interface cues and metadata before scoring, and document residual cues.
6. Specify repeated runs or matched cases where nondeterminism matters. Record failures, refusals, latency and cost without inventing unavailable measurements.
7. Plan analysis using per-criterion results, uncertainty intervals or ranges, scorer agreement and failure-slice comparisons. State when the sample is too small for stable conclusions.
8. Add a debrief that asks learners to distinguish observed results from explanations, identify validity threats and propose the next discriminating test.
## Output contract: Blind Comparison Lab Pack
Return:
1. **Learning and evaluation brief**: audience, learning outcomes, bounded question, scope and non-claims.
2. **Case manifest**: case ID, stratum, permission status, held-out status, reference evidence and contamination concern.
3. **Reproducible run protocol**: unit of comparison, comparison regime, blinded system labels, provider, product, model and version details, task packet, enabled tools and retrieval, settings, order randomization, privacy controls and failure capture.
4. **Anchored scoring guide**: criterion, observable evidence, scale anchors, non-applicable rule and example.
5. **Scoring and adjudication sheets**: blinded result ID, independent judgments, rationale, disagreement and resolution status.
6. **Analysis plan**: comparisons, uncertainty treatment, scorer agreement, failure slices, cost/latency fields and claims that would be unsupported.
7. **Facilitator debrief**: questions that test interpretation, limitations and transfer rather than model preference.
8. **Decision boundary**: what this lab can support, what remains unknown and which owner may act on the result.
## Verification and completion
The design is complete only when the cases are authorized and task-relevant; lab-held-out status and contamination uncertainty are documented; the unit of comparison and comparison regime are explicit; systems receive inputs and capabilities that are comparable under that regime; scoring anchors are observable; identity masking and disagreement handling are feasible; analysis retains uncertainty; and the debrief tests learning.
For each completion condition, record the expected design evidence, the actual supplied evidence, `Met`, `Not met`, or `Blocked` status, and the unresolved organizer action.
If cases, permissions, candidate-system access or scoring capacity are missing, produce a provisional lab plan and a focused acquisition list. Reject requests to fabricate runs, scores, endorsements or a claim that one model is universally best.
Variables to Replace
Replace each listed value in the Prompt with information relevant to your task.
- audience_and_learning_outcomes
- task_and_candidate_systems
- cases_and_reference_evidence
- lab_constraints
How to Use This Prompt
Run this prompt in any capable AI assistant while planning the lab. Paste the learning outcome and task, then attach sanitized candidate cases, reference evidence, available model/version information and constraints. Do not upload protected coursework or personal data. Have the facilitator test the blinding and rubric on synthetic examples before learners use the pack.
Example Use Case
A university AI club compares three assistants on extracting claims from openly licensed or otherwise authorized abstracts. The pack tests the masking, randomizes labelled outputs, uses held-out abstracts, records scorer disagreement and concludes only which system performed better under that specific protocol.
Was this useful?