Experiment Readout Prompt
Produce an evidence-grounded experiment readout covering design validity, metric effects, statistical uncertainty, caveats, segment findings, and an approval-ready recommendation.
Use in AI
Choose an AI tool to copy the current Prompt with a short usage note. Nothing is sent to that tool.
Prepare a decision-grade readout for the experiment described below. Experiment inputs - Objective and hypothesis: [Experiment objective] - Design, assignment method, variants, population, dates, stopping rule, and decision thresholds: [Experiment design and decision rule] - Metric results, sample sizes, analysis tables, statistical outputs, queries, or exports: [Results data and analysis outputs] - Metric definitions, exposure rules, exclusions, instrumentation details, and known data-quality issues: [Data definitions and quality notes] - Product, business, financial, and operational context: [Business and operational context] - Privacy restrictions, permitted recommendations, required approvers, and rollout authority: [Privacy constraints and approval policy] Input gate Treat the objective, experiment design, decision rule, and results as blocking prerequisites for a final readout. If any are absent or materially contradictory, return an Input Gaps section listing what is missing, why it affects validity, and the minimum evidence needed; then provide only a clearly labeled provisional readout where bounded analysis remains safe. Treat contextual details, prior tests, and implementation notes as useful but non-blocking unless they alter the estimand, eligibility, exposure, or decision rule. Never infer missing values, convert an unreported result into zero, or silently resolve conflicting definitions. Claude’s operating boundary Use Claude to inspect and synthesize only the text, files, tables, and analysis outputs actually supplied in this conversation. You may recompute simple quantities such as absolute lift, relative lift, rates, and interval interpretations when the required values are present; show the formula, inputs, and rounding. Do not claim access to an experimentation platform, warehouse, notebook, dashboard, source system, or live deployment unless access and resulting evidence are explicitly provided. Do not execute queries, alter experiment allocation, stop a test, ship a variant, change tracking, contact participants, or approve a rollout. Recommendations are proposals until the named human owner authorizes them. Evidence rules - Label each material statement as one of: supplied fact, derived calculation, assumption, hypothesis, conflict, or unknown. - Cite the supporting file, table, chart, query output, or input section for every headline number and integrity conclusion. If source locations are unavailable, identify the evidence descriptively rather than inventing a citation. - Preserve the analysis framework actually used. Do not translate a frequentist result into a Bayesian probability, or vice versa. Report confidence intervals and p-values for frequentist analyses, or credible intervals and posterior decision probabilities for Bayesian analyses, only when supplied or validly derivable. - Distinguish statistical significance, practical significance, and decision-threshold attainment. Do not treat a non-significant result as proof of no effect. - Mark post hoc metrics, changed stopping rules, unplanned exclusions, and segment analyses as exploratory. Do not present them as preregistered confirmatory evidence. - State when the available material is insufficient to reproduce or independently verify a result. Analysis sequence 1. Reconstruct the experiment contract: hypothesis, unit of randomization, unit of analysis, eligibility, exposure definition, control and treatment experiences, analysis window, primary metric, guardrails, minimum detectable effect, power or sample-size basis, stopping rule, and decision thresholds. Record omissions and any changes made after launch. 2. Check design and execution validity. Evaluate, where evidence permits: randomization and allocation, sample-ratio mismatch, crossover or contamination, repeated units, exposure logging, pre-experiment imbalance, novelty or learning effects, seasonality, concurrent launches, attrition, missingness, bot or fraud filtering, metric-definition drift, delayed outcomes, and interference between units. 3. Reconcile populations and counts from assignment through eligibility, exposure, inclusion, and analysis. Explain exclusions and denominator changes by variant. Flag unexplained losses, incompatible totals, or analysis populations that differ from the stated estimand. 4. Evaluate the primary metric using the declared method. Report control and treatment values, absolute and relative effect, uncertainty interval, test statistic or posterior quantity when available, sample sizes, threshold comparison, and practical impact. Identify peeking, optional stopping, underpowering, multiplicity, clustering, variance reduction, or model adjustments that affect interpretation. 5. Review guardrail and secondary metrics without allowing favorable secondary outcomes to override a failed primary decision rule. State multiplicity controls, if any, and separate confirmatory metrics from diagnostics. 6. Analyze segments only when sample sizes, definitions, and outputs support it. Report interaction evidence rather than relying solely on within-segment significance. Flag sparse cells, unstable estimates, multiple comparisons, and segments defined after observing outcomes. Do not recommend targeting a subgroup from directional noise alone. 7. Translate the evidence into a recommendation of launch, do not launch, continue collecting data, rerun, investigate instrumentation, or inconclusive. Compare expected benefit with downside severity, reversibility, operational cost, and guardrail impact. If the prespecified rule and business recommendation differ, show both and explain why. 8. Define the smallest safe follow-up, its owner, required approval, prerequisite evidence, monitoring signals, and rollback or recovery trigger. Keep all actions proposed unless execution evidence proves they occurred. Privacy and safety controls Use aggregate results wherever possible. Do not reproduce personal identifiers, credentials, access tokens, confidential row-level records, or sensitive attributes unnecessary to the decision. If supplied material exposes such data, stop using or quoting it, identify the issue, and request a redacted or aggregated replacement. Do not infer sensitive traits or propose discriminatory targeting. Escalate for human review when the recommendation could materially affect customers, regulated outcomes, finances, production reliability, or contractual commitments. A rollout recommendation must include authorization, monitoring, and rollback requirements; it is not approval to act. Required deliverable 1. Readout status - Status: final, provisional, blocked, or inconclusive - Decision requested and named decision owner - One-sentence hypothesis and experiment outcome - Recommended disposition with confidence level expressed in terms supported by the analysis - Most important caveat 2. Experiment contract table Columns: field; prespecified definition; observed implementation; evidence source; discrepancy; decision impact. Include population, variants, randomization unit, analysis unit, dates, exposure, primary metric, guardrails, stopping rule, minimum detectable effect, and decision threshold. 3. Population reconciliation table Columns: stage; control count; treatment count; exclusion or loss reason; expected observation; actual observation; evidence; status. Cover assigned, eligible, exposed, retained, and analyzed populations where available. 4. Integrity assessment For each applicable check, provide: check; method or diagnostic; expected condition; actual observation; evidence; pass, fail, unresolved, or not applicable; consequence. Include sample-ratio mismatch, contamination, missingness, instrumentation, denominator consistency, timing, and concurrent-change checks. Do not mark a check passed without an observed result. 5. Metric results Create separate primary, guardrail, and secondary metric tables. Use columns: metric; status as confirmatory or exploratory; control; treatment; absolute effect; relative effect; uncertainty interval; p-value or posterior quantity; sample size; practical threshold; threshold met; method; evidence; interpretation. Use not reported or not derivable rather than fabricated values. 6. Segment findings Columns: segment; rationale; sample sizes; effect and uncertainty; interaction evidence; multiplicity treatment; stability concern; classification as confirmatory, exploratory, or unsupported; implication. Omit unsupported personalization recommendations. 7. Evidence and uncertainty register Columns: ID; statement; classification as supplied fact, derived calculation, assumption, hypothesis, conflict, or unknown; source; consequence if wrong; resolution needed. Include formulas for derived headline values. 8. Decision analysis Compare the prespecified rule, statistical evidence, practical effect, guardrail outcomes, downside severity, reversibility, implementation cost, and operational readiness. State the recommended disposition, credible alternatives, trade-offs, and conditions that would change the recommendation. 9. Verification and acceptance matrix Columns: acceptance check; expected evidence; actual evidence observed; status as pass, fail, unresolved, or unavailable; owner; required resolution. At minimum verify source-to-readout number reconciliation, population totals, metric definitions, analysis method, uncertainty reporting, decision-rule application, exploratory labeling, privacy review, approver identification, and monitoring or rollback readiness. 10. Authorized handoff List proposed action, owner, prerequisite, required human approver, monitoring metric, alert threshold, rollback or recovery trigger, and current state as proposed, approved, executed, verified, blocked, or unavailable. Use approved, executed, or verified only when the supplied evidence demonstrates that state. End with the smallest safe next action.
Variables to Replace
Replace each listed value in the Prompt with information relevant to your task.
- Experiment objective
- Experiment design and decision rule
- Results data and analysis outputs
- Data definitions and quality notes
- Business and operational context
- Privacy constraints and approval policy
How to Use This Prompt
In Claude, replace every bracketed variable with the corresponding experiment materials. Provide the experiment plan or preregistration, assignment and exposure counts, metric definitions, result tables or statistical outputs, data-quality notes, decision thresholds, business context, and approval policy. Redact credentials and unnecessary personal data, upload or paste the evidence, then run the prompt. If blocking inputs are unavailable, use Claude’s provisional output to request them rather than treating it as a final readout.
Example Use Case
A product team has completed a checkout A/B test and must decide whether to roll out a new flow. Provide Claude with the hypothesis, randomization and exposure design, conversion and revenue outputs, guardrail metrics, sample-size plan, exclusion rules, and rollout approvals. The result is a readout that reconciles populations, checks sample-ratio mismatch and instrumentation, distinguishes confirmatory from exploratory findings, and gives an evidence-bounded rollout or rerun recommendation.
Was this useful?