Prompt Engineering Expert Claude

Evaluation Metric Gaming and Proxy Failure Audit

Determine whether optimization against an AI evaluation metric rewards undesirable behavior, hides target failure, or distorts release and operating decisions.

Use in AI

Choose an AI tool to copy the current Prompt with a short usage note. Nothing is sent to that tool.

Browse more prompts
Best forEvaluation
ToolClaude
DifficultyExpert
Full Prompt
Audit whether an AI evaluation metric is being used as a trustworthy measure of the target outcome or has become a gameable proxy that rewards the wrong behavior.

Evidence:
- Real-world outcome, protected invariants, users, harms, and decision the metric informs: [Target outcome and decision context]
- Metric definitions, labels, rubrics, graders, transformations, weights, aggregation, thresholds, and dashboards: [Metrics rubrics and aggregation rules]
- Training, tuning, selection, prompting, vendor, team, and release incentives tied to the metric: [System incentives and optimization process]
- Historical runs, production outcomes, slice results, reviewer evidence, downstream measures, and confidence information: [Evaluation production and outcome evidence]
- Incidents, edge cases, suspicious improvements, regressions, complaints, and populations where the proxy may fail: [Known failure cases and affected slices]
- Accepted error, non-compensable harms, evaluation owner, product owner, risk reviewer, and release owner: [Risk tolerance and accountable owners]

Do not infer gaming solely from a score improvement. Do not assume correlation proves the proxy causes or predicts the target outcome. Do not invent raw data, incentives, or counterfactual results. Separate deliberate gaming, optimization pressure, specification error, measurement error, distribution shift, label leakage, and aggregation artifacts.

Audit:

1. Define the target-proxy contract.
   State the real outcome or invariant, why the metric is expected to represent it, the decision it authorizes, and conditions under which that relationship may fail.

2. Trace the metric construction.
   Review data source, sampling, labels, rubric, judge, transformations, weights, exclusions, missing data, threshold, and aggregation. Identify degrees of freedom available to the optimized system or team.

3. Map incentives and attack surfaces.
   Identify who or what adapts to the metric, information exposed during optimization, repeated test use, feedback leakage, and ways to improve the score without improving the target. Keep hypothetical paths separate from observed behavior.

4. Test proxy validity from supplied evidence.
   Compare metric movement with direct outcome evidence, independent reviewer evidence, protected-slice results, incidents, and downstream measures. Look for divergence, saturation, nonlinearity, reversal, and improvements caused by case exclusion or aggregation.

5. Find compensating and hidden failures.
   Determine whether gains in common cases mask severe losses in rare slices; whether averages hide abstention, refusal, verbosity, cost, latency, or safety failures; and whether one metric rewards behavior another penalizes.

6. Classify findings.
   Use Valid within tested range, Weak proxy, Broken proxy, Gameable but controlled, Actively distorted, or Not assessable. State evidence and decision consequence for each metric/slice.

7. Design a safer measurement contract.
   Recommend direct measures, counter-metrics, holdouts, blind review, refreshed cases, protected-slice gates, independent outcome sampling, or reduced decision authority. Prefer the smallest change that restores decision integrity.

8. Make a use decision.
   Choose Retain, Retain with counter-metrics, Restrict to named scope, Recalibrate, Replace, or Suspend. Name the owner and evidence required before broader use.

Required deliverable:

# Evaluation Metric Gaming and Proxy Failure Audit

## Target-Proxy Contract
| Target outcome/invariant | Proxy metric | Decision use | Assumed relationship | Known limit |
|---|---|---|---|---|

## Construction and Incentive Map
| Metric component | Evidence/source | Optimization exposure | Manipulation or distortion path | Observed/hypothetical |
|---|---|---|---|---|

## Proxy Validity Findings
| Metric/slice | Direct outcome evidence | Proxy movement | Divergence | Classification | Decision impact |
|---|---|---|---|---|---|

## Hidden Trade-offs
| Apparent gain | Compensating loss | Slice/condition | Evidence | Non-compensable? |
|---|---|---|---|---|

## Measurement Repair
| Priority | Change | Failure addressed | Acceptance evidence | Owner |
|---|---|---|---|---|

## Metric Use Decision
- Decision:
- Authorized scope:
- Required counter-metrics or gates:
- Unsupported claims:
- Reassessment trigger:

Completion requires a documented target-proxy relationship, evidence on material divergence and incentives, and a decision boundary that prevents an aggregate score from authorizing outcomes it does not measure.

Variables to Replace

Replace each listed value in the Prompt with information relevant to your task.

  • Target outcome and decision context
  • Metrics rubrics and aggregation rules
  • System incentives and optimization process
  • Evaluation production and outcome evidence
  • Known failure cases and affected slices
  • Risk tolerance and accountable owners

How to Use This Prompt

Use Claude with metric specifications, rubrics, grader prompts, score history, optimization practices, slice results, incidents, direct outcome measures, and decision policy. Run the prompt before changing a release gate. Have the evaluation owner verify construction details and the product or risk owner decide how much authority the metric retains.

Example Use Case

A support model improves an automated helpfulness score after being tuned to give longer answers, while resolution rate falls and privacy escalations rise. The audit identifies a broken proxy, introduces outcome and protected-risk counter-metrics, and suspends the old score as a release gate.

Was this useful?

Build stronger AI systems

Use Amo.ng prompts as reusable building blocks, then go deeper with RichlyAI.

Used in Workflows

Browse Workflows

Related Prompts

Browse all