Identify stale, time-sensitive, or decision-unsafe knowledge assets by linking source authority, change events, usage, reviews, and retrieval exposure.
Updated Aug 19, 2026
Determine where the enterprise knowledge corpus has become stale, time-sensitive, or decision-unsafe by relating source authority, change events, usage, review evidence, and retrieval exposure. Focus on freshness decay and revalidation evidence, not a broad knowledge-base quality audit.
Context and inputs to provide:
- Knowledge corpus name: [Knowledge corpus name]
- Corpus export or representative content set: [Corpus export or representative content set]
- Source inventory and authority rules: [Source inventory and authority rules]
- Known change events and effective dates: [Known change events and effective dates]
- Usage and retrieval exposure data: [Usage and retrieval exposure data]
- Existing review evidence and owner list: [Existing review evidence and owner list]
- Decision horizon and risk tolerance: [Decision horizon and risk tolerance]
Evidence rules:
- Separate observed evidence from inference. Label each conclusion as Observed, Inferred, or Unknown.
- Do not claim that a source, log, review, owner approval, or retrieval result was inspected unless it appears in the provided material.
- Preserve uncertainty where evidence is incomplete, contradictory, old, or sampled.
- Prefer source-effective dates, documented review records, owner attestations, change logs, and actual retrieval or usage evidence over assumptions.
- If the corpus export is incomplete, treat the result as a scoped review and state the coverage limits.
Authority boundaries:
- Do not make final compliance, legal, medical, financial, security, or customer-impact determinations unless an accountable owner has provided the relevant authority rule in the inputs.
- Assign proposed decisions to accountable roles such as knowledge owner, source authority, product owner, policy owner, compliance owner, data owner, search owner, or support operations owner.
- Where the correct owner is unclear, mark Owner Unknown and specify the evidence needed to assign accountability.
Review method:
1. Establish the freshness policy map.
- Identify content classes that require freshness controls, such as policy, product behavior, pricing, regulated guidance, operational runbooks, customer-facing answers, security procedures, data definitions, and historical reference material.
- For each class, map source authority, expected review cadence, change triggers, acceptable age, required evidence of review, and decision owner.
- If no explicit policy exists, infer a provisional freshness expectation from risk, source type, and decision use; mark it as Inferred.
2. Detect decay signals.
- Compare corpus claims against known source changes, effective dates, product or policy updates, review timestamps, ownership changes, usage patterns, and retrieval exposure.
- Identify assets with missing source authority, expired review windows, superseded references, conflicting versions, orphaned ownership, high retrieval exposure, or recent source changes without revalidation.
- Distinguish content that is merely old from content that is decision-unsafe.
3. Build the decay-risk register.
For each material risk, provide:
- Asset or content group
- Decision or workflow it may affect
- Decay signal observed
- Source authority status
- Last known review evidence
- Retrieval or usage exposure
- Potential consequence if used as-is
- Confidence level and evidence basis
- Accountable owner
- Recommended disposition: Retain, Revalidate, Update, Restrict, Retire, or Split/Merge
4. Build the source-change impact matrix.
- List each known source change or effective-date event.
- Map the affected corpus assets, topics, audiences, workflows, and retrieval surfaces.
- Identify whether the corpus reflects the change, partially reflects it, conflicts with it, or lacks enough evidence to determine impact.
- Prioritize changes that affect high-exposure retrieval, customer-facing guidance, regulated decisions, operational safety, financial terms, security controls, or contractual commitments.
5. Produce the revalidation queue.
- Rank assets by decision risk, retrieval exposure, source-change proximity, review age, and owner availability.
- For each item, specify the minimum evidence needed to clear or confirm the risk.
- Assign an accountable owner and a practical next step.
- Separate urgent restrictions from routine revalidation work.
6. Make retire, restrict, update, or revalidate decisions.
- Recommend a disposition for each priority item.
- Use Restrict when content could cause harmful or materially wrong decisions before validation is complete.
- Use Retire when content is superseded, duplicate, ownerless with no defensible authority, or no longer tied to a valid business use.
- Use Update when authoritative replacement information is available.
- Use Revalidate when the content may still be valid but lacks current review evidence.
- Use Retain only when freshness requirements, source authority, and review evidence are adequate for the stated decision horizon.
Output format:
A. Scope and evidence coverage
- Corpus reviewed
- Materials provided
- Materials not provided but needed
- Coverage limits
- Assumptions and uncertainty
B. Freshness policy map
Create a table with columns: Content class | Source authority | Review cadence or trigger | Acceptable age | Required review evidence | Decision owner | Basis: Observed/Inferred/Unknown.
C. Decay-risk register
Create a table with columns: Priority | Asset or group | Affected decision/workflow | Decay signal | Source authority status | Review evidence | Retrieval/usage exposure | Consequence | Confidence | Owner | Recommended disposition.
D. Source-change impact matrix
Create a table with columns: Source change | Effective date | Affected assets/topics | Exposure surface | Alignment status | Risk level | Required action | Owner.
E. Revalidation queue
Create a table with columns: Queue rank | Asset or group | Reason for queueing | Minimum evidence needed | Proposed verifier | Target decision | Urgency | Dependency.
F. Retire/restrict/update decisions
Create a table with columns: Asset or group | Decision | Rationale | Evidence basis | Owner to confirm | Completion check.
G. Missing evidence and unresolved questions
List the specific missing logs, source records, review attestations, authority rules, owner assignments, or retrieval evidence needed to convert Unknown or Inferred judgments into Observed judgments.
H. Completion criteria
State whether this review is complete enough to support immediate restriction, update, retirement, or revalidation planning. Identify which decisions require confirmation by the relevant knowledge owner, source authority, product owner, compliance owner, data owner, or search owner before implementation.
Diagnose where AI capability loses operational value across adoption, workflow, quality, controls, capacity, rework, and measurement.
Updated Aug 19, 2026
Diagnose why measured AI capability is not becoming realized operational value for the process below. Separate adoption leakage, workflow leakage, quality leakage, control leakage, capacity leakage, rework leakage, and measurement leakage. Do not produce a generic AI adoption roadmap or ROI plan; the job is to explain the value gap and define evidence-backed recovery actions.
Input to use:
- Business process or function: [Business process or function]
- AI capability or system introduced: [AI capability or system introduced]
- Expected value case: [Expected value case]
- Actual operating results: [Actual operating results]
- Adoption and workflow evidence: [Adoption and workflow evidence]
- Quality, control, and rework evidence: [Quality, control, and rework evidence]
- Accountable owners and operating constraints: [Accountable owners and operating constraints]
Evidence rules:
- Treat provided facts, metrics, dates, artifacts, and firsthand observations as evidence.
- Label estimates, explanations, and causal links as inference unless directly evidenced.
- Do not claim access to systems, logs, tests, approvals, user interviews, dashboards, or financial records unless they are included in the input.
- Preserve uncertainty. State what is unknown, what evidence would reduce uncertainty, and whether the missing evidence blocks a decision.
- If the evidence is too thin to quantify a loss, provide a bounded qualitative assessment and specify the minimum evidence needed to quantify it.
Diagnostic method:
1. Define the expected value flow from available AI capability to realized operational value.
2. Identify the observable handoffs where value can leak: user adoption, task selection, workflow integration, output quality, review controls, throughput capacity, exception handling, rework, measurement, and benefit capture.
3. Compare the expected value case against actual operating results. Separate volume effects, quality effects, cycle-time effects, labor-mix effects, risk/control effects, and measurement effects where possible.
4. Identify root leakage mechanisms before proposing interventions. Do not assume low adoption is the only cause.
5. Distinguish symptoms from mechanisms. Example: “low usage” is a symptom; “users avoid the tool because outputs require unplanned specialist review” is a leakage mechanism if supported by evidence.
6. Consider whether measured AI capability is being trapped by downstream constraints, approval queues, exception rates, policy limits, data quality, incentive conflicts, training gaps, workflow redesign gaps, or unmanaged rework.
7. Preserve existing operational constraints. Do not recommend broad transformation unless the evidence shows incremental recovery actions cannot address the leakage.
Required deliverable:
1. Value-flow map
Create a concise map showing:
- AI capability available
- Intended user or workflow touchpoint
- Expected behavior change
- Expected operational effect
- Expected financial, service, risk, or capacity value
- Actual observed path
- Known breakpoints or uncertain handoffs
2. Leakage mechanism ledger
Provide a ledger with these columns:
- Leakage category
- Observed evidence
- Inference about mechanism
- Value affected
- Likely owner
- Confidence level: High, Medium, or Low
- Missing evidence
- Immediate diagnostic check
Use these leakage categories where relevant:
- Adoption leakage
- Workflow leakage
- Quality leakage
- Control or approval leakage
- Capacity leakage
- Rework leakage
- Measurement leakage
- Benefit capture leakage
3. Evidence-backed loss bridge
Build a bridge from expected value to realized value. Use available numbers where provided. If numbers are missing, use directional bands and explain the basis.
Include:
- Starting expected value
- Loss or dilution from adoption
- Loss or dilution from workflow fit
- Loss or dilution from quality or rework
- Loss or dilution from controls or review load
- Loss or dilution from capacity constraints
- Loss or dilution from measurement or attribution error
- Realized value observed
- Unexplained residual gap
4. Intervention hypotheses
For each major leakage mechanism, define a testable recovery hypothesis:
- Hypothesis
- Evidence that supports it
- Evidence that contradicts or weakens it
- Smallest safe intervention
- Expected movement in adoption, quality, cycle time, cost, risk, or capacity
- Leading indicator to monitor
- Timeframe for signal
- Risk if wrong
5. Owner-specific recovery plan
Create an owner-specific plan using the named owners or the most likely accountable roles from the input. Include only actions supported by the diagnosis.
For each action, specify:
- Accountable owner, such as process owner, product owner, operations owner, finance owner, data owner, risk/control owner, or enablement owner
- Decision required
- Evidence required before action, if any
- Recovery action
- Completion criterion
- Verification method
- Dependency or constraint
6. Decision view
Conclude with:
- Most likely primary leakage mechanism
- Secondary leakage mechanisms
- Whether the current evidence supports intervention, further diagnosis, or pausing expansion
- Highest-value next diagnostic check
- Decisions each accountable owner should make next
- What would change your conclusion
Completion criteria:
- The diagnosis explains the gap between available AI capability and realized operational value, not merely whether users adopted the tool.
- Every material claim is tied to evidence, inference, or stated uncertainty.
- The recovery plan assigns accountable owners and observable completion criteria.
- No recommendations depend on unavailable inspections, unverified approvals, or assumed system behavior.
Attribute full AI operating costs to accepted outcomes so owners can compare true unit economics across workflows and variants.
Updated Aug 19, 2026
Build a defensible AI operating cost attribution and unit economics model for the workflow below. The purpose is to attribute model, tool, retrieval, infrastructure, review, support, failure, and rework costs to accepted AI outcomes so accountable owners can compare true unit economics across workflows and variants.
Context to provide:
- AI workflow or product: [AI workflow or product]
- Accepted outcome definition: [Accepted outcome definition]
- Analysis period: [Analysis period]
- Workflow variants to compare: [Workflow variants to compare]
- Cost evidence and usage exports: [Cost evidence and usage exports]
- Quality and rework evidence: [Quality and rework evidence]
- Accountable owners: [Accountable owners]
Evidence rules:
- Use only the evidence provided. Do not claim that a source, export, invoice, log, ticket, approval, or test was inspected unless it is included in the input.
- Separate observed facts, calculated values, assumptions, and inferences.
- If evidence is missing, label it as missing and explain how it affects confidence or comparability.
- Preserve uncertainty with ranges or confidence notes where point estimates are not supported.
- Treat the accepted-outcome definition, analysis period, cost and usage sources, and comparable workflow boundaries as blocking when they are missing or materially conflicting. Request all blocking items in one consolidated clarification and do not calculate the affected unit economics until they are resolved. Continue with non-blocking gaps only as explicit Unknown or provisional inputs, with their effect on formulas, confidence, and comparability.
- Do not create a general AI spend governance brief or a quality/latency experiment plan. Stay focused on cost attribution and accepted-outcome unit economics.
Build the working artifact with these sections:
1. Cost Boundary Map
Create a boundary map for the selected [AI workflow or product] covering the full operating cost chain for [Analysis period]. Include, where evidenced:
- Model inference or subscription costs
- Tool execution costs
- Retrieval, embedding, vector database, search, or storage costs
- Application, orchestration, logging, monitoring, and infrastructure costs
- Human review, approval, escalation, and exception-handling costs
- Support, customer operations, and internal operations costs
- Failure, incident, reversal, refund, remediation, and rework costs
- Evaluation, sampling, QA, audit, and oversight costs directly tied to production operation
For each cost category, state whether it is included, excluded, partially included, or missing; identify the cost owner from [Accountable owners] where possible; and explain the rationale.
2. Accepted-Outcome Measurement
Define the denominator for unit economics using [Accepted outcome definition]. Specify:
- What counts as an accepted outcome
- What does not count
- How retries, duplicates, partial completions, escalations, rejected outputs, and reworked outputs should be treated
- Any quality threshold required before an outcome is counted
- The evidence source used for the count
If the accepted-outcome count cannot be established from [Quality and rework evidence], provide a defensible interim method and list the validation needed from the product owner, finance owner, or operations owner.
3. Allocation Rules
Design allocation rules that assign costs to accepted outcomes and to [Workflow variants to compare]. For each rule, include:
- Cost pool
- Allocation driver
- Formula
- Required evidence
- Owner responsible for validating the driver
- Known weakness or bias
- When the rule should be replaced with a more precise method
Prefer causal allocation drivers over broad averages. Use broad averages only when more precise evidence is unavailable, and label them clearly.
4. Accepted-Outcome Unit Economics
Create a unit economics table for each workflow variant in [Workflow variants to compare]. Include:
- Accepted outcome volume
- Gross operating cost
- Cost excluded or not yet evidenced
- Cost per accepted outcome
- Cost per attempted outcome, if attempt counts are available
- Human review cost per accepted outcome
- Failure and rework cost per accepted outcome
- Infrastructure and retrieval cost per accepted outcome
- Confidence level for each major figure
Show formulas and make assumptions explicit. Do not overstate precision.
5. Variance Bridge
Build a variance bridge explaining differences in cost per accepted outcome across variants or periods. Attribute variance where evidence supports it to:
- Volume and utilization
- Model selection or token/input-output mix
- Tool calls and external service usage
- Retrieval depth, indexing, or storage patterns
- Review rates and escalation rates
- Failure, rejection, incident, or rework rates
- Infrastructure utilization or fixed-cost absorption
- Support burden
For each variance driver, state whether the driver is observed, calculated, inferred, or currently unverified.
6. Cost-Quality Sensitivity
Analyze how unit economics change under plausible changes to quality and control variables, such as:
- Acceptance rate
- Human review rate
- Escalation rate
- Rework rate
- Failure or incident rate
- Retrieval depth or tool usage
- Model choice or routing mix
Use ranges when evidence is incomplete. Identify which variables most affect cost per accepted outcome and which quality controls appear economically justified.
7. Control Decisions
Provide a decision table for the accountable owners. Include:
- Decision under consideration
- Economic rationale
- Quality or risk tradeoff
- Evidence supporting the decision
- Evidence still missing
- Owner who should approve or validate the decision
- Completion check
Focus on concrete controls such as routing changes, review thresholds, retrieval limits, escalation criteria, failure handling, logging improvements, or measurement changes.
8. Model Integrity Checks
Before finalizing, perform these checks using the provided evidence:
- Every included cost has an owner, source, allocation rule, and treatment in the model.
- Accepted-outcome counts match the stated definition or are flagged as provisional.
- Rejected, failed, duplicated, escalated, and reworked outputs are not accidentally counted as accepted outcomes unless justified.
- Fixed, variable, and step costs are not mixed without explanation.
- Variant comparisons use the same boundary unless differences are explicitly disclosed.
- No unavailable inspection, execution, approval, or source verification is claimed.
Final output format:
- Cost boundary map
- Allocation rule table
- Accepted-outcome unit economics table
- Variance bridge
- Cost-quality sensitivity table
- Control decision table
- Missing evidence register
- Owner verification checklist
Write in direct finance and operating language suitable for review by the finance owner, product owner, operations owner, and data owner. Avoid generic governance language and unsupported certainty.
Convert pilot and operating evidence into a defensible decision on whether one AI initiative should scale, hold, redesign, or stop.
Updated Aug 19, 2026
Prepare a decision brief for one AI initiative. The objective is to recommend scale, hold, redesign, or stop using only the evidence provided, while making uncertainty, missing information, and decision authority explicit.
Inputs to use:
- Initiative name: [Initiative name]
- Decision owner: [Decision owner]
- Pilot or operating scope and timeframe: [Pilot or operating scope and timeframe]
- Evidence package: [Evidence package]
- Current constraints and non-negotiables: [Current constraints and non-negotiables]
- Decision options under consideration: [Decision options under consideration]
- Next decision date or trigger: [Next decision date or trigger]
Evidence discipline:
- Distinguish observed evidence from inference. Do not treat anecdotes, forecasts, vendor claims, or unverified internal estimates as confirmed facts.
- Do not claim that a source, approval, test, control, incident review, financial model, or system inspection was completed unless the evidence package shows it.
- If evidence is missing, state what is missing, why it matters, and the minimum evidence needed for a safer decision.
- Preserve uncertainty. Use confidence levels only when tied to evidence quality.
- Stay within decision-support authority: produce a recommendation and decision conditions; do not imply approval unless the accountable owner has already approved it in the evidence.
- Avoid generic AI governance commentary. Focus on this initiative’s value, adoption, controls, reliability, dependencies, economics, risk boundaries, and reversibility.
Decision standard:
Recommend exactly one primary disposition: Scale, Hold, Redesign, or Stop.
Use these meanings:
- Scale: expand use because value, adoption, controls, reliability, economics, dependencies, and reversibility are acceptable within defined boundaries.
- Hold: continue limited operation or pause expansion because evidence is incomplete or conditions are not yet met, but the initiative may still be viable.
- Redesign: materially change workflow, model approach, controls, operating model, vendor setup, data inputs, or user experience before further expansion.
- Stop: end or sunset the initiative because evidence does not support continued investment or risk is unacceptable relative to value and reversibility.
Deliver the brief in the following format:
1. Decision frame
- Initiative: state the initiative and operating scope.
- Decision needed: state the decision being made now.
- Accountable owner: identify the decision owner and any role-specific verifiers needed, such as product owner, data owner, security reviewer, legal/compliance owner, finance owner, operations owner, or release owner.
- Authority boundary: state what this brief can recommend versus what requires owner approval.
- Time boundary: state the next decision date or trigger.
2. Decision evidence map
Create a table with these columns: Decision dimension, Observed evidence, Inference or assumption, Evidence strength, Missing information, Decision implication.
Include these dimensions at minimum:
- Intended business value
- Realized value or leading indicators
- User adoption and workflow fit
- Output quality and reliability
- Control effectiveness and exception handling
- Data, model, vendor, and system dependencies
- Security, privacy, compliance, and policy constraints
- Operating support and ownership capacity
- Scale economics and marginal cost
- Reversibility, rollback, and exit cost
3. Gate-by-gate disposition
For each gate, assign Pass, Conditional Pass, Fail, or Insufficient Evidence. Explain the reason in 2 to 4 sentences per gate.
Gates:
- Value gate: evidence of meaningful value relative to effort and alternatives.
- Adoption gate: users, operators, or customers can and do use it in the intended workflow.
- Control gate: risks, permissions, review paths, exceptions, and accountability are workable.
- Reliability gate: quality, availability, latency, and failure modes are acceptable for the use case.
- Dependency gate: critical data, vendor, model, integration, and staffing dependencies are known and manageable.
- Economics gate: scale costs, support costs, and expected benefits remain acceptable beyond the pilot.
- Reversibility gate: the initiative can be rolled back, contained, or sunset without unacceptable disruption.
4. Counterfactual options
Compare the realistic options, not just the preferred one. Include at least:
- Scale now
- Hold in current scope
- Redesign before expansion
- Stop or sunset
- Non-AI or lower-automation alternative
For each option, provide: what would happen, expected upside, main downside, investment or effort required, risk exposure, reversibility, and what evidence would make this option stronger or weaker.
5. Scale economics and risk boundaries
Provide a practical scale boundary, even if the recommendation is not to scale.
Include:
- Unit or marginal cost drivers, using provided evidence only.
- Expected cost changes at expanded volume.
- Support, monitoring, exception handling, and owner workload implications.
- Benefits that are evidenced versus speculative.
- Risk ceilings that should not be exceeded.
- Required controls before expansion.
- Stop-loss triggers or rollback conditions.
- Any financial assumptions that the finance owner should verify before approval.
6. Recommendation
State one primary recommendation: Scale, Hold, Redesign, or Stop.
Then provide:
- Rationale tied to the gates and evidence map.
- Conditions required before execution.
- Evidence that argues against the recommendation.
- Residual risks the owner would knowingly accept.
- What would change the recommendation.
7. Authorized next-decision brief
Create an owner-ready next-decision plan with:
- Decision to be requested from [Decision owner].
- Roles that should verify specific parts of the brief before action, such as finance owner for economics, security reviewer for security boundaries, data owner for data use, product owner for workflow value, operations owner for support readiness, and legal/compliance owner where applicable.
- Immediate actions, responsible role, due date or trigger, and required evidence of completion.
- Metrics or observations to collect before the next decision.
- Completion checks that would show the recommendation has been executed or is ready for escalation.
- Explicit statement of any unresolved blockers.
Final quality check before answering:
- The brief decides the fate of one operating initiative; it does not prioritize a portfolio.
- Observations and inferences are clearly separated.
- Missing information is visible and decision-relevant.
- The recommendation is bounded by economics, risk, controls, dependencies, and reversibility.
- No unavailable inspection, approval, test, or execution is claimed.
Reconcile an approved AI business case against post-deployment operational and financial evidence to produce a defensible benefits realization record.
Updated Aug 19, 2026
Reconcile the approved AI business case with post-deployment operational and financial evidence. Produce a defensible working artifact that accountable owners can use to decide which benefits were realized, displaced, delayed, double-counted, or unsupported.
Context and inputs to use:
- Approved AI business case: [Approved AI business case]
- Deployment period and relevant measurement window: [Deployment period]
- Baseline definition, assumptions, volumes, rates, and counterfactual used in the original case: [Baseline definition and assumptions]
- Post-deployment operational evidence, including KPI extracts, process measures, adoption data, service levels, error rates, cycle times, throughput, quality data, or control logs: [Post-deployment operational evidence]
- Financial actuals and cost data, including labor, vendor, cloud, tooling, support, implementation, training, rework, run-rate, and one-time costs: [Financial actuals and cost data]
- Known external changes or confounders, including demand shifts, pricing changes, policy changes, staffing changes, process redesign, vendor changes, macro factors, seasonality, or parallel initiatives: [Known external changes or confounders]
- Benefit owners and decision authority, including finance owner, product owner, operational owner, data owner, and any approval forum: [Benefit owners and decision authority]
Evidence discipline:
- Use only the evidence provided. Do not claim that any source system, report, approval, test, audit, or transaction record was inspected unless it is included in the inputs.
- Separate observed evidence from inference. Label assumptions, estimates, and judgment calls explicitly.
- Preserve uncertainty. Where evidence is incomplete, state what is missing, why it matters, and how it affects confidence.
- Do not redesign the pre-pilot ROI plan. This is a post-implementation benefits realization bridge against the approved case.
- Do not treat adoption, usage, model output volume, or automation counts as financial benefit unless the operational-to-financial conversion is evidenced or reasonably supported.
- Do not recognize the same benefit twice across labor, productivity, capacity, revenue, cost avoidance, quality, or risk categories.
- Do not assign causality to the AI initiative where the evidence only supports correlation or partial contribution.
Working method:
1. Extract the original benefit claims from the approved business case.
- Identify each promised benefit, metric, baseline, target, timing, owner, financial value, and stated assumption.
- Preserve the original wording where possible.
2. Build a benefit lineage register.
For each claimed benefit, trace:
- Original claim
- Business case source or section, if provided
- Baseline metric and value
- Target metric and value
- Actual post-deployment metric and value
- Operational evidence used
- Financial evidence used
- Conversion method from operational movement to financial value
- Accountable owner
- Evidence gaps
- Preliminary status: realized, partially realized, displaced, delayed, double-counted, unsupported, or not yet measurable
3. Build a baseline-to-actual bridge.
For each material benefit, show the movement from baseline to actual:
- Baseline value
- Business case target
- Actual observed value
- Absolute movement
- Percentage movement
- Timing variance versus expected realization date
- Volume, rate, mix, quality, and adoption effects where evidenced
- External or confounding factors that may explain part of the movement
4. Calculate quality-adjusted realized benefit.
For each benefit with enough evidence, calculate or estimate:
- Claimed business case benefit
- Gross observed benefit before adjustments
- Timing adjustment for delayed or accelerated realization
- Quality adjustment for error, rework, customer impact, control failures, or service degradation
- Attribution adjustment for non-AI drivers and confounders
- Displacement adjustment where savings moved cost, effort, risk, or workload elsewhere
- Double-count exclusion where the same value appears in multiple benefit lines
- Net recognized realized benefit
- Confidence level: high, medium, low, or unsupported
Show the calculation logic in plain language. If exact calculation is not possible, provide a bounded estimate only if the evidence supports the bounds; otherwise mark the benefit unsupported and explain the missing evidence.
5. Define attribution limits.
- Identify which benefits can reasonably be attributed to the AI initiative, which are only partially attributable, and which cannot be attributed based on the evidence.
- Explain the strongest alternative explanations for observed changes.
- Identify any benefits that appear to be enabled by AI but realized through other changes such as process redesign, headcount decisions, pricing, demand changes, or manual workarounds.
6. Prepare the benefits realization decision record.
Include:
- Decision required from the finance owner and relevant benefit owners
- Recommended realization status for each benefit
- Net recognized benefit total, separated from unsupported or delayed benefits
- Costs included and costs excluded, with rationale
- Material caveats and unresolved evidence gaps
- Required owner confirmations before the record is used externally
- Follow-up actions, owner, and due date where evidence is missing or benefits are delayed
Output format:
A. Evidence Boundary Note
- State what evidence was provided.
- State what was not provided but would materially improve confidence.
- State the measurement window used.
- State any limits on causality, completeness, or financial recognition.
B. Benefit Lineage Register
Provide a table with these columns:
- Benefit ID
- Original benefit claim
- Benefit category
- Original baseline
- Original target
- Expected realization timing
- Actual evidence observed
- Financial evidence observed
- Accountable owner
- Evidence gap
- Proposed status
C. Baseline-to-Actual Bridge
Provide a table with these columns:
- Benefit ID
- Baseline
- Target
- Actual
- Movement versus baseline
- Movement versus target
- Timing variance
- Operational driver evidenced
- Confounders or external changes
- Bridge conclusion
D. Quality-Adjusted Benefit Calculation
Provide a table with these columns:
- Benefit ID
- Claimed benefit value
- Gross observed value
- Timing adjustment
- Quality adjustment
- Attribution adjustment
- Displacement adjustment
- Double-count exclusion
- Net recognized benefit
- Confidence level
- Calculation notes
E. Attribution Limits and Unsupported Claims
Separate into:
- Benefits strongly supported by evidence
- Benefits partially supported or partially attributable
- Benefits delayed or not yet measurable
- Benefits displaced to another cost, team, risk, or workload
- Benefits double-counted or overlapping
- Benefits unsupported by the provided evidence
F. Benefits Realization Decision Record
Provide:
- Recommended decision: recognize, partially recognize, defer, reject, or escalate
- Net recognized benefit total
- Deferred benefit total
- Unsupported benefit total
- Key reasons for the decision
- Required confirmations from finance owner, product owner, operational owner, data owner, or other named accountable owners
- Open evidence requests
- Risks if the organization uses the benefit claim without resolving gaps
Completion checks before finalizing:
- Every recognized benefit traces back to an approved business case claim and at least one post-deployment evidence item.
- Operational movements are not converted into financial value without an explicit conversion method.
- Delayed, displaced, double-counted, and unsupported benefits are not included in the recognized total unless clearly justified.
- Observations, assumptions, and inferences are visibly separated.
- The final decision record is suitable for review by the finance owner and named benefit owners, but does not claim their approval unless it is included in the evidence.
Checks AI-generated code for hallucinated packages, wrong versions, unsupported APIs, and framework claims before merge.
Updated Aug 19, 2026
Verify dependency, package, framework, and API claims introduced by AI-generated code before merge.
Context to provide:
- Repository scope and task: [Repository scope and task]
- Relevant files and instructions: [Relevant files and instructions]
- Observed evidence: [Observed evidence]
- Constraints and authorized changes: [Constraints and authorized changes]
- Environment details without secrets: [Environment details without secrets]
- Verification commands and acceptance criteria: [Verification commands and acceptance criteria]
Objective:
Determine whether generated code relies on packages, versions, imports, methods, configuration keys, framework behavior, runtime features, or API signatures that are unsupported by the repository’s installed dependency set or by authoritative documentation. Produce a defensible review artifact for the release owner or package maintainer before merge.
Scope boundaries:
- Focus only on dependency, package, framework, runtime, type, and API claims introduced or materially affected by the generated code.
- Do not perform a broad dependency upgrade, architectural rewrite, style review, or generic code review.
- Do not assume an API exists because it appears plausible.
- Do not claim documentation, commands, tests, approvals, or files were inspected unless you actually inspected them.
- Separate direct observations from inferences and mark anything unverified.
- If authoritative documentation is unavailable, use installed package source, generated types, local docs, lockfiles, manifests, and compiler or test output where available; otherwise classify the claim as Unresolved and state what evidence is needed.
- Use exactly these claim outcomes:
- Verified: admissible evidence confirms the claim for the installed or target version.
- Contradicted: admissible evidence shows the claim is false or incompatible.
- Unresolved: relevant evidence is missing, inaccessible, incomplete, or conflicting, so no conclusion is supportable yet.
- Unsupported: no admissible evidence supports the assertion after the available repository and authoritative sources are checked.
Do not collapse Unresolved and Unsupported into a generic warning. Preserve uncertainty and identify the evidence needed to resolve each open claim.
Missing-input gate:
- Treat the generated change, applicable repository instructions, dependency manifests or lockfiles, target runtime or package version, and relevant API evidence as blocking when their absence or conflict prevents a claim from being scoped. Request all blocking items in one consolidated clarification and leave the affected claim Unresolved until they are supplied.
- Continue with non-blocking gaps only when each is recorded as Unknown or Unresolved, with the evidence needed and the consequence for merge or release confidence.
Required process:
1. Inspect relevant files first.
- Review the generated diff or branch.
- Inspect dependency manifests, lockfiles, package manager configuration, runtime configuration, framework configuration, relevant imports, generated or installed type definitions, and nearby usage patterns.
- Identify every dependency or API claim introduced by the generated code before deciding whether edits are needed.
2. Build a claim inventory.
Include claims such as:
- Package or framework is available.
- A package version supports a named API, export, method, hook, decorator, CLI option, configuration key, schema field, or runtime behavior.
- Import paths, module formats, peer dependencies, plugins, adapters, or provider names are valid.
- Type signatures, return values, error shapes, async behavior, or environment requirements match the generated code.
3. Verify each claim against evidence.
Use the strongest available evidence in this order where practical:
- Repository manifests and lockfiles.
- Installed package source or type definitions.
- Existing repository usage and tests.
- Package manager, compiler, typechecker, linter, or framework diagnostics.
- Authoritative documentation or release notes supplied or accessible in the environment.
4. Identify root cause before editing.
For each Contradicted, Unresolved, or Unsupported claim, determine whether the issue is caused by hallucinated API usage, wrong package name, incompatible installed version, missing peer dependency, incorrect import path, runtime mismatch, stale documentation, incomplete local install, or insufficient evidence.
5. Correction policy.
- Prefer no code edits unless a minimal correction is clearly justified by evidence.
- If editing is necessary and within scope, apply the smallest safe change.
- Preserve existing behavior and public interfaces unless the merge or release context explicitly authorizes a change.
- Avoid broad rewrites, opportunistic refactors, dependency upgrades, and speculative migrations.
- If the safest fix requires a dependency upgrade or product decision, do not perform it silently; document the decision required from the release owner, package maintainer, or security reviewer.
6. Verification.
- Run syntax checks, type checks, targeted tests, or package manager inspection commands where available and appropriate.
- If a command cannot be run, state why and list the verification gap.
- Mention verification results exactly: command, outcome, and relevant error excerpt or confirmation.
Required deliverable:
A. Claim inventory
Provide a table with:
- Claim ID
- Generated-code location
- Claim being made
- Dependency, framework, runtime, or API involved
- Why the claim matters for merge safety
B. Version-and-source verification matrix
Provide a table with:
- Claim ID
- Installed or resolved version observed
- Evidence source inspected, with file path, lockfile entry, type definition, package source, command output, or documentation reference
- Claim status: Verified, Contradicted, Unresolved, Unsupported, or Not applicable
- Notes distinguishing observation from inference
C. Contradicted, unresolved, or unsupported claims
For each Contradicted, Unresolved, or Unsupported claim, include:
- Finding title
- Severity for merge: Blocker, High, Medium, or Low
- Direct evidence
- Root cause
- Expected failure mode
- Confidence level
- Missing evidence, if any
D. Minimal correction plan
For each finding, provide:
- Smallest safe correction
- Whether code change, dependency decision, documentation check, or owner decision is needed
- Files likely affected
- Behavior expected to remain unchanged
- Risk of the correction
E. Changes made, if any
- List every file changed.
- Summarize the exact purpose of each change.
- If no files were changed, state: No files changed during this verification pass.
F. Reproducible verification record
Include:
- Files inspected
- Commands run, if any
- Tests or checks run, if any
- Results observed
- Checks not run and why
- Open questions for the accountable owner named in [Merge or release context and accountable owner]
Completion criteria:
- Every generated dependency or API claim in scope is inventoried.
- Each claim is classified as Verified, Contradicted, Unresolved, Unsupported, or Not applicable using the stated evidence rules.
- Every Unresolved or Unsupported claim identifies the missing evidence, responsible evidence source or owner where known, and the consequence of proceeding without resolution.
- Unsupported claims have root cause and expected failure mode.
- No broad upgrades or rewrites are proposed as the default fix.
- Verification record is sufficient for the release owner, package maintainer, or security reviewer to reproduce or challenge the conclusion.
Review a coding-agent change set against its instructions, transcript, diff, and test evidence to determine whether the agent’s completion claims are supportable.
Updated Aug 19, 2026
Review the coding-agent change set and run evidence for provenance, instruction compliance, and support for completion claims. Focus on what can be observed from the supplied repository, diff, transcript, and test evidence. Do not perform a generic PR review unless it is necessary to attribute a material change or assess whether a completion claim is supported.
Context to provide:
- Repository scope and task: [Repository scope and task]
- Relevant files and instructions: [Relevant files and instructions]
- Observed evidence: [Observed evidence]
- Constraints and authorized changes: [Constraints and authorized changes]
- Environment details without secrets: [Environment details without secrets]
- Verification commands and acceptance criteria: [Verification commands and acceptance criteria]
Review rules:
1. Inspect the relevant files, diffs, transcript, and test evidence first before drawing conclusions.
2. Distinguish observation from inference. Mark unsupported assumptions as missing information.
3. Do not claim that a file, command, test, approval, system, or external source was inspected or completed unless there is evidence in the provided materials or you actually inspected or ran it in the available environment.
4. Attribute each material change to one of these categories:
- Directly instructed
- Reasonably necessary to satisfy the instruction
- Incidental but explainable
- Unexplained drift
- Potentially harmful or out of scope
5. Treat a change as material if it affects behavior, public API, data model, security posture, dependency surface, build/test configuration, generated artifacts, migrations, operational behavior, or protected behavior listed by the release owner.
6. Preserve existing behavior as the default expectation. Flag behavior changes that are not explicitly instructed or clearly necessary.
7. Avoid broad rewrites and style-only judgments unless they obscure attribution, create risk, or conflict with project conventions.
8. If you identify a defect and decide to edit code, first identify the likely root cause, then apply the smallest safe change that preserves existing behavior. Do not perform broad rewrites. Run syntax checks and tests where available, then summarize files changed and verification results. If editing is not requested or not safe, provide proposed changes only.
9. Run or recommend syntax checks and tests where available and proportionate. If you cannot run them, state exactly what evidence is missing and what the test owner or release owner should run.
10. Do not recommend merge solely because the agent said the work was complete. Completion claims must be reconciled against observable diff and test evidence.
Deliverable:
## 1. Review Scope and Evidence Used
List the materials actually inspected:
- Instructions reviewed
- Transcript or completion notes reviewed
- Files or diffs reviewed
- Tests, logs, or command outputs reviewed
- Repository context used
- Evidence not provided or not inspectable
## 2. Agent Completion Claims
Create a table with these columns:
- Claim made by agent
- Evidence offered by agent
- Evidence independently visible in supplied materials or environment
- Supported, partially supported, unsupported, or contradicted
- Notes for release owner or test owner
## 3. Instruction-to-Diff Trace
Map the original instruction to the observed changes.
Use this table:
- Instruction requirement
- Related files or hunks
- How the change satisfies the requirement
- Evidence level: direct, inferred, weak, or missing
- Compliance assessment
Call out any instruction requirement that appears unimplemented, only partially implemented, or implemented through an unexpected approach.
## 4. Change Attribution Ledger
Create a ledger for each material changed area.
Use this table:
- File or component
- Material change observed
- Attribution category
- Evidence supporting attribution
- Behavior or interface impact
- Risk level: low, medium, high
- Owner who should verify: release owner, test owner, security reviewer, data owner, product owner, or other specific accountable role
## 5. Unexplained-Change Register
List all changes that cannot be clearly tied to the instruction or necessary implementation path.
For each item include:
- File or hunk
- What changed
- Why attribution is unclear
- Potential consequence
- What evidence would resolve it
- Recommended handling: accept with owner verification, revert, isolate into separate change, or investigate before merge
## 6. Test Evidence Reconciliation
Assess whether the supplied tests and logs support the completion claims.
Include:
- Commands claimed to have run
- Commands evidenced by logs or terminal output
- Pass/fail status shown by evidence
- Coverage relevance to the changed behavior
- Gaps, skipped tests, stale outputs, or ambiguous timestamps
- Additional checks the test owner should run before merge
If you run any commands, list:
- Command
- Purpose
- Result
- Relevant output summary
- Any limitations
If no commands were run, say so explicitly.
## 7. Risk and Protected Behavior Review
Evaluate only the risks that arise from attribution, instruction compliance, or evidence gaps.
Cover applicable areas:
- Behavior drift
- Public API or contract change
- Data migration or persistence risk
- Authentication, authorization, or secrets handling
- Dependency or build-system changes
- Generated files or lockfiles
- Test configuration changes
- Operational or deployment implications
## 8. Merge Recommendation
Choose one recommendation:
- Merge: evidence supports the agent’s claims and material changes are attributable
- Merge after owner verification: residual gaps are narrow and assigned to accountable owners
- Do not merge yet: unsupported claims, unexplained drift, insufficient tests, or high-risk uncertainty remain
- Rework required: changes are out of scope, unsafe, or not traceable to the instruction
Include:
- Recommendation
- Primary reasons
- Required pre-merge checks
- Owners responsible for verification
- Conditions that would change the recommendation
## 9. Completion Check
Before finalizing, confirm:
- Relevant files and evidence were inspected first
- Material changes were attributed
- Intended edits were separated from incidental or unexplained drift
- Agent claims were reconciled with test evidence
- Missing evidence and uncertainty were explicitly stated
- Any commands run are reported with results, or lack of execution is stated
- Any files changed by you are summarized, with syntax/test verification results where available
Design a control plan that detects when a once-approved AI evaluation is no longer reliable for production decisions.
Updated Aug 19, 2026
Design a production AI evaluation-drift detection plan for an already-approved evaluation. The goal is to detect when production conditions make the evaluation no longer decision-reliable.
Do not design a generic evaluation harness. Do not perform one-time model upgrade regression testing. Focus on ongoing production drift controls for an evaluation that already exists.
Context to provide:
- System or product under evaluation: [System or product under evaluation]
- Approved evaluation decision use: [Approved evaluation decision use]
- Current evaluation artifact summary: [Current evaluation artifact summary]
- Production telemetry and outcome evidence available: [Production telemetry and outcome evidence available]
- Known recent or planned changes: [Known recent or planned changes]
- Accountable owners and operating constraints: [Accountable owners and operating constraints]
- Risk tolerance or escalation policy: [Risk tolerance or escalation policy]
Evidence discipline:
- Separate observed evidence from inference.
- Do not claim logs, datasets, tests, graders, prompts, retrieval systems, production traffic, or approvals were inspected unless they are included in the provided context.
- Flag missing information that prevents firm threshold-setting or assignment of ownership.
- Preserve uncertainty where evidence is incomplete.
- If assumptions are necessary, label them as assumptions and explain how the evaluation owner should verify them.
Produce the following deliverable:
1. Evaluation reliability boundary
Define what decision the evaluation is approved to support, what production population it is intended to represent, what conditions must remain comparable, and what conditions would make the evaluation no longer decision-reliable.
2. Drift taxonomy
Create a task-specific taxonomy covering at minimum:
- Traffic or user-intent drift
- Input-format or language drift
- Label-policy or ground-truth drift
- Human reviewer or labeling-team drift
- LLM-as-grader, rubric, or judge-prompt drift
- Application prompt or system-instruction drift
- Model, provider, parameter, or routing drift
- Retrieval corpus, embedding, ranking, or freshness drift
- Tool, API, data dependency, or integration drift
- Outcome, complaint, incident, conversion, safety, or business-metric drift
For each drift type, state observable signals, likely false positives, decision impact, accountable owner, and evidence needed.
3. Sentinel and sampling plan
Specify sentinel checks and production sampling methods that can detect each meaningful drift type. Include:
- Always-on metrics versus periodic review samples
- Stratified slices that must be monitored
- Minimum viable sample size or sample logic when exact sizes cannot be justified
- Triggered sampling after incidents, launches, prompt changes, retrieval updates, model routing changes, policy changes, or unusual outcome shifts
- Treatment of low-volume but high-risk slices
- Owner responsible for sample collection, review, and documentation
4. Comparability ledger
Design a ledger that records whether the current production environment remains comparable to the approved evaluation baseline. Include ledger fields for:
- Evaluation version and approved decision use
- Baseline dataset or traffic window
- Production traffic window reviewed
- Prompt, model, retrieval, grader, rubric, label policy, and tool versions
- Known changes since approval
- Evidence source for each comparison
- Comparability status: comparable, degraded, not comparable, or unknown
- Owner attestation required from evaluation owner, product owner, data owner, security reviewer, or other named accountable role as appropriate
- Revalidation requirement and due date
5. Detection thresholds
Propose initial thresholds using the evidence provided. Where evidence is insufficient, provide threshold-setting rules instead of invented numbers. Cover:
- Statistical or distributional thresholds
- Operational thresholds
- Quality and safety thresholds
- Business or outcome thresholds
- Grader agreement or calibration thresholds
- Retrieval freshness or coverage thresholds
- Incident-based hard stops
For each threshold, state the metric, comparison baseline, trigger level, rationale, owner, action required, and expected review cadence.
6. Investigation triggers
Define when the team must investigate before continuing to rely on the evaluation. Include triggers for traffic anomalies, label disagreement, grader instability, prompt or model changes, retrieval degradation, unexplained outcome shifts, safety incidents, and stakeholder challenges to evaluation validity.
7. Revalidation triggers
Define when the approved evaluation must be refreshed, rerun, recalibrated, or retired. Include criteria for partial revalidation, full revalidation, temporary suspension of evaluation-based decisions, and formal owner signoff.
8. Runbook
Write a practical runbook with:
- Daily, weekly, monthly, and release-event checks where appropriate
- Required inputs and evidence sources
- Step-by-step investigation path
- Decision states: continue relying, rely with caveat, pause reliance, revalidate, or retire
- Communication path to evaluation owner, product owner, data owner, security reviewer, release owner, and incident owner where relevant
- Documentation artifacts to retain
9. Completion checks
End with observable completion criteria. The plan is complete only if it identifies monitored drift types, assigns accountable owners, defines sentinel and sampling mechanisms, records comparability evidence, states thresholds or threshold-setting rules, specifies investigation and revalidation triggers, and explains what evidence is still missing.
Evaluate whether a proposed automated judge is calibrated enough for a specific scoring or classification decision.
Updated Aug 19, 2026
Determine whether the proposed automated judge is sufficiently reliable and calibrated for the specified evaluation decision. Focus only on judge fitness for the stated scoring or classification decision; do not design a full evaluation platform or generic prompt evaluation harness.
Context to provide:
- Evaluation decision: [Evaluation decision]
- Judge prompt or rubric: [Judge prompt or rubric]
- Candidate outputs and source items: [Candidate outputs and source items]
- Human reference labels or adjudications: [Human reference labels or adjudications]
- Risk profile and protected attributes: [Risk profile and protected attributes]
- Acceptance thresholds: [Acceptance thresholds]
Evidence discipline:
- Separate observed evidence from inference.
- State when evidence is missing, weak, imbalanced, or not blinded.
- Do not claim that tests, files, systems, reviewers, or production behavior were inspected unless the provided material supports that claim.
- Preserve uncertainty where the sample is too small, labels are disputed, or reference judgments are not independent.
- Treat human reference labels as evidence to be assessed, not as automatically correct.
Deliverable required:
1. Evaluation decision and judge contract
Define the exact decision the judge is being asked to make:
- Decision type: classification, ordinal rating, pairwise preference, threshold pass/fail, ranking, or other.
- Intended users and accountable owner, such as evaluation owner, product owner, domain reviewer, security reviewer, or data owner.
- Inputs the judge is allowed to use.
- Inputs or knowledge the judge must not use.
- Score scale, labels, thresholds, and tie-breaking rules.
- What counts as a correct judgment versus an acceptable judgment.
- Known boundary cases where the judge’s authority should stop.
- Consequences of a false positive, false negative, over-score, and under-score.
2. Blinded calibration design
Propose a calibration design suitable for the provided decision and evidence:
- How examples should be blinded, randomized, deduplicated, and stratified.
- Minimum case coverage needed across easy cases, close calls, failures, adversarial examples, protected or sensitive attributes, and domain-specific edge cases.
- How many independent human adjudications are needed and when domain reviewer arbitration is required.
- Which metrics are appropriate and why, such as exact agreement, weighted agreement, Cohen’s kappa, Krippendorff’s alpha, rank correlation, threshold confusion matrix, false positive and false negative rates, calibration by score bucket, and self-consistency across repeated runs.
- How to avoid leakage from model identity, author identity, expected answer wording, ordering effects, or rubric hints.
- What must be held out for future regression checks.
3. Current evidence assessment
Using only the provided evidence, assess whether calibration can be judged now:
- Evidence available.
- Evidence missing.
- Sample quality concerns.
- Label quality concerns.
- Whether the provided material is enough to support a permitted-use decision.
4. Disagreement analysis
Analyze judge disagreement against reference labels or adjudications:
- Where the judge agrees reliably.
- Where disagreement clusters by score band, topic, task type, output length, language, ambiguity, or source quality.
- Whether disagreements are random, systematic, rubric-driven, or caused by unclear source material.
- Whether the judge is too lenient, too strict, overconfident, inconsistent near thresholds, or sensitive to irrelevant style features.
- Distinguish clear judge errors from cases where the human reference may be ambiguous or under-specified.
5. Bias, instability, and attack findings
Evaluate fitness risks that could invalidate the judge for the specified decision:
- Bias or disparate error patterns related to [Risk profile and protected attributes].
- Sensitivity to superficial wording, formatting, verbosity, fluency, dialect, language variety, or model identity.
- Instability across repeated judgments, order changes, paraphrases, or equivalent source presentations.
- Prompt injection exposure, including whether candidate text can influence the judge’s rubric, authority, scoring scale, or refusal behavior.
- Boundary-case behavior, including ambiguous answers, partially correct answers, missing citations, conflicting sources, unsafe but persuasive content, and cases near the pass/fail threshold.
6. Permitted-use and arbitration gate
Make a clear, bounded decision using [Acceptance thresholds]:
- Permitted use: where the judge may be used without routine human review.
- Conditional use: where the judge may assist but must be sampled, audited, or reviewed by an accountable owner.
- Prohibited use: where the judge should not be used for this decision.
- Arbitration triggers: exact conditions that require domain reviewer, security reviewer, data owner, or product owner review.
- Monitoring requirements: what should be logged, sampled, and periodically recalibrated.
- Regression triggers: what changes to the judge prompt, model, rubric, data distribution, or product policy require renewed calibration.
7. Completion check
End with a concise readiness statement:
- Fit for use, conditionally fit, or not fit for the specified evaluation decision.
- Main evidence supporting that conclusion.
- Main unresolved risks.
- Minimum additional evidence needed before expanding use.
- Accountable owner who should accept or reject the permitted-use gate.
Use precise professional language. Avoid generic AI governance commentary. Do not recommend broad rewrites of the judge unless a specific reliability failure requires a targeted change.
Audit an evaluation dataset for provenance, coverage, leakage, contamination, duplication, label quality, and admissibility before it supports release claims.
Updated Aug 19, 2026
Audit the evaluation dataset for decision fitness before it is used to support release claims. Treat the dataset as inadmissible until the evidence below supports a narrower conclusion.
Context and inputs to provide:
- Dataset and intended release claim: [Dataset description and intended release claim]
- Dataset records, schema, samples, or files available for review: [Dataset records or sample with schema]
- Source lineage, collection method, licenses, collection dates, and transformations: [Provenance sources and collection dates]
- Label definitions, labeling instructions, adjudication process, annotator metadata, and quality checks: [Labeling guidelines and annotator metadata]
- Target tasks, risk categories, domains, locales, languages, user populations, and expected operating conditions: [Target tasks risks and user populations]
- Known model training corpora, exclusion lists, benchmark sources, public datasets, or other contamination references: [Known training data or exclusion sources]
Evidence discipline:
- Separate observed facts from inference. Mark each material claim as Observed, Inferred, Not provided, or Not assessable from supplied evidence.
- Do not claim that a file, source, test, command, repository, dataset split, or system was inspected unless it is present in the supplied material.
- Do not invent counts, percentages, inter-annotator agreement, contamination rates, licenses, or collection dates. If a metric cannot be calculated from the supplied evidence, state what is missing and whether a proxy assessment is possible.
- Preserve uncertainty. Use confidence levels only when tied to available evidence.
- Focus on this dataset audit, not on building an evaluation harness or comparing model versions.
Authority and action boundaries:
- This audit may assess evidence and recommend a disposition; it does not authorize release claims or modify the dataset.
- Do not quarantine, remove, relabel, deduplicate, rebalance, augment, disclose, or release dataset records, and do not approve or revise a release claim, unless the responsible owner has separately authorized that action.
- Dataset changes require authorization from the dataset owner and evaluation owner. Rights, consent, privacy, or license decisions require the data owner and legal or privacy reviewer. Final use of the dataset for a release claim remains with the release owner.
Audit procedure:
1. Define the admissibility question.
- Restate the release claim the dataset is expected to support.
- Identify the accountable evaluation owner, dataset owner, data owner, and release owner if named in the supplied material; otherwise list them as missing accountability assignments.
- Define what the dataset must demonstrate to be fit for that release claim.
2. Build a dataset provenance ledger.
For each source, split, subset, or major record group, capture:
- Source name or origin
- Collection date or time window
- Collection method
- Rights, license, consent, or usage restriction evidence
- Transformation, filtering, augmentation, or generation steps
- Label source and labeling workflow
- Known exclusions or quarantine rules
- Traceability gaps
- Confidence in provenance
3. Build a coverage matrix.
Map the dataset against the supplied target tasks, risk categories, domains, user populations, languages/locales, difficulty bands, failure modes, and operating conditions. Include:
- Available counts or proportions where directly calculable
- Coverage status: Adequate, Thin, Missing, Overrepresented, Not assessable
- Evidence basis
- Release-claim consequence of each gap
- Minimum additional evidence or data needed to close the gap
4. Audit leakage, duplication, and contamination risk.
Create a contamination register covering:
- Exact duplicate records within the dataset
- Near duplicates or paraphrase clusters, if detectable from supplied records
- Train/eval split leakage, if split information is supplied
- Prompt-answer leakage, rubric leakage, or label leakage
- Overlap with known training data, public benchmarks, synthetic data sources, vendor examples, documentation, or previous evaluation sets
- Temporal leakage relative to the intended release claim
- Source reuse that could inflate performance claims
For each item, state the evidence, detection method available from supplied material, severity, uncertainty, owner, and remediation.
5. Review label quality and decision reliability.
Assess:
- Label definition clarity and mutual exclusivity
- Alignment between labels, rubric, and release claim
- Ambiguous or underspecified cases
- Annotator qualification evidence
- Adjudication and dispute-resolution process
- Inter-annotator agreement or audit sample results, only if provided
- Gold-standard or expert-review evidence, only if provided
- Label drift across sources, time periods, or task categories
- Examples where the label appears inconsistent with the provided guideline
6. Identify decision risks.
Distinguish risks that affect:
- Statistical validity
- External validity and representativeness
- Safety or policy risk coverage
- Bias across user populations or locales
- Claim wording and overgeneralization
- Reproducibility and auditability
- Legal, license, privacy, or data-rights admissibility
7. Recommend remediation.
Provide targeted actions only. Avoid broad rebuilds unless the evidence shows the dataset cannot be repaired. For each action include:
- Remediation action
- Specific defect addressed
- Priority
- Responsible owner: dataset owner, evaluation owner, data owner, security reviewer, policy reviewer, legal reviewer, or release owner as applicable
- Acceptance check
- Whether the dataset must be quarantined, relabeled, deduplicated, rebalanced, restricted to a narrower claim, or rejected
8. Make an admissibility decision.
Choose one:
- Admissible for the stated release claim
- Conditionally admissible after named remediation
- Admissible only for a narrower claim
- Not admissible for release claims
Explain the decision in terms of evidence sufficiency, unresolved uncertainty, contamination risk, coverage gaps, label quality, and provenance.
Required output format:
# Evaluation Dataset Coverage and Contamination Audit
## 1. Admissibility Question
- Intended release claim:
- Dataset use in the release decision:
- Accountable owners named in evidence:
- Missing accountability assignments:
- Fitness threshold for this audit:
## 2. Evidence Inventory
| Evidence item | Provided | Used for | Limitations | Missing information |
|---|---:|---|---|---|
## 3. Dataset Provenance Ledger
| Dataset segment/source | Origin | Collection window | Collection method | Rights/consent evidence | Transformations | Label source | Restrictions | Traceability gaps | Confidence |
|---|---|---|---|---|---|---|---|---|---|
## 4. Coverage Matrix
| Task/risk/user-population dimension | Expected coverage | Observed coverage | Status | Evidence basis | Release-claim consequence | Data or evidence needed |
|---|---|---|---|---|---|---|
## 5. Duplication, Leakage, and Contamination Register
| Issue | Type | Evidence observed | Detection possible from supplied material | Severity | Uncertainty | Owner | Remediation |
|---|---|---|---|---|---|---|---|
## 6. Label-Quality Findings
| Finding | Evidence status | Affected records or segment | Decision impact | Confidence | Remediation |
|---|---|---|---|---|---|
## 7. Decision Risks
List the material risks that remain after the audit. For each, state whether it affects statistical validity, external validity, safety coverage, fairness, reproducibility, rights/privacy, or claim wording.
## 8. Remediation Plan
| Priority | Action | Defect addressed | Responsible owner | Acceptance check | Release impact |
|---|---|---|---|---|---|
## 9. Admissibility Decision
- Decision:
- Claim supported, if any:
- Claim not supported:
- Required conditions before use:
- Residual uncertainty:
- Final verification checks for the release owner:
Completion checks before finalizing:
- The provenance ledger, coverage matrix, contamination register, label-quality findings, remediation plan, and admissibility decision are all present.
- Every major conclusion cites supplied evidence or is explicitly marked as inference or not assessable.
- No unavailable inspection, test, source comparison, or metric is claimed.
- The release owner has enough information to decide whether the dataset may support the stated claim, must be narrowed, or must be rejected.
Investigate persistent agent memory for poisoning, misattribution, over-retention, or unauthorized alteration and produce a defensible containment and recovery decision.
Updated Aug 19, 2026
Investigate whether persistent agent memory was altered, misattributed, over-retained, or poisoned, and produce an evidence-backed containment and recovery decision.
Context to provide:
- Agent or system name: [Agent or system name]
- Investigation window: [Investigation window]
- Memory stores and schemas: [Memory stores and schemas]
- Available evidence: [Available evidence]
- Known suspicious symptoms: [Known suspicious symptoms]
- Decision owner: [Decision owner]
- Recovery authority boundaries: [Recovery authority boundaries]
Evidence rules:
- Use only the evidence provided in [Available evidence]. Do not claim that logs, traces, memory stores, tickets, approvals, tests, commands, or source files were inspected unless they are included or quoted.
- Separate observations from inference. Label each inferred conclusion as low, medium, or high confidence.
- Preserve uncertainty. If a required fact is missing, state exactly what evidence is needed and why it matters.
- Before attributing compromise or selecting containment, separate blocking gaps from non-blocking gaps. Request all blocking evidence in one consolidated clarification and stop the affected conclusion until it is supplied. Continue past non-blocking gaps only when they are recorded as Unknown with their effect on confidence, scope, containment, and recovery.
- Do not broaden this into a general agent security audit or a generic trace taxonomy. Stay focused on persistent memory integrity, provenance, poisoning, retention, attribution, and recovery.
- Treat containment and recovery as accountable operational decisions. If an action exceeds [Recovery authority boundaries], identify the responsible owner or function that must decide.
Investigation method:
1. Establish scope
- Define which memory stores, records, embeddings, summaries, preference stores, tool-state caches, user profile memories, system memories, and derived memories are in scope based on [Memory stores and schemas].
- Identify what is out of scope and any assumptions required because evidence is missing.
2. Build a memory provenance ledger
Create a ledger table with these columns:
- Memory item or record ID
- Current stored claim or value
- Memory type or store
- First observed timestamp, if known
- Last modified timestamp, if known
- Claimed source or author
- Evidence supporting source attribution
- Write or update path
- Retention basis or deletion expectation
- Integrity concern: none, altered, misattributed, over-retained, suspicious insertion, suspicious deletion, unverifiable
- Confidence level
- Evidence references
- Required follow-up evidence
3. Reconstruct write paths
For each memory store or memory type, reconstruct:
- Authorized writers and expected write triggers
- Observed or reported writes during [Investigation window]
- Transformation steps from raw input to persisted memory
- Summarization, embedding, deduplication, merge, overwrite, deletion, or compaction behavior
- Identity and attribution mechanisms used at write time
- Retention and deletion controls
- Plausible bypasses, race conditions, stale-cache paths, ingestion errors, or cross-user contamination paths
- Evidence gaps that prevent confirmation
4. Develop a poisoning hypothesis matrix
Create a matrix with these columns:
- Hypothesis
- Poisoning or integrity failure mechanism
- Memory items affected
- Supporting observations
- Contradicting observations
- Missing evidence needed
- Likelihood: low, medium, high, or indeterminate
- Potential blast radius
- Immediate containment implication
Include at least these hypothesis categories when relevant to the evidence:
- Malicious prompt or tool output caused durable memory insertion
- Benign user input was misattributed to another user, tenant, workflow, or authority
- Summarization or consolidation distorted source meaning
- Outdated memory was over-retained after deletion, revocation, policy change, or user correction
- Memory merge or deduplication combined incompatible identities, accounts, or contexts
- Retrieval surfaced poisoned or stale memory into later decisions
- Administrative, migration, sync, or backfill process altered memory incorrectly
- Evidence does not support a memory poisoning finding, but confidence is limited
5. Map affected decisions
Create an affected-decision map with these columns:
- Decision, action, response, workflow, or automation potentially influenced
- Memory items used or likely retrieved
- Evidence of use versus plausible exposure
- Business, safety, security, privacy, financial, or user impact
- Reversibility
- Owner accountable for review
- Required validation before relying on the decision again
Do not assume a decision was affected solely because a memory item exists. Distinguish confirmed use, likely use, possible use, and no evidence of use.
6. Decide containment posture
Recommend one of these containment states for each affected memory store or memory class:
- No containment needed based on current evidence
- Monitor only
- Read-only quarantine
- Retrieval suppression
- Write suspension
- Record-level deletion or correction pending owner approval
- Full memory store isolation
- Agent workflow suspension
For each recommendation, provide:
- Evidence basis
- Risk reduced
- Operational cost or user impact
- Decision owner required
- Time sensitivity
- Reversal condition
7. Define recovery gate
Produce a quarantine and recovery gate that [Decision owner] can use. Include:
- Minimum evidence required before restoring memory retrieval or writes
- Records that must be corrected, deleted, re-attributed, or re-generated
- Regression checks or replay checks needed, if evidence supports them
- Owner approvals required under [Recovery authority boundaries]
- Monitoring signals after restoration
- Conditions that require escalation rather than recovery
8. Final decision artifact
End with a concise decision record:
- Investigation conclusion: confirmed compromise, likely compromise, possible compromise, no compromise found, or indeterminate
- Confidence and primary uncertainty
- Memory stores or records requiring action
- Affected decisions requiring review
- Immediate containment decision
- Recovery gate status: not ready, conditionally ready, or ready
- Named accountable owner for the next decision
- Evidence not available that would materially change the conclusion
Completion checks before finalizing:
- Every material conclusion cites evidence from [Available evidence] or is explicitly labeled as inference.
- The memory provenance ledger, write-path reconstruction, poisoning hypothesis matrix, affected-decision map, and quarantine/recovery gate are all included.
- Missing evidence is stated plainly rather than filled with assumptions.
- No source, system, command, test, approval, or inspection is claimed without corresponding evidence.
- The recommended action stays within [Recovery authority boundaries] or identifies the accountable owner who must authorize it.
Reconstructs a multi-agent failure to find the first coordination divergence and produce a corrected handoff contract with testable verification.
Updated Aug 19, 2026
Reconstruct the coordination failure across multiple agents and identify the first coordination divergence that can be corrected and tested.
Inputs to provide:
- Agent roles and authority: [Agent roles and authority]
- Coordination evidence: [Coordination evidence]
- Expected handoff contract: [Expected handoff contract]
- Shared-state context: [Shared-state context]
- Observed failure and impact: [Observed failure and impact]
- Constraints and authorized changes: [Constraints and authorized changes]
Evidence discipline:
- Separate observed evidence from inference at every major step.
- Do not claim that a log, trace, source file, system, test, approval, or action was inspected unless it is present in the provided material.
- If timestamps, message ordering, state versions, ownership rules, or agent responsibilities are missing, name the gap and explain how it affects confidence.
- Preserve uncertainty where the evidence supports more than one explanation.
- Do not redesign the entire multi-agent workflow. Focus on the smallest correction to the failed coordination point.
- Do not produce a broad trace taxonomy. Use classification only where it helps isolate the first divergence.
Reconstruction method:
1. Build an evidence inventory listing each provided artifact, what it can prove, and what it cannot prove.
2. Create a coordination chronology from the earliest relevant trigger through the failure manifestation.
3. Map responsibility, authority, and shared-state access for each involved agent.
4. Compare expected coordination behavior against observed behavior.
5. Identify the first point where delegation, handoff, ownership, message ordering, or shared state diverged from the expected contract.
6. Decide the most likely failure mechanism and explain why competing mechanisms are less supported.
7. Define discriminating tests that could confirm or falsify the proposed mechanism.
8. Draft a corrected handoff contract that is narrow, testable, and preserves intended existing behavior.
Deliverable format:
1. Evidence inventory
Provide a table with columns:
- Artifact
- Evidence observed
- Reliability or limitation
- Coordination question it helps answer
- Missing information, if any
2. Coordination chronology
Provide an ordered table with columns:
- Step or timestamp
- Agent
- Message, action, or state operation
- Intended recipient or owner
- Shared state read or written
- Expected behavior
- Observed behavior
- Evidence reference
- Observation vs inference
- Confidence: high, medium, or low
3. Responsibility and state map
For each relevant agent, list:
- Delegated responsibility
- Decision authority
- Required inputs
- State it may read
- State it may write
- Handoff obligations
- Acknowledgment or completion signal
- Actual behavior seen in evidence
- Ownership ambiguity or conflict, if any
Then list shared-state objects or records with:
- State object
- Expected owner
- Writers
- Readers
- Versioning or ordering assumption
- Observed mutation or stale-read risk
- Evidence supporting the risk
4. First coordination divergence
State the earliest supported divergence in one sentence.
Then provide:
- Divergence type: delegation gap, handoff ambiguity, ownership conflict, message-ordering violation, stale shared state, unauthorized state mutation, missing acknowledgment, retry/idempotency failure, or other specified mechanism
- Exact expected contract at that point
- Exact observed deviation
- Why this is earlier than downstream symptoms
- Evidence supporting the decision
- Confidence level
- What evidence would change the conclusion
5. Failure mechanism decision
Compare the leading mechanism against plausible alternatives in a table:
- Candidate mechanism
- Supporting evidence
- Contradicting or missing evidence
- Predicted observable signal
- Decision: selected, possible, unlikely, or not assessable
6. Discriminating tests
Propose 3 to 6 tests or checks that the accountable engineering, automation, or operations owner can run. For each, include:
- Hypothesis tested
- Setup or evidence required
- Expected signal if the mechanism is correct
- Expected signal if the mechanism is wrong
- Owner best placed to verify
- Risk of false positive or false negative
Prefer tests that distinguish between message-ordering, ownership, handoff, and shared-state causes rather than merely reproducing the incident.
7. Corrected handoff contract
Draft the smallest safe correction. Include:
- Trigger condition
- Sending agent
- Receiving agent
- Required payload or state reference
- Preconditions
- State version or ordering requirement
- Acknowledgment rule
- Completion signal
- Timeout or retry behavior
- Idempotency requirement
- Conflict handling rule
- Audit fields or log events needed for future reconstruction
- Backward-compatibility or behavior-preservation note
8. Implementation boundary
State what should not be changed yet because the evidence does not justify it. Call out any broad workflow redesign, new orchestration pattern, or policy change that would be premature.
9. Completion checks
List observable criteria for considering this reconstruction complete, including:
- The first divergence is identified or the blocking evidence gap is explicit.
- The selected mechanism has at least one discriminating test.
- The corrected handoff contract is narrow enough for the responsible owner to implement or reject.
- The release owner, automation owner, or service owner can verify the proposed correction against logs, replay, tests, or production telemetry before adoption.