Compare overlapping SaaS tools against validated capabilities, dependencies, contracts, controls, switching costs, and migration risk to support a defensible consolidation decision.
Updated Aug 19, 2026
Prepare a defensible SaaS portfolio redundancy and consolidation analysis for a consequential portfolio decision.
Context to provide:
- SaaS applications in scope: [SaaS applications in scope]
- Validated business capabilities and requirements: [Validated business capabilities and requirements]
- Known integrations, data flows, and dependencies: [Known integrations data flows and dependencies]
- Contract, commercial, renewal, termination, and exit terms: [Contract commercial and exit terms]
- Control, compliance, security, privacy, audit, and risk requirements: [Control compliance and risk requirements]
- Migration constraints, timing limits, business calendar constraints, and decision horizon: [Migration constraints and decision horizon]
Working rules:
- Use only the evidence provided. Do not claim that systems, contracts, logs, security controls, usage reports, or files were inspected unless the supplied evidence shows that.
- Distinguish observed facts, stated assumptions, reasoned inference, and missing information.
- Do not make seat utilization the primary decision basis. Consider usage only if supplied and only as supporting evidence.
- Do not treat this as a one-vendor renewal recommendation or an initial procurement evaluation. The job is portfolio-level redundancy and consolidation across tools already in scope.
- Preserve legitimate differences between tools. Do not collapse tools as redundant merely because they share a category label.
- Identify where accountable owners must verify facts before execution: product owner, business capability owner, data owner, security reviewer, finance owner, procurement or vendor manager, legal reviewer, and migration owner as applicable.
- If evidence is insufficient for a final decision, provide a provisional decision and list the specific evidence needed to confirm or change it.
Authority and execution boundaries:
- This analysis may recommend retain, consolidate, retire, or defer; it does not authorize contract renewal or termination, remove user access, migrate or delete data, communicate a final decision, or initiate implementation.
- Portfolio commitment requires approval from the accountable portfolio or process owner and finance owner. Contract actions require the procurement owner and legal reviewer. Data movement or deletion requires the data and privacy owners. Control changes and production transition require the security, service, migration, and release owners as applicable.
- Do not present a recommendation as approved or executable until the named owners have verified the supporting evidence, accepted the residual risks, and authorized the next action.
Decision criteria:
Evaluate each tool and overlap area against:
1. Coverage of validated business capabilities and must-have requirements.
2. Unique capability, workflow, data, user group, regional, or regulatory dependencies.
3. Integration, identity, reporting, automation, API, and data-flow dependencies.
4. Data ownership, retention, portability, residency, privacy, and deletion constraints.
5. Contractual constraints, renewal timing, notice periods, termination rights, minimum commitments, price protections, and exit assistance.
6. Switching costs, migration effort, change management burden, retraining, process disruption, and operational risk.
7. Security, privacy, compliance, audit, continuity, and administrative control coverage.
8. Total-cost scenarios, including parallel-run periods, implementation effort, migration support, internal labor, vendor services, penalties, and stranded commitments where evidence allows.
9. Reversibility of the decision and the risk of losing hard-to-recover data, integrations, controls, or institutional workflow knowledge.
Required deliverable:
1. Decision boundary
- State the applications in scope.
- State the decision horizon.
- State what is explicitly out of scope.
- Identify the decision owners and verification owners needed from the evidence provided.
2. Evidence register
Create a compact table with columns:
- Evidence item
- Source or supplied artifact
- What it supports
- Reliability level: high / medium / low
- Gaps or verification needed
3. Capability-to-tool map
Create a table with columns:
- Business capability or requirement
- Current tool or tools supporting it
- Requirement criticality: must-have / important / optional
- Coverage quality: strong / partial / weak / unknown
- Unique differentiators or constraints
- Evidence basis
- Consolidation implication
4. Redundancy and dependency analysis
For each meaningful overlap area, assess:
- Apparent redundancy
- True redundancy after requirements and dependencies are considered
- Non-obvious dependencies
- Integration and data-flow impact
- Business process impact
- Control or compliance impact
- Users, teams, regions, or workflows likely affected if supplied
- Whether the overlap supports consolidation, coexistence, or retirement
5. Total-cost scenarios
Compare at least three scenarios where evidence permits:
- Retain current portfolio
- Consolidate into one or more existing tools
- Retire one or more tools after migration or exit
For each scenario, include:
- Cost components supported by evidence
- Cost components that are likely but not quantified
- Contract timing and stranded-cost exposure
- Migration and change costs
- Operational risk cost drivers
- Confidence level
- Missing financial inputs needed for a firmer estimate
6. Migration and exit constraints
Create a table with columns:
- Tool or capability affected
- Exit or migration constraint
- Source of constraint: contract / data / integration / control / process / people / timing / unknown
- Severity: high / medium / low
- Mitigation option
- Owner to verify
- Decision impact
7. Control coverage and risk comparison
Assess whether consolidation would weaken, maintain, or improve:
- Identity and access administration
- Role-based access or privilege controls
- Audit logging and evidence retention
- Security monitoring and incident response support
- Privacy, residency, retention, deletion, or export obligations
- Business continuity and disaster recovery expectations
- Regulatory or contractual control obligations
Do not infer control sufficiency from vendor reputation. Tie each point to supplied evidence or mark it as unknown.
8. Retain / consolidate / retire decision record
Provide a decision record with:
- Recommended decision for each application: retain / consolidate / retire / defer pending evidence
- Rationale tied to capability coverage, dependencies, cost, contract terms, controls, and migration risk
- Conditions that must be true for the recommendation to remain valid
- Required owner verifications before commitment
- Risks accepted by the decision
- Risks requiring mitigation before execution
- Earliest safe decision point and earliest safe execution point if determinable
- Reversal difficulty: low / medium / high
9. Portfolio recommendation
Give a concise portfolio-level recommendation that states:
- Preferred consolidation path
- Tools to retain and why
- Tools to retire or phase out and why
- Tools that should remain temporarily during transition
- Sequencing logic
- No-regret next steps that do not prematurely commit the organization
10. Completion checks
End with a checklist confirming whether the analysis includes:
- Capability-to-tool map
- Redundancy and dependency analysis
- Total-cost scenarios
- Migration and exit constraints
- Control coverage comparison
- Retain / consolidate / retire decision record
- Named owner verifications
- Explicit missing information
- Clear distinction between evidence and inference
Decide whether an existing systematic review needs an update and define the safest bounded scope before committing review resources.
Updated Aug 19, 2026
Produce a Systematic Review Update Decision Brief that determines whether the existing review should be updated now, deferred, monitored, or amended with a bounded scope.
Do not redo initial screening, extraction, or full evidence synthesis unless the update decision cannot be made without a narrowly defined check. Focus on update triggers, conclusion stability, and the minimum defensible update scope.
Inputs:
- Review title and citation: [Review title and citation]
- Review question or PICO: [Review question or PICO]
- Original search end date: [Original search end date]
- Original conclusions and certainty ratings: [Original conclusions and certainty ratings]
- Surveillance sources and date range: [Surveillance sources and date range]
- Current decision need: [Current decision need]
- Known methodological or data issues: [Known methodological or data issues]
Evidence discipline:
- Separate observed evidence from inference.
- Cite sources for every factual claim about new studies, retractions, corrections, guidelines, methods changes, or decision context.
- State the date accessed for web sources when available.
- If a source is mentioned but not accessible, mark it as “not inspected” and do not rely on it as evidence.
- Do not claim that searches, screening, extraction, statistical checks, author contact, or protocol approval were completed unless the supplied evidence supports that claim.
- Identify missing information that could materially change the decision.
- Preserve uncertainty where evidence is incomplete or surveillance is limited.
Decision criteria:
Assess whether any of the following update triggers are present:
1. New eligible studies plausibly change effect direction, effect size, precision, heterogeneity, applicability, or certainty.
2. New evidence affects outcomes prioritized for the current decision need.
3. New safety, equity, subgroup, implementation, cost, or long-term outcome evidence changes practical interpretation.
4. Retractions, expressions of concern, corrected data, duplicate publications, or serious integrity concerns affect included or influential studies.
5. Changes in methods, risk-of-bias expectations, outcome standards, reporting norms, or certainty assessment make the original conclusions unreliable for current use.
6. New guidelines, regulatory decisions, policy needs, payer decisions, or clinical practice changes create a time-sensitive decision requirement.
7. The original review has become stale because the evidence base, intervention, comparator, population, or context has materially changed.
Deliverable format:
1. Decision Snapshot
- Existing review: cite the review.
- Original evidence currency: state the original search end date.
- Current decision need: summarize the decision the update must support.
- Recommended decision: choose one: No update needed now; Continue surveillance; Focused update; Full update; Protocol amendment before update; Defer because evidence is insufficient or inaccessible.
- Confidence in decision: High / Moderate / Low, with a short reason.
- Accountable verification owner: name the role that should verify the decision, such as review lead, methods lead, clinical content owner, guideline chair, research director, or data owner.
2. Surveillance Evidence Log
Create a table with these columns:
- Evidence item
- Source and citation
- Date range or publication date
- Why it may matter
- Status: observed / inferred / not inspected
- Link to original review question or PICO
- Potential update trigger
- Limitations or uncertainty
Include new studies, trial registrations, preprints when relevant, retractions, corrections, expressions of concern, updated guidelines, methodological changes, and decision-context changes. Do not include marginal items unless they could affect the update decision.
3. New-Evidence Impact Matrix
Create a matrix with these rows where applicable:
- Population
- Intervention or exposure
- Comparator
- Primary outcomes
- Harms or safety outcomes
- Equity or subgroup effects
- Certainty of evidence
- Risk of bias or study integrity
- Applicability to the current decision need
- Methods or reporting expectations
For each row, provide:
- New or changed evidence observed
- Likely impact on original conclusion: none / minor / moderate / major / unknown
- Basis for judgment
- What would need to be checked in a formal update
4. Conclusion-Stability Assessment
For each major original conclusion, assess:
- Original conclusion and certainty rating
- New evidence or issue relevant to that conclusion
- Whether the conclusion is likely stable, possibly unstable, likely unstable, or cannot be assessed
- Reasoning, distinguishing evidence from inference
- Consequence if the conclusion changed
5. Update Trigger Decision
State whether the threshold for an update is met. Use this structure:
- Trigger status: met / not met / unclear
- Triggering evidence or issue
- Decision risk if no update is performed
- Decision risk if an update is performed too broadly
- Recommended update type and scope
- Explicit exclusions from the update scope
- Minimum evidence checks required before work begins
6. Bounded Update Scope
Define the smallest responsible update scope:
- Review question or PICO elements to retain unchanged
- Elements to amend, if any
- Outcomes to prioritize
- Study designs to include or exclude
- Date range to search
- Databases, registries, and targeted sources to search
- Whether re-screening of prior included studies is needed and why
- Whether reanalysis is likely needed and why
- Whether certainty assessment must be revised
- Known exclusions that prevent scope creep
7. Protocol Amendment and Work Plan
If any update or amendment is recommended, draft a practical work plan:
- Required protocol amendments
- Review team roles: review lead, information specialist, methods lead, content reviewer, statistician if needed, and data owner if corrected data are involved
- Priority tasks in sequence
- Evidence checks before extraction or synthesis
- Estimated level of effort: low / medium / high, with rationale
- Decision points where the update may be narrowed, expanded, paused, or stopped
- Approval or sign-off needed from the accountable review, guideline, policy, or research owner
If no update is recommended, provide a surveillance plan instead:
- Conditions that would trigger reconsideration
- Sources to monitor
- Suggested surveillance interval
- Owner responsible for monitoring
8. Missing Information and Verification Checks
List only missing information that could materially alter the update decision. Then provide completion checks:
- Are all relied-upon sources cited?
- Are inaccessible sources clearly marked as not inspected?
- Are observations separated from inference?
- Is the update decision tied to explicit triggers?
- Is the recommended scope bounded enough to prevent unnecessary full-review work?
- Is owner verification clearly assigned?
Keep the brief practical, evidence-bound, and suitable for review lead or policy owner decision-making. Avoid generic literature-summary prose.
Audit consequential claims from current citation back to original source evidence, documenting where provenance breaks, mutates, or overstates authority.
Updated Aug 19, 2026
Objective: Trace each material claim through its full citation chain to the earliest available original source, dataset, record, or primary evidence. Identify where citations merely repeat intermediaries, where meaning changes across the chain, where authority is unsupported, and what release or correction decision is justified.
Inputs to use:
- Material claims to audit: [Material claims to audit]
- Cited sources and intermediaries: [Cited sources and intermediaries]
- Decision context: [Decision context]
- Source access constraints: [Source access constraints]
- Accountable release owner: [Accountable release owner]
Instructions:
1. Treat each claim as a separate audit unit. Do not collapse similar claims unless the evidence chain is demonstrably identical.
2. For every claim, start with the provided citation or cited authority, then trace backward through cited papers, reports, datasets, webpages, legal records, interviews, statistics, or quoted statements until you reach the earliest accessible source that materially supports, weakens, contradicts, or fails to support the claim.
3. Use Perplexity search to locate originals and intermediaries where available. Prefer primary sources, original datasets, official records, archived pages, full papers, methodological appendices, and author-provided data over summaries, media reports, abstracts, AI-generated pages, and circular citations.
4. Do not assume that a cited source supports the claim because it is referenced. Verify the exact passage, table, figure, dataset field, legal clause, quote, date range, population, geography, method, and scope relevant to the claim.
5. Distinguish observations from inference:
- Observation: what the source explicitly says or contains.
- Inference: what must be concluded from comparing sources, methods, dates, or wording.
- Unknown: what cannot be verified from accessible evidence.
6. Identify transformation points in the citation chain, including paraphrase drift, scope expansion, causal overstatement, missing denominators, outdated data reuse, quote truncation, statistical re-labeling, citation laundering, broken links, inaccessible originals, and claims supported only by secondary repetition.
7. Preserve uncertainty. If the original source is paywalled, missing, archived only, unavailable, or ambiguous, state that limitation and evaluate the chain only from available evidence.
8. Do not perform a generic source credibility ranking. Focus on whether the cited chain actually carries the audited claim from origin to current use without unsupported transformation.
9. Do not synthesize competing positions unless needed to evaluate whether a claim’s provenance is distorted. The central deliverable is claim-to-origin traceability.
10. Do not state that a source, dataset, archive, document, or full text was inspected unless you can cite the accessible evidence used.
Source-protection and action boundaries:
- Use only supplied or lawfully accessible material. Do not bypass access controls, disclose confidential or personal information, or reproduce more licensed text than is needed to document the finding. Record access, privacy, confidentiality, or reuse constraints in the evidence limits and refer uncertain legal questions to the legal or privacy reviewer.
- This audit may recommend wording, citation, hold, removal, or correction decisions; it does not publish, remove, correct, contact source owners, alter records, or authorize release.
- Publication, withholding, correction, and removal decisions require authorization from the accountable release owner. Legal, privacy, or data-rights consequences require confirmation from the relevant legal, privacy, or data owner before action.
- Stop and request an owner decision when completing the trace would require unauthorized access, disclosure, or use of protected material.
Deliver the audit in this format:
A. Audit scope and evidence limits
- Claims audited:
- Sources or citation chains checked:
- Sources not accessible or not located:
- Material assumptions:
- Consequence of unresolved gaps for [Decision context]:
B. Claim-to-origin provenance graph
For each claim, provide a compact chain using this structure:
- Claim ID:
- Current claim wording:
- Current cited authority:
- Intermediary 1:
- Intermediary 2, if any:
- Earliest accessible original source:
- Evidence type at origin:
- Direct support status: Supported / Partially supported / Not supported / Contradicted / Unverifiable
- Provenance confidence: High / Medium / Low, with reason
C. Citation transformation ledger
Create a table with columns:
- Claim ID
- Chain step
- Source and link
- What this source actually states
- How the next source or current claim changes it
- Transformation type
- Materiality of change: Low / Medium / High
- Evidence quoted or summarized
D. Source authority findings
For each original or near-original source, assess only authority factors relevant to the claim:
- Source identity and role in producing the evidence
- Whether it is primary, near-primary, or secondary
- Method, dataset, sample, jurisdiction, date range, or evidentiary basis
- Known limitations stated by the source
- Whether the current claim exceeds those limitations
- Citation needed for the authority finding
E. Broken-chain register
List every provenance failure or unresolved gap:
- Claim ID
- Break or defect type
- Location in chain
- Why it matters
- What evidence would repair the chain
- Whether repair is feasible before release
- Responsible owner: [Accountable release owner] or another named owner if more appropriate
F. Release and correction decision
For each claim, recommend one decision:
- Release as written
- Release with citation upgrade
- Release with narrowed wording
- Hold pending source repair
- Remove claim
- Issue correction or clarification if already published
For each decision, include:
- Required wording change, if any
- Required citation replacement, if any
- Residual risk
- Verification step the [Accountable release owner] should complete before release
Completion checks:
- Every material claim has a traced chain or an explicit unresolved-chain entry.
- Every releaseable claim links to an original or defensible near-original source.
- Every unsupported transformation is recorded in the ledger or broken-chain register.
- No unavailable source is described as inspected.
- The final decision is usable by [Accountable release owner] for publication, release, correction, or withholding.
Validate AI-generated SQL and its reported results before they are used for a consequential decision.
Updated Aug 19, 2026
Verify the AI-generated SQL and its claimed result before anyone relies on it for the stated decision. Treat this as a defensible review artifact, not a general SQL improvement exercise.
Context to provide:
- Repository scope and task: [Repository scope and task]
- Relevant files and instructions: [Relevant files and instructions]
- Observed evidence: [Observed evidence]
- Constraints and authorized changes: [Constraints and authorized changes]
- Environment details without secrets: [Environment details without secrets]
- Verification commands and acceptance criteria: [Verification commands and acceptance criteria]
Working rules:
1. Inspect the relevant files, saved queries, schemas, models, tests, lineage notes, and query artifacts first where they are available.
2. Do not claim that a table, file, source, command, test, query execution, permission boundary, or approval was inspected unless there is direct evidence in the provided materials or accessible workspace.
3. Identify the likely root cause of any discrepancy before editing SQL or proposing a corrected query.
4. Apply the smallest safe change needed to correct the query. Preserve existing intended behavior unless the evidence shows it is wrong.
5. Avoid broad rewrites, style-only edits, or migration to a different modeling pattern unless required to remove a proven defect.
6. Distinguish observations from inference. Mark unsupported assumptions explicitly.
7. Respect access boundaries. Do not suggest bypassing row-level security, protected schemas, production safeguards, or data owner controls.
8. Run syntax checks, query compilation, dry runs, unit tests, dbt tests, warehouse explain plans, or limited validation queries where available and appropriate. If execution is unavailable, state exactly what could not be run and what evidence substitutes for it.
9. Do not present a corrected result as final unless it is reproducible from inspected SQL and accessible evidence.
Missing-input gate:
- Treat the exact SQL, intended metric or decision and grain, relevant schema or lineage, and comparison source or acceptance threshold as blocking when their absence or conflict prevents semantic review or reconciliation. Request all blocking items in one consolidated clarification and stop the affected conclusion until they are supplied.
- Continue with non-blocking gaps only as explicit Unknowns with their decision impact.
Review procedure:
A. Build the query intent contract
- Restate the decision or metric question in operational terms.
- Define the expected grain of the result.
- Identify required dimensions, filters, joins, exclusions, deduplication rules, aggregation logic, time window, timezone, and freshness expectations.
- Identify the authoritative definition owner when evident, such as data owner, analytics owner, finance owner, product owner, or compliance owner.
- List evidence supplied versus evidence missing.
B. Review SQL semantics before result claims
- Check whether selected columns, grouping, joins, filters, CTEs, window functions, null handling, distinct logic, and aggregation match the query intent contract.
- Check for join multiplication, accidental inner joins, incomplete predicates, slowly changing dimension issues, many-to-many joins, late-arriving data, timezone drift, partition filters, and snapshot-versus-current-state confusion.
- Check whether access constraints or row-level filters could change the observed result.
- Identify whether the SQL answers the stated question, a narrower question, a broader question, or a different question.
C. Review execution and reproducibility evidence
- Determine whether the claimed result can be traced to a specific SQL text, execution environment, parameters, source versions, run time, and data freshness state.
- If query execution is available, run the safest reproducible check permitted by the environment and record the command/query used, result shape, row counts, and relevant totals.
- If execution is not available, perform static validation and specify the minimum evidence needed from the data owner or analytics owner to complete reconciliation.
D. Reconcile against authoritative totals
- Compare the claimed result to authoritative totals, source-of-truth reports, known control totals, prior certified extracts, ledger totals, or approved metric definitions where supplied.
- Reconcile at the most useful level: total, time bucket, segment, source system, account, customer, product, or other relevant dimension.
- Explain every material variance using evidence where possible. Separate confirmed causes from plausible causes.
E. Correct only what is justified
- If the SQL is wrong and enough evidence exists, provide a corrected query or minimal patch.
- If there is not enough evidence to correct the query safely, provide a bounded correction proposal and list the exact evidence required before use.
- Preserve naming, output shape, filters, permissions, and downstream expectations unless the discrepancy requires a specific change.
Deliverable format:
1. Query intent contract
- Decision or metric question:
- Intended grain:
- Required population and exclusions:
- Required time logic and timezone:
- Required joins and source precedence:
- Authoritative definitions or totals used:
- Accountable owner to verify final use:
- Missing context that limits certainty:
2. Semantic and execution review
Create a table with columns:
- Review area
- Evidence inspected
- Observation
- Risk to claimed result
- Status: Pass / Fail / Unverified
- Notes or required follow-up
Include at minimum: grain, joins, filters, aggregation, deduplication, time logic, null handling, access boundaries, source freshness, result reproducibility, and output shape.
3. Reconciliation table
Create a table with columns:
- Reconciliation item
- AI-claimed value
- Verified or control value
- Difference
- Materiality assessment
- Evidence source
- Explanation
- Status: Reconciled / Variance explained / Unresolved
4. Unsupported-result register
Create a table with columns:
- Unsupported claim or result component
- Why it is unsupported
- Evidence needed
- Accountable owner or source to confirm
- Decision impact
5. Corrected-query and acceptance record
- Root cause summary before any SQL change:
- Minimal corrected SQL or patch, if justified:
- Behavior intentionally preserved:
- Behavior intentionally changed:
- Files changed, if any:
- Syntax checks, tests, dry runs, or executions performed:
- Verification results:
- Remaining uncertainty:
- Acceptance decision: Accept / Accept with caveats / Reject / Cannot determine
- Acceptance criteria met or unmet:
- Owner verification required before decision use:
Completion check:
- The claimed result is either reconciled, corrected and reproducible, or explicitly rejected as unsupported.
- All material assumptions are labeled.
- Any changed files are summarized.
- Verification results and unavailable checks are stated plainly.
- The final acceptance decision is tied to the acceptance threshold or materiality standard.
Trace sensitive context across agent handoffs, memory, retrieval, tools, logs, and shared workspaces to identify unauthorized propagation and contamination risk.
Updated Aug 19, 2026
Conduct an evidence-based audit of sensitive context propagation across the specified agent workflow. Trace actual context movement through handoffs, memory, retrieval, tools, logs, outputs, and shared workspaces. Do not produce a generic sensitive-data checklist. Separate observed evidence from inference, expose missing information, and preserve uncertainty.
Context to provide:
- System or workflow under review: [System or workflow under review]
- Authorized purposes and tenant boundaries: [Authorized purposes and tenant boundaries]
- Sensitive context categories: [Sensitive context categories]
- Evidence package: [Evidence package]
- Known incidents or concerns: [Known incidents or concerns]
- Accountable owners: [Accountable owners]
- Review date range: [Review date range]
Evidence rules:
- Use only the evidence provided in [Evidence package] and clearly identified user-supplied context.
- Do not claim that a source, system, log, tool, workspace, approval, command, or deletion was inspected or completed unless evidence is present.
- Label each material statement as one of: Observed, Inferred, Not evidenced, or Requires owner confirmation.
- Distinguish sensitive context from ordinary workflow context.
- Distinguish authorized propagation from unauthorized propagation, excessive retention, purpose drift, and cross-tenant or cross-task contamination.
- If evidence is incomplete, state the missing artifact and why it matters.
Audit scope:
Trace sensitive context across these surfaces where evidence exists:
1. User inputs, uploaded files, tickets, records, conversations, or task instructions.
2. Agent-to-agent handoffs, delegation messages, intermediate reasoning summaries, or task state objects.
3. Short-term memory, long-term memory, vector stores, retrieval indexes, embeddings, caches, and session stores.
4. Tool calls, API payloads, browser sessions, database queries, SaaS integrations, webhooks, automations, and background jobs.
5. Logs, traces, analytics events, evaluation datasets, transcripts, error reports, monitoring records, and support workspaces.
6. Shared folders, project workspaces, collaboration tools, exported artifacts, generated documents, and downstream notifications.
7. Human review queues, escalation paths, approval records, and operator notes.
Required deliverable:
1. Audit Boundary and Evidence Inventory
Create a concise table with:
- Evidence item
- Source or owner if known
- Date range covered
- Context surfaces covered
- Reliability limits
- Material gaps
2. Context Lineage Map
Create a lineage table that traces each sensitive context category through the workflow:
- Context item or category
- Origin
- Initial authorized purpose
- Receiving agent, service, tool, memory store, log, workspace, or person
- Transfer mechanism
- Transformation or summarization performed
- Retention location and retention duration if evidenced
- Tenant, customer, project, task, or workspace boundary crossed
- Evidence reference
- Status: authorized, questionable, unauthorized, excessive, contaminated, or not evidenced
Then provide a short narrative explaining the highest-risk propagation paths. Do not infer a path merely because it is technically possible; identify it as a hypothesis if not evidenced.
3. Sensitivity and Purpose Register
Create a register with:
- Sensitive context category
- Sensitivity rationale
- Data subject, tenant, customer, project, or task boundary affected
- Authorized purpose from [Authorized purposes and tenant boundaries]
- Actual observed use
- Purpose alignment: aligned, narrowed, expanded, drifted, unrelated, or not evidenced
- Minimum context needed for the task
- Excess context observed
- Owner accountable for purpose decision from [Accountable owners], or owner not identified
4. Cross-Agent and Cross-Boundary Contamination Findings
For each finding, include:
- Finding title
- Evidence basis
- Contamination type: cross-agent, cross-tenant, cross-task, cross-customer, cross-project, memory reuse, retrieval bleed, logging exposure, tool propagation, workspace exposure, or purpose drift
- Affected context
- Affected boundary
- How the propagation occurred or is suspected to occur
- Impact on confidentiality, integrity, compliance, customer trust, operational safety, or decision quality
- Likelihood rating: evidenced, plausible, weakly supported, or unknown
- Severity rating: critical, high, medium, low, or informational
- Confidence level and reason
- Missing evidence that would change the rating
5. Minimization and Containment Controls
Propose controls tied to the observed lineage, not generic policy slogans. For each control, include:
- Propagation point addressed
- Control objective
- Specific change to inputs, prompts, handoff schema, memory policy, retrieval filtering, tool payloads, logging, access control, workspace permissions, retention, or operator procedure
- Expected reduction in sensitive context exposure
- Owner responsible for implementation
- Verification method
- Residual risk
Prioritize smallest effective controls that reduce propagation without breaking the authorized workflow. Avoid broad rewrites unless the evidence shows the workflow design itself is unsafe.
6. Deletion, Quarantine, and Revalidation Plan
Create an action plan with:
- Artifact or location requiring deletion, quarantine, redaction, re-indexing, access review, or retention change
- Reason action is needed
- Required owner approval: data owner, security reviewer, legal/privacy owner, product owner, platform owner, or customer account owner as appropriate
- Preconditions before action
- Execution evidence needed after action
- Revalidation test or sampling method
- Rollback or exception handling if deletion would impair legal hold, auditability, customer support, or service reliability
Do not state that deletion, quarantine, redaction, or re-indexing has been completed. State only the plan and the evidence needed to verify completion.
7. Open Questions and Owner Decisions
List unresolved questions that materially affect risk or remediation. For each, identify:
- Question
- Why it matters
- Evidence needed
- Accountable owner from [Accountable owners], or owner not identified
- Decision deadline if inferable from [Known incidents or concerns]
8. Completion Check
End with a completion check stating whether the audit is ready for owner review. Include:
- Whether every sensitive context category in [Sensitive context categories] was traced or marked not evidenced
- Whether every material propagation path has an evidence reference or uncertainty label
- Whether contamination findings are tied to actual lineage evidence
- Whether minimization controls map to specific propagation points
- Whether deletion and revalidation actions identify accountable owners and verification evidence
- Remaining blockers before security reviewer, data owner, product owner, or platform owner decision
Investigate wrong-principal agent actions and delegated authority failures across agent, tool, and downstream-system boundaries.
Updated Aug 19, 2026
Review the supplied evidence to determine which principal acted, under which delegated authority, at each agent, tool, and downstream-system boundary, and where that principal-authority binding failed.
Focus on agent identity spoofing, confused-deputy behavior, stale delegation, wrong-principal actions, and authorization-context loss. Do not perform a broad agent security audit. Do not infer inspection, execution, approval, containment, or remediation unless the supplied evidence supports it.
Context to provide:
- Incident or workflow name: [Incident or workflow name]
- Time window: [Time window]
- Agent and tool boundary description: [Agent and tool boundary description]
- Evidence bundle: [Evidence bundle]
- Delegation policy sources: [Delegation policy sources]
- Accountable owners: [Accountable owners]
Evidence discipline:
- Separate observed facts from inference, hypothesis, and missing information.
- Treat logs, traces, request IDs, token claims, policy excerpts, configuration snapshots, audit events, tickets, and owner statements as evidence only when supplied.
- Preserve uncertainty when timestamps conflict, identities are aliased, policies are incomplete, or downstream enforcement behavior is not evidenced.
- Do not assume that the initiating user, agent runtime identity, tool credential, service account, and downstream effective principal are the same.
- If a claim depends on unavailable evidence, state what evidence would be needed and which owner should provide or verify it.
Authority boundaries:
- Do not declare the issue fixed, contained, approved, or closed without evidence tied to a security reviewer, system owner, data owner, release owner, or service owner named in the supplied context.
- Do not recommend bypassing authorization checks, expanding privileges for convenience, or relying on logging as a substitute for enforcement.
- Prefer the smallest containment and repair steps that restore correct principal binding while preserving legitimate workflow behavior.
Produce the following deliverable.
1. Review scope and evidence inventory
- State the workflow, systems, principals, tools, downstream services, and time window actually covered by the evidence.
- List supplied evidence by type and relevance.
- List important evidence not provided, including why it matters.
- State confidence level for the review and the main reasons for that confidence.
2. Principal and delegation chain
Create a step-by-step chain showing how authority was carried or transformed across boundaries. Use a table with these columns:
- Step
- Timestamp or sequence marker
- Boundary crossed
- Request, trace, job, or session identifier
- Initiating principal
- Agent runtime principal
- Tool or connector principal
- Downstream effective principal
- Delegated authority claimed
- Delegation source or policy reference
- Token, credential, session, or capability involved
- Observed action
- Evidence citation from supplied material
- Confidence
- Missing evidence or ambiguity
Call out every point where identity was asserted, translated, delegated, cached, proxied, impersonated, or lost.
3. Authorization failure chronology
Build a concise chronology of the failure path. For each event, identify:
- What happened
- Which principal appeared to act
- Which principal should have been authoritative, if determinable
- What delegated authority was present, stale, missing, overbroad, or misapplied
- Whether the event indicates spoofing, confused deputy behavior, stale delegation, wrong-principal action, authorization-context loss, or another clearly named failure mode
- What evidence supports the classification
- What remains uncertain
4. Action-to-authority matrix
Create a matrix mapping each consequential action to the authority required and the authority actually observed. Use these columns:
- Action or operation
- Resource or data scope
- Required principal or role
- Required delegation condition, scope, audience, tenant, time limit, or consent
- Observed principal
- Observed delegated authority
- Enforcement point expected
- Enforcement point evidenced
- Decision observed, such as allowed, denied, skipped, inherited, cached, or unknown
- Mismatch type
- Potential impact
- Evidence citation
Highlight actions where the system accepted an authority context that was missing, stale, meant for another audience, meant for another tenant, too broad, inherited from the wrong actor, or not revalidated downstream.
5. Affected scope
Separate confirmed, likely, possible, and not evidenced impact. Cover:
- Users, tenants, accounts, workspaces, repositories, environments, datasets, secrets, transactions, or external systems affected
- Time interval of exposure
- Data or operations reachable under the mistaken authority
- Whether unauthorized read, write, execute, approve, delete, publish, spend, or disclose actions are evidenced
- Whether lateral movement, replay, token reuse, cached delegation, or downstream propagation is evidenced or merely plausible
- Constraints that limit impact
6. Containment and identity-control repair gates
Define repair gates that accountable owners can use to decide whether the workflow may resume or continue operating. Include immediate containment and durable control gates.
For each gate, provide:
- Gate name
- Failure it addresses
- Required control or change
- Owner accountable for verification
- Evidence required to pass
- Pass criteria
- Fail or block condition
- Residual risk if accepted
Include gates where relevant for:
- Principal provenance at agent start and tool-call time
- Delegation freshness and revocation handling
- Audience, tenant, resource, and scope binding
- Prevention of confused-deputy delegation reuse
- Downstream reauthorization rather than blind trust in upstream context
- Service account and connector credential separation
- Audit correlation across user, agent, tool, and downstream identifiers
- Token, session, capability, or cached grant invalidation
- Least-privilege restoration without breaking legitimate workflow paths
7. Owner decision record
Prepare a decision-ready summary for the named accountable owners. Include:
- Most likely root cause, stated with confidence and evidence basis
- Highest-risk unresolved uncertainty
- Minimum containment required before further use
- Minimum identity-control repair required before normal operation
- Verification tasks by owner role
- Open questions that block a defensible decision
- Items that can be safely deferred, with rationale
Completion checks:
- Confirm that the principal and delegation chain covers each supplied boundary or states why a boundary could not be assessed.
- Confirm that every consequential action is mapped to required and observed authority, or marked unknown with missing evidence.
- Confirm that affected scope distinguishes observed impact from plausible but unproven exposure.
- Confirm that each repair gate has an accountable owner, evidence requirement, and pass/fail criterion.
- Do not present remediation as complete unless evidence supplied in the prompt demonstrates completion.
Design a controlled retrieval experiment that compares chunking and metadata choices without turning into a full RAG redesign.
Updated Aug 19, 2026
Design a controlled retrieval experiment to isolate how chunking choices and metadata choices affect retrieval and answer quality. Keep the scope limited to corpus configuration decisions: do not redesign the full RAG architecture, diagnose the whole pipeline, change model selection, or introduce unrelated retrieval changes unless they are explicitly controlled constants.
Context to use:
- Corpus description and sample documents: [Corpus description and sample documents]
- Current retrieval setup: [Current retrieval setup]
- Candidate chunking options: [Candidate chunking options]
- Candidate metadata fields: [Candidate metadata fields]
- Query logs or benchmark questions: [Query logs or benchmark questions]
- Answer evaluation standard: [Answer evaluation standard]
- Decision owners and constraints: [Decision owners and constraints]
Evidence discipline:
- Separate observations supported by the provided inputs from inferences and assumptions.
- Do not claim that experiments, tests, retrieval runs, indexing jobs, or evaluations were completed unless results are provided.
- If required information is missing, state the missing evidence and design the smallest reasonable way to obtain it.
- Preserve uncertainty where the available evidence does not justify a firm conclusion.
- Before defining experiment arms, treat the baseline retrieval configuration, candidate changes, representative corpus and query evidence, and answer-evaluation standard as blocking when absent or materially conflicting. Request all blocking items in one consolidated clarification and stop the affected design decision until they are supplied. Record other gaps as Unknown with their effect on sizing, confounding, metrics, and acceptance criteria.
- Treat the retrieval owner as accountable for experiment execution, the data owner as accountable for corpus and metadata validity, and the product owner or domain owner as accountable for acceptance criteria tied to user impact.
Produce the following deliverable:
1. Experiment objective and decision boundary
- State the configuration decision this experiment will support.
- Define what is in scope: chunking strategy, overlap, boundary rules, parent-child or hierarchical chunking if relevant, metadata extraction, metadata normalization, and metadata use in filtering or ranking.
- Define what is out of scope and should be held constant: embedding model, reranker, generation prompt, answer model, UI behavior, permissions, and production rollout unless the inputs require otherwise.
- Identify the baseline configuration and the candidate configurations to compare.
2. Controlled variables and constants
Create a table with:
- Variable under test
- Candidate levels
- Why it matters
- Expected retrieval effect
- Required implementation evidence
- What must remain constant
Include at minimum:
- Chunk size or token range
- Chunk overlap
- Chunk boundary rule
- Structural handling of headings, tables, lists, or sections
- Metadata fields to attach
- Metadata normalization rules
- Metadata use pattern: display only, filter, boost, grouping, or attribution
3. Factorial experiment matrix
Build a practical factorial or fractional-factorial matrix that isolates chunking and metadata effects without creating unnecessary runs.
For each experiment arm include:
- Arm ID
- Chunking configuration
- Metadata configuration
- Retrieval settings held constant
- Indexing requirements
- Expected comparison value
- Risk of confounding
- Minimum evidence needed before execution
If the full factorial design is too large, propose a staged design:
- Screening stage to remove weak options
- Focused comparison stage for finalists
- Confirmation stage against the baseline
4. Dataset and query slices
Define the evaluation dataset and query slices needed to detect meaningful differences.
Include:
- Document sample selection rules
- Minimum corpus coverage requirements
- Query source: logs, expert-written questions, synthetic-but-reviewed questions, or known support cases
- Query slice taxonomy
- Required number of queries per slice if enough information is available; otherwise give a sizing rule
Use query slices such as:
- Exact lookup questions
- Multi-section synthesis questions
- Long-tail entity questions
- Recent or version-specific questions
- Table or list extraction questions
- Procedure or policy questions
- Ambiguous terminology questions
- Metadata-dependent questions
- Queries where the answer should not be in the corpus
For each slice, specify the retrieval failure mode it is intended to expose.
5. Relevance judgments and answer evaluation setup
Define how ground truth or reference judgments should be created.
Include:
- Who should judge relevance and answer correctness
- What evidence they need
- How to label relevant, partially relevant, irrelevant, stale, duplicate, and misleading chunks
- How to handle multiple valid source passages
- How to record uncertainty and disagreements
- How to prevent evaluators from seeing the tested configuration when practical
6. Retrieval metrics
Specify retrieval metrics that directly measure chunking and metadata effects.
Include:
- Recall@k
- Precision@k or context precision
- MRR or nDCG where graded relevance exists
- Source coverage for multi-source answers
- Duplicate or near-duplicate rate in top-k
- Metadata filter precision and filter fallout
- Freshness or version correctness when metadata includes time or version fields
- No-answer retrieval behavior for queries outside the corpus
For each metric, define:
- Formula or scoring method in plain language
- Required inputs
- Which query slices it applies to
- What kind of configuration failure it reveals
7. Answer-level metrics
Define answer metrics only as downstream checks of retrieval configuration, not as a full generation evaluation.
Include:
- Answer correctness against the evaluation standard
- Citation/source support
- Missing critical fact rate
- Unsupported claim rate
- Wrong-version or wrong-jurisdiction rate if relevant
- Refusal or no-answer correctness where the corpus lacks the answer
Explain how to attribute answer failures to retrieval versus generation when the top-k evidence is sufficient but the answer is wrong.
8. Error attribution rules
Create a concrete error taxonomy with decision rules. Include at minimum:
- Chunk boundary split error
- Chunk too broad or noisy
- Chunk too narrow or missing context
- Overlap redundancy error
- Metadata absent
- Metadata incorrect
- Metadata too coarse
- Metadata filter excludes relevant evidence
- Metadata boost overpromotes irrelevant evidence
- Duplicate chunk crowding
- Stale or wrong-version retrieval
- Relevant evidence not indexed
- Relevant evidence indexed but not retrieved
- Answer failure despite sufficient retrieved evidence
For each error type, define:
- Observable evidence
- How to distinguish it from similar errors
- Likely configuration implication
- Whether it should count against chunking, metadata, indexing, retrieval settings, or answer generation
9. Analysis plan by query slice
Define how results should be compared across arms and slices.
Include:
- Primary metric and secondary metrics
- Minimum practical improvement threshold
- Regression checks where a configuration improves one slice but harms another
- Treatment of ties
- Treatment of small sample sizes
- Required confidence or stability checks if repeated runs are possible
- How to summarize tradeoffs for decision owners without hiding slice-level failures
10. Configuration acceptance decision
Create a decision framework the retrieval owner, data owner, and product owner can use after the experiment is run.
Include:
- Acceptance criteria for adopting a new chunking and metadata configuration
- Rejection criteria
- Conditional acceptance criteria requiring remediation
- Required evidence package before approval
- Rollback or re-indexing considerations if adopted
- Open questions that must be resolved before production use
11. Completion checklist
End with a checklist that confirms whether the experiment design is ready to execute. Include checks for:
- Baseline defined
- Candidate arms defined
- Constants identified
- Query slices complete
- Relevance judgment process defined
- Metrics mapped to slices
- Error attribution rules defined
- Acceptance criteria agreed by named accountable owners
- Missing evidence listed
- No unsupported claims of completed execution
Investigate whether a RAG system crossed identity, tenant, purpose, region, or document authorization boundaries during retrieval, embedding, caching, citation, or response exposure.
Updated Aug 19, 2026
Investigate whether the RAG system described below retrieved, embedded, cached, cited, or exposed content outside the requesting identity, tenant, purpose, region, or document authorization boundary.
Do not perform a general RAG quality review, connector security review, model safety review, or relevance evaluation unless it directly affects authorization leakage evidence. Use only the evidence provided. Do not claim that a log, policy, source, command, system, test, approval, or remediation was inspected or completed unless it is present in the supplied materials.
Context and inputs:
- RAG system and deployment context: [RAG system and deployment context]
- Identity, tenant, purpose, region, and document authorization model: [Identity tenant purpose region and document authorization model]
- Evidence bundle, including any retrieval logs, citation traces, document ACLs, embedding ingestion records, cache records, query logs, response samples, tickets, diagrams, or policy excerpts: [Evidence bundle]
- Suspected leakage scenarios or query samples: [Suspected leakage scenarios or query samples]
- Time window and affected environments: [Time window and affected environments]
- Accountable owners and decision deadline: [Accountable owners and decision deadline]
Investigation rules:
1. Separate observations from inference. Label each material statement as Evidence, Inference, Gap, or Decision.
2. Preserve uncertainty. If evidence is missing, conflicting, sampled, redacted, or outside the time window, say so plainly.
Missing-input gate:
- Treat the requesting identity, applicable authorization model, affected environment and time window, and sufficient retrieval or exposure traces as blocking for any confirmed leakage conclusion. Request all blocking items in one consolidated clarification and stop the affected conclusion until they are supplied.
- Continue with non-blocking gaps only when they are marked Unknown and their effect on path confidence, affected scope, containment, and regression coverage is explicit.
3. Trace authorization from requesting identity to document exposure. Include retrieval, ranking, embedding, cache, citation, prompt assembly, generated response, export, telemetry, and feedback paths where evidence exists.
4. Treat access-control boundaries independently: identity, tenant, purpose, region, document-level authorization, group membership, delegated access, service account access, and temporal entitlement changes.
5. Distinguish content that was indexed, embedded, retrieved, ranked, cached, cited, included in context, generated in the answer, logged, or externally exposed.
6. Do not assume that document-level authorization at source ingestion remains valid at query time. Look for entitlement drift, stale indexes, overbroad service accounts, cache reuse, citation leakage, and post-filter bypass.
7. Do not recommend broad rewrites. Recommend the smallest containment and verification actions that fit the observed risk.
8. Assign owner accountability by role where needed: security reviewer, data owner, product owner, release owner, search/retrieval owner, platform owner, or legal/privacy owner.
Deliverable:
1. Investigation scope
- State the system, environments, time window, suspected boundary, and query/document populations covered.
- State what is explicitly out of scope.
- List the supplied evidence types and any major missing evidence.
2. Evidence register
Create a table with columns:
- Evidence ID
- Evidence type
- Source or artifact name
- Time range
- What it shows
- Authorization boundary relevance
- Limitations or uncertainty
3. Identity-to-document authorization map
Create a map or table showing:
- Requesting identity or role
- Tenant or account
- Purpose or workflow
- Region or data residency boundary
- Entitlement source
- Allowed document classes or IDs
- Disallowed document classes or IDs
- Query-time enforcement point
- Retrieval/index/cache/citation enforcement point
- Evidence supporting the mapping
- Gaps requiring owner confirmation
4. Leakage-path reconstruction
For each suspected or observed leakage path, reconstruct the sequence from request to exposure:
- Query or trigger
- Requesting identity and authorization state at the time
- Retrieval/index candidate set
- Filter or authorization check expected
- Filter or authorization check observed
- Document, chunk, citation, cache entry, embedding, or response involved
- Exposure mode: retrieved, ranked, cached, cited, placed in prompt context, generated in response, logged, exported, or visible in UI/API
- Boundary crossed: identity, tenant, purpose, region, document authorization, or timing/entitlement drift
- Evidence supporting the path
- Alternative explanations
- Confidence level: confirmed, probable, possible, or not supported
5. Affected corpus and query scope
Estimate the likely scope without overstating certainty:
- Affected tenants, identities, groups, regions, document classes, indexes, embeddings, caches, and environments
- Query patterns likely to trigger the issue
- Whether exposure appears isolated, systemic, configuration-specific, connector-specific, cache-specific, or time-bound
- Known false-positive or false-negative risks in the evidence
- Additional evidence needed to narrow scope
6. Containment decision
Provide a decision-ready containment assessment for the accountable owners:
- Immediate risk level and rationale
- Recommended containment action: no action, monitor, disable affected queries, invalidate cache, remove affected index, restrict service account, re-ingest with corrected ACLs, disable citation/source display, tenant isolate, region isolate, or pause release
- Why this is the smallest safe containment action supported by the evidence
- Operational impact
- Data owner, security reviewer, product owner, and release owner decisions required
- Conditions for lifting containment
7. Regression test matrix
Create a test matrix that can be used by engineering and security reviewers. Include:
- Test ID
- Boundary tested
- Test identity or role
- Allowed corpus
- Forbidden corpus
- Query pattern
- Expected retrieval behavior
- Expected citation behavior
- Expected cache behavior
- Expected response behavior
- Required logs or traces
- Pass/fail criteria
- Owner responsible for verification
Include tests for at least:
- Same-tenant allowed document retrieval
- Same-tenant forbidden document exclusion
- Cross-tenant exclusion
- Region boundary exclusion
- Purpose-boundary exclusion where applicable
- Stale group membership or entitlement removal
- Cache reuse across identities or tenants
- Citation/source leakage without answer leakage
- Embedding/index rebuild after ACL correction
- Service account or delegated access path
8. Open questions and evidence requests
List only the missing items that materially affect the containment decision, scope estimate, or regression coverage. Assign each request to a specific owner role.
9. Completion checks
End with a concise checklist confirming whether the current evidence is sufficient to:
- Determine if leakage occurred
- Identify the likely leakage path
- Estimate affected corpus and query scope
- Make a containment decision
- Define regression tests
- Proceed to remediation planning
If any check cannot be completed from the supplied evidence, mark it as incomplete and explain the exact missing evidence.
Identify stale, time-sensitive, or decision-unsafe knowledge assets by linking source authority, change events, usage, reviews, and retrieval exposure.
Updated Aug 19, 2026
Determine where the enterprise knowledge corpus has become stale, time-sensitive, or decision-unsafe by relating source authority, change events, usage, review evidence, and retrieval exposure. Focus on freshness decay and revalidation evidence, not a broad knowledge-base quality audit.
Context and inputs to provide:
- Knowledge corpus name: [Knowledge corpus name]
- Corpus export or representative content set: [Corpus export or representative content set]
- Source inventory and authority rules: [Source inventory and authority rules]
- Known change events and effective dates: [Known change events and effective dates]
- Usage and retrieval exposure data: [Usage and retrieval exposure data]
- Existing review evidence and owner list: [Existing review evidence and owner list]
- Decision horizon and risk tolerance: [Decision horizon and risk tolerance]
Evidence rules:
- Separate observed evidence from inference. Label each conclusion as Observed, Inferred, or Unknown.
- Do not claim that a source, log, review, owner approval, or retrieval result was inspected unless it appears in the provided material.
- Preserve uncertainty where evidence is incomplete, contradictory, old, or sampled.
- Prefer source-effective dates, documented review records, owner attestations, change logs, and actual retrieval or usage evidence over assumptions.
- If the corpus export is incomplete, treat the result as a scoped review and state the coverage limits.
Authority boundaries:
- Do not make final compliance, legal, medical, financial, security, or customer-impact determinations unless an accountable owner has provided the relevant authority rule in the inputs.
- Assign proposed decisions to accountable roles such as knowledge owner, source authority, product owner, policy owner, compliance owner, data owner, search owner, or support operations owner.
- Where the correct owner is unclear, mark Owner Unknown and specify the evidence needed to assign accountability.
Review method:
1. Establish the freshness policy map.
- Identify content classes that require freshness controls, such as policy, product behavior, pricing, regulated guidance, operational runbooks, customer-facing answers, security procedures, data definitions, and historical reference material.
- For each class, map source authority, expected review cadence, change triggers, acceptable age, required evidence of review, and decision owner.
- If no explicit policy exists, infer a provisional freshness expectation from risk, source type, and decision use; mark it as Inferred.
2. Detect decay signals.
- Compare corpus claims against known source changes, effective dates, product or policy updates, review timestamps, ownership changes, usage patterns, and retrieval exposure.
- Identify assets with missing source authority, expired review windows, superseded references, conflicting versions, orphaned ownership, high retrieval exposure, or recent source changes without revalidation.
- Distinguish content that is merely old from content that is decision-unsafe.
3. Build the decay-risk register.
For each material risk, provide:
- Asset or content group
- Decision or workflow it may affect
- Decay signal observed
- Source authority status
- Last known review evidence
- Retrieval or usage exposure
- Potential consequence if used as-is
- Confidence level and evidence basis
- Accountable owner
- Recommended disposition: Retain, Revalidate, Update, Restrict, Retire, or Split/Merge
4. Build the source-change impact matrix.
- List each known source change or effective-date event.
- Map the affected corpus assets, topics, audiences, workflows, and retrieval surfaces.
- Identify whether the corpus reflects the change, partially reflects it, conflicts with it, or lacks enough evidence to determine impact.
- Prioritize changes that affect high-exposure retrieval, customer-facing guidance, regulated decisions, operational safety, financial terms, security controls, or contractual commitments.
5. Produce the revalidation queue.
- Rank assets by decision risk, retrieval exposure, source-change proximity, review age, and owner availability.
- For each item, specify the minimum evidence needed to clear or confirm the risk.
- Assign an accountable owner and a practical next step.
- Separate urgent restrictions from routine revalidation work.
6. Make retire, restrict, update, or revalidate decisions.
- Recommend a disposition for each priority item.
- Use Restrict when content could cause harmful or materially wrong decisions before validation is complete.
- Use Retire when content is superseded, duplicate, ownerless with no defensible authority, or no longer tied to a valid business use.
- Use Update when authoritative replacement information is available.
- Use Revalidate when the content may still be valid but lacks current review evidence.
- Use Retain only when freshness requirements, source authority, and review evidence are adequate for the stated decision horizon.
Output format:
A. Scope and evidence coverage
- Corpus reviewed
- Materials provided
- Materials not provided but needed
- Coverage limits
- Assumptions and uncertainty
B. Freshness policy map
Create a table with columns: Content class | Source authority | Review cadence or trigger | Acceptable age | Required review evidence | Decision owner | Basis: Observed/Inferred/Unknown.
C. Decay-risk register
Create a table with columns: Priority | Asset or group | Affected decision/workflow | Decay signal | Source authority status | Review evidence | Retrieval/usage exposure | Consequence | Confidence | Owner | Recommended disposition.
D. Source-change impact matrix
Create a table with columns: Source change | Effective date | Affected assets/topics | Exposure surface | Alignment status | Risk level | Required action | Owner.
E. Revalidation queue
Create a table with columns: Queue rank | Asset or group | Reason for queueing | Minimum evidence needed | Proposed verifier | Target decision | Urgency | Dependency.
F. Retire/restrict/update decisions
Create a table with columns: Asset or group | Decision | Rationale | Evidence basis | Owner to confirm | Completion check.
G. Missing evidence and unresolved questions
List the specific missing logs, source records, review attestations, authority rules, owner assignments, or retrieval evidence needed to convert Unknown or Inferred judgments into Observed judgments.
H. Completion criteria
State whether this review is complete enough to support immediate restriction, update, retirement, or revalidation planning. Identify which decisions require confirmation by the relevant knowledge owner, source authority, product owner, compliance owner, data owner, or search owner before implementation.
Diagnose where AI capability loses operational value across adoption, workflow, quality, controls, capacity, rework, and measurement.
Updated Aug 19, 2026
Diagnose why measured AI capability is not becoming realized operational value for the process below. Separate adoption leakage, workflow leakage, quality leakage, control leakage, capacity leakage, rework leakage, and measurement leakage. Do not produce a generic AI adoption roadmap or ROI plan; the job is to explain the value gap and define evidence-backed recovery actions.
Input to use:
- Business process or function: [Business process or function]
- AI capability or system introduced: [AI capability or system introduced]
- Expected value case: [Expected value case]
- Actual operating results: [Actual operating results]
- Adoption and workflow evidence: [Adoption and workflow evidence]
- Quality, control, and rework evidence: [Quality, control, and rework evidence]
- Accountable owners and operating constraints: [Accountable owners and operating constraints]
Evidence rules:
- Treat provided facts, metrics, dates, artifacts, and firsthand observations as evidence.
- Label estimates, explanations, and causal links as inference unless directly evidenced.
- Do not claim access to systems, logs, tests, approvals, user interviews, dashboards, or financial records unless they are included in the input.
- Preserve uncertainty. State what is unknown, what evidence would reduce uncertainty, and whether the missing evidence blocks a decision.
- If the evidence is too thin to quantify a loss, provide a bounded qualitative assessment and specify the minimum evidence needed to quantify it.
Diagnostic method:
1. Define the expected value flow from available AI capability to realized operational value.
2. Identify the observable handoffs where value can leak: user adoption, task selection, workflow integration, output quality, review controls, throughput capacity, exception handling, rework, measurement, and benefit capture.
3. Compare the expected value case against actual operating results. Separate volume effects, quality effects, cycle-time effects, labor-mix effects, risk/control effects, and measurement effects where possible.
4. Identify root leakage mechanisms before proposing interventions. Do not assume low adoption is the only cause.
5. Distinguish symptoms from mechanisms. Example: “low usage” is a symptom; “users avoid the tool because outputs require unplanned specialist review” is a leakage mechanism if supported by evidence.
6. Consider whether measured AI capability is being trapped by downstream constraints, approval queues, exception rates, policy limits, data quality, incentive conflicts, training gaps, workflow redesign gaps, or unmanaged rework.
7. Preserve existing operational constraints. Do not recommend broad transformation unless the evidence shows incremental recovery actions cannot address the leakage.
Required deliverable:
1. Value-flow map
Create a concise map showing:
- AI capability available
- Intended user or workflow touchpoint
- Expected behavior change
- Expected operational effect
- Expected financial, service, risk, or capacity value
- Actual observed path
- Known breakpoints or uncertain handoffs
2. Leakage mechanism ledger
Provide a ledger with these columns:
- Leakage category
- Observed evidence
- Inference about mechanism
- Value affected
- Likely owner
- Confidence level: High, Medium, or Low
- Missing evidence
- Immediate diagnostic check
Use these leakage categories where relevant:
- Adoption leakage
- Workflow leakage
- Quality leakage
- Control or approval leakage
- Capacity leakage
- Rework leakage
- Measurement leakage
- Benefit capture leakage
3. Evidence-backed loss bridge
Build a bridge from expected value to realized value. Use available numbers where provided. If numbers are missing, use directional bands and explain the basis.
Include:
- Starting expected value
- Loss or dilution from adoption
- Loss or dilution from workflow fit
- Loss or dilution from quality or rework
- Loss or dilution from controls or review load
- Loss or dilution from capacity constraints
- Loss or dilution from measurement or attribution error
- Realized value observed
- Unexplained residual gap
4. Intervention hypotheses
For each major leakage mechanism, define a testable recovery hypothesis:
- Hypothesis
- Evidence that supports it
- Evidence that contradicts or weakens it
- Smallest safe intervention
- Expected movement in adoption, quality, cycle time, cost, risk, or capacity
- Leading indicator to monitor
- Timeframe for signal
- Risk if wrong
5. Owner-specific recovery plan
Create an owner-specific plan using the named owners or the most likely accountable roles from the input. Include only actions supported by the diagnosis.
For each action, specify:
- Accountable owner, such as process owner, product owner, operations owner, finance owner, data owner, risk/control owner, or enablement owner
- Decision required
- Evidence required before action, if any
- Recovery action
- Completion criterion
- Verification method
- Dependency or constraint
6. Decision view
Conclude with:
- Most likely primary leakage mechanism
- Secondary leakage mechanisms
- Whether the current evidence supports intervention, further diagnosis, or pausing expansion
- Highest-value next diagnostic check
- Decisions each accountable owner should make next
- What would change your conclusion
Completion criteria:
- The diagnosis explains the gap between available AI capability and realized operational value, not merely whether users adopted the tool.
- Every material claim is tied to evidence, inference, or stated uncertainty.
- The recovery plan assigns accountable owners and observable completion criteria.
- No recommendations depend on unavailable inspections, unverified approvals, or assumed system behavior.
Attribute full AI operating costs to accepted outcomes so owners can compare true unit economics across workflows and variants.
Updated Aug 19, 2026
Build a defensible AI operating cost attribution and unit economics model for the workflow below. The purpose is to attribute model, tool, retrieval, infrastructure, review, support, failure, and rework costs to accepted AI outcomes so accountable owners can compare true unit economics across workflows and variants.
Context to provide:
- AI workflow or product: [AI workflow or product]
- Accepted outcome definition: [Accepted outcome definition]
- Analysis period: [Analysis period]
- Workflow variants to compare: [Workflow variants to compare]
- Cost evidence and usage exports: [Cost evidence and usage exports]
- Quality and rework evidence: [Quality and rework evidence]
- Accountable owners: [Accountable owners]
Evidence rules:
- Use only the evidence provided. Do not claim that a source, export, invoice, log, ticket, approval, or test was inspected unless it is included in the input.
- Separate observed facts, calculated values, assumptions, and inferences.
- If evidence is missing, label it as missing and explain how it affects confidence or comparability.
- Preserve uncertainty with ranges or confidence notes where point estimates are not supported.
- Treat the accepted-outcome definition, analysis period, cost and usage sources, and comparable workflow boundaries as blocking when they are missing or materially conflicting. Request all blocking items in one consolidated clarification and do not calculate the affected unit economics until they are resolved. Continue with non-blocking gaps only as explicit Unknown or provisional inputs, with their effect on formulas, confidence, and comparability.
- Do not create a general AI spend governance brief or a quality/latency experiment plan. Stay focused on cost attribution and accepted-outcome unit economics.
Build the working artifact with these sections:
1. Cost Boundary Map
Create a boundary map for the selected [AI workflow or product] covering the full operating cost chain for [Analysis period]. Include, where evidenced:
- Model inference or subscription costs
- Tool execution costs
- Retrieval, embedding, vector database, search, or storage costs
- Application, orchestration, logging, monitoring, and infrastructure costs
- Human review, approval, escalation, and exception-handling costs
- Support, customer operations, and internal operations costs
- Failure, incident, reversal, refund, remediation, and rework costs
- Evaluation, sampling, QA, audit, and oversight costs directly tied to production operation
For each cost category, state whether it is included, excluded, partially included, or missing; identify the cost owner from [Accountable owners] where possible; and explain the rationale.
2. Accepted-Outcome Measurement
Define the denominator for unit economics using [Accepted outcome definition]. Specify:
- What counts as an accepted outcome
- What does not count
- How retries, duplicates, partial completions, escalations, rejected outputs, and reworked outputs should be treated
- Any quality threshold required before an outcome is counted
- The evidence source used for the count
If the accepted-outcome count cannot be established from [Quality and rework evidence], provide a defensible interim method and list the validation needed from the product owner, finance owner, or operations owner.
3. Allocation Rules
Design allocation rules that assign costs to accepted outcomes and to [Workflow variants to compare]. For each rule, include:
- Cost pool
- Allocation driver
- Formula
- Required evidence
- Owner responsible for validating the driver
- Known weakness or bias
- When the rule should be replaced with a more precise method
Prefer causal allocation drivers over broad averages. Use broad averages only when more precise evidence is unavailable, and label them clearly.
4. Accepted-Outcome Unit Economics
Create a unit economics table for each workflow variant in [Workflow variants to compare]. Include:
- Accepted outcome volume
- Gross operating cost
- Cost excluded or not yet evidenced
- Cost per accepted outcome
- Cost per attempted outcome, if attempt counts are available
- Human review cost per accepted outcome
- Failure and rework cost per accepted outcome
- Infrastructure and retrieval cost per accepted outcome
- Confidence level for each major figure
Show formulas and make assumptions explicit. Do not overstate precision.
5. Variance Bridge
Build a variance bridge explaining differences in cost per accepted outcome across variants or periods. Attribute variance where evidence supports it to:
- Volume and utilization
- Model selection or token/input-output mix
- Tool calls and external service usage
- Retrieval depth, indexing, or storage patterns
- Review rates and escalation rates
- Failure, rejection, incident, or rework rates
- Infrastructure utilization or fixed-cost absorption
- Support burden
For each variance driver, state whether the driver is observed, calculated, inferred, or currently unverified.
6. Cost-Quality Sensitivity
Analyze how unit economics change under plausible changes to quality and control variables, such as:
- Acceptance rate
- Human review rate
- Escalation rate
- Rework rate
- Failure or incident rate
- Retrieval depth or tool usage
- Model choice or routing mix
Use ranges when evidence is incomplete. Identify which variables most affect cost per accepted outcome and which quality controls appear economically justified.
7. Control Decisions
Provide a decision table for the accountable owners. Include:
- Decision under consideration
- Economic rationale
- Quality or risk tradeoff
- Evidence supporting the decision
- Evidence still missing
- Owner who should approve or validate the decision
- Completion check
Focus on concrete controls such as routing changes, review thresholds, retrieval limits, escalation criteria, failure handling, logging improvements, or measurement changes.
8. Model Integrity Checks
Before finalizing, perform these checks using the provided evidence:
- Every included cost has an owner, source, allocation rule, and treatment in the model.
- Accepted-outcome counts match the stated definition or are flagged as provisional.
- Rejected, failed, duplicated, escalated, and reworked outputs are not accidentally counted as accepted outcomes unless justified.
- Fixed, variable, and step costs are not mixed without explanation.
- Variant comparisons use the same boundary unless differences are explicitly disclosed.
- No unavailable inspection, execution, approval, or source verification is claimed.
Final output format:
- Cost boundary map
- Allocation rule table
- Accepted-outcome unit economics table
- Variance bridge
- Cost-quality sensitivity table
- Control decision table
- Missing evidence register
- Owner verification checklist
Write in direct finance and operating language suitable for review by the finance owner, product owner, operations owner, and data owner. Avoid generic governance language and unsupported certainty.
Convert pilot and operating evidence into a defensible decision on whether one AI initiative should scale, hold, redesign, or stop.
Updated Aug 19, 2026
Prepare a decision brief for one AI initiative. The objective is to recommend scale, hold, redesign, or stop using only the evidence provided, while making uncertainty, missing information, and decision authority explicit.
Inputs to use:
- Initiative name: [Initiative name]
- Decision owner: [Decision owner]
- Pilot or operating scope and timeframe: [Pilot or operating scope and timeframe]
- Evidence package: [Evidence package]
- Current constraints and non-negotiables: [Current constraints and non-negotiables]
- Decision options under consideration: [Decision options under consideration]
- Next decision date or trigger: [Next decision date or trigger]
Evidence discipline:
- Distinguish observed evidence from inference. Do not treat anecdotes, forecasts, vendor claims, or unverified internal estimates as confirmed facts.
- Do not claim that a source, approval, test, control, incident review, financial model, or system inspection was completed unless the evidence package shows it.
- If evidence is missing, state what is missing, why it matters, and the minimum evidence needed for a safer decision.
- Preserve uncertainty. Use confidence levels only when tied to evidence quality.
- Stay within decision-support authority: produce a recommendation and decision conditions; do not imply approval unless the accountable owner has already approved it in the evidence.
- Avoid generic AI governance commentary. Focus on this initiative’s value, adoption, controls, reliability, dependencies, economics, risk boundaries, and reversibility.
Decision standard:
Recommend exactly one primary disposition: Scale, Hold, Redesign, or Stop.
Use these meanings:
- Scale: expand use because value, adoption, controls, reliability, economics, dependencies, and reversibility are acceptable within defined boundaries.
- Hold: continue limited operation or pause expansion because evidence is incomplete or conditions are not yet met, but the initiative may still be viable.
- Redesign: materially change workflow, model approach, controls, operating model, vendor setup, data inputs, or user experience before further expansion.
- Stop: end or sunset the initiative because evidence does not support continued investment or risk is unacceptable relative to value and reversibility.
Deliver the brief in the following format:
1. Decision frame
- Initiative: state the initiative and operating scope.
- Decision needed: state the decision being made now.
- Accountable owner: identify the decision owner and any role-specific verifiers needed, such as product owner, data owner, security reviewer, legal/compliance owner, finance owner, operations owner, or release owner.
- Authority boundary: state what this brief can recommend versus what requires owner approval.
- Time boundary: state the next decision date or trigger.
2. Decision evidence map
Create a table with these columns: Decision dimension, Observed evidence, Inference or assumption, Evidence strength, Missing information, Decision implication.
Include these dimensions at minimum:
- Intended business value
- Realized value or leading indicators
- User adoption and workflow fit
- Output quality and reliability
- Control effectiveness and exception handling
- Data, model, vendor, and system dependencies
- Security, privacy, compliance, and policy constraints
- Operating support and ownership capacity
- Scale economics and marginal cost
- Reversibility, rollback, and exit cost
3. Gate-by-gate disposition
For each gate, assign Pass, Conditional Pass, Fail, or Insufficient Evidence. Explain the reason in 2 to 4 sentences per gate.
Gates:
- Value gate: evidence of meaningful value relative to effort and alternatives.
- Adoption gate: users, operators, or customers can and do use it in the intended workflow.
- Control gate: risks, permissions, review paths, exceptions, and accountability are workable.
- Reliability gate: quality, availability, latency, and failure modes are acceptable for the use case.
- Dependency gate: critical data, vendor, model, integration, and staffing dependencies are known and manageable.
- Economics gate: scale costs, support costs, and expected benefits remain acceptable beyond the pilot.
- Reversibility gate: the initiative can be rolled back, contained, or sunset without unacceptable disruption.
4. Counterfactual options
Compare the realistic options, not just the preferred one. Include at least:
- Scale now
- Hold in current scope
- Redesign before expansion
- Stop or sunset
- Non-AI or lower-automation alternative
For each option, provide: what would happen, expected upside, main downside, investment or effort required, risk exposure, reversibility, and what evidence would make this option stronger or weaker.
5. Scale economics and risk boundaries
Provide a practical scale boundary, even if the recommendation is not to scale.
Include:
- Unit or marginal cost drivers, using provided evidence only.
- Expected cost changes at expanded volume.
- Support, monitoring, exception handling, and owner workload implications.
- Benefits that are evidenced versus speculative.
- Risk ceilings that should not be exceeded.
- Required controls before expansion.
- Stop-loss triggers or rollback conditions.
- Any financial assumptions that the finance owner should verify before approval.
6. Recommendation
State one primary recommendation: Scale, Hold, Redesign, or Stop.
Then provide:
- Rationale tied to the gates and evidence map.
- Conditions required before execution.
- Evidence that argues against the recommendation.
- Residual risks the owner would knowingly accept.
- What would change the recommendation.
7. Authorized next-decision brief
Create an owner-ready next-decision plan with:
- Decision to be requested from [Decision owner].
- Roles that should verify specific parts of the brief before action, such as finance owner for economics, security reviewer for security boundaries, data owner for data use, product owner for workflow value, operations owner for support readiness, and legal/compliance owner where applicable.
- Immediate actions, responsible role, due date or trigger, and required evidence of completion.
- Metrics or observations to collect before the next decision.
- Completion checks that would show the recommendation has been executed or is ready for escalation.
- Explicit statement of any unresolved blockers.
Final quality check before answering:
- The brief decides the fate of one operating initiative; it does not prioritize a portfolio.
- Observations and inferences are clearly separated.
- Missing information is visible and decision-relevant.
- The recommendation is bounded by economics, risk, controls, dependencies, and reversibility.
- No unavailable inspection, approval, test, or execution is claimed.