AI Trace Review and Failure Taxonomy
Use ChatGPT to reconstruct evidence-supported AI and agent traces, govern failure classifications, analyze recurring patterns within sampling limits, identify observability gaps, and propose sanitized regression cases and measurable prevention work.
Analyze the supplied AI-system evidence to reconstruct observable traces, classify supported failure mechanisms, measure recurring patterns within the sampling design, identify observability gaps, and propose prevention and regression-evaluation work. Inputs - Review scope and decisions: [Review scope and decisions] - System and telemetry map: [System and telemetry map] - Population and sampling design: [Population and sampling design] - Sanitized trace evidence: [Sanitized trace evidence] - Outcome and incident evidence: [Outcome and incident evidence] - Taxonomy and measurement definitions: [Taxonomy and measurement definitions] - Governance and privacy constraints: [Governance and privacy constraints] - Acceptance criteria and authorized analyses: [Acceptance criteria and authorized analyses] Input handling 1. First check whether the review objective, analysis unit, review period, eligible population, sample construction, trace identifiers, environment and component versions, outcome criteria, privacy boundary, and authorized analyses are sufficiently defined. 2. If a missing input blocks reconstruction, classification, denominator selection, or safe handling, ask for all blocking items in one consolidated list and stop before drawing conclusions. 3. If a gap is non-blocking, continue only after labeling the relevant statement as an assumption, hypothesis, unknown, unavailable measurement, or owner decision. 4. Preserve conflicting timestamps, versions, statuses, labels, and outcomes. Do not resolve a conflict by preference; identify the evidence needed to discriminate between alternatives. 5. Use only sanitized, access-approved material supplied in the conversation. Do not request secrets, credentials, unrestricted customer content, confidential prompts, private endpoints, or unnecessary tool payloads. ChatGPT operating boundary ChatGPT may organize, compare, classify, and calculate from the supplied material. It must not claim to have queried telemetry stores, opened unavailable files, inspected production, contacted providers, executed tools, confirmed external state, changed sampling, altered retention, deployed a fix, or run an evaluation unless direct evidence of that action is supplied. Do not request or infer hidden chain-of-thought. Use these work-status terms consistently: - Requested: an action or artifact was sought, but completion evidence was not supplied. - Proposed: a recommendation, taxonomy change, check, or evaluation that has not been authorized or performed. - Executed: contemporaneous evidence shows that an operation was attempted; this does not by itself prove its intended effect. - Unavailable: the required source or field was not supplied or accessible for this review. - Unverified: a claim, result, or side effect lacks independent confirmation. - Verified: the supplied authoritative source and acceptance evidence support the claim. Evidence rules - Separate supplied facts, direct observations, derived calculations, assumptions, hypotheses, unknowns, unsupported claims, conflicts, recommendations, and verified results. - Cite each material source by its supplied identifier, environment, time window, component version, schema or semantic-convention version, and sampling status where available. - Treat a missing span as missing evidence, not proof that an operation did not happen. - Account for parent-child links, trace links, retries, fallbacks, parallel branches, queues, asynchronous work, clock skew, and out-of-order arrival before asserting chronology. - Distinguish tool availability, selection, argument construction, approval request, approval result, invocation attempt, returned response or error, independently confirmed side effect, compensation, and final reconciliation. Never collapse these into a single tool-success claim. - Minimize sensitive content. Prefer identifiers, hashes, classifications, derived attributes, and redacted excerpts when full content is unnecessary. - Apply supplied tenant, purpose, region, access, retention, and reviewer restrictions. Refer high-impact privacy, security, safety, financial, legal, regulatory, or contractual judgments to accountable human owners. Focused workflow 1. Define the review boundary: objective, decisions, analysis unit, period, population, sample, exclusions, privacy limits, environments, versions, and definition of done. 2. Inventory telemetry by system stage, including identifiers, schemas, sampling, retention, content-capture policy, owners, authoritative outcome sources, and blind spots. 3. Validate trace integrity and reconstruct the smallest evidence-supported chronology for each authorized case. Mark recorded events, derived timings, inferred transitions, contradictions, and missing evidence separately. 4. Determine the expected outcome, observed outcome, first supported divergence, initiating failure, propagation, visible symptom, detection event, recovery attempt, and final outcome. 5. Draft or refine a hierarchical, versioned taxonomy. Keep outcome status, failure stage, mechanism, symptom, contributing factors, severity, recoverability, owner, and confidence as separate dimensions. 6. Calibrate the taxonomy on a representative pilot when reviewer labels are supplied. Record disagreements and ambiguous rules; do not invent reviewer consensus. 7. Classify the authorized sample. Assign one primary mechanism only when evidence discriminates it from alternatives; otherwise mark it unresolved and retain competing hypotheses. 8. Analyze recurring patterns using explicit numerators and denominators. Reflect sampling, trace loss, missingness, version coverage, detectability, duplicates, retries, and weighting in every rate or comparison. 9. Prioritize patterns using severity, affected scale, exposure, reversibility, detectability, recurrence, and confidence rather than frequency alone. 10. Convert supported findings into proposed observability changes, remediations, and sanitized evaluation cases. Keep historical evidence separate from held-out evaluation data when contamination would invalidate measurement. 11. Finish with the smallest safe next action that reduces material uncertainty or recurring risk without changing production. Classification model Use these outcome states where applicable: Successful, Partially successful, Failed safely, Failed unsafely, Abandoned or timed out, Escalated, Outcome unknown, and Not assessable. Locate the first observable failure stage, such as input or intent, routing, context or memory, retrieval, model generation, validation, safeguard, tool authorization, tool execution, downstream side effect, human escalation, delivery, infrastructure, observability, or user outcome. Define a mechanism narrowly enough to guide prevention. Examples may include retrieval miss, stale context, unsupported model claim, format validation failure, policy false positive, invalid tool arguments, authorization failure, timeout, retry defect, non-idempotent execution, reconciliation failure, queue failure, configuration error, or incomplete telemetry. Treat these as candidate mechanisms, not findings, until supported. Use confidence levels: - High: direct and consistent evidence demonstrates the mechanism. - Moderate: evidence supports the mechanism, but a material gap remains. - Low: multiple mechanisms remain plausible. - Unresolved: available evidence cannot discriminate between competing mechanisms. Every taxonomy label must include a stable code, parent, name, definition, inclusion criteria, exclusions, positive example, near-miss, default owner, related evaluation, introduction version, and retirement or replacement mapping. Measurement rules For each metric, state the numerator, denominator, unit, authoritative source, derivation, missing-data treatment, sampling limitation, and whether the value is observed, calculated, estimated, unavailable, or unverified. Report reviewed-set counts or proportions rather than production rates when the eligible population or sampling probability is unknown. Use percentile, cost, version-comparison, or causal claims only when the supplied sample and definitions support them. Treat correlations as hypotheses and name the cheapest safe discriminating check. Required deliverable Return concise markdown with the following task-specific sections. 1. Boundary and evidence status State the objective, decisions, analysis unit, period, population, sampling method, environments, versions, supplied sources, privacy limits, missing inputs, conflicts, assumptions, authorized analyses, and definition of done. Include a source ledger with evidence status and limitations. 2. Telemetry coverage map Provide a table with: stage or component; owner; telemetry source; identifiers and links; schema or convention version; sampling and retention; sensitive-content handling; authoritative outcome source; known gap. 3. Sample and denominator ledger Provide a table with: population or stratum; eligible cases; reviewed cases; exclusions or missing cases; sampling method; weighting; denominator; permitted inference; limitation. Do not estimate unavailable counts unless an approved method is supplied. 4. Trace reconstruction register Provide a table with: case; ordered stage or event; time or duration; parent, link, branch, or retry; observable evidence; derived or inferred detail; tool-operation status where relevant; outcome; contradiction or missing evidence. Include a trace-completeness assessment for each case. 5. Versioned failure taxonomy Provide a table with: code; parent; stage; mechanism; definition; inclusion criteria; exclusions; positive example; near-miss; default owner; related evaluation; version status. 6. Classified case register Provide a table with: case; outcome status; first observable divergence; primary mechanism or unresolved hypotheses; contributing factors; symptoms; severity; recoverability; detection source; confidence; reviewer status; supporting evidence. Do not report calibration or consensus unless reviewer evidence is supplied. 7. Pattern and impact findings For each material pattern, report observed count, valid denominator, supported rate, severity, affected segments and versions, latency or cost evidence, detection and recovery evidence, sampling and observability limitations, confidence, alternative explanations, and discriminating check. Distinguish association from causation. 8. Observability gap register Provide a table with: priority; blind spot; affected cases or stages; diagnostic consequence; privacy risk; minimum necessary telemetry; proposed owner; approval required; observable acceptance condition. 9. Prevention and evaluation backlog Provide a table with: priority; supporting finding and cases; proposed intervention; expected mechanism; proposed owner; sanitized evaluation case; telemetry required; success and failure measures; regression risk; approval gate; rollback or retirement path. Label all unperformed work Proposed. 10. Governance and verification plan State the proposed taxonomy version, owner and reviewers, calibration requirement, unresolved disagreements, mapping and backfill policy, historical-label preservation rule, privacy and access controls, review cadence, drift triggers, and retest plan. End with: - the smallest safe next action; - accountable owner or owner decision required; - evidence expected; - completion condition; - current status using the defined work-status terms. Acceptance checks Before finalizing, verify from the supplied material that: - the analysis unit, population, sample, exclusions, and denominator are explicit; - sampling, missingness, and trace loss constrain every rate; - schemas, environments, and component versions are recorded when available; - chronology preserves links, branches, retries, fallbacks, asynchronous work, and clock uncertainty; - observations are separated from inference and hidden reasoning is neither requested nor inferred; - tool response, external side effect, and reconciliation are distinct; - stage, mechanism, symptom, impact, owner, and recovery are not conflated; - unresolved mechanisms remain unresolved; - taxonomy codes have inclusion and exclusion rules and a version history; - severe rare cases are not hidden by aggregate frequency; - each recommendation cites affected cases, an owner, approval gate, verification method, observable acceptance condition, and rollback or retirement path; - no inspection, calculation, classification, test, approval, deployment, fix, or outcome is described as completed without supplied evidence. If any acceptance check cannot be met, mark it Unavailable, Unverified, or Owner decision required and explain the narrowest evidence needed.