Reusable AI capability
Diagnose RAG Grounding and Retrieval Failures
Evaluate a retrieval-augmented generation system by separating corpus, retrieval, context assembly, answer generation, citation, and abstention failures, then define reproducible regression criteria.
This Skill packages a reusable way to use the linked Prompt or Workflow; Amo.ng does not run it for you.
# Diagnose RAG Grounding and Retrieval Failures Skill ID: AMO-S-000007 Skill URL: https://amo.ng/skills/diagnose-rag-grounding-and-retrieval-failures Purpose: Produce an evidence-grounded RAG quality assessment that attributes failures to the responsible system layer and turns representative examples into measurable remediation and release criteria. Required inputs: - RAG use case, intended users, and decision or answer types - Representative query set, including difficult, ambiguous, unanswerable, and access-sensitive cases - Corpus scope, source versions, and any known coverage limitations - Captured retrieval results, assembled context, generated answers, citations, and abstention behavior - Expected answers or human relevance judgments where available - Current quality metrics, thresholds, and known failure reports - Applicable privacy, access-control, safety, latency, and release constraints How to use: When to use: - Evaluating a RAG system before release or after a model, corpus, embedding, chunking, or reranking change - Investigating unsupported answers, irrelevant retrieval, incomplete citations, or failures to abstain - Comparing retrieval or generation configurations using a stable evaluation set - Converting observed RAG failures into regression cases and acceptance thresholds When not to use: - Designing an entire RAG architecture from scratch - Treating model fluency or user preference as proof of factual grounding - Evaluating a system without inspectable queries, retrieved context, answers, citations, or reference evidence - Authorizing production deployment or accepting residual security, privacy, or compliance risk Instructions: 1. Use the linked AMO-P-000242 asset as the evaluation framework: supply the query set, corpus description, captured RAG outputs, expected behavior, and current acceptance thresholds rather than asking for a generic RAG review. 2. Define the evaluation unit and success criteria before assessing results. Keep corpus coverage, retrieval relevance, context sufficiency, generation groundedness, citation support, citation completeness, and abstention behavior as separate dimensions. 3. Create a traceable record for each evaluated case, linking the query, retrieved evidence, answer claims, citations, expected behavior, observed failure, and proposed classification. 4. Label every conclusion as a supplied fact, evaluator observation, assumption, inference, missing information, or uncertainty. Do not infer successful retrieval or citation support from an answer that merely sounds plausible. 5. Check whether each material answer claim is supported by the cited passage and whether the citation identifies the correct source. Record partial support, contradictory support, missing support, and inaccessible evidence separately. 6. Attribute failures conservatively. Distinguish absent corpus evidence from poor retrieval, truncated or polluted context, unsupported generation, incorrect citation attachment, and inappropriate refusal or non-abstention. 7. Define remediation hypotheses and bounded tests for the responsible layer, such as corpus repair, chunking changes, metadata filters, reranking, context assembly, prompting, citation rules, or abstention thresholds. 8. Propose a versioned regression set with measurable pass criteria and non-compensable gates for severe grounding, access-control, privacy, or safety failures. 9. Keep technical recommendations separate from release authorization. Require review by the release owner and relevant data, security, privacy, or domain reviewers before production changes, use of sensitive evaluation data, or acceptance of consequential residual risk. Expected output: A RAG evaluation brief containing the evaluation scope, case-level evidence table, layer-specific failure taxonomy, aggregate metrics with uncertainty, corpus and observability gaps, ranked remediation hypotheses, regression suite, acceptance thresholds, and a conditional release recommendation for the release owner and relevant data, security, or privacy reviewers. Constraints and boundaries: - Do not claim that queries were run, sources were retrieved, or configurations were tested unless corresponding execution records are supplied. - Do not use answer plausibility, citation presence, or aggregate scores alone as evidence of correctness. - Preserve source provenance and access boundaries; redact or minimize sensitive query, document, and user data. - Do not collapse retrieval, generation, and citation errors into a single generic hallucination category. - Treat benchmark composition, judge reliability, sampling limits, and missing reference answers as sources of uncertainty. - Production release remains with the release owner; permission changes and security or privacy risk acceptance remain with the designated data, security, and privacy owners. Powered by Prompt: RAG Retrieval and Citation Quality Evaluation Lab Source ID: AMO-P-000242 https://amo.ng/prompts/rag-retrieval-citation-quality-evaluation-lab Completion criteria: Complete when: - Every evaluated finding references a specific query, retrieved passage, answer claim, citation, or supplied execution record. - Material claims are classified as supported, partially supported, contradicted, unsupported, or unverifiable. - Corpus, retrieval, context, generation, citation, and abstention failures are scored or recorded separately. - Aggregate results identify sample size, sampling method, missing data, and uncertainty or evaluator-disagreement limits. - Each recommended change has a corresponding test, expected result, and regression or safety guardrail. - The release recommendation clearly distinguishes observed evidence from unrun tests and unresolved gaps. Use this Amo.ng Skill with your preferred AI tool. Supply the required inputs and follow the usage instructions. # Diagnose RAG Grounding and Retrieval Failures Skill ID: AMO-S-000007 Skill URL: https://amo.ng/skills/diagnose-rag-grounding-and-retrieval-failures Purpose: Produce an evidence-grounded RAG quality assessment that attributes failures to the responsible system layer and turns representative examples into measurable remediation and release criteria. Required inputs: - RAG use case, intended users, and decision or answer types - Representative query set, including difficult, ambiguous, unanswerable, and access-sensitive cases - Corpus scope, source versions, and any known coverage limitations - Captured retrieval results, assembled context, generated answers, citations, and abstention behavior - Expected answers or human relevance judgments where available - Current quality metrics, thresholds, and known failure reports - Applicable privacy, access-control, safety, latency, and release constraints How to use: When to use: - Evaluating a RAG system before release or after a model, corpus, embedding, chunking, or reranking change - Investigating unsupported answers, irrelevant retrieval, incomplete citations, or failures to abstain - Comparing retrieval or generation configurations using a stable evaluation set - Converting observed RAG failures into regression cases and acceptance thresholds When not to use: - Designing an entire RAG architecture from scratch - Treating model fluency or user preference as proof of factual grounding - Evaluating a system without inspectable queries, retrieved context, answers, citations, or reference evidence - Authorizing production deployment or accepting residual security, privacy, or compliance risk Instructions: 1. Use the linked AMO-P-000242 asset as the evaluation framework: supply the query set, corpus description, captured RAG outputs, expected behavior, and current acceptance thresholds rather than asking for a generic RAG review. 2. Define the evaluation unit and success criteria before assessing results. Keep corpus coverage, retrieval relevance, context sufficiency, generation groundedness, citation support, citation completeness, and abstention behavior as separate dimensions. 3. Create a traceable record for each evaluated case, linking the query, retrieved evidence, answer claims, citations, expected behavior, observed failure, and proposed classification. 4. Label every conclusion as a supplied fact, evaluator observation, assumption, inference, missing information, or uncertainty. Do not infer successful retrieval or citation support from an answer that merely sounds plausible. 5. Check whether each material answer claim is supported by the cited passage and whether the citation identifies the correct source. Record partial support, contradictory support, missing support, and inaccessible evidence separately. 6. Attribute failures conservatively. Distinguish absent corpus evidence from poor retrieval, truncated or polluted context, unsupported generation, incorrect citation attachment, and inappropriate refusal or non-abstention. 7. Define remediation hypotheses and bounded tests for the responsible layer, such as corpus repair, chunking changes, metadata filters, reranking, context assembly, prompting, citation rules, or abstention thresholds. 8. Propose a versioned regression set with measurable pass criteria and non-compensable gates for severe grounding, access-control, privacy, or safety failures. 9. Keep technical recommendations separate from release authorization. Require review by the release owner and relevant data, security, privacy, or domain reviewers before production changes, use of sensitive evaluation data, or acceptance of consequential residual risk. Expected output: A RAG evaluation brief containing the evaluation scope, case-level evidence table, layer-specific failure taxonomy, aggregate metrics with uncertainty, corpus and observability gaps, ranked remediation hypotheses, regression suite, acceptance thresholds, and a conditional release recommendation for the release owner and relevant data, security, or privacy reviewers. Constraints and boundaries: - Do not claim that queries were run, sources were retrieved, or configurations were tested unless corresponding execution records are supplied. - Do not use answer plausibility, citation presence, or aggregate scores alone as evidence of correctness. - Preserve source provenance and access boundaries; redact or minimize sensitive query, document, and user data. - Do not collapse retrieval, generation, and citation errors into a single generic hallucination category. - Treat benchmark composition, judge reliability, sampling limits, and missing reference answers as sources of uncertainty. - Production release remains with the release owner; permission changes and security or privacy risk acceptance remain with the designated data, security, and privacy owners. Powered by Prompt: RAG Retrieval and Citation Quality Evaluation Lab Source ID: AMO-P-000242 https://amo.ng/prompts/rag-retrieval-citation-quality-evaluation-lab Completion criteria: Complete when: - Every evaluated finding references a specific query, retrieved passage, answer claim, citation, or supplied execution record. - Material claims are classified as supported, partially supported, contradicted, unsupported, or unverifiable. - Corpus, retrieval, context, generation, citation, and abstention failures are scored or recorded separately. - Aggregate results identify sample size, sampling method, missing data, and uncertainty or evaluator-disagreement limits. - Each recommended change has a corresponding test, expected result, and regression or safety guardrail. - The release recommendation clearly distinguishes observed evidence from unrun tests and unresolved gaps.Copy skill copies the Skill details. Use with AI adds a short instruction for your preferred AI tool; neither action runs the Skill.
Purpose
Produce an evidence-grounded RAG quality assessment that attributes failures to the responsible system layer and turns representative examples into measurable remediation and release criteria.
Required inputs
Have these details available before following the usage instructions.
- RAG use case, intended users, and decision or answer types
- Representative query set, including difficult, ambiguous, unanswerable, and access-sensitive cases
- Corpus scope, source versions, and any known coverage limitations
- Captured retrieval results, assembled context, generated answers, citations, and abstention behavior
- Expected answers or human relevance judgments where available
- Current quality metrics, thresholds, and known failure reports
- Applicable privacy, access-control, safety, latency, and release constraints
How to use this Skill
When to use:
- Evaluating a RAG system before release or after a model, corpus, embedding, chunking, or reranking change
- Investigating unsupported answers, irrelevant retrieval, incomplete citations, or failures to abstain
- Comparing retrieval or generation configurations using a stable evaluation set
- Converting observed RAG failures into regression cases and acceptance thresholds
When not to use:
- Designing an entire RAG architecture from scratch
- Treating model fluency or user preference as proof of factual grounding
- Evaluating a system without inspectable queries, retrieved context, answers, citations, or reference evidence
- Authorizing production deployment or accepting residual security, privacy, or compliance risk
Instructions:
1. Use the linked AMO-P-000242 asset as the evaluation framework: supply the query set, corpus description, captured RAG outputs, expected behavior, and current acceptance thresholds rather than asking for a generic RAG review.
2. Define the evaluation unit and success criteria before assessing results. Keep corpus coverage, retrieval relevance, context sufficiency, generation groundedness, citation support, citation completeness, and abstention behavior as separate dimensions.
3. Create a traceable record for each evaluated case, linking the query, retrieved evidence, answer claims, citations, expected behavior, observed failure, and proposed classification.
4. Label every conclusion as a supplied fact, evaluator observation, assumption, inference, missing information, or uncertainty. Do not infer successful retrieval or citation support from an answer that merely sounds plausible.
5. Check whether each material answer claim is supported by the cited passage and whether the citation identifies the correct source. Record partial support, contradictory support, missing support, and inaccessible evidence separately.
6. Attribute failures conservatively. Distinguish absent corpus evidence from poor retrieval, truncated or polluted context, unsupported generation, incorrect citation attachment, and inappropriate refusal or non-abstention.
7. Define remediation hypotheses and bounded tests for the responsible layer, such as corpus repair, chunking changes, metadata filters, reranking, context assembly, prompting, citation rules, or abstention thresholds.
8. Propose a versioned regression set with measurable pass criteria and non-compensable gates for severe grounding, access-control, privacy, or safety failures.
9. Keep technical recommendations separate from release authorization. Require review by the release owner and relevant data, security, privacy, or domain reviewers before production changes, use of sensitive evaluation data, or acceptance of consequential residual risk.
Expected output:
A RAG evaluation brief containing the evaluation scope, case-level evidence table, layer-specific failure taxonomy, aggregate metrics with uncertainty, corpus and observability gaps, ranked remediation hypotheses, regression suite, acceptance thresholds, and a conditional release recommendation for the release owner and relevant data, security, or privacy reviewers.
Constraints and boundaries:
- Do not claim that queries were run, sources were retrieved, or configurations were tested unless corresponding execution records are supplied.
- Do not use answer plausibility, citation presence, or aggregate scores alone as evidence of correctness.
- Preserve source provenance and access boundaries; redact or minimize sensitive query, document, and user data.
- Do not collapse retrieval, generation, and citation errors into a single generic hallucination category.
- Treat benchmark composition, judge reliability, sampling limits, and missing reference answers as sources of uncertainty.
- Production release remains with the release owner; permission changes and security or privacy risk acceptance remain with the designated data, security, and privacy owners.
Powered by an Amo.ng Prompt
RAG Retrieval and Citation Quality Evaluation Lab
Open the linked prompt to use the instructions that power this Skill.
Completion criteria
Complete when:
- Every evaluated finding references a specific query, retrieved passage, answer claim, citation, or supplied execution record.
- Material claims are classified as supported, partially supported, contradicted, unsupported, or unverifiable.
- Corpus, retrieval, context, generation, citation, and abstention failures are scored or recorded separately.
- Aggregate results identify sample size, sampling method, missing data, and uncertainty or evaluator-disagreement limits.
- Each recommended change has a corresponding test, expected result, and regression or safety guardrail.
- The release recommendation clearly distinguishes observed evidence from unrun tests and unresolved gaps.
Was this useful?
Explore related Workflows
Browse WorkflowsImplement a Retrieval-Grounded Support Triage Pilot
Implement an approved knowledge-assistant architecture, integrate it into bounded support triage in sandbox or shadow mode, evaluate retrieval/citation/abstention quality, and prepare a controlled change decision.
Related Prompts
Browse PromptsAutomate KPI Reporting from Approved Metric Definitions
Implement a repeatable KPI-reporting process from approved metric definitions, with traceable calculations, authoritative reconciliation, failure and rerun tests, disabled delivery, and recovery evidence.
Analytical Conclusion Sensitivity Review
Test whether a consequential analytical conclusion survives plausible changes to data, cohort, definitions, assumptions, model choices, and missing-information treatment.
Data Lineage Break Investigation
Reconstruct where a data product diverged from authoritative lineage, bound affected outputs and decisions, and define safe repair and reprocessing.
Forecast Assumption and Model Drift Challenge
Challenge forecast assumptions, structural stability, backtest evidence, scenario sensitivity, and decision thresholds before relying on projected outcomes.
Knowledge Entitlement Drift Review
Reconcile authoritative access policy with effective permissions across source, ingestion, index, cache, retrieval, citation, and response layers.
AI-Generated SQL Result Verification and Reconciliation
Validate AI-generated SQL and its reported results before they are used for a consequential decision.