Reusable AI capability

Diagnose RAG Grounding and Retrieval Failures

Evaluate a retrieval-augmented generation system by separating corpus, retrieval, context assembly, answer generation, citation, and abstention failures, then define reproducible regression criteria.

This Skill packages a reusable way to use the linked Prompt or Workflow; Amo.ng does not run it for you.

Skill ID
AMO-S-000007
Powered by
Prompt
Published

Copy skill copies the Skill details. Use with AI adds a short instruction for your preferred AI tool; neither action runs the Skill.

Purpose

Produce an evidence-grounded RAG quality assessment that attributes failures to the responsible system layer and turns representative examples into measurable remediation and release criteria.

Required inputs

Have these details available before following the usage instructions.

  • RAG use case, intended users, and decision or answer types
  • Representative query set, including difficult, ambiguous, unanswerable, and access-sensitive cases
  • Corpus scope, source versions, and any known coverage limitations
  • Captured retrieval results, assembled context, generated answers, citations, and abstention behavior
  • Expected answers or human relevance judgments where available
  • Current quality metrics, thresholds, and known failure reports
  • Applicable privacy, access-control, safety, latency, and release constraints

How to use this Skill

When to use:
- Evaluating a RAG system before release or after a model, corpus, embedding, chunking, or reranking change
- Investigating unsupported answers, irrelevant retrieval, incomplete citations, or failures to abstain
- Comparing retrieval or generation configurations using a stable evaluation set
- Converting observed RAG failures into regression cases and acceptance thresholds

When not to use:
- Designing an entire RAG architecture from scratch
- Treating model fluency or user preference as proof of factual grounding
- Evaluating a system without inspectable queries, retrieved context, answers, citations, or reference evidence
- Authorizing production deployment or accepting residual security, privacy, or compliance risk

Instructions:
1. Use the linked AMO-P-000242 asset as the evaluation framework: supply the query set, corpus description, captured RAG outputs, expected behavior, and current acceptance thresholds rather than asking for a generic RAG review.
2. Define the evaluation unit and success criteria before assessing results. Keep corpus coverage, retrieval relevance, context sufficiency, generation groundedness, citation support, citation completeness, and abstention behavior as separate dimensions.
3. Create a traceable record for each evaluated case, linking the query, retrieved evidence, answer claims, citations, expected behavior, observed failure, and proposed classification.
4. Label every conclusion as a supplied fact, evaluator observation, assumption, inference, missing information, or uncertainty. Do not infer successful retrieval or citation support from an answer that merely sounds plausible.
5. Check whether each material answer claim is supported by the cited passage and whether the citation identifies the correct source. Record partial support, contradictory support, missing support, and inaccessible evidence separately.
6. Attribute failures conservatively. Distinguish absent corpus evidence from poor retrieval, truncated or polluted context, unsupported generation, incorrect citation attachment, and inappropriate refusal or non-abstention.
7. Define remediation hypotheses and bounded tests for the responsible layer, such as corpus repair, chunking changes, metadata filters, reranking, context assembly, prompting, citation rules, or abstention thresholds.
8. Propose a versioned regression set with measurable pass criteria and non-compensable gates for severe grounding, access-control, privacy, or safety failures.
9. Keep technical recommendations separate from release authorization. Require review by the release owner and relevant data, security, privacy, or domain reviewers before production changes, use of sensitive evaluation data, or acceptance of consequential residual risk.

Expected output:
A RAG evaluation brief containing the evaluation scope, case-level evidence table, layer-specific failure taxonomy, aggregate metrics with uncertainty, corpus and observability gaps, ranked remediation hypotheses, regression suite, acceptance thresholds, and a conditional release recommendation for the release owner and relevant data, security, or privacy reviewers.

Constraints and boundaries:
- Do not claim that queries were run, sources were retrieved, or configurations were tested unless corresponding execution records are supplied.
- Do not use answer plausibility, citation presence, or aggregate scores alone as evidence of correctness.
- Preserve source provenance and access boundaries; redact or minimize sensitive query, document, and user data.
- Do not collapse retrieval, generation, and citation errors into a single generic hallucination category.
- Treat benchmark composition, judge reliability, sampling limits, and missing reference answers as sources of uncertainty.
- Production release remains with the release owner; permission changes and security or privacy risk acceptance remain with the designated data, security, and privacy owners.

Powered by an Amo.ng Prompt

RAG Retrieval and Citation Quality Evaluation Lab

Open the linked prompt to use the instructions that power this Skill.

Open prompt

Completion criteria

Complete when:
- Every evaluated finding references a specific query, retrieved passage, answer claim, citation, or supplied execution record.
- Material claims are classified as supported, partially supported, contradicted, unsupported, or unverifiable.
- Corpus, retrieval, context, generation, citation, and abstention failures are scored or recorded separately.
- Aggregate results identify sample size, sampling method, missing data, and uncertainty or evaluator-disagreement limits.
- Each recommended change has a corresponding test, expected result, and regression or safety guardrail.
- The release recommendation clearly distinguishes observed evidence from unrun tests and unresolved gaps.

Browse Prompts
Data Analysis Expert ChatGPT

AI Trace Review and Failure Taxonomy

Use ChatGPT to reconstruct evidence-supported AI and agent traces, govern failure classifications, analyze recurring patterns within sampling limits, identify observability gaps, and propose sanitized regression cases and measurable prevention work.

Updated Aug 13, 2026

View prompt Verified ✓ 181 views · 6 copies
Data Analysis Expert ChatGPT

Multi-Currency Revenue Reconciliation Model

Reconcile multi-currency revenue from source transactions through recognition, exchange-rate conversion, payments, settlements, fees, taxes, journals, and ledger reporting while preserving timing, policy, and currency differences.

Updated Aug 5, 2026

View prompt Verified ✓ 119 views · 11 copies

Was this useful?