Skip to main content
RAG evaluation metricsLLM evaluationAI quality

RAG Evaluation Metrics That Catch Real Retrieval Failures

12 September 2026 · OpenCrevo

RAG evaluation metrics only matter if they fail a pull request when retrieval breaks. A weekly notebook score that nobody gates is theatre. Use three numbers on a frozen 20-row golden set: citation hit, faithfulness, and answer completeness. If you need the surrounding discipline, read AI QA vs quality engineering (/blog/ai-quality-assurance-vs-quality-engineering/) and wire the same fixtures into an LLM evaluation framework (/blog/llm-evaluation-framework-in-ci/).

The failure: retrieval silently swapped the source

A support bot answers "What is the refund window?" with 30 days. The correct policy chunk says 14 days after a docs move. The chat still sounds confident. BLEU and "thumbs up" miss it. Citation hit would have gone to zero on that row.

Three RAG evaluation metrics you can compute in CI

  • Citation hit — retrieved chunk IDs must include the gold chunk_id. This is the first gate.
  • Faithfulness — every claim in the answer must be supportable by retrieved text. Unsupported claims fail the row.
  • Answer completeness — required entities from gold ("14 days", "store credit") must appear.
{
  "id": "refund-window",
  "question": "What is the refund window?",
  "gold_chunk_id": "policy-refunds-v3",
  "must_include": ["14 days"],
  "forbidden": ["30 days"]
}
type Row = {
  gold_chunk_id: string;
  must_include: string[];
  forbidden: string[];
};

export function scoreRagRow(
  row: Row,
  retrievedIds: string[],
  answer: string
) {
  const citationHit = retrievedIds.includes(row.gold_chunk_id) ? 1 : 0;
  const completeness = row.must_include.every((s) =>
    answer.toLowerCase().includes(s.toLowerCase())
  )
    ? 1
    : 0;
  const clean = row.forbidden.every(
    (s) => !answer.toLowerCase().includes(s.toLowerCase())
  )
    ? 1
    : 0;
  return { citationHit, completeness, clean };
}

export function gate(rows: ReturnType<typeof scoreRagRow>[]) {
  const citation = rows.reduce((a, r) => a + r.citationHit, 0) / rows.length;
  if (citation < 0.85) throw new Error(`citation hit ${citation} < 0.85`);
}

Run this on every index or prompt change. Do not wait for a monthly eval day. Not sure where your RAG quality stands? Start with the free QA maturity assessment. Productionise it through AI quality engineering services or a self-hosted control plane on the Enterprise page.

Pass, fail, cleanup

  • Pass: Moving the refund chunk out of the index drops citation hit below 0.85 and CI fails.
  • Fail: You only store an LLM-as-judge score with no gold chunk_id. Judges will agree with a fluent wrong answer.
  • Cleanup: Keep the golden set under 50 rows until it is stable. Delete unused judge prompts that are not gated.
START YOUR QUALITY JOURNEY

Your next chapter starts with a conversation.

Book a free quality audit. We'll review your AI system, identify the highest-risk failure modes, and map a quality roadmap tailored to your stack.