Skip to main content

Command Palette

Search for a command to run...

Common RAG Failure Cases

Published
•3 min read•View as Markdown

Retrieval-Augmented Generation (RAG) has become the backbone of modern AI applications—from chatbots and search assistants to internal knowledge copilots. While RAG significantly reduces hallucinations and improves factual grounding, it still fails in subtle but costly ways.

This article breaks down the most common RAG failure cases and offers practical mitigations you can apply immediately.

1. Poor Recall: When Relevant Data Isn’t Retrieved

What goes wrong

The retriever fails to fetch the most relevant documents, even though they exist in the knowledge base. This leads the LLM to answer with partial or incorrect information.

Common causes

  • Weak or generic embeddings

  • Low top-k retrieval

  • Over-compressed chunks

  • Mismatch between query language and document language

Symptoms

  • Answers feel vague

  • Important facts are missing

  • Model “guesses” instead of citing sources

Quick mitigations

  • Increase top-k (e.g., from 3 → 10)

  • Use domain-specific embedding models

  • Apply hybrid search (vector + keyword)

  • Normalize queries (expand acronyms, rephrase)

2. Bad Chunking: Context Without Meaning

What goes wrong

Documents are split poorly—breaking logical units like paragraphs, tables, or code blocks. The retriever fetches fragments that lack context.

Symptoms

  • Answers reference incomplete sentences

  • Definitions are cut off

  • Tables or lists lose structure

Quick mitigations

  • Chunk by semantic boundaries, not fixed tokens

  • Use overlapping chunks (10–20%)

  • Preserve metadata (headings, sections)

  • Chunk differently for prose vs code vs tables

Rule of thumb: A chunk should answer one question clearly.

3. Query Drift: The Retriever Loses the User’s Intent

What goes wrong

As queries become longer or conversational, embeddings drift away from the core intent. The retriever focuses on irrelevant keywords.

Symptoms

  • Retrieved documents are “related but wrong”

  • Follow-up questions break retrieval

  • Multi-intent queries confuse results

Quick mitigations

  • Rewrite queries using an LLM (“query condensation”)

  • Extract the core intent before retrieval

  • Use conversation memory selectively

  • Run retrieval on multiple query variants

4. Outdated or Stale Indexes

What goes wrong

The knowledge base is correct—but no longer current. The retriever returns obsolete documents, leading to confidently wrong answers.

Symptoms

  • Answers contradict recent updates

  • Deprecated APIs or policies are suggested

  • Time-sensitive facts are incorrect

Quick mitigations

  • Schedule automatic re-indexing

  • Store timestamps in metadata

  • Prefer newer documents during retrieval

  • Add “last updated” constraints in prompts

5. Hallucinations from Weak or Thin Context

What goes wrong

Even when documents are retrieved, they don’t contain enough information to answer the question. The LLM fills the gaps with fabricated reasoning.

Symptoms

  • Fluent but unverifiable answers

  • No citations or weak grounding

  • Overconfident tone despite missing data

Quick mitigations

  • Enforce “answer only from context” prompts

  • Set fallback responses (“I don’t have enough info”)

  • Require citation-based answering

  • Add a relevance check before generation

Bonus: Silent Failure — When Everything Looks Fine

Sometimes RAG fails without obvious errors: latency spikes, subtle inaccuracies, or degraded answer quality over time.

Mitigations

  • Log retrieved documents

  • Evaluate retrieval separately from generation

  • Use RAG-specific evaluation metrics (Recall@K, MRR)

  • Add human or LLM-based answer reviewers

Final Thoughts

Most RAG failures are retrieval problems, not model problems. Before switching LLMs or increasing parameters, inspect:

  • What was retrieved?

  • Why was it retrieved?

  • Was it enough?

Strong RAG systems are built on good data hygiene, smart retrieval, and defensive prompting—not just bigger models.