Common RAG Failure Cases
Retrieval-Augmented Generation (RAG) has become the backbone of modern AI applications—from chatbots and search assistants to internal knowledge copilots. While RAG significantly reduces hallucinations and improves factual grounding, it still fails in subtle but costly ways.
This article breaks down the most common RAG failure cases and offers practical mitigations you can apply immediately.
1. Poor Recall: When Relevant Data Isn’t Retrieved
What goes wrong
The retriever fails to fetch the most relevant documents, even though they exist in the knowledge base. This leads the LLM to answer with partial or incorrect information.
Common causes
Weak or generic embeddings
Low
top-kretrievalOver-compressed chunks
Mismatch between query language and document language
Symptoms
Answers feel vague
Important facts are missing
Model “guesses” instead of citing sources
Quick mitigations
Increase
top-k(e.g., from 3 → 10)Use domain-specific embedding models
Apply hybrid search (vector + keyword)
Normalize queries (expand acronyms, rephrase)
2. Bad Chunking: Context Without Meaning
What goes wrong
Documents are split poorly—breaking logical units like paragraphs, tables, or code blocks. The retriever fetches fragments that lack context.
Symptoms
Answers reference incomplete sentences
Definitions are cut off
Tables or lists lose structure
Quick mitigations
Chunk by semantic boundaries, not fixed tokens
Use overlapping chunks (10–20%)
Preserve metadata (headings, sections)
Chunk differently for prose vs code vs tables
Rule of thumb: A chunk should answer one question clearly.
3. Query Drift: The Retriever Loses the User’s Intent
What goes wrong
As queries become longer or conversational, embeddings drift away from the core intent. The retriever focuses on irrelevant keywords.
Symptoms
Retrieved documents are “related but wrong”
Follow-up questions break retrieval
Multi-intent queries confuse results
Quick mitigations
Rewrite queries using an LLM (“query condensation”)
Extract the core intent before retrieval
Use conversation memory selectively
Run retrieval on multiple query variants
4. Outdated or Stale Indexes
What goes wrong
The knowledge base is correct—but no longer current. The retriever returns obsolete documents, leading to confidently wrong answers.
Symptoms
Answers contradict recent updates
Deprecated APIs or policies are suggested
Time-sensitive facts are incorrect
Quick mitigations
Schedule automatic re-indexing
Store timestamps in metadata
Prefer newer documents during retrieval
Add “last updated” constraints in prompts
5. Hallucinations from Weak or Thin Context
What goes wrong
Even when documents are retrieved, they don’t contain enough information to answer the question. The LLM fills the gaps with fabricated reasoning.
Symptoms
Fluent but unverifiable answers
No citations or weak grounding
Overconfident tone despite missing data
Quick mitigations
Enforce “answer only from context” prompts
Set fallback responses (“I don’t have enough info”)
Require citation-based answering
Add a relevance check before generation
Bonus: Silent Failure — When Everything Looks Fine
Sometimes RAG fails without obvious errors: latency spikes, subtle inaccuracies, or degraded answer quality over time.
Mitigations
Log retrieved documents
Evaluate retrieval separately from generation
Use RAG-specific evaluation metrics (Recall@K, MRR)
Add human or LLM-based answer reviewers
Final Thoughts
Most RAG failures are retrieval problems, not model problems. Before switching LLMs or increasing parameters, inspect:
What was retrieved?
Why was it retrieved?
Was it enough?
Strong RAG systems are built on good data hygiene, smart retrieval, and defensive prompting—not just bigger models.