When RAG gives a wrong answer, there is one line you hear first: "Let us change the embeddings."
Sometimes that is the answer. But in my experience, the culprit is somewhere else more often: document parsing, chunking, the retrieval method, and whether the system knows how to say "I don't know." This post splits the RAG pipeline into 5 stages and gives you an order for finding the breakage.
The basics live in the What is RAG post and the Classic vs Graph vs Agentic RAG post. This is the repair manual.
The full map: 5 gates before an answer comes out
[1. Parsing] Reading docs (PDFs, tables, scans)
→ [2. Chunking] Splitting (units, overlap, metadata)
→ [3. Retrieval] Finding (vector, keyword, hybrid)
→ [4. Reranking] Choosing (reordering top candidates)
→ [5. Generation] Answering (with citations and refusal)
Users only see step 5: "The answer is wrong." But the cause can sit in any of steps 1–4. The key is suspecting them in order. Fixing from the back wastes money and time.
Stage 1, parsing: reading breaks before search breaks
Borrow the key line from the Mistral OCR post: search can break only after document parsing already broke.
Typical symptoms:
| Symptom | Parsing warning sign |
|---|---|
| Table content never shows in answers | Tables arrive as broken text |
| Frequent "it is not in the documents" | Scanned and image pages arrive empty |
| Page and clause citations are off | Headers, footers, and footnotes mix into the body |
The check is simple. Skip search and read the raw parsed output with your own eyes. Search cannot find what a human cannot find in the document.
Check 1: Ctrl+F the source pages of 5 questions in the parsed raw text
→ If they are missing here, parsing is confirmed. Do not touch the embeddings.
Stage 2, chunking: the split decides the answer unit
Even clean parsing fails with bad splits:
Chunks too big: retrieval was right, but junk rides along and blurs the answer
Chunks too small: context is cut, leaving a "so what?" state
No overlap: content on the boundary disappears from both sides
No metadata: "2024 policy" and "2026 policy" get mixed up
Just remember three working rules:
Rule 1: Split along document structure (clause, section, table units)
Rule 2: Add overlap (usually 10–20%, prevents boundary loss)
Rule 3: Attach metadata (doc name, date, version, page)
Version mixing is especially common. For policies, manuals, and terms, the date is part of the answer. Trusting vector similarity alone with no date metadata happily returns the old version at rank 1.
Stage 3, retrieval: some games you lose with vectors alone
Vector search finds similar meanings well. But keyword search (like BM25) finds exact matches better — proper nouns, product codes, clause numbers, figures.
Vector wins: meaning questions like "tell me the refund policy"
Keyword wins: "Article 14, paragraph 3", "model XB-200", "March 2026 notice"
Real questions: both mixed → hybrid as the default
The hybrid structure looks like this:
Question ──┬──→ vector search top 20 ──┐
└──→ keyword search top 20 ─┴──→ merge → [stage 4 rerank] → top 5 → generate
Sketched as a concept, the flow is this:
# hybrid_search.py — concept sketch
def hybrid_search(query, top_n=5):
dense_hits = vector_search(query, k=20) # meaning-based
sparse_hits = keyword_search(query, k=20) # exact-match based
merged = reciprocal_rank_fusion(dense_hits, sparse_hits)
reranked = rerank(query, merged[:20]) # stage 4
return reranked[:top_n]
The merge method (weighted sum, RRF, and so on) and its weights are decided on your data. Do not copy numbers from someone else's blog. Reproduce them on a 30-question eval set — the same idea as the 5-question check in the mini eval post.
Stages 4–5, reranking and generation: "I don't know" must exist at the end
Reranking re-sorts 20 candidates by relevance to the question and passes only the top 5 to generation. It costs a little, but the answer quality jumps enough to make it good value.
At generation time, the principle from the hallucination vs abstention post applies: if there is no evidence, make it say "I don't know." Recent measurements point the same way:
- Gemini 4 Argon: 15% hallucination rate, lowest among leading models (accuracy a bit lower, overall on par)
- GPT-6 Astra: hallucination rate improved from 92% to 51%, accuracy up too
→ "Models that refuse well" are gaining the edge in knowledge work
There is also the Hebbia case: record citation recall on Opus 5.5 while keeping token efficiency. Forcing citations is a working device against hallucination:
Three lines for your generation prompt:
1. Mark the source chunk number on every answer sentence
2. If it is not in the sources, answer "not confirmed in the documents"
3. Never write uncertain numbers or dates
CodeBridge Mini Lab: classify 20 failures
1. Collect 20 wrong questions (from user logs and tests)
2. Find the stage, not the answer:
Parsing? → Is the evidence in the parsed raw text (if not, parsing)
Chunking? → Does reading the evidence chunk answer it (if context is cut, chunking)
Retrieval? → Is the evidence in the top 20 (if not, retrieval)
Reranking? → Is it in the 20 but not the 5 (if so, reranking)
Generation? → Is it in the 5 but the answer is wrong (if so, generation/prompt)
3. Count the distribution:
e.g. parsing 8, chunking 5, retrieval 4, reranking 2, generation 1
→ Fix the winner first. Touch embeddings only when retrieval wins.
Once this classification is done, vague prescriptions like "change the embeddings" disappear. The numbers decide what to do next.
Conclusion: suspect from the front
The order, summarized:
Check parsed raw text → inspect chunking and metadata → hybrid search → rerank → force citations and allow refusal
RAG is a pipeline, not a model. Overlay this checklist on the architecture sketch in the internal-docs chatbot post and the gaps show immediately. And remember: "I don't know" is performance too.
Further reading
- What is RAG? Retrieval to generation in one go
- Mistral OCR 4.1 and RAG: parsing can break first
- Is a model bad if it says it does not know?
References
- Artificial Analysis: Gemini 4 Argon
- Artificial Analysis: Benchmarking GPT-6 Astra
- Anthropic: Introducing Claude Opus 5.5
Go deeper with a course
If you want to tie parsing, chunking, retrieval, and generation into one structure and build a working chatbot, a course that grows from Classic RAG into GraphRAG and Agentic RAG follows directly from this 5-stage checklist.