Say you build a chatbot for company policy. A language model can answer general questions, but it cannot magically know the policy that changed yesterday. RAG (Retrieval-Augmented Generation) searches related sources before answering and passes them to the model.

In short: search finds evidence, and the language model writes sentences from that evidence. You do not retrain the model's knowledge every time.

What actually happens

Imagine a user asks, "What are this year's remote-work request rules?" The system does roughly three things:

  1. Search: finds related passages in company policy docs.
  2. Build context: puts the found passages plus the user question into the model input.
  3. Generate: the model reads the material and answers. Showing doc titles or links lets users check the sources.

If search returns "up to twice per week with manager approval," the model can answer from that. If nothing relevant is found, saying "I don't know" beats inventing an answer.

Question: What are the remote-work request rules?
Retrieved doc: HR policy 4.2 — up to twice per week, manager approval required
Answer: Per HR policy 4.2, you can apply up to twice per week with manager approval.

This example is simplified to show the idea. Real results change with doc quality, search quality, and model interpretation.

Search and generation play different roles

The retriever decides what to read. The language model decides how to phrase what it read. If search brings back the wrong docs, even great phrasing cannot save accuracy. If search finds the right docs but the model adds facts outside them, the answer can still go wrong.

There is more than one way to find docs. You can match exact words, match similar meanings, or combine both. For meaning-based search, embeddings that turn sentences into numeric vectors are common.

That split helps you debug. Bad sources point to retrieval. Faithful sources with wrong claims point to generation.

When RAG is useful

RAG shines when docs change often or answers must show sources. Product manuals, internal knowledge docs, and policy guides are classic cases. Updating the source lets the next search use the new doc. Still, an update alone does not guarantee a correct answer.

Question What to check in RAG
Where is the evidence? Do retrieved docs actually connect to the answer
Is the material current? Are latest docs reflected in search targets
What if no material exists? Does it explain limits instead of guessing

Common myth: does RAG remove hallucinations?

No. Search results can be wrong, and the model can misread sources or invent facts beyond them. RAG is a structure for providing sources, not an automatic fact guarantee.

Another myth is that RAG always needs a vector database. For a small doc set, simple search can start the work. The right search method depends on your docs and questions.

Treat RAG as risk reduction, not risk removal. Evals and source display still do heavy lifting.

Conclusion: what you fetch matters as much as what the model knows

RAG finds outside material, adds it to the input, then generates an answer. Search finds evidence, and the model reads it to respond. So "which sources you fetch and pass" matters as much as "what the model knows." For how this differs from tuning the model itself, continue with our RAG vs fine-tuning comparison.

References

Go deeper with a course

If you want to design retrieval, chunking, and evals for your own document project, learn by building a working system.