Learn / RAG in 7 lessons / What RAG is, and when not to use it
What RAG is, and when not to use it
Retrieval-augmented generation in one picture, why it exists, and the cases where it's the wrong tool.
The problem RAG solves
A large language model’s knowledge is frozen at training time and general-purpose. Ask it about your company’s internal runbook, a document uploaded five minutes ago, or a fact that changed last week, and it will either say it doesn’t know or - worse - guess fluently and sound confident while doing it.
Retraining or fine-tuning the model every time your data changes is slow and expensive. Retrieval-augmented generation (RAG) sidesteps that: keep the model frozen, and instead fetch the relevant pieces of your own data at the moment of the question, then hand them to the model as context.
question --> [retriever: search your documents] --> top-k relevant chunks
|
v
prompt = instructions + chunks + question --> LLM --> answer
That’s the whole idea. Everything else in this track - chunking, embeddings, hybrid search, evaluation - is about making each of those two steps (retrieve, then generate) actually work well, because each one fails in its own specific ways.
Why not just fine-tune instead?
Fine-tuning bakes facts into model weights. It’s the right tool when you want to teach a model a style, a format, or a skill (how to structure an answer, a house tone, a narrow classification task). It’s usually the wrong tool for teaching a model facts that change, because:
- Every data update means a new training run.
- You can’t cite a source - the fact is now diffused into weights, not traceable to a document.
- It’s easy to overfit and quietly degrade the model’s general ability.
RAG keeps facts as retrievable, swappable text, so updating your data is as simple as re-indexing a document - no retraining.
When RAG is the wrong choice
RAG is popular, which makes it tempting to reach for by default. Don’t, when:
- The whole knowledge base fits in the context window. If you have 20 short FAQ entries, just put them all in the system prompt. A retriever adds failure modes (a bad chunk boundary, a missed match) for no benefit.
- You need an exact, guaranteed answer, like “what is this user’s current account balance.” That’s a database query, not a semantic search - RAG returns probably relevant text, not a guaranteed-correct fact.
- The task is really about reasoning or generation, not lookup. RAG doesn’t make a model better at multi-step arithmetic or code review; it only gives it better source material to read from.
- Latency budget is very tight. A retrieval step (even a fast one) adds a network hop and compute before the model even starts generating.
The rest of this track assumes you’ve decided RAG is the right tool - and shows what it takes to make it reliable in production, not just in a demo. See lesson 7 and the article Why Your RAG Pipeline Is Confidently Wrong for what “reliable” actually requires.
Key takeaways
- RAG = retrieve relevant text at request time, then put it in the prompt so the model answers from it instead of from memory.
- It exists because models can't know your private/recent data and shouldn't have to be retrained to learn it.
- It is not a fix for a model that reasons badly, and it is not free - it adds latency, infra, and a whole new failure surface.
- Skip RAG when the answer set is small enough to fit in the prompt directly, or when you need guaranteed-exact lookups (use a database call, not a retriever).
Quick check
2 questions - see how much stuck.