In short
An LLM knows what it learned during training. Past that date, it knows nothing about what happened in the world — and it never knew anything about your internal documents. RAG (Retrieval-Augmented Generation) solves this problem without modifying the model: instead of memorizing everything, the model learns to go and fetch.
The principle is simple. When you ask a question, a search engine identifies the most relevant passages in an external knowledge base, then slips them directly into the model’s context before it generates its answer. The model answers by leaning on those passages, like a doctor checking a file before talking to the patient.
Explanation
The problem RAG solves
An LLM’s weights are frozen at training time. GPT-4 does not know today’s news. Claude does not know your company’s internal documentation. Fine-tuning (retraining) a model on new data is expensive, takes time, and has to be redone at every update. RAG is an alternative: plug an external memory onto the model at inference time.
The four-step architecture
1. Embeddings — turning text into coordinates
Before you can search, you have to index. Each document in the knowledge base is split into pieces (chunks), then each chunk is converted into a numeric vector — an embedding. This embedding represents the text’s semantics: two sentences saying the same thing with different words will have close vectors. This computation is done once, at indexing time.
2. The vector store — a library of coordinates
All these vectors are stored in a vector database (vector store). Unlike a SQL database that looks for exact words, the vector store searches by geometric proximity — it finds the vectors “closest” to a given query.
3. Retrieval — fetching the relevant passages
When the user asks a question, that question is also converted into an embedding. The vector store then computes the distance between this query vector and every stored vector (usually via cosine similarity), and returns the top N semantically closest chunks — even if no exact word matches.
4. Augmented generation — answering based on the sources
The retrieved chunks are injected into the prompt sent to the LLM, right before the user’s question. The model now has an enriched context. It generates its answer based on these excerpts, and can cite their source.
[Question] → [Query embedding] → [Vector search]
↓
[Relevant chunks retrieved]
↓
[Prompt = question + chunks] → [LLM] → [Sourced answer]
In short: a RAG system is a librarian who runs through the stacks before answering. It has four mechanical stages — turn books into codes (embeddings), shelve them (vector store), retrieve the right books (retrieval), answer while citing pages (generation). Each stage has its own failure modes.
Why it is useful
RAG makes it possible to use a generalist model on a proprietary knowledge base without retraining it. A company can plug its model into its technical documentation, its contracts, its support tickets — and get an assistant that knows its context without having exposed its data to a training provider.
| If your concern is… | RAG is… | Why |
|---|---|---|
| Changing knowledge (case law, prices, stock, internal docs) | Suitable | Updating an index is less costly than re-training. |
| Sourced and auditable answers (medical, legal, compliance) | Highly suitable | Each answer can cite the exact source passage. |
| High question volume on small stable corpus | Oversized | Direct long context or fine-tuning may suffice and cost less per query. |
| Specific style/tone (fixed terminology, very strict JSON format) | Unsuitable | RAG injects context, does not change response form — fine-tuning preferable. |
| Ultra-sensitive data, mandatory on-premise hosting | Suitable | Embeddings and vector store can stay internal, local model optional. |
| Critical latency (< 200 ms p99) | Questionable | Retrieval adds 200-500 ms — measure before committing. |
Concrete examples
Up-to-date data
An LLM trained in 2024 does not know the court decisions of 2025. With RAG, a law firm can index case law on the fly and query a base that is always fresh. No retraining — just an update to the document base.
Reducing hallucinations
Without RAG, the model generates by drawing on its training probabilities. It can “invent” a plausible but false reference. With RAG, it relies on concrete passages provided in context: the answer is anchored in real documents. Hallucinations do not disappear, but their frequency drops significantly on factual questions.
Traceable sources
Every assertion from the model can be linked back to the source chunk it came from. The user can verify. That is a major advantage in contexts where trust has to be audited — medical, legal, regulatory.
The chunking limit
Splitting documents into chunks is a critical decision. Too small, and the chunks lose their context: a passage that only makes sense inside the previous paragraph becomes unusable. Too big, and the chunks consume a significant share of the model’s context window, diluting relevant information in noise. There is no universal size — it is a parameter to calibrate depending on the type of documents and the nature of the queries.
Another documented limit: even with the right documents retrieved, the model may fail to use them correctly if they sit in the middle of a long context. Research by Liu et al. (Stanford, 2024) showed that LLM performance follows a U-shaped curve depending on where the relevant information sits in the context — information at the beginning and the end is better exploited than information in the middle.
The hidden cost
Every RAG query involves a vector search plus the injection of N chunks into the prompt. This increases the number of tokens sent to the model on every call — and billing follows. For high-volume applications, the cost of retrieval becomes an architectural parameter in its own right.
In short: RAG does not eliminate hallucinations — it relocates them. If retrieval returns a bad chunk, the model still invents, but this time leaning on text that seems relevant. RAG quality ≈ chunking quality × retriever quality × LLM quality. A bad link in the middle breaks everything.
What to remember
- RAG does not train the model — it gives it access to an external memory at inference time.
- The flow is: indexing (embeddings + vector store) → retrieval (similarity search) → augmented generation (enriched prompt).
- Main advantages: up-to-date data without retraining, fewer hallucinations, citable sources, lower cost than fine-tuning.
- Main limits: retrieval quality (garbage in, garbage out), chunking sensitivity, “lost in the middle” effect, per-query cost.
- The quality of the embedding model and the document-chunking strategy matter as much as the choice of LLM.