In Brief
Fine-tuning and RAG are the two main paths for adapting an LLM to a specific domain. The first modifies the model’s weights — the equivalent of memorization. The second gives the model access to a searchable library at response time. Since 2023, the literature has accumulated to compare both approaches — and the conclusion is neither simple nor definitive.
The Metaphor That Holds
Imagine an expert who must answer pointed questions about tax law.
With fine-tuning, you put them through intensive training: they read thousands of documents, assimilate the concepts, and integrate this knowledge into long-term memory. Questioned six months later with no reference materials, they can answer from memory. The problem: if legislation changes, the entire training process must start over.
With RAG (Retrieval-Augmented Generation), you give them access to an up-to-date library they can consult before responding. Each answer draws on dynamically retrieved documents. The library can be updated without touching the expert. However, the quality of the response depends on what the library contains — and the relevance of the retrieved documents.
Both approaches are valid. They don’t answer the same questions.
In short: fine-tuning = continuing education in the expert’s head. RAG = library they consult before answering. The choice rarely depends on raw performance — it mostly depends on what changes faster: the knowledge, or the tone.
Fine-Tuning: What It Really Means
Full fine-tuning modifies all weights of a model on a targeted dataset. In 2022, LoRA (Low-Rank Adaptation) changed the game: instead of adjusting all parameters, it trains low-rank matrices (~1/1000 of the parameters) inserted into the transformer layers. Results on MMLU are comparable to full fine-tuning, at a fraction of the cost.
QLoRA went further in 2023: 4-bit NF4 quantization + LoRA. Result: fine-tuning a 65-billion-parameter model on a single 48GB GPU, with 99.3% of ChatGPT’s performance on tested benchmarks. What required a GPU cluster became accessible on a single high-end machine.
Variants continue to multiply: AdaLoRA (adaptive rank budget allocation), DoRA (direction/magnitude decomposition), CLoQ (constrained quantization). The ecosystem is fragmenting, complicating direct comparisons.
What fine-tuning does well: adapting response style and format, memorizing invariant structures (business ontology, stable terminology), improving performance on well-defined tasks when labeled data exists.
What it does less well: updating rapidly changing factual knowledge, handling rare entities poorly represented in the training data.
In short: QLoRA (2023) democratized fine-tuning — a 65-billion-parameter model now fits on a single 48 GB GPU, where 8 used to be needed. But this accessibility does not change the fundamental observation: re-training to update one fact is demolishing and rebuilding a wall to change a screw.
RAG: From Naive to Agentic
RAG was formalized by Lewis et al. in 2020 (NeurIPS). In its naive version, the system retrieves the K passages most similar to the query and injects them into the prompt. Simple, but fragile: response quality depends entirely on retrieval quality.
Since then, architectures have evolved across several generations:
- Advanced RAG: query rewriting before retrieval, result reranking. A well-calibrated reranker can increase MRR@5 by +59%.
- Hybrid search: combination of BM25 (lexical) and vector search (semantic). Complementary: BM25 excels on exact terms, vector search on meaning.
- Modular/Agentic RAG: iterative retrieval loops, self-correction (Self-CRAG). In one study, Self-CRAG achieves +0.456 FactScore compared to the standard RAG baseline.
What RAG does well: up-to-date knowledge, long documents, infrequent entities (arXiv 2403.01432 confirms RAG’s superiority on rare entities).
What it does less well: latency performance (retrieval doubles time-to-first-token: 495ms → 965ms in one published measurement), reduced throughput (-78% in the same study), dependence on the quality of the indexed corpus.
In short: RAG pays for its flexibility in latency and retrieval complexity. When the corpus is dirty or poorly indexed, the agent answers poorly — not because it lacks knowledge, but because it consults the wrong pages. Quality of RAG = quality of corpus + quality of retriever.
What Direct Comparisons Show
Several published empirical results allow for more refined reasoning.
Factual knowledge injection: Soudani et al. (EMNLP 2024) show that RAG systematically outperforms unsupervised fine-tuning. The nuance is critical: unsupervised FT (on raw documents) ≠ supervised FT (on labeled question-answer pairs). This confusion feeds many poorly framed comparisons in the literature.
Agricultural domain (arXiv 2401.08406): FT contributes +6 points, RAG added on top contributes +5 additional points. The combination is cumulative.
Multi-hop QA (arXiv 2601.07054): supervised FT is better than RAG — but only when labeled data exists. That’s not always the case.
RAG+FT combination: Tencent publishes +7.79% Exact Match by combining both compared to RAG alone (53.76% vs 44.20% for FT alone). The combination wins — but it also accumulates costs and operational complexity.
The Decision Tree
| If your main concern is… | Prefer | Why |
|---|---|---|
| Changing knowledge (regulation, prices, stock) | RAG | Updating an index is less costly than re-training. |
| Specific response style/format (tone, JSON structure, fixed terminology) | Fine-tuning (LoRA/QLoRA) | RAG does not change style, FT bakes it into weights. |
| Maximum performance on bounded task with labeled data | Supervised fine-tuning | On labeled QA, supervised FT beats RAG (multi-hop QA arXiv 2601.07054). |
| Rare entities or under-represented in training corpus | RAG | RAG is documented superior on infrequent entities (arXiv 2403.01432). |
| Performance + flexibility, large budget | Combined RAG + FT | +7.79% EM on Tencent benchmark, but accumulates complexity. |
| Quick start, existing documentary base, non-ML team | RAG | Mature stack (LangChain, LlamaIndex), no dedicated GPU required. |
| Production model exposed without knowledge drift | RAG | No re-training = no drift to manage. |
Here’s how to approach the choice in practice.
Start with the data question: do you have labeled question-answer pairs that are sufficiently representative? If not, supervised fine-tuning is not a realistic option. RAG or unsupervised FT remain possible, but with different limitations.
Then, the temporality question: do your knowledge requirements change frequently? If so, RAG is structurally better suited — updating an index is less costly than re-fine-tuning.
Then, the style vs. content question: are you trying to change the tone and format of responses (RAG doesn’t help much here) or inject new facts (RAG is made for that)?
Finally, the resources question: do you have a GPU available for fine-tuning, retrieval infrastructure for RAG, or neither? QLoRA has made fine-tuning more accessible, but RAG remains simpler to deploy on existing bases.
If you have the resources and data, the combination almost always wins on benchmarks. The real question is: does the marginal gain justify the additional complexity?
Debates That Remain Open
Several questions are unresolved in the literature.
Does fine-tuning truly encode knowledge, or mostly style? Models fine-tuned on medical corpora produce responses that sound medical, but they also hallucinate on specific facts. The style/knowledge distinction is not trivial to establish.
Does RAG amplify hallucinations via retrieval noise? This is documented. A BadRAG study shows that 0.04% of a poisoned corpus is sufficient to redirect 98.2% of responses toward malicious content. RAG is not a protection against misinformation — it can be its vector.
Are comparison baselines fair? Often not. Comparing RAG to unsupervised FT favors RAG. Comparing RAG to supervised FT changes the result. Papers don’t always pose the same question.
Is the combination always justified? On benchmarks, often yes. In production, the answer depends on operational cost, acceptable latency, and the maintenance complexity the team can absorb.
Key Takeaways
- RAG and fine-tuning are not competing on the same register: one adapts the model’s memory, the other gives it access to an external library.
- RAG systematically outperforms unsupervised fine-tuning for fact injection — but this comparison is often poorly framed in the literature.
- The RAG+FT combination wins on benchmarks, but accumulates costs and complexity: the marginal gain must be evaluated case by case.
- QLoRA has made fine-tuning accessible on limited hardware (65B on 1 48GB GPU), but RAG remains simpler to update.
- Unresolved debates (knowledge vs. style, retrieval noise, fair baselines) must be kept in mind before concluding on the superiority of one approach.