In brief
We call RAG (Retrieval-Augmented Generation) the family of techniques that consist in having a language model search relevant documents in a knowledge base and read them before producing an answer. As of April 2026, three large families coexist: dense vector RAG (texts are turned into lists of numbers and we look for the closest ones), vectorless, sparse or hybrid RAG (combining keyword search and semantic search), and emerging approaches (GraphRAG, Self-RAG, Adaptive, Agentic). The arrival of very long contexts (1 million tokens at Gemini, 200,000 at Claude Opus) has not killed RAG: it has redrawn its territory.
Understanding the problem: why search a base before answering
Imagine a librarian who knows by heart everything they read during their studies, but nothing published since. If you ask “what does the latest quarterly report from such-and-such company say?”, they cannot answer. Large language models (LLMs) are in this situation: they learned once and for all, and do not know your internal documents, recent updates, or the details of a specific business corpus.
RAG solves this problem with a librarian’s reflex: go look before speaking. The principle decomposes into three roles found in nearly every variant:
- The retriever (the searcher): receives the question, dives into the base, and brings back a small bundle of candidate documents. The librarian who runs through the stacks.
- The reranker (the critic): receives the candidate stack and reorders it from most to least relevant. A more careful but slower reader, used only on the pre-selected pile.
- The generator (the writer): receives the question and the retained documents, and produces the final answer. The LLM that puts it into words.
In plain terms: a RAG pipeline is “question → search → sort → write”. Each step can be implemented with very different techniques, and the choice of techniques at each step defines the RAG family.
Before comparing families, a transversal vocabulary:
Base vocabulary
- Corpus: the set of documents we search through (knowledge base).
- Chunking: each document is cut into digestible pieces, typically 200 to 800 tokens. A chunk is a piece. We search on chunks, not entire documents, because an LLM has a limited context window and a relevant paragraph drowned in 100 pages dilutes the information.
- Token: a piece of a word (about 0.75 word in English, slightly less in French). It is the unit LLMs manipulate.
- Query: the question as sent to the retriever.
Why long context does not kill RAG
In 2024-2026, the argument “models now take 1 million tokens of context, RAG becomes useless” circulated. Reality is more nuanced. Sending the entire corpus as context for every request raises four concrete problems:
- Cost: billing 1 million input tokens is expensive. On Claude Opus in 2026, loading 1 million tokens per question costs several dollars per request; a RAG pipeline that retains only the 10 most relevant chunks (≈ 10,000 tokens) costs a hundred times less.
- Latency: the more tokens we send, the longer the model takes to start answering (the metric is called TTFT, time-to-first-token). Long context adds seconds.
- Freshness: an index (a search structure) is updated document by document. Re-sending a million tokens for every request means reloading everything for each new piece.
- “Lost in the middle”: even when context fits, LLMs struggle to extract a precise piece of information when it is perched at the center of a very long text (documented by Liu et al. 2023, “Lost in the Middle”; confirmed by RecaLLM in 2026). Concretely, the same question is answered better with 5 targeted chunks than 500 raw pages.
A partial workaround, available at Anthropic and Google, is prompt caching: a stable context block (rules, reference documents) is cached, and subsequent queries pay only a fraction (about 10%) of the normal rate on that block. Useful for stable corpora queried often. But it does not replace an index when the corpus moves or exceeds the cache.
In plain terms: long context is excellent for reading one or two documents end to end; RAG remains indispensable as soon as the corpus is large, changes often, or requires precise source citation.
Family 1 — Dense vector RAG (the classical 2020+ approach)
The principle: turning text into geometry
The founding idea of vector RAG: convert each piece of text into a list of numbers (for instance 1024 or 1536 values), so that two texts about the same thing yield two close lists in space. This list is called a dense embedding. “Dense” because nearly all the numbers are non-zero, in contrast with the “sparse” (mostly empty) representations we will see later.
Vocabulary
- Embedding: numerical representation of a text as a vector (list of numbers). “the capital of France” and “Paris, French political center” yield two vectors pointing in very close directions.
- Cosine similarity: measures the angle between two vectors. Small angle → close vectors → semantically close texts. The most common calculation for “which document is closest to this question”.
- Dimension: number of values in the vector (often 768, 1024, 1536 or 3072). Higher dimension distinguishes more nuances but is more expensive to store.
An embedding model (OpenAI text-embedding-3, Cohere embed-v4.0, Voyage voyage-4-large, open-source BGE or mxbai) does the text → vector conversion. The model is applied to all corpus chunks (once), then to each question (per request). The retriever simply returns the chunks whose vectors are closest to the question’s vector.
Technical building blocks
The embedder: the model that turns text into a vector. 2025-2026 models incorporate two trends. First, Matryoshka Representation Learning (MRL), which allows a reduced dimension to be used without retraining — a 1536-vector can be truncated to 512 depending on the desired tradeoff. Second, context extension: Cohere embed-v4.0 embeds chunks up to 128,000 tokens, against 512 for the previous generation. On the open-source side, mxbai-embed-large-v1 and BGE-large reach and surpass proprietary ones on the MTEB benchmark (a public ranking of embedding models on 56 tasks).
The vector database: a structure optimized to quickly find the vectors closest to a target vector among millions. The most widespread relies on the HNSW algorithm (Hierarchical Navigable Small World), which builds a multi-level graph easing navigation. Players in 2026: Qdrant (open-source leader, $50M raise in March 2026), Pinecone (managed enterprise reference), pgvector (PostgreSQL extension for teams already on Postgres), Weaviate and Milvus (cloud-native), Chroma (local prototyping), Turbopuffer (serverless backed on S3).
The reranker: once the retriever has brought back 50 to 200 candidates, a heavier and more precise model (a cross-encoder) reads the question AND each candidate together, assigning a relevance score. We keep only the top 5 to 10 before sending to the generator.
Vocabulary — bi-encoder vs cross-encoder
- A bi-encoder encodes question and document separately, in two independent vectors, and compares the vectors. Very fast (document vectors can be precomputed) but less precise.
- A cross-encoder reads question and document together and produces a single score. Much more precise but cannot be prepared in advance: it must be run at every request on every pair. Hence the use: bi-encoder for massive retrieval, cross-encoder for targeted reranking on a small pile.
References 2026: Voyage rerank-2.5 and Cohere Rerank 3 on the proprietary side, bge-reranker-v2-m3 on the open-source side.
Structural limitations documented in 2025-2026
Several papers converge on weaknesses of the dense vector paradigm:
- Systematic biases of dense retrievers (Goyal et al. 2026): they favor short documents (brevity bias), information at the start of the chunk (position bias), literal matches, and repeated terms. Rewriting the query masks but does not fix these biases.
- Query-document asymmetry (Maekawa et al. 2026): the question is typically a complex sentence with instructions, while the document is a flat fragment. Resulting vectors do not live in the same “region” of space, requiring dedicated adapters.
- LLM passivity (Sun et al. 2026, “Don’t Retrieve, Navigate”): the generator model receives a stack of chunks without ever seeing the corpus organization. It cannot backtrack or combine scattered evidence.
- Degradation on near-duplicate corpora (CHOP, Park et al. 2026): when the base contains many similar documents, retrieval errs more often and injects noise, generating hallucinations.
These limits do not disqualify vector RAG, which remains the dominant production pattern, but they explain renewed interest for other approaches.
Family 2 — Vectorless, sparse, hybrid: the keyword revenge
Before embeddings, classical document search relied on so-called sparse methods, based on the words themselves. In 2024-2026, these methods come back in force, alone or combined with dense.
BM25: the trusty keyword relevance score
BM25 (for Best Matching 25, invented in the 1990s by Robertson and Zaragoza) is a relevance formula that orders documents according to word-by-word match with the query. It is the technique behind Elasticsearch’s, Lucene’s, and most classical search engines’ search bars.
In plain terms: BM25 asks three questions of each document.
- Do the query words appear in the document? (term frequency, TF)
- Are these words rare in the corpus as a whole? (inverse document frequency, IDF — a rare matched word is worth more than a banal one)
- Is the document of reasonable size for this number of matches? (length normalization — a very long document matching 3 times is not more relevant than a short one matching 3 times) The combined answer gives a score, we sort, we return the top-k.
BM25 is “stupid” in the sense that it ignores semantics: “cat” and “feline” are two different words to it. But it is robust, fast, training-free, drift-free, and explainable (one sees exactly which terms contributed to the match).
Pedagogical aside — “On TREC Robust04, no neural reranker improves MAP by more than 0.05 over BM25”: what does this mean?
To objectively evaluate search systems, the IR (information retrieval) community uses benchmark corpora: a frozen set of documents, a list of queries, and human judgments indicating which documents are relevant for which query.
- TREC Robust04 (Text REtrieval Conference, Robust 2004 edition) is one of the oldest: 528,155 news articles, 250 queries, thousands of human “query-relevant document” pairings. It is a public standard for system comparison.
- MAP (Mean Average Precision) is the main metric. For a given query, Average Precision measures at which positions in the ordered list the truly relevant documents appear: if relevant ones are at the top, score approaches 1; if scattered far down, the score collapses. Average over all queries gives the “Mean”. MAP = 0.3 is typical; MAP = 0.5 is excellent.
- A baseline is a reference system any new technique is compared against. BM25 has been the historical baseline of textual search for 30 years.
Translation: in a study comparing 443 systems on TREC Robust04 (Chatterjee 2026), no neural reranker (even the most sophisticated) improved MAP by more than 0.05 points relative to BM25 in open-world evaluation. In other words: gains announced by modern techniques are often less than 5% on the main metric, compared to a 1990s system. BM25 is not a relic to replace — it is a competitive reference.
In practice in 2026, the BM25 + cross-encoder reranker pipeline remains a strong baseline often competitive with much heavier architectures.
SPLADE: when sparse learns semantics
SPLADE (SParse Lexical AnD Expansion, Formal et al., Naver Labs, 2021) bridges BM25 and dense embeddings. The model, trained on BERT, encodes each text as a sparse vector over the tokenizer’s vocabulary: typically 30,000 dimensions, of which only 50 to 200 values are non-zero.
How to read it: for a document about cats, SPLADE activates not only the dimensions “cat” and “feline” present in the text, but also “meow”, “whiskers”, “kitten” — semantically related terms. Result: we keep the efficiency of inverted indices (the classical search-engine structure that retrieves all documents containing a given term in microseconds) while capturing part of the semantics.
Variants 2025-2026: ESPLADE (extended vocabulary), SPLADE-Doc (productionized version), Sparton (GPU kernel optimized via Triton).
ColBERT: one vector per token
ColBERT (Contextualized Late Interaction over BERT, Khattab & Zaharia, 2020) takes a third path: instead of one vector per document, ColBERT encodes each token as a vector. To score a question-document pair, the similarity of each question token with the best matching document token is computed, then summed. This operation is called MaxSim (late interaction), hence the “L” in the name.
In plain terms: a classical dense embedding crushes the entire document into a single point; ColBERT keeps a cloud of points (one per token). The cloud is finer but takes more space.
ColBERTv2 (2022) adds compression to make the approach viable in production. ColPali (Faysse et al. 2024) extends ColBERT to images of PDF pages, allowing retrieval on visual documents without an OCR step.
Limit 2026 (Ghosh et al.): ColBERT degrades sharply (drop of 86 to 97%) on long narrative queries beyond ~20 words, due to the intrinsic ceiling of MaxSim.
Hybrid search: combining both worlds
The practical discovery of 2023-2025: BM25 and dense complement each other. BM25 captures exact matches (proper nouns, identifiers, precise technical terms); dense embeddings capture semantics and paraphrases. Doing both in parallel and merging results beats each one taken alone.
The most used fusion is called RRF (Reciprocal Rank Fusion). Principle: for each document, look at its rank in each list (e.g. rank 3 in BM25, rank 7 in dense) and assign it a score 1/(k + rank) per list, then sum. The k is a small constant (typically 60) preventing rank 1 from crushing all. No weights to calibrate, robust by default.
Concrete example — RRF on 3 documents
Document BM25 rank Dense rank RRF score (k=60) A 1 15 1/61 + 1/75 ≈ 0.029 B 8 2 1/68 + 1/62 ≈ 0.031 C 3 3 1/63 + 1/63 ≈ 0.032 C, first nowhere but well-placed everywhere, comes out winner. Exactly the kind of document hybrid aims to surface.
Reference result 2026: on TREC-COVID (benchmark of 171,000 COVID articles), RRF(SPLADE + BGE) reaches nDCG@10 = 0.828, surpassing dense alone (Prajapati 2026). In production, Elasticsearch 8.14+, Qdrant, Weaviate and Vespa offer hybrid natively.
Vocabulary — other metrics you’ll encounter
- nDCG@k (normalized Discounted Cumulative Gain): like MAP but weights top-k ranks with a logarithm. Favors systems placing good documents at the very top. “@10” = only the first 10 results are considered.
- MRR (Mean Reciprocal Rank): mean of
1/rank of the first relevant document. If the right document is always first, MRR = 1.- recall@k: proportion of truly relevant documents appearing in the top-k. Measures coverage, not precision.
The full vectorless RAG case
Some teams push the logic to the end: no vector database, no external embedding provider. Everything relies on Markdown-structured claims, indexed by wikilinks and keywords, searched by grep lexical and TF-IDF similarity (an earlier formulation than BM25) in Python stdlib. Gain: no external dependency, no infrastructure cost, no reindexation, total explainability. Cost: synonyms captured by embeddings are missed. Relevant for bases of 1,000 to 10,000 documents with stable technical vocabulary, with a small number of human operators. Does not scale to consumer-facing systems.
Family 3 — Emerging approaches (post-2024)
The three previous approaches share the same structure: a mechanical retriever followed by a generator. Emerging approaches disrupt this scheme by entrusting the LLM itself with part of the reasoning about what to search, when, and how.
GraphRAG: building a knowledge graph from the corpus
GraphRAG (Microsoft, Edge et al., April 2024) replaces chunk indexing with a more ambitious upstream step: traverse the corpus, extract entities (people, organizations, places, concepts) and the relationships between them, then build a knowledge graph. Over this graph, a community detection algorithm (Leiden) groups strongly connected entities into communities, and an LLM summarizes each community.
Two usage modes:
- Local search: targeted factual question → identify the relevant subgraph → answer.
- Global search: holistic question (“what are the main trends of the corpus?”) → traverse community summaries → synthesis.
GraphRAG shines on multi-hop questions (requiring linking pieces of information scattered across the corpus) and global thematic questions. Cost: building the graph is heavy, each new document requires an update. LightRAG (He et al., October 2024) offers a lighter variant.
Paper 2026 (Fan et al., “Do We Still Need GraphRAG?”): GraphRAG remains superior on multi-hop queries, but the gap narrows with more capable LLMs that can recompose from flat chunks.
Self-RAG: the model decides when to retrieve
Self-RAG (Asai et al., 2023) introduces a paradigm shift: rather than calling the retriever systematically, fine-tune the LLM so it generates reflection tokens during its own inference:
[Retrieve]: the model decides it needs to search before answering.[IsRel]: it judges if the retrieved passage is relevant.[IsSup]: it judges if the answer is well supported by the passage.[IsUse]: it judges if the answer is ultimately useful.
Advantage: useless retrievals on questions the model already knows are skipped, and the relevance judgment is made explicit. Cost: the generator model must be fine-tuned, which is heavy and specific.
CRAG (Corrective RAG, Shi et al., January 2024) offers a lighter version: a T5 evaluator classes retrieved passages into {correct, ambiguous, incorrect}, and triggers a backup web search if all are incorrect.
Adaptive RAG: choosing the strategy per query
Adaptive RAG (Jeong et al., March 2024) routes each question to a different strategy via a lightweight classifier:
- Simple question (direct factual) → no retrieval, direct LLM answer.
- Single-document question → classical single-step retrieval.
- Complex multi-hop question → iterative multi-step retrieval.
The gain is not on the quality of each isolated type, but on global economy: complexity is paid only where it adds value.
Agentic RAG: an agent that plans its searches
Agentic RAG abandons the fixed “retrieve then generate” sequence in favor of a loop where an LLM agent decomposes the question into sub-questions, launches several successive searches, reads its own intermediate results, and dynamically decides when it has enough information. It is slower and more expensive per request, but handles problems that sequential approaches miss (cross-investigation, multi-step reasoning).
LangGraph and LlamaIndex are the dominant frameworks in 2026. Recent extensions: Doctor-RAG (Jiao et al., April 2026) adds a failure-aware repair layer; “Don’t Retrieve, Navigate” (Sun et al. 2026) proposes an agent that learns the corpus structure rather than retrieving passively.
Other useful variants to know
- LongRAG: much larger chunks (~6,000 tokens), to exploit extended context windows.
- RAPTOR: hierarchical tree of recursive summaries — indexed at the paragraph, section, and document levels simultaneously.
- HyDE (Hypothetical Document Embeddings): instead of embedding the question, ask the LLM to produce a hypothetical answer, then embed that answer to look for close documents. Useful when the question is short and ambiguous.
- SAGE (Wang et al., 2026): selective extraction reducing tokens transmitted to the generator by 60%, without losing answer quality.
Which family to choose?
There is no universal winner. The choice depends on the corpus, question type, and operational constraints.
Decision matrix
| Context | Recommendation | Why |
|---|---|---|
| Corpus < 10,000 chunks, factual questions | Classical dense vector RAG | Mature tooling, fast time-to-production, sufficient in the vast majority of cases. |
| Technical vocabulary, proper nouns, exact identifiers | Hybrid BM25 + dense (RRF) | Robustness on exact keywords + semantics. Best production tradeoff in 2026. |
| Multi-hop questions, cross-document reasoning | GraphRAG or Agentic RAG | Explicit structure (graph) or planning (agent) to recompose scattered evidence. |
| Small stable corpus (< 200,000 tokens), varied questions | Direct long context + prompt caching | Maximum simplicity, no index to maintain, cache absorbs cost. |
| High explainability requirement (legal, medical, audit) | BM25 or hybrid | Clean, traceable citations, no opaque cosine distance. |
| Need to avoid useless retrievals | Self-RAG or Adaptive RAG | The model decides itself when to search. |
| Visual corpus (scanned PDFs, diagrams) | ColPali | Page image retrieval, no OCR. |
Some heuristic rules
- Start simple: a mature dense vector RAG (embedder + Qdrant + Voyage reranker) covers 80% of cases. Add sophistication only when the success metric stagnates.
- Measure before optimizing: without an evaluation set (questions + expected answers), any optimization is blind. RAGAS (open-source framework) automates part of the scoring.
- Hybrid by default in production: except for tiny corpora, the gain of hybrid over dense alone is documented and the overhead is modest.
- GraphRAG and Agentic are expensive: reserve for cases where flat approaches clearly fail.
What we keep in 2026
RAG is no longer a technique, it is a family. Several architectures coexist and the choice is guided by the profile of the corpus and questions, not by a universal hierarchy. Long context has not eliminated RAG: it has absorbed small stable corpora and left RAG with the large, dynamic, and source-citing ones. BM25 remains a competitive baseline thirty years after its invention — it has become elegant again to combine it with dense rather than try to replace it. Finally, emerging approaches (GraphRAG, Self-RAG, Agentic) do not impose themselves as replacements: they add to the toolbox for problems the classical sequential pipeline does not solve.
Going further
- fine-tuning-vs-rag article: when to choose one or the other, and why both are complementary.
- rag article: step-by-step introduction to the simplest pipeline.
- context-engineering article: what happens before retrieval (selection, compression, ordering of context).
- agentic-ai article: the agent loops underlying Agentic RAG.