In short
When you type “cheap car to buy” into a search engine, you do not want a page that contains those exact words — you want results about buying used vehicles on a small budget. Semantic search makes that distinction. It relies on embeddings: numerical representations that encode the meaning of a text, not just its words. This article explains what these vectors are, why they require specialized indexing tools, and what this means concretely inside a RAG system.
In short: an embedding is a postal address for the meaning of a text. Instead of writing “12 Lilac Street” to locate a house, you write “(longitude=-1.234, latitude=48.567, …)” with a thousand coordinates. Two texts close in meaning live a stone’s throw from each other in this giant city. The search engine no longer reads the words — it computes the distance between addresses.
From word to vector: a brief history
The idea of encoding text as numbers is not new. The earliest methods — TF-IDF, bag-of-words — treated a document as a bag of words, with no order and no meaning. They worked well for exact matching, not for understanding.
In 2013-2014, Word2Vec and GloVe laid the foundation of modern embeddings. These models assign each word a fixed vector learned from vast corpora. The remarkable property: words used in similar contexts end up close together in vector space. “King” and “queen” are neighbors. “Paris” and “France” share a geometric relationship similar to the one between “Berlin” and “Germany”.
But these models had a fundamental flaw: each word had a single vector, regardless of context. “Bank” (the financial institution) and “bank” (the river edge) received the same representation.
BERT, released in 2018 by Devlin et al. at Google, solved this problem with contextualized representations: a word’s vector depends on its neighbors in the sentence. A decisive breakthrough — but one that created a new problem for large-scale search.
In short: Word2Vec gave each word a single address, like a paper dictionary. BERT gives each word a different address depending on context, as if “bank” in “river bank” and “bank” in “investment bank” were located in two different neighborhoods of the semantic city. The price: you have to re-read the entire sentence each time.
Sentence-BERT: the practical turning point
To compare two sentences with the original BERT, they had to be encoded together in a single pass. On 10,000 sentences, that means about 50 million pairs to process — roughly 65 hours of computation on a standard GPU. Unusable in production.
Reimers and Gurevych (EMNLP 2019) sidestepped this problem with Sentence-BERT (SBERT). The architecture uses two identical BERT encoders working in parallel: each sentence is encoded separately, then the representations are compared. Processing time for 10,000 comparisons drops to 5 seconds.
This architectural shift set the standard for all modern semantic search: encode separately, store, compare. It is the foundation of the RAG pipeline.
How a sentence embedding works
An embedding model takes a text as input and produces a vector of floating-point numbers — typically between 384 and 4,096 dimensions depending on the model. This vector is a coordinate in a mathematical space where the proximity between points expresses kinship of meaning.
The metric used is cosine similarity: it measures the angle between two vectors, independent of their size. Two semantically close sentences will have a cosine close to 1. Two unrelated sentences, close to 0. Sentences with opposite meaning, negative.
What makes this mechanism powerful: a query worded differently from a document can still retrieve that document, if the two have nearby vectors. “Cheap car” and “buying a vehicle on a tight budget” can land in the same neighborhood.
What makes it fragile: the vector mostly captures surface similarity — theme, lexical register, syntax. It struggles with implicit semantics: conversational undertones, reasoning by analogy, pragmatics. A rhetorical question and a sincere question can produce very similar vectors while calling for opposite answers.
Another structural bias is worth mentioning: embedding models project texts into a narrow cone of vector space, using only a fraction of the available dimensions. This phenomenon, called anisotropy, causes artificial compression of semantic diversity.
Why SQL databases are no longer enough
A classic database excels at one thing: looking up an exact value. “Give me every customer whose name is Smith.” The query compares strings or numbers.
Vector similarity search is a different problem. The question is no longer “are these two values identical?” but “among a million vectors, which are closest to this query vector?”. This is the k-nearest neighbor problem (kNN).
Exact search has a complexity proportional to the number of vectors times their dimension. On a billion vectors in 768 dimensions, it is infeasible in real time.
The solution: approximate nearest neighbor (ANN) algorithms. They give up the guarantee of finding the exact neighbors in exchange for finding very close neighbors in a fraction of the time.
The two main families of indexing
Two approaches dominate ANN indexing in vector stores:
IVF (Inverted File Index) — Partition vector space into clusters via k-means. At query time, only the closest clusters are searched (the nprobes parameter). Product quantization (PQ) further compresses memory: each vector is split into sub-vectors, each approximated by a code of a few bits. This is the algorithm at the core of FAISS, the library from Meta AI (Johnson et al., 2017), an industrial and academic reference.
HNSW (Hierarchical Navigable Small World) — A graph-based approach. Each vector is a node connected to its closest neighbors. A hierarchy of layers enables fast navigation: you enter the graph through the sparsest layers (long range), then descend toward the dense layers (local precision). Search complexity is logarithmic in the number of vectors. HNSW generally offers a better recall/latency trade-off than IVF, at the cost of a higher memory footprint. Formalized by Malkov and Yashunin (2016).
| Criterion | IVF-PQ | HNSW |
|---|---|---|
| Memory footprint | Low (PQ compression) | High |
| Query latency | Variable (depends on nprobes) | Stable, logarithmic |
| Recall at iso-budget | Slightly lower | Generally higher |
| Typical use case | Billions of vectors, memory-constrained | Millions of vectors, latency-critical |
The vector store ecosystem
From indexing libraries, the field has evolved toward full-featured vector databases, with persistence, metadata filters, and standardized APIs. A few representatives:
- FAISS (Meta AI, open-source): the reference library for raw performance, CPU and GPU, up to billions of vectors. Not a standalone database — it plugs into an architecture.
- Milvus (open-source, LF AI): distributed architecture, designed for massive volumes, supports multiple index types.
- Weaviate (open-source, Go): object-oriented with schema, GraphQL API, built-in embedding modules.
- Chroma (open-source, Python): embedded, zero-configuration, aimed at local development.
- Pinecone (proprietary, cloud): fully managed, automatic scaling, latency guaranteed by contract.
None of these solutions is universally superior. The choice depends on data volume, latency constraints, available infrastructure, and budget.
The RAG pipeline and its weak points
Embeddings and vector stores are the central components of a RAG pipeline (Retrieval-Augmented Generation). The principle: instead of memorizing knowledge in the model’s weights, you index documents in a vector store. At each query, you retrieve the most relevant passages and inject them into the LLM’s context.
The full pipeline:
- Splitting documents into passages (chunking)
- Embedding each passage
- Indexing in the vector store
- At inference time: embed the query, run ANN search, retrieve the top-k passages
- Inject into the LLM’s prompt
This pipeline is conceptually simple. It is fragile in practice. Gao et al. (2023) documented systematic failures in a literature review, and Barnett et al. (2024) identified seven recurring failure points:
- Chunking that is too fine loses context. Too wide introduces noise.
- The embedding model may not capture the nuances of the business domain.
- The query and the documents may belong to different semantic distributions, even when they cover the same topic.
- Cosine similarity does not capture confidence: an “ambiguous” vector and a “precise” vector can be treated as equivalent.
The last point is documented in Semantics at an Angle (2025): cosine similarity ignores vector magnitude, erasing potentially useful information.
Dense, sparse, hybrid: an unsettled debate
The community contrasts two major families of retrieval methods. Dense methods (embeddings) capture semantics but can miss exact matches. Sparse methods (BM25, TF-IDF) catch precise lexical matches but ignore meaning.
The practical answer — combining both in hybrid retrieval — is empirically superior on most tasks. It does introduce a weighting parameter that is hard to calibrate without validation data. Recent work (DAT, 2025) explores per-query dynamic tuning, but this approach is not yet standardized.
On some tasks — code retrieval, highly specialized domains, out-of-domain — BM25 systematically outperforms dense methods. This result, highlighted on the BEIR benchmark, remains hard to predict a priori.
Which vector store to choose by volume and need
| Context | Recommendation | Why |
|---|---|---|
| < 100K chunks, rapid prototype | FAISS in-memory or Chroma | Zero infra, direct Python integration, sub-ms latency. Ideal for iterating. |
| 100K - 10M chunks, cloud production | Pinecone, Weaviate, Qdrant Cloud | Managed, scalable, rich metadata. Pinecone = leader, Weaviate/Qdrant = mature open source. |
| > 10M chunks, on-prem or sovereignty | Milvus, Vespa | Distributed, horizontally scalable, optimized for industrial volumes. Vespa = Yahoo!, Milvus = Zilliz. |
| Technical vocabulary, exact identifiers | BM25 + dense (hybrid RRF) | Dense alone misses proper nouns or exact identifiers. Hybrid combines strengths of both. |
| Critical latency (< 50 ms p99) | HNSW in-memory (FAISS, Qdrant configured HNSW) | IVF is less memory-costly but slower at lookup. HNSW dominates on latency. |
| Corpus evolution (frequent insertion) | Qdrant or Weaviate | HNSW structures support efficient incremental insertion. FAISS rebuild required at each significant batch. |
What to remember
- An embedding is a vector representation of a text in a space where geometric proximity expresses kinship of meaning. This is the mechanism that lets a system understand the intent of a query, not just its words.
- Vector stores are databases optimized for large-scale similarity search. The two dominant algorithms — IVF and HNSW — offer different trade-offs between memory, latency, and precision.
- RAG relies entirely on retrieval quality: a bad embedding or bad document chunking produces mediocre results, regardless of the power of the LLM.
- Cosine similarity has structural limits: it ignores vector magnitude and struggles with implicit semantics. These limits are active in every RAG pipeline.
- Combining dense + sparse (hybrid retrieval) is empirically superior to either method alone on most tasks, but it introduces calibration complexity.