Vector search connects information retrieval, machine learning, and database systems. Their vocabularies overlap, but a shared word can describe a different operation: recall can mean agreement with exact neighbors or recovery of documents judged relevant. That difference changes how you evaluate a retrieval system.
This glossary follows the pipeline through embedding models, ANN algorithms, distance and similarity, hybrid retrieval, reranking, and operations. Read the families together when selecting a database or investigating a missed answer. A graph index cannot fix a missing source document, and a reranker cannot judge candidates that the retriever never returned.
The definitions and implementation references were checked on October 4, 2026. Formulas and the fictional worked example explain the concepts; they are not performance benchmarks. Start with a versioned corpus and query set, then compare candidate systems against the same relevance judgments and operating constraints.
- 01Representation, index, and reranker solve different problems.Embeddings represent content; an ANN index searches candidates; a reranker orders the candidates it receives. Identify the failing stage before changing the stack.
- 02Unit-length normalization makes cosine and dot-product scores equal.Match the embedding model's scoring contract, keep distance and similarity directions explicit, and define what happens to zero vectors.
- 03ANN recall and relevance recall have different denominators.Agreement with exact neighbors evaluates the index. Recovery of judged relevant documents evaluates the retrieval task. Neither is answer accuracy.
- 04Fusion and compression need corpus-specific evaluation.RRF combines ranks without score normalization; quantization changes the representation used for search. Their costs and quality effects depend on the workload.
- 05Filters, revisions, and measurement boundaries belong in the result.Record permissions, dataset and model versions, cutoffs, load, and latency scope. Without them, a leaderboard or vendor result cannot predict your application.
01 — RepresentationEmbedding models.
Models turn text or other inputs into representations that a retriever can compare. Separate the representation, the model architecture, and the training objective: none of these alone tells you whether a document answers a question.
Embedding. A numerical representation of an input. A useful retrieval embedding puts related queries and documents near each other under the model’s intended similarity function; similarity is not a guarantee of factual relevance.
Dense embedding. A vector whose dimensions usually carry learned, distributed information. Its coordinates are not normally readable keyword labels. Dense retrieval can match paraphrases without requiring the same words in the query and document.
Sparse embedding. A representation with relatively few nonzero coordinates. Lexical and learned sparse methods can use vocabulary-aligned dimensions. Sparse does not mean BM25: BM25 is a scoring method, while sparsity describes a representation.
Bi-encoder. An architecture that encodes queries and documents separately. Document representations can be computed before search. Query and document encoders may share weights or differ; both must produce compatible representations.
Cross-encoder. An architecture that processes a query and document together to score their relationship. It does not produce the reusable independent document vectors used by a conventional bi-encoder index.
Late interaction. A retrieval design that independently encodes inputs but postpones richer interaction until scoring. ColBERT compares query-token and document-token representations rather than reducing each passage to a single vector.
ColBERT. A late-interaction text retrieval model described by Khattab and Zaharia. Its token-level scoring allows fine-grained matches; it does not imply a universal storage multiplier or a guaranteed improvement on long documents.
SPLADE. Sparse Lexical and Expansion Model: a learned sparse retrieval approach that can expand a representation with related vocabulary terms. It connects neural representation learning to inverted-index retrieval.
ColPali. A document retrieval model that represents document-page images and uses late interaction. This can preserve visual layout information that plain extracted text loses; document type and evaluation determine its usefulness.
Embedding dimension. The length of a vector. More coordinates can increase storage and arithmetic, but dimension is not a standalone quality score. Compare models on the task before trading dimensions for cost.
Token. A model-specific input unit. Tokenization governs input length and truncation. A character limit, word count, and token limit are different quantities; check the encoder’s tokenizer and model documentation.
Pooling. Combining token-level representations into a representation for a span or document. The pooling strategy is part of the model contract, not an interchangeable detail to change after indexing.
Chunk. The indexed unit cut from a source document. Chunk boundaries affect what evidence can be retrieved. Preserve the source ID, location, and permissions so a retrieved vector can lead back to usable evidence.
MTEB. Massive Text Embedding Benchmark: an evaluation framework spanning multiple embedding tasks. Compare task, language, dataset, model revision, and evaluation settings rather than treating an aggregate leaderboard position as your retrieval result.
Matryoshka embedding / MRL. Matryoshka Representation Learning trains nested prefixes of a representation to remain useful. Truncation is appropriate only when the model supports it; indexed and query vectors must use the same supported dimension.
Multimodal embedding. A representation that supports inputs such as text and images. Verify whether a model supports cross-modal matching, joint inputs, or separate encoders; the word multimodal alone does not specify that contract.
Architecture references: Sentence-BERT explains independently encoded sentence representations; ColBERT specifies late interaction and MaxSim; SPLADE describes learned sparse expansion; ColPali covers document-page images; Matryoshka Representation Learning establishes supported nested representations; MTEB describes task-specific embedding evaluation.
For the document-unit decision, use our RAG chunking playbook. Chunking and encoder choice should be evaluated together: a suitable model cannot supply evidence that was cut away from the indexed unit.
Bi-encoder
Precompute document representations, then compare them to an encoded query. Verify model-specific scoring and input formatting.
Cross-encoder
Score each candidate with the query. Measure the ranking gain and cost on a fixed candidate set.
Late interaction
Preserve finer-grained representations and interact at scoring time. Evaluate storage, compression, and query cost.
02 — Candidate searchANN algorithms.
An index accelerates candidate search. Its job is to approximate the neighbors defined by a metric; semantic relevance still depends on the representation and documents. There is no universal collection-size threshold that selects an algorithm.
ANN. Approximate nearest neighbor search. It trades exact agreement with nearest neighbors for a resource or latency advantage. The size and direction of that trade-off depend on the index, data, parameters, and hardware.
Exact kNN. Exact k-nearest-neighbor search computes the nearest results under the selected metric. A flat scan provides an evaluation baseline; selective filtering or suitable hardware may also make exact search practical in an application.
HNSW. Hierarchical Navigable Small World: a graph index with multiple layers used to navigate toward near neighbors. The HNSW paper describes the algorithm; a database implementation adds its own storage, filtering, and update behavior.
NSG. Navigating Spreading-out Graph: another graph-based ANN method. It is distinct from HNSW, not its predecessor; the original NSG paper was first submitted after the original HNSW paper.
Vamana / DiskANN. Vamana is a graph construction approach associated with DiskANN. DiskANN targets efficient search with disk-based storage. Check the implementation and deployment constraints rather than assuming disk residency removes every memory requirement.
IVF. Inverted file indexing partitions vectors into lists, then searches selected lists near the query. A coarse quantizer assigns vectors to those lists. Searching more lists changes the speed and recall trade-off.
IVFFlat. An IVF index that stores the vectors without product quantization. It needs representative data for training its partitions. In pgvector it is a separate index type from HNSW, with different build and query behavior.
IVF-PQ. An IVF index with product-quantized vector codes. Partition selection limits candidates and compressed codes approximate distances. This combines two sources of approximation, which should be evaluated separately where possible.
Faiss. An open-source library for similarity search and clustering. It offers multiple index families, including exact flat indexes, IVF, quantization, and HNSW. It is a library, not a complete database access-control or persistence service.
ScaNN. A Google similarity-search library that combines search-space pruning and scoring optimizations, including quantization. Its suitability depends on the workload and supported execution environment, not a fixed minimum vector count.
Annoy. Spotify’s approximate-neighbor library using random projection trees and file-backed indexes. Its built index is designed for read-only lookup; that lifecycle matters when assessing continuously changing collections.
M / m. An HNSW connectivity parameter. More connections increase graph storage and can alter search quality. The exact meaning, capitalization, supported range, and defaults belong to the implementation.
ef_construction. The candidate-list size used while constructing an HNSW graph. It controls construction search effort; it is not the graph-degree parameter. Evaluate build time, insert cost, and resulting retrieval behavior together.
ef_search / ef. The search candidate-list size for HNSW queries. pgvector exposes hnsw.ef_search. Increasing search effort can improve agreement with exact neighbors at additional cost; confirm behavior with filters and the chosen library.
nlist / lists; nprobe / probes. IVF partition count and query-time partition search count. Faiss uses nlist and nprobe; pgvector uses lists and ivfflat.probes. Similar concepts do not make configuration syntax portable between products.
ANN recall@k. The fraction of the exact top-k neighbors recovered by approximate search for the same query, metric, corpus, and filters. Define tie handling and missing results. This measures index fidelity, not human relevance.
Implementation references: the HNSW paper describes multilayer graph search; the NSG paper describes a separate graph method; DiskANN, ScaNN, and Annoy document their respective implementations. The Faiss index reference distinguishes flat, IVF, graph, and quantized indexes.
For a PostgreSQL implementation, read the self-hosted RAG tutorial alongside the pgvector documentation. In pgvector, HNSW construction effort, connectivity, and query effort are separate settings. Record the installed extension version before using examples or copying settings from another system.
03 — GeometryDistance and similarity.
Use the similarity contract of the embedding model and the operators supported by the index. Keep score direction explicit: larger similarity is usually better, while smaller distance is better. Comparing two numbers without their definitions invites ranking mistakes.
Cosine similarity. The dot product divided by the product of vector lengths. It compares direction, not magnitude. For nonzero real vectors it lies between −1 and 1. The zero vector has no defined cosine similarity.
Cosine distance. Often defined as 1 minus cosine similarity. Confirm the API definition. pgvector’s cosine-distance operator returns distance, so ascending order retrieves the most similar items; subtract from 1 to report similarity.
Dot product / inner product. The sum of coordinate-wise products. For unit-length vectors it equals cosine similarity. Without normalization, magnitude affects the score; that can be intentional for a model trained with inner-product scoring.
L2 distance / Euclidean distance. The square root of the sum of squared coordinate differences. Some libraries return squared L2 instead: ranking is preserved, but thresholds and displayed values are different.
L1 distance / Manhattan distance. The sum of absolute coordinate differences. Its geometry differs from L2. Use it only with a representation and search implementation that support the intended comparison.
Hamming distance. The number of differing positions in equal-length binary strings or vectors. It can compare binary codes; it is not the cosine similarity of the original floating-point vectors.
Jaccard similarity. Set intersection size divided by union size. For binary representations, active coordinates can be treated as sets. Define the empty-union convention; libraries may expose Jaccard distance rather than similarity.
L2 normalization. Dividing a nonzero vector by its Euclidean norm to give it unit length. Cosine can be calculated directly without storing normalized vectors. Normalize for dot-product equivalence only when compatible with the model.
MaxSim. ColBERT’s aggregation matches each query-token representation to its most similar document-token representation, then sums the query-token contributions. It is not simply the best similarity between any pair of tokens.
Quantization. Encoding vectors or their components with a compact representation. The encoding changes storage and scoring behavior. Evaluate approximation errors on queries; a smaller index alone does not establish better end-to-end value.
Scalar quantization / SQ. Quantizing coordinates individually, for example to lower-precision numeric values. Calibration and rescaling rules depend on the implementation. Coordinate-level precision reduction is different from learning codebooks for groups of coordinates.
Product quantization / PQ. Splitting vectors into subvectors and representing each with a learned codebook entry. Distance calculations can use those codes. Codebook training and the chosen partitioning affect the approximation.
Binary quantization. Encoding vector information as bits. Binary candidate search may be followed by rescoring with higher-precision vectors. Retaining original vectors means total system storage differs from the binary code size alone.
A fictional vector calculation
For hand-picked vectors q = [1, 0], a = [3, 4], and b = [1, 1], dot(q, a) = 3 and dot(q, b) = 1. Cosine(q, a) = 0.6 while cosine(q, b) is approximately 0.7071. Dot product ranks a first; cosine ranks b first because a has greater magnitude but a less aligned direction. These are arithmetic examples, not learned embeddings or relevance judgments.
The Python example below rejects zero vectors, unequal dimensions, and nonfinite inputs before computing cosine. It also shows rank fusion with a fictional constant of 2. Production systems should use their documented distance operators and numerical implementations.
import math
def cosine(a, b):
if not a or len(a) != len(b):
raise ValueError("nonempty equal dimensions required")
if not all(math.isfinite(x) for x in [*a, *b]):
raise ValueError("finite coordinates required")
sa, sb = max(map(abs, a)), max(map(abs, b))
if sa == 0 or sb == 0:
raise ValueError("cosine undefined for zero vectors")
ua, ub = [x / sa for x in a], [y / sb for y in b]
na, nb = math.hypot(*ua), math.hypot(*ub)
return math.fsum((x / na) * (y / nb)
for x, y in zip(ua, ub))
def rrf(rankings, constant=2):
if not math.isfinite(constant) or constant < 0:
raise ValueError("finite nonnegative constant required")
scores = {}
for ranking in rankings:
seen = set()
for rank, doc in enumerate(ranking, start=1):
if not isinstance(doc, str):
raise ValueError("string document IDs required")
if doc in seen:
raise ValueError("duplicate ID within a ranking")
seen.add(doc)
scores[doc] = scores.get(doc, 0) + 1 / (constant + rank)
return sorted(scores.items(), key=lambda item: (-item[1], item[0]))
print(round(cosine([1, 0], [3, 4]), 4)) # 0.6
print(rrf([["A", "B", "C"], ["B", "D", "A"]]))
# [("B", 0.583333...), ("A", 0.533333...),
# ("D", 0.25), ("C", 0.2)]The RRF code expects string document IDs and ranks starting at 1, rejects repeated IDs within a list, gives an absent document no contribution, and breaks equal scores by ID. It is a small teaching implementation for finite inputs, not a search service or a benchmark of database performance.
04 — Rank fusionHybrid retrieval.
Dense, lexical, and learned sparse retrieval provide different signals. A hybrid pipeline combines them, but adding a retriever is a hypothesis to test. A rare product identifier and a broad paraphrase may benefit from different candidate sources.
Hybrid retrieval. Combining results from multiple retrieval signals, commonly lexical or sparse search and dense vectors. Both retrievers must use compatible document identifiers, visibility rules, and document revisions before their results are merged.
BM25. A lexical scoring function using term-frequency saturation, inverse document frequency, and document-length normalization. It does not produce only integer scores. Tokenization, stemming, and implementation details affect matching and score values.
k1 and b. BM25 parameters controlling term-frequency saturation and document-length normalization. They are not embedding dimensions or ANN parameters. Tune them on development queries and reserve separate queries for final evaluation.
BM25F. A field-aware extension of BM25 that combines evidence from fields such as title and body. Field weights and length normalization influence scoring; matching a field does not automatically establish document relevance.
TF-IDF. Term frequency–inverse document frequency weighting. It combines within-document occurrence with collection-wide rarity. It is a useful lexical concept, but should not be treated as an exact synonym for BM25.
Reciprocal rank fusion / RRF. Combining ranked lists by adding reciprocal contributions from each document’s rank. Raw score scales are not needed. The rank constant and candidate windows still affect results; fusion is not parameter-free.
Score normalization. Transforming different retriever scores before combining them. Min-max and z-score methods depend on the observed score distribution. Normalized values should not be interpreted automatically as calibrated relevance probabilities.
Convex combination. A weighted combination with nonnegative weights summing to one. With normalized retriever scores, it offers a tunable fusion rule. Select weights with evaluation, not by assuming equal numerical scores carry equal evidence.
Query expansion / HyDE. Query expansion adds terms or alternative representations. HyDE generates a hypothetical document and embeds it for retrieval. The generated text is a search aid, not verified evidence to cite in an answer.
Multi-query retrieval. Running reformulations of a query and merging their candidates. Reformulations may improve coverage or drift away from the original request. Deduplicate by document ID and measure the added retrieval cost.
Metadata filter. Restricting eligible results using attributes such as tenant, date, or category. Evaluate candidate search under the actual filter distribution. A highly selective filter can change recall, result count, and the best search strategy.
Pre-filter / post-filter. Filtering before candidate search or after candidate generation. Neither label promises ANN recall. In pgvector approximate-index scans, filtering can leave too few rows; iterative scans can search further within configured limits.
Read the fusion score as a rank calculation
The Cormack, Clarke, and Büttcher RRF paper defines the score as the sum of 1 / (c + rank) over retriever lists. Here c names the rank constant so it is not confused with the top-k cutoff. The paper used c = 60 in its experiments; that experimental choice is not evidence that it is optimal for your corpus.
In the fictional Python example, the lexical ranking is A, B, C and the dense ranking is B, D, A. With c = 2, B receives 1/4 + 1/3 = 7/12, while A receives 1/3 + 1/5 = 8/15. D receives 1/4 and C receives 1/5. The fused order is B, A, D, C. Those values describe only these invented rankings; no query relevance labels were used to select them.
See the hybrid search reference for pipeline vocabulary. The Stanford information retrieval text explains BM25 saturation and length normalization. Keep the fusion rule explicit, then compare lexical-only, dense-only, and combined retrieval on held-out queries.
05 — Final orderingReranking.
Candidate retrieval decides which documents a reranker can see. Reranking then changes order within that candidate set. It cannot recover a document that was never included, and retrieval quality does not directly measure answer truthfulness.
Reranker. A scoring stage that reorders retrieved candidates. It may use a cross-encoder, late interaction, or another model. Compare its gain on judged queries with its latency and resource cost.
Cross-encoder reranker. A reranker that jointly processes the query and each candidate. Joint interaction allows different relevance features from an independently encoded vector comparison, but does not guarantee a quality gain for every query type.
Late-interaction reranker. A reranking stage that compares independently computed token-level representations. Candidate selection can happen with another retriever first. Storage and scoring costs depend on the representations and the compression used.
Commercial and open-weight rerankers. Cohere and Voyage offer reranking services; BAAI publishes the BGE reranker family. Validate current model names, accepted inputs, limits, terms, and data handling before selecting a specific service or deployment.
Multimodal reranker. A reranker accepting representations or inputs beyond plain text, such as document images. Confirm whether query and candidate modalities are supported; support for image retrieval does not imply support for every reranking task.
LLM-as-judge reranker. Prompting a language model to assess candidate relevance. Define the rubric and test order sensitivity and repeatability. Retrieved instructions are untrusted content; a scoring prompt does not make them authoritative.
Top-k / top-n. The candidate cutoff and final retained-result cutoff in a pipeline. These are local naming conventions, not universal parameter meanings. State the cutoff at each retrieval, fusion, reranking, and context-selection stage.
Judgments / qrels. Document relevance labels for each evaluation query. Record how labels were produced, which documents were assessed, and whether labels are binary or graded. Unjudged documents and irrelevant documents are not necessarily equivalent.
Relevance recall@k. The relevant documents retrieved in the first k results divided by all relevant documents in the evaluation set for that query. A hit-rate metric asking whether any relevant document appears has a different denominator.
Precision@k. The relevant documents in the first k results divided by k. Specify the convention when fewer than k results are returned and keep it fixed between experiments.
NDCG@k. Normalized discounted cumulative gain: graded relevance gains discounted by rank, divided by the ideal ranking’s gain at the same cutoff. Specify gain and discount formulas and how queries with no relevant documents are handled.
MRR. Mean reciprocal rank: average the reciprocal of the first relevant result’s rank across queries, assigning zero when none is retrieved. State any cutoff; MRR does not measure how many relevant documents were recovered.
A fictional evaluation with explicit denominators
Suppose the exact nearest neighbors at cutoff 3 are A, B, C, and the approximate search returns A, D, B. ANN recall@3 is 2/3 because two exact neighbors were recovered. Separately, suppose the documents judged relevant are A and C. Relevance recall@3 is 1/2, precision@3 is 1/3, and reciprocal rank is 1 because A is first. A hit-rate test would count this query as a hit despite the missed relevant document C.
This invented single-query calculation explains why the metrics cannot be swapped. To report MRR, average reciprocal ranks across the stated query set. To report relevance recall, disclose the judgments and their completeness. For rank-aware graded evaluation, see the Stanford evaluation chapter. Measure final answer correctness separately from retrieval quality.
06 — OperationsOperational metrics and infrastructure.
The vocabulary of a database is also an operating contract. Evaluate indexing, updates, deletion, isolation, and restore procedures alongside query quality. A promising benchmark cannot resolve a missing permission check or stale document revision.
Vector database / vector search extension. A database or extension providing vector storage and similarity queries. Pinecone, Qdrant, Weaviate, Milvus, and Chroma have different service and deployment models. pgvector adds vector types and search indexes to PostgreSQL.
Index. A data structure used to accelerate a search operation. An index is not the source document or an access-control rule. Keep source content, identifiers, and authorization available when vectors are retrieved.
Collection / namespace. Logical groupings exposed by a product. Their schemas, limits, and isolation semantics differ by implementation. A namespace name alone does not prove that every query enforces the required tenant boundary.
Sharding. Partitioning data across storage or compute units. Routing, balancing, and cross-shard merging affect behavior. Benchmark with the actual shard arrangement rather than extrapolating an isolated local-index result.
Replication. Maintaining copies of data for availability or read capacity. State consistency and failure behavior. Replication is different from a backup because a deletion or corruption may propagate to the copies.
Index build time. Elapsed time to create the search structure from a specified dataset. Report preprocessing and training inclusion, hardware, parallelism, and data size. Rebuilding every vector is a different job from inserting a new document.
Index size. The storage footprint of an index. Separate graph or code storage from source vectors, document content, metadata, replicas, and temporary build memory. A compression ratio needs an explicit numerator and denominator.
QPS / throughput. Completed queries per second under a specified workload and concurrency. Report errors and latency with throughput; successful performance at low concurrency does not establish behavior under a burst or sustained load.
P50 / P95 / P99 latency. Percentiles of a measured latency distribution. Declare the measured boundary, time window, sample size, load, and percentile calculation. A server search timer excludes network, embedding generation, reranking, and answer generation unless explicitly included.
Cold start / warm cache. Initial or uncached work versus a warmed execution path. State which resource is cold: a process, model, index, or storage cache. A warm-cache benchmark should not be presented as cold-start user experience.
Freshness / deletion. How quickly a document change or deletion becomes visible to retrieval. Version identifiers help detect mismatches between source content and vectors. Test the behavior required by the application rather than assuming updates are immediate.
Evaluation set. A versioned collection of queries, corpus documents, and relevance judgments. Record language, query categories, exclusions, cutoff, filters, model revision, and index settings so another run can be compared meaningfully.
SLO. Service-level objective for the behavior that users depend on. A latency objective and a retrieval-quality objective answer different questions. Report both with their measurement windows and test realistic failure and overload paths.
Embedding migration. Moving to a new representation model or revision. Re-embed documents as needed, version the index, and route queries to a compatible representation. Mixing incompatible embedding spaces can silently change results even when dimensions match.
A decision table for the first investigation
| Observed problem | Inspect first | Evidence to collect |
|---|---|---|
| A relevant document never appears | Source coverage, chunks, filters, embedding compatibility | Judged query, eligible document IDs, corpus and model revision |
| Exact search works; ANN misses neighbors | Index parameters and filtered candidate search | Exact top-k and ANN top-k under identical filters |
| Good candidates rank too low | Fusion and reranking | Candidate lists, final ranking, graded judgments |
| Compressed search loses quality | Quantization and rescoring | Quality and total storage for each representation |
| Latency spikes under load | Queues, caches, concurrency, remote calls | End-to-end percentiles, throughput, failures, hardware and time window |
| A tenant sees unauthorized content | Authorization and filter enforcement | Denied-access tests at retrieval and source-document boundaries |
The table is an engineering diagnostic, not a measured ranking of techniques. Fix the observed failure, then rerun the same evaluation before increasing architectural complexity.
07 — Putting it togetherUse vocabulary to make testable choices.
Write down the representation, ranking rule, and evaluation contract.
A retrieval design becomes easier to assess when each stage has a clear responsibility. Identify the indexed unit, embedding revision, similarity function, candidate-search algorithm, filters, fusion rule, reranker, and final context cutoff. These are separate choices, even when a vendor wraps them in a single query API.
Then preserve a small, representative evaluation set and record its limitations. Exact-neighbor agreement, human relevance, answer correctness, storage, and latency each describe a different outcome. A glossary supplies the language; repeatable evaluation supplies the reason to choose a configuration.
For help applying that process to a business workflow, explore our AI and digital transformation services. Start with evidence about the retrieval failure and operational requirement you need to resolve.