Glossary
Hybrid search
Hybrid search is a retrieval method that runs a keyword search and a vector search over the same content and merges their two ranked result lists into one. Keyword search matches the exact words in a query, semantic search matches its meaning, and each finds relevant results the other misses.
Why hybrid search matters for enterprise retrieval
Enterprise content is full of strings that carry meaning only as exact matches: part numbers, error codes, contract references, SKUs, function names, people's names. Vector search handles paraphrase well and these strings badly, because an embedding model places similar-looking identifiers close together whether or not they refer to the same thing. Keyword search has the opposite profile. In a RAG system a passage that neither method returns is a passage the model never sees, so the gap between the two shows up directly as wrong or missing answers.
Keyword and vector retrieval
Keyword retrieval uses an inverted index, a map from each term to the documents that contain it. The standard scoring function, BM25, ranks documents higher when a query term appears often in them, when the term is rare across the collection, and when the document is not padded with unrelated text. It finds "ERR-4402" in a support ticket every time.
Vector retrieval converts queries and passages into embeddings and ranks by vector similarity, typically in a vector database. It matches "staff leaving" to a passage on "employee attrition" with no shared words, and it is weaker on identifiers and on vocabulary the model rarely saw in training.
Merging the results
BM25 scores and vector similarities sit on unrelated scales, so they cannot simply be added. Two methods are common.
Reciprocal rank fusion ignores the scores and uses only positions. Each document earns 1 ÷ (60 + rank) from each list, and the totals are summed. A document ranked first by keyword search and fifth by vector search scores 1 ÷ 61 + 1 ÷ 65 ≈ 0.0318; one ranked second by both scores 1 ÷ 62 + 1 ÷ 62 ≈ 0.0323 and comes out on top. Consistent agreement beats a single first place.
Weighted score fusion rescales each list's scores to a common range and combines them with a tuned weight. It keeps information about how far apart results are, at the cost of sensitivity to how the scores are rescaled.
| Property | Reciprocal rank fusion | Weighted score fusion |
|---|---|---|
| Uses | Rank positions only | Rescaled scores |
| Tuning | One constant | A weight and a rescaling method |
| Sensitivity to score distributions | None | High |
Reranking and filtering
Fusion is often followed by reranking. A reranking model reads the query and each candidate passage together and scores their relevance. It is more accurate than either first-stage method and far too slow to run across a whole collection, so it is applied only to the fused top candidates, for instance cutting 50 down to 5.
Metadata filters (date range, document type, business unit, access permissions) apply to both searches. A filter applied to one side and not the other lets the unfiltered side return documents the first excludes, which is one route to an assistant quoting content a user is not entitled to read.
What hybrid search means for search and storage teams
Running hybrid search means running two indexes built from one corpus, and keeping them consistent is the operational burden. Every new, changed or deleted document has to reach both. A document deleted from the vector index but still in the keyword index keeps turning up in answers, and the reverse produces results that differ depending on how a question is phrased. Pipelines that update both indexes from the same change feed, keyed on the source object and its version, avoid that drift.
Capacity and cost split the same way. The vector index is the expensive part: it is usually held in memory or on flash and grows with the number of chunks and the vector size. The inverted index is typically smaller and sits comfortably on disk. When the chunking scheme or embedding model changes, both are rebuilt from the corpus, which is why the corpus itself stays on durable storage that can be read end to end at speed.
Latency adds up across stages: the slower of the two searches, then fusion, then reranking, then fetching the chosen chunks and any source objects shown as citations. Teams tuning a hybrid system tend to find that the reranker and the final fetches, more than either index, set the response time users experience.
Hybrid search and the storage layer
Both indexes in a hybrid system are derived from one corpus and rebuilt from it when chunking or models change. Scality RING holds that corpus as S3 objects on standard x86 servers, scaling to 300 billion objects in a single RING, with versioning available so every chunk in both indexes can name the exact object version it came from.














