Glossary

Semantic search

Semantic search finds content by meaning instead of exact wording. A trained model converts documents and queries into vectors, called embeddings, and the search returns the content whose vectors lie closest to the query's, so a document can match a question it shares no words with.

Why semantic search matters for enterprise data

Large organizations describe the same thing in many ways. Engineering writes "node eviction", support writes "server dropped out of the cluster", and a customer writes "one of the boxes went offline". Keyword search treats those as unrelated, and synonym lists only patch the gaps someone thought to list. Semantic search compares meaning directly, across departments, writing styles and, with multilingual models, languages. It is also the default retriever in retrieval-augmented generation, so its recall sets a ceiling on how good an AI assistant's answers can be.

How semantic search works

  1. Documents are split into passages by document chunking.
  2. An embedding model turns each passage into a fixed-length vector.
  3. The vectors are stored in an approximate nearest-neighbour index, usually in a vector database.
  4. At query time the same model turns the query into a vector.
  5. The index returns the passages whose vectors are closest to it.
  6. Optionally, a slower reranking model reorders the top candidates.

Steps 1 to 3 run once per passage at indexing time. Steps 4 to 6 run for every query and set search latency.

Retrieval and reranking models

Two model arrangements are used, with very different costs. A bi-encoder encodes each text on its own, so passage vectors are computed once and compared with any number of later queries by simple arithmetic. A cross-encoder reads the query and a passage together and scores the pair, which is more accurate but has to run once for every pair. Comparing every pair in a collection of 10,000 passages means 10,000 × 9,999 ÷ 2, about 50 million model runs; at enterprise scale that is out of reach. Production systems therefore use a bi-encoder to search the whole collection and, where accuracy justifies the added latency, a cross-encoder to rerank the top few dozen results.

Strengths and limits

Handles wellHandles poorly
Paraphrase and synonymsExact identifiers: part numbers, error codes, names
Natural-language questionsVocabulary unlike the model's training data
Cross-language matching (multilingual models)Long passages squeezed into one vector
Images, audio and text in one shared space (multimodal models)Fixed relevance thresholds, since scores differ by model

Similarity scores are relative. A score of 0.8 from one model means something different from 0.8 from another, so semantic search returns the top results by rank, and any cut-off is calibrated per model. The weakness on identifiers is the main reason many systems pair semantic search with keyword search in hybrid search.

What semantic search means at enterprise scale

The index is the expensive tier. Its size follows the number of passages, and it is usually held in memory or on fast flash to keep query latency low. Growth in the corpus, and especially the move from text to images, video frames and audio transcripts, translates directly into index memory, which is why compression of vectors and careful choice of what to index become budget decisions as much as quality decisions.

The corpus stays in the picture after indexing. Every result points back to a source object, which is fetched to show a citation or a preview, so source storage sits on the query path as well as the ingestion path. And every vector is tied to the model that produced it: adopting a better embedding model means re-reading the corpus and re-embedding all of it, with the old index serving traffic until the new one is complete.

Query load scales separately from corpus size. An assistant rolled out to tens of thousands of employees, or a fleet of agents issuing several searches per task, turns the index into a shared service with its own capacity plan: replicas sized for query throughput, and an ingestion schedule that makes new documents searchable within hours of their creation. A stale index produces answers that are confidently out of date.

Permissions travel with the data only when they are copied. Semantic search will find the most relevant passage in the collection regardless of who wrote it or who may read it, so access lists stored with each passage and applied as filters at query time are what keep a company-wide search from becoming a company-wide leak.

Scality and semantic search

Scality RING stores the source objects that semantic search indexes and cites, over S3 on standard x86 servers, with a ceiling of 300 billion objects in one RING. The vectors themselves are indexed in a separate vector database. Scality's ADI page lists "Indexed and vectorized enterprise data with GPU-direct access" among the capabilities of its Enterprise AI for Core and Edge use case.