Glossary

Embeddings

Embeddings are lists of numbers, produced by a trained model, that represent a piece of content (a sentence, a passage, an image, a block of code) as a point in a vector space. Content with related meaning ends up close together, which lets software compare meaning by measuring distance.

They are the representation behind semantic search, recommendation systems and the retrieval step of retrieval-augmented generation.

Why embeddings matter for AI data infrastructure

Embeddings are how unstructured content becomes searchable by AI systems. Every document that enters an AI knowledge base is cut into passages and each passage is embedded, so the embedding set becomes a derived dataset that grows in step with the corpus. Unlike most derived data, it is also disposable on a schedule: whenever a better model arrives, every vector is regenerated from the source. For infrastructure teams that makes embeddings a recurring compute job, a capacity line in the most expensive storage tier, and a dependency on keeping the original content readable.

How embedding models produce vectors

Current text embedding models are neural networks that read a passage and output one vector for it. They are trained on pairs of texts that belong together, such as a question and the passage that answers it, or a title and its article. During training the model learns to place each pair close together and to push unrelated texts apart. Training with "hard negatives", passages that look relevant but are not, teaches it finer distinctions.

Because each text is encoded on its own, a passage is embedded once and its vector can be compared with any number of later queries. That reuse is what makes search over hundreds of millions of passages affordable: the expensive model runs at ingestion, and query time needs only one embedding plus vector arithmetic.

Dimensions, precision and size

An embedding's size is its number of dimensions multiplied by the bytes per value.

Dimensions32-bit float8-bit integer1-bit binary
3841,536 bytes384 bytes48 bytes
7683,072 bytes768 bytes96 bytes
1,0244,096 bytes1,024 bytes128 bytes
3,07212,288 bytes3,072 bytes384 bytes

Lower precision trades some retrieval accuracy for size: 8-bit quantization is a quarter of the float size and binary a thirty-second. Some models are trained so that the first few hundred dimensions of each vector work as a shorter embedding on their own, which allows a coarse, cheap first search followed by a precise one on fewer candidates.

Types of embeddings

  • Passage embeddings: one vector per chunk of text, the standard unit for RAG.
  • Multilingual embeddings: texts with the same meaning in different languages land close together.
  • Multimodal embeddings: images, video frames or audio share a space with text, so a caption sits near its image.
  • Code embeddings: trained on source code and documentation for code search.
  • Sparse embeddings: long vectors with few non-zero weights tied to vocabulary terms, searchable with an inverted index.
  • Multi-vector embeddings: one vector per token, compared token by token, more accurate and many times larger.

What embeddings mean for storage and AI platform teams

Capacity planning starts from chunk count. Half a billion chunks at 1,024 dimensions in 32-bit floats is 500,000,000 × 4,096 bytes, about 2 TB of raw vectors, before the index structures that make them searchable. At 8-bit precision the same set is about 512 GB. Because the index usually sits in memory or on flash, the choice of dimensions and precision is a hardware budget decision as much as a quality one, and it is worth measuring on the organization's own queries before it is fixed.

The model is the schema. Vectors from two models, or two versions of one model, are not comparable even when their dimensions match, so changing models means re-embedding everything. At 2,000 passages per second, half a billion chunks take 250,000 seconds, close to 70 hours of embedding compute, plus the time to read the source text and rebuild the index. During the switch the old index keeps serving traffic, so for a period two complete vector sets exist side by side. Recording the model identifier with every vector is what prevents a mixed index.

Embeddings also inherit the sensitivity of the text they encode. They are derived from the source content and sit in the same governance scope: access rules, residency requirements and deletion requests that apply to a document apply to its vectors too.

Embeddings and Scality storage

An embedding is always derived from a source item, and every re-embedding pass reads the corpus again from start to finish. Scality RING holds that corpus as S3 objects on standard x86 servers, so the original documents, images and transcripts stay available to each new model, while the vectors are stored and searched in a vector database. Erasure coding schemes in RING are defined per storage class, so source content and exported vector snapshots can carry different protection settings.