Glossary

RAG storage

RAG storage is the storage that holds everything a retrieval-augmented generation (RAG) system reads and produces: the source documents, the text and chunks extracted from them, the embeddings and search indexes built from those chunks, and the records of questions and answers. These layers differ in size, speed and value, and in most deployments they sit on different systems.

Why RAG storage matters at enterprise scale

A RAG pilot built on a few thousand documents fits comfortably on a single server. Production is a different problem. The corpus grows to tens of millions of documents drawn from many departments, it is kept for years, and it is reprocessed every time the team adopts a new parser, chunking scheme or embedding model. Storage decisions made at the pilot stage end up deciding how quickly the system can change, whether every answer can be traced back to a source, and what it costs to keep all of it.

The layers also have very different value. The source corpus is the only part that cannot be regenerated. Everything else is derived from it by repeatable processing, so it can be rebuilt at a price in compute and time.

The layers of a RAG system

LayerWhat it holdsRebuildableTypical home
Source corpusOriginal PDFs, office files, web pages, images, transcripts, exported recordsNoObject storage, file shares, data lake
Extracted textText and tables recovered by parsers and OCRYes, by re-parsingObject storage, as JSON or Parquet
Chunks and metadataPassages with source reference, version, permissions and datesYes, by re-chunkingObject storage or a database
Vector indexEmbeddings plus a nearest-neighbour search structureYes, by re-embeddingA vector database in memory or on flash
Keyword indexInverted index of termsYesA search engine
Query and answer recordsPrompts, retrieved chunk IDs, answers, feedbackNoObject storage or a log platform

Two layers are irreplaceable: the corpus, and the record of what the system told people. The rest is a cache of processing work.

Capacity and reprocessing

The relative sizes follow from a few parameters. Ten million documents averaging 200 KB make a 2 TB corpus. If they yield 50 million chunks, and each chunk becomes a 1,024-dimension vector of 4-byte values, the raw vectors take 50,000,000 × 4,096 bytes, about 205 GB. The vectors are a tenth of the corpus, but they live on the most expensive tier, memory or flash, while the corpus can sit on dense, lower-cost capacity. Corpora heavy in scanned pages, images or video tilt the ratio further, because most of their bytes carry no text.

Reprocessing reads the layers in order. A new parser or OCR engine reads the entire corpus again. A new chunking scheme re-reads the extracted text and re-embeds every chunk. A new embedding model re-embeds everything, because vectors from different models cannot be compared. Read throughput from the corpus store sets the floor on how long any of these takes.

Access patterns at query time

A question passes through the layers in sequence: the query is embedded, the vector index is searched, the matching chunks are fetched by ID, and the source object is often fetched as well so the answer can show a citation. Chunk and source fetches are small random reads, so per-request storage latency counts for more than throughput. When an answer waits for ten parallel reads, the slowest of the ten sets the pace, which makes the tail of the latency distribution more important than its median.

What RAG storage means for AI infrastructure teams

In practice RAG storage splits into two tiers with opposite economics. A small, hot tier holds the indexes and is sized for query rate and latency. A large, durable tier holds the corpus, the extracted text and the logs, and it grows every time another department or data source is onboarded. Sizing the durable tier only for capacity leaves out the reprocessing case: a 2 PB corpus read at 10 GB/s takes about 56 hours to scan once, so the throughput of that tier decides whether a model upgrade is a weekend job or a month-long project.

Governance spans every layer. Personal data in one source document also exists in its extracted text, its chunks, its vectors and the logs of answers that quoted it, so an erasure request under GDPR Article 17 reaches all of them. The same applies to access control: a chunk cut from a restricted document carries the restricted content but none of the original permissions unless they are copied into its metadata. Chunks that record the source object's key and version ID can be traced to the exact revision they came from, which is what makes deletion, audit and citation workable at scale.

Placement matters for multi-site and sovereign deployments. Derived copies are still copies of the data, so residency rules that apply to the corpus apply to the chunks, vectors and logs built from it.

RAG storage on Scality RING and ADI

Scality RING is software-defined object and file storage on standard x86 servers, scaling to 300 billion objects in a single RING. In a RAG system it holds the layers stored as objects: the corpus, extracted text, chunk files and query records, all over S3. Erasure coding schemes in RING are defined per storage class, so a corpus and a log archive can sit under different schemes. RING supports S3 Object Lock on versioned buckets: in compliance mode no user, including the account root, can delete a locked version before its retain-until date, while governance mode can be bypassed by an identity holding s3:BypassGovernanceRetention.

The vector index itself runs in a vector database alongside. Scality's ADI is positioned for "petabyte to exabyte AI workloads", the range where the durable tier of a RAG system ends up.