Glossary
Retrieval-augmented generation (RAG)
Retrieval-augmented generation (RAG) is a technique in which a language model answers a question using passages retrieved from an external collection at query time, in addition to what it learned during training. The retrieved text is inserted into the prompt, so answers can draw on current, private and citable sources.
Knowledge frozen at a training cutoff
A large language model knows what its training data contained, up to a cutoff date, and nothing held inside the organization. Enterprises want answers built on their own contracts, manuals, tickets, research and records, material that changes every day and carries access restrictions. Retraining on every document change is out of reach. RAG keeps the model as it is and changes what it reads at the moment of the question, which moves the hard part from model training into data infrastructure: answer quality now rests on how enterprise content is stored, prepared, indexed and kept current.
Retrieved knowledge has properties that trained knowledge lacks. It updates when a document is added, edited or deleted and the index catches up. Its answers trace back to specific passages. It can be filtered per user at retrieval time, where weights answer every user the same way. Changing a model's style or behaviour is a separate job, closer to fine-tuning than to retrieval.
RAG architecture in two paths
RAG architecture is the arrangement of components and data flows that turns the technique into a running service. It has two paths on different schedules: the offline path prepares content ahead of time, and the online path serves each query. They meet at the index and at the source store.
| Layer | Path | Job | Common components |
|---|---|---|---|
| Source store | Both | Holds original documents and media | Object storage, file shares, databases, SaaS systems |
| Ingestion | Offline | Turns sources into indexed chunks | Connectors, parsers, chunkers, embedding models |
| Index | Both | Similarity and keyword search with filters | Vector database, search engine, metadata store |
| Retrieval | Online | Finds and ranks candidate passages | Query encoder, search service, reranker |
| Orchestration | Online | Builds prompts, calls tools, tracks conversation state | Application framework or agent runtime |
| Generation | Online | Writes the answer | Language model on inference servers |
On the offline side, connectors detect new, changed and deleted items in each source along with their permissions. Parsers recover text and structure, with layout analysis for PDFs, OCR for scans and transcription for audio. The text is split by document chunking, tagged with metadata, encoded into embeddings and written to a vector database or search index.
On the online side, the question is encoded the same way, and an approximate nearest-neighbour search finds the closest chunks while examining only a small fraction of stored vectors, trading some recall for speed. Results are filtered by metadata, including the user's permissions, often merged with keyword matches through hybrid search, then rescored by a reranker that reads query and passage together. Orchestration assembles the prompt within a token budget: with 8,000 tokens reserved for context and chunks averaging 500 tokens, no more than 16 chunks fit.
Permissions follow one of two designs. Labels copied into chunk metadata at ingestion are fast to filter on, but only as correct as the last sync with the source. Checking each candidate against the source system's live permissions is exact and adds latency to every query.
Recall as the ceiling on answer quality
No answer is better than the passages it was given. When the passage holding the answer misses the shortlist, the model falls back on weaker passages or on its weights, a frequent cause of RAG hallucination. Variants of the basic pattern exist mostly to raise that ceiling. Hybrid retrieval catches exact names, part numbers and codes that vectors blur. GraphRAG searches a graph of entities and relationships extracted from the collection. Agentic RAG lets the model decide when and what to retrieve, issuing several queries and calling other tools before answering. Multimodal RAG indexes images, tables, audio and video next to text.
RAG evaluation measures whether those choices work, by running a fixed set of test questions against each configuration. Recall at k, the share of relevant passages that land in the top k results, scores retrieval and caps everything downstream. Faithfulness, the share of claims in an answer that the retrieved text supports, scores generation, and an answer can be faithful to an outdated passage and still be wrong. Reproducing a score later takes the same corpus revision, index build, prompts and judge model, so evaluation runs record the versions of the data they ran against.
Re-indexing, permission lag and agentic query load
The corpus dwarfs everything derived from it. Source documents, scans, recordings and their extracted text far outweigh the vector index, run to hundreds of terabytes or petabytes at enterprise scale, and are read in bulk every time the collection is reprocessed, which makes object storage their natural home. Index size can be worked out in advance: ten million chunks at 1,024 dimensions in 32-bit floats occupy 10,000,000 × 1,024 × 4 bytes, about 41 GB, before index structures, metadata and chunk text. Replacing the embedding model re-encodes every chunk, turning a model upgrade into a full corpus read and a full index rewrite while the old index keeps serving. RAG storage covers capacity and throughput for these layers.
Freshness and permission changes are storage-layer problems. Between a document changing and its re-indexing, answers can rest on stale content, and a file deleted or restricted at the source stays retrievable until the index reflects it. Change notifications, versioning and access metadata on the storage layer set how narrow those windows get.
Load climbs with adoption. A pilot for a few analysts becomes a service handling thousands of queries an hour, and agentic designs retrieve several times per answer, multiplying traffic on every online layer. Inference servers keep a key-value cache of attention state, and sharing it across servers for common prefixes, such as a system prompt or a frequently retrieved document, shortens time to first token. Multimodal collections grow the footprint again: originals plus transcripts, extracted frames, captions and embeddings, derived data that can approach the source in size and needs equal protection, since rebuilding it means reprocessing everything.
Scality ADI and retrieval-augmented generation
Two layers of the architecture above map to Scality ADI: the source store, through high-concurrency S3 and an S3 over RDMA data path, and the inference-state layer, through a centralized cache for distributed inference (KV cache). Among its target workloads Scality lists "Multimodal Agentic RAG, VSS, Deep Research Agents".
Index, reranker and orchestration run above the storage, so retrieval quality and permission filtering stay with the RAG application.














