Glossary

RAG architecture

RAG architecture is the arrangement of components and data flows in a retrieval-augmented generation system. An ingestion pipeline turns source content into a searchable index, a retrieval layer queries it, and a generation layer composes answers from what comes back.

Why RAG architecture matters beyond the model

Prototype RAG systems fit on a laptop: a folder of PDFs, an embedding model, an in-memory index and a hosted language model. Production systems serving an enterprise look very different. They ingest millions of documents from dozens of sources, enforce each source's permissions, re-index continuously, serve many concurrent users and keep logs for audit. Most of the engineering effort, and most of the failures, sit in the data layers around the model. The technique itself is described under retrieval-augmented generation; the architecture is what makes it work at scale.

Offline and online paths

A RAG system runs two paths on different schedules. The offline path prepares content ahead of time; the online path serves each query. They meet at the index and the source store.

LayerPathFunctionTypical components
Source storeBothHolds original documents and mediaObject storage, file shares, databases, SaaS systems
IngestionOfflineTurns sources into indexed chunksConnectors, parsers, chunkers, embedding models
IndexBothSimilarity and keyword search with filtersVector database, search engine, metadata store
RetrievalOnlineFinds and ranks candidate passagesQuery encoder, search service, reranker
OrchestrationOnlineBuilds prompts, calls tools, manages conversation stateApplication framework or agent runtime
GenerationOnlineProduces the answerLanguage model on inference servers

Ingestion, retrieval and generation

Ingestion connects to source systems and detects new, changed and deleted items along with their permissions. Parsers extract text and structure from each format, including layout analysis for PDFs, OCR for scans and transcription for audio. Text is split by document chunking, tagged with metadata and encoded into vectors. Deletions and permission changes are propagated so removed content stops being retrievable. It is a specific form of AI data pipeline.

Retrieval uses approximate nearest-neighbour indexes that examine a small fraction of stored vectors, trading a little recall for speed. Results are filtered by metadata, including the user's permissions, often merged with keyword results in hybrid search, and reranked by a model that reads query and passage together.

Generation works within a token budget. If 8,000 tokens are reserved for context and chunks average 500 tokens, at most 16 chunks fit. Inference servers keep a key-value (KV) cache of attention state for tokens already processed, and reusing it for shared prefixes, such as a common system prompt or a frequently retrieved document, shortens time to first token. In agentic designs the orchestrator lets the model retrieve several times before answering, which multiplies the load on every online layer.

Access control follows one of two designs. In the first, permission labels are copied into chunk metadata at ingestion and applied as a filter at query time, so correctness depends on how quickly permission changes in the source propagate. In the second, candidates are checked against the source system's live permissions before use, which is exact but adds latency to every query. Freshness works the same way: the interval between a document changing and its re-indexing is the window in which answers can draw on outdated content.

What RAG architecture means for AI storage design

Each layer holds a different dataset with a different access pattern, and treating them as one storage problem leads to either overspending or bottlenecks.

  • Source content is the largest dataset, often petabytes, read in bulk during ingestion and individually when a citation is opened. Capacity, durability and parallel read throughput matter most.
  • Intermediate artefacts, the extracted text and chunks, are written once per ingestion run and kept so the index can be rebuilt without re-parsing every source.
  • The index is read randomly at low latency on every query and usually lives in memory or on flash.
  • Inference state in the KV cache sits in GPU memory first and, in larger deployments, overflows to shared storage, where reuse across servers saves recomputation.
  • Logs and evaluation data, the queries, retrieved passages and answers, accumulate steadily and are kept for audit and for RAG evaluation.

Re-indexing is the event that stresses the design. A new embedding model or chunking strategy means reading the whole corpus, re-encoding it and rewriting the index, while the old index keeps serving queries. At petabyte scale that is a sustained, high-throughput job, and the storage beneath the source and intermediate layers determines whether it takes days or weeks. Capacity and performance for these layers are treated under RAG storage.

Scality ADI in a RAG architecture

Scality ADI addresses the source store and inference-state layers. Scality lists a "Centralized Cache for Distributed Inference (KV cache)" among ADI's capabilities, alongside high-concurrency S3 and S3 over RDMA data paths, and its MultiScale design lets capacity and throughput grow independently of each other.