Glossary

Agentic RAG

Agentic RAG is a form of retrieval-augmented generation in which the language model acts as an agent: it decides whether to search, which sources to query and how to phrase each query, runs those searches in a loop, and checks what it found before answering.

A standard RAG pipeline runs one fixed search per question. Agentic RAG turns retrieval into a series of decisions taken at run time.

Why agentic RAG matters at scale

Enterprise questions rarely live in one place. Working out why a customer's renewal stalled can take the contract from a document system, the support history from a ticketing tool, and usage figures from a database. A fixed pipeline embeds the question, searches one index and hands back the closest passages, which works when the answer sits in one or two passages and fails when it has to be assembled. Agentic RAG handles the assembly, and in doing so it changes the load on every system it touches: one user question becomes many searches, across many sources, in sequence.

How the agent loop works

  1. The model reads the request and, for compound questions, splits it into sub-questions.
  2. It picks a tool: a vector index, a keyword index, a database query, an internal API or a web search.
  3. It writes a query for that tool, often worded differently from the user's question.
  4. The tool runs and returns results.
  5. The model judges whether the results are relevant and sufficient.
  6. It repeats from step 2 with a revised query or a different tool, or moves on.
  7. It writes the answer from the evidence gathered, with citations.

The loop ends when the model judges it has enough, or when it hits a limit on steps or tokens.

Common agentic retrieval patterns

PatternWhat the agent does
Query rewritingRephrases or expands the question before searching, for instance spelling out acronyms or adding product names
DecompositionSplits a multi-part question into sub-questions, retrieves for each, and combines the results
RoutingChooses among sources by question type: policy documents, tickets, code, structured data
Multi-hop retrievalUses the result of one lookup as the query for the next
Self-checkingGrades its own retrieved passages and searches again when they are weak
Multi-agent retrievalGives each source or sub-task to a specialist agent and merges their findings

Cost, latency and failure modes

Every turn of the loop adds a model call and a tool call, and the context grows as evidence accumulates. With 1.5 seconds per model call and 0.2 seconds per retrieval, a single-pass pipeline answers in 1.7 seconds and a four-step agent run in 6.8 seconds plus a final generation. Tokens grow faster than steps: if each step adds 2,000 tokens and the full history is resent, four steps consume 2,000 + 4,000 + 6,000 + 8,000 = 20,000 input tokens, ten times a single pass.

Agent loops also fail in ways fixed pipelines do not. They can keep searching without converging. An early wrong conclusion steers every later query. Text inside a retrieved document phrased as an instruction can redirect the agent, a risk that grows with the number of sources it reads. And the agent acts with the credentials of its tools, so it retrieves whatever those credentials can reach unless permissions are enforced per user. Evaluation therefore covers the path the agent took as well as its answer, as described under RAG evaluation.

What agentic RAG means for data infrastructure

The first consequence is read amplification. A chat assistant that triggers one retrieval per question becomes, once agentic, a system that triggers five to ten, each fetching several chunks and often the source objects behind them. Multiply that by a fleet of agents running unattended research or monitoring tasks and the corpus store sees a sustained stream of small, concurrent reads that a pilot never produced. Because the steps run in sequence, storage latency adds up across the run: 50 ms of slow reads per step becomes half a second over ten steps.

The second is identity. An agent with a broad service account can read anything that account can, so a question from a junior employee can surface a board document. Passing the requesting user's identity through to the storage and index layers, and enforcing access there, keeps the agent inside the same boundaries as the person who asked.

The third is record keeping. Auditing an agent means storing its trajectory: every query, tool call, retrieved passage and intermediate judgement. Those logs run many times larger than chat transcripts. Agents that write back, saving notes or summaries for later reuse, add model-generated content to the corpus, and that content needs provenance metadata so it can be told apart from human-authored sources, filtered or removed.

Agentic RAG and Scality ADI

Scality's ADI page lists "Multimodal Agentic RAG, VSS, Deep Research Agents" among the capabilities of its Enterprise AI for Core and Edge use case, in which ADI manages bi-directional AI data flows between core and edge environments.

The corpus such agents draw on can be held in Scality RING as S3 objects on standard x86 servers. For the sequential small reads an agent loop produces, Scality publishes measured RING latencies of 511 µs for a GET and 741 µs for a PUT, figures that belong to the configuration on which they were measured.