Glossary

RAG evaluation

RAG evaluation is the measurement of how well a retrieval-augmented generation system finds the right information and turns it into correct answers supported by that information. It scores the retrieval step, the generated answer and the system end to end, usually by running a fixed set of test questions against each configuration.

Why RAG evaluation matters for enterprise AI

Every component of a RAG system gets replaced over its life: the embedding model, the chunk size, the reranker, the prompt, the language model itself. Each change improves some answers and quietly breaks others. Without a repeatable measurement, a team has no way to tell whether a change helped, and regressions surface as complaints from users weeks later.

Evaluation also justifies spending. Re-embedding a corpus of hundreds of millions of chunks costs days of GPU time. A measured gain in retrieval quality on representative questions is the evidence that the cost is worth it, and a flat result is the evidence that it is not.

What gets measured

A wrong answer can start at any stage, so evaluation is organized by stage to locate the failure.

StageWhat is measuredTypical failure
IngestionWhether content was parsed and chunked intactTables flattened, text lost in OCR, passages cut mid-sentence
RetrievalWhether the relevant passages reached the top resultsThe passage with the answer is missing from the top k
RerankingWhether the best passages were placed firstThe right passage retrieved but cut off below the limit
GenerationWhether the answer sticks to the retrieved textClaims the context does not support
End to endWhether the final answer is correct and usefulA confident answer to a slightly different question

For retrieval, the most watched measure is recall at k: the share of relevant passages that appear in the top k results. The model can only use what reaches its context, so recall sets the ceiling for everything after it. Precision and ranking measures matter more when k is small. For generation, the core measures are faithfulness (the share of claims in the answer that the retrieved text supports), answer relevance, correctness against a reference answer, and citation accuracy. Faithfulness and correctness are independent: an answer can be faithful to an outdated passage, and correct while ignoring the context entirely. Grounding failures are covered under RAG hallucination.

Test sets and model judges

A test set pairs questions with what is needed to score them: the IDs of relevant passages, reference answers, or both. Questions come from subject experts, from real user queries, or from a language model that writes questions from passages in the corpus. Generated questions are cheap but tend to reuse the passage's own wording, which flatters keyword retrieval and understates how hard real questions are.

Labelling by hand is slow, so many teams use a language model as the judge. A model judge has its own errors and biases, including a preference for longer answers, so its scores are checked against human labels on a sample, and the judge model and prompt are recorded with every score so runs stay comparable.

Offline and online evaluation

Offline evaluation runs the fixed test set against a configuration and reports the metrics. It is repeatable and is how configurations are compared before release. Online evaluation watches the live system: explicit feedback, users rephrasing the same question, escalations to a human, clicks on citations. Operational measures sit alongside quality: end-to-end latency at the median and 95th percentile, retrieval latency, tokens per answer and cost per query. Online signals reflect the real mix of questions, which any offline set only approximates.

What RAG evaluation means for AI platform teams

Reproducing an evaluation score is a data management problem as much as a machine learning one. A score is only meaningful against the exact inputs that produced it: the corpus snapshot, the chunk set, the index, the test set, the judge prompt. Relevance labels point at specific chunk IDs, so re-chunking the corpus invalidates them unless they are remapped. Teams that cannot recover the corpus as it stood on the day of a run cannot explain why a score moved.

Full copies are rarely affordable at this scale. Copying a 2 PB corpus for every experiment doubles capacity each time, while object versioning keeps only what changed and lets a run refer to a precise revision. Each run also produces its own outputs (retrieved IDs, answers, scores), and keeping them is what turns a regression from a mystery into a diff.

Online evaluation draws on query logs, which hold real user questions and often personal data. Those logs inherit retention and access rules, and they grow with traffic. Storage changes show up in the operational metrics too: moving chunks or source objects to a slower tier moves retrieval latency, which the same evaluation harness will measure.

Scality and RAG evaluation

Scality RING stores test sets, corpus snapshots and run outputs as S3 objects on standard x86 servers. With versioning enabled, each object keeps its prior versions, so a labelled test set can be matched to the corpus revision it was written against without duplicating the corpus.

Where evaluation records have to stay unaltered for an audit period, RING applies S3 Object Lock retention. Compliance mode prevents deletion by any user, including the account root, until the retain-until date; governance mode can be overridden by an identity holding s3:BypassGovernanceRetention that sends x-amz-bypass-governance-retention:true. After the retain-until date the locked versions can be deleted like any others.