Glossary
GraphRAG
GraphRAG is a retrieval-augmented generation method that uses a language model to turn a document collection into a knowledge graph of entities and relationships, groups closely connected entities into communities with written summaries, and retrieves from the graph and those summaries when answering questions.
The term also covers, more loosely, any RAG system that retrieves through a graph structure alongside or in place of a vector index of text chunks.
Why GraphRAG matters for enterprise questions
Chunk-based retrieval answers local questions well: the renewal date of a contract, the default value of a setting. The answer sits in a few passages that resemble the question. Two other kinds of question defeat it.
- Global questions concern the collection as a whole, such as the recurring causes across a year of incident reports. No single passage resembles the answer, and the material is spread over thousands of documents.
- Multi-hop questions connect facts held in different documents, such as which suppliers two product lines share. Each fact is retrievable on its own, but the link between them appears nowhere.
Similarity search explains the second failure. The question's vector lands near passages that resemble the whole question, while each intermediate fact sits in a passage that resembles only part of it, so those passages rank low. A graph records the links explicitly, so retrieval can follow them.
How a graph index is built
- Source documents are split into chunks.
- A language model reads every chunk and extracts entities (people, organizations, systems, places, concepts), the relationships between them, and short descriptions of each.
- Duplicate entities across chunks are merged, producing a graph of nodes and weighted edges.
- A community detection algorithm partitions the graph into a hierarchy of groups of closely connected entities.
- The model writes a summary of each community at each level of the hierarchy.
The graph is only as good as the extraction. The prompt sets which entity types count, the model decides what is a relationship, and merging by name can turn one company spelled two ways into two nodes, or two people with the same name into one. The output is a set of tables (entities, relationships, communities, summaries and the chunks behind each) kept alongside or instead of a vector index.
Local and global search
| Mode | Starting point | Suited to |
|---|---|---|
| Local search | Entities matching the question, then their neighbours, relationships and source chunks | Questions about specific entities and how they connect |
| Global search | Community summaries at a chosen level of the hierarchy | Questions about themes and patterns across the whole collection |
Global search works in two passes. Each community summary produces a partial answer on its own, and the partial answers are then merged into one response. Because the summaries were written at indexing time, the collection never has to fit into one context window.
Indexing cost and updates
Building a vector index takes one embedding pass. Building a graph index takes language model calls on every chunk and every community, so its cost scales with the corpus in generated and read tokens. A one-million-token corpus cut into 600-token chunks gives about 1,667 chunks; with 2,000 tokens of instructions per extraction call, extraction alone reads about 1,667 × 2,600 ≈ 4.3 million tokens. At an enterprise corpus of ten billion tokens the same ratio means roughly 43 billion tokens of model input before summaries are written.
Updates are heavier too. A new document can add entities and edges that shift community membership, which forces the affected summaries to be regenerated, so graph indexes are usually refreshed in batches.
What GraphRAG means for AI infrastructure teams
A graph index build is an inference job running on GPUs, closer in scale to a batch analytics workload than to an embedding pass. That pushes most teams to scope GraphRAG to the parts of the corpus where cross-document questions carry real value (contracts, incident histories, research archives) and keep plain hybrid search for the rest.
Each build also leaves a set of artefacts: entity and relationship tables, community structures and summaries, often written as Parquet files. Keeping each run's output as its own versioned set lets a team compare graphs built with different prompts or models and roll back a bad build. The graph remains derived data; the corpus is the source of truth it is regenerated from.
Deletion and access control are harder than with vectors. A community summary blends content from many documents, so removing one source can mean regenerating every summary it fed. The same blending crosses permission boundaries: a global answer built from summaries can carry facts from documents the user cannot open. Graphs built per permission domain, or summaries filtered by the access lists of their source chunks, are the usual responses, and both depend on chunk metadata that traces every entity back to a specific source object and version.
GraphRAG and Scality RING
Scality RING stores the source corpus and the tables produced by each indexing run as S3 objects on standard x86 servers, up to 300 billion objects in a single RING. Erasure coding schemes in RING are defined per storage class, which lets the irreplaceable corpus and the regenerable graph artefacts sit under different protection schemes in one system.














