Glossary

Document chunking

Document chunking is the process of splitting documents into smaller passages, called chunks, before they are embedded and indexed for retrieval. Each chunk becomes one unit a search can return, so the chunking scheme decides what a RAG system can find, what it passes to the language model and how precisely it can cite a source.

Why chunking matters for RAG at scale

Four constraints make chunking unavoidable. Embedding models accept a limited number of tokens, and text past the limit is dropped. One vector summarizes its whole input, so a vector for a 40-page document averages many topics and matches none of them well. The language model's context holds a limited amount of text, and every irrelevant sentence in a retrieved chunk crowds out a relevant one. And an answer can only cite as precisely as the chunk it came from.

At enterprise scale chunking is also a multiplier. The number of chunks sets the number of vectors, the size of the index and the cost of every future re-embedding pass, so a decision made in a notebook during a pilot carries through to the hardware bill.

Chunking methods

MethodSplit ruleCharacteristics
Fixed-sizeEvery N tokens or charactersSimple and predictable; cuts through sentences and sections
RecursiveParagraphs first, then lines, then sentences, until each piece fitsKeeps natural units where possible; a common default
Structure-awareHeadings, list items, tables, HTML or PDF layout blocksFollows how the author organized the content
SemanticWhere the topic shifts between adjacent sentencesTopic-aligned boundaries; adds an embedding pass at ingestion
Content-specificFunctions in code, clauses in contracts, row groups in tablesMatches the logical unit of each document type

Large collections mix document types, so production pipelines usually route each type to its own chunker.

Chunk size and overlap

Chunk size trades precision against completeness. Small chunks match specific questions well but can separate a statement from the context needed to read it. Large chunks keep context together but dilute the vector and use more of the prompt. Many systems settle between 200 and 1,000 tokens.

Overlap repeats the end of each chunk at the start of the next so that a sentence on a boundary appears whole at least once. It has a direct cost. A 10,000-token document cut into 500-token chunks gives 20 chunks with no overlap; with 100 tokens of overlap the step between chunks falls to 400 tokens and the count rises to 25. That is 25% more vectors, index memory and embedding compute across the entire corpus.

Parsing and context enrichment

Chunking works on extracted text and can be no better than the extraction. PDFs store positioned characters instead of paragraphs, so reading order across columns has to be rebuilt. Headers and page numbers interrupt sentences. Tables lose their structure when flattened, which is why tables split across chunks commonly repeat their header row in each. Scanned pages need OCR before any text exists.

A chunk removed from its document also loses what the document supplied: which customer, which product, which year. Enrichment adds it back, by prefixing the title and heading path to each chunk, by having a model write a one-line description of where the chunk sits, or by indexing small chunks for matching and returning the larger section around a match to the model.

What document chunking means for AI infrastructure teams

Chunking is rarely decided once. Teams revise it as evaluation exposes weak spots, and each revision means re-reading the extracted text and re-embedding every chunk. Keeping the extracted text as a stored intermediate layer means a chunking change never forces a re-parse of the source, which matters when the source is millions of scanned pages and OCR is the slowest step in the pipeline.

Chunk metadata is what makes a large index governable. Each chunk carries the source object key and version, its offsets in the document, its heading path, the document date, the access list inherited from the source and the embedding model used. Those fields support citation, permission filtering at query time, and the removal of every chunk derived from a document when it is updated or deleted. Without them, a superseded policy stays answerable long after it was replaced, as described under RAG hallucination.

The sizing consequences compound. Chunk size, overlap and enrichment together set chunk count, and chunk count sets the vector database footprint and the duration of every reindexing job. A change from 1,000-token to 250-token chunks roughly quadruples all three.

Chunking and source objects in RING

Chunks are rebuilt from source documents whenever the parser, chunking scheme or embedding model changes, so the originals stay in place long after ingestion. Scality RING stores them as S3 objects with versioning, and each chunk can record the version ID of the object it was cut from. Where specific source versions have to be held unchanged, RING applies S3 Object Lock: in compliance mode no user, including the account root, can delete a locked version before its retain-until date, while governance mode can be bypassed by an identity holding s3:BypassGovernanceRetention.