Glossary

Large language model (LLM)

A large language model (LLM) is a neural network trained on very large amounts of text to predict the next token, a word or fragment of a word, in a sequence. Answering questions, summarizing documents and writing code all come from repeating that one prediction many times.

For the teams that build and run AI platforms, an LLM is also a set of very large files and a set of data flows: the corpus it learns from, the checkpoints saved while it trains, the weights it is served from and the documents it retrieves at query time.

Why large language models matter for AI infrastructure

Most discussion of LLMs centres on GPUs, yet much of the operational effort around them is data handling. A training cluster sits idle whenever it is waiting for data to load or for a checkpoint to finish writing, and GPU time is the most expensive resource in the building. Once a model is in production, the load shifts to inference: weights loaded onto many servers at once, retrieval collections queried continuously, and caches holding the working state of thousands of conversations.

Many organizations also want to run models on their own documents, inside their own data centres or jurisdiction, which places the corpus, the model versions and the retrieval data on infrastructure the platform team operates.

How an LLM is trained and served

Text is split into tokens, and each token is turned into a vector of numbers. Almost all current LLMs use the transformer architecture, a stack of layers in which every token is compared with the tokens before it (attention) and then transformed by a small network. The learned numbers in those layers, the parameters or weights, range from a few billion to several hundred billion in current models.

  • Pretraining adjusts the weights over a very large text corpus until predictions improve. It consumes nearly all of the compute.
  • Fine-tuning continues training on curated instructions and answers, or on an organization's own domain data.
  • Inference serves requests. The prompt is processed in one pass, then the answer is generated one token at a time, with intermediate results held in a key-value (KV) cache so they are not recomputed at each step.

An LLM's knowledge is fixed at the end of training. Retrieval-augmented generation supplies current documents at query time, usually found through a vector database of embeddings, so answers can rest on sources the organization controls.

The data an LLM depends on

DataWhen it is usedHow it behaves on storage
Training corpusPretraining and fine-tuningVery large, read repeatedly and in parallel by every GPU node; throughput bound
CheckpointsSaved at intervals during trainingLarge bursts written by many GPUs at once; read back after a failure
Model weights and versionsDeployment and rollbackRead by many inference servers simultaneously; older versions retained
Retrieval collectionsEvery query in a RAG systemDocuments, embeddings and indexes updated continuously, read at low latency
KV cacheDuring inferencePer-request working state that outgrows GPU memory under long contexts

Size follows directly from parameter count. A 70-billion-parameter model stored at 16 bits per weight is 70 × 2 = 140 GB. Training state is far larger, because each parameter carries its weight, gradient and optimizer values: at a commonly cited 16 bytes per parameter, one checkpoint of the same model is about 1.12 TB.

What LLMs mean for AI infrastructure teams

Checkpoint time is GPU time. Writing 1.12 TB at 10 GB/s takes about 112 seconds, during which the training job typically pauses; at 50 GB/s it takes about 22 seconds. Over a run that checkpoints every hour for several weeks, the difference adds up to days of idle accelerators, so storage write throughput shows up directly in the cost of training.

Capacity grows in layers that are easy to underestimate. A team that keeps twenty checkpoints of one run holds over 22 TB for that run alone, before counting the corpus, the cleaned and tokenized copies of it, the fine-tuning sets and every model version kept for rollback or audit. Most of that data is hot for a short period and cold afterwards, which makes placement across media a recurring cost decision rather than a one-time sizing exercise.

Inference creates a different pattern. Scaling a service from ten to a hundred replicas means a hundred servers reading the same 140 GB of weights at the same moment, and a retrieval layer that has to stay current as source documents change. Where models run on proprietary or regulated documents, those documents, their embeddings and the prompts that reference them stay inside the organization's control, which ties the AI platform to the same sovereignty and retention rules as the rest of the data estate. For a lean platform team, the practical consequence is that training, inference and retrieval all draw on one data foundation, and the more of it that lives on a single scalable platform, the fewer copies and pipelines there are to keep consistent.

Scality and LLM workloads

Scality's Autonomous Data Infrastructure (ADI) is built for AI data at petabyte to exabyte scale. It lists training, inference, retrieval and agent traffic as distinct access patterns, provides high-concurrency S3 and an S3 over RDMA data path for GPU servers, offers a centralized cache for distributed inference (KV cache), and applies power-aware media tiering across hot, warm and cold AI data. Scality RING, software-defined object and file storage on standard x86 servers, holds up to 300 billion objects in a single RING, enough headroom for corpora, checkpoints and model versions to share one namespace as they accumulate.