Table of Contents
2. SCALITY Software and Services
3.1. What type of information we collect
3.2. What information we do not collect
Specific Provisions — United States
1.3. Purpose of use and legal basis
1.6. California residents' privacy rights
1.7. Change to the Privacy Policy (Specific Provisions)
Specific Provisions — United Kingdom
3.4. How is your Personal Data collected?
3.5. How we use your personal data
3.6. Purpose of use and legal basis
3.9. Withdrawal of your consent
Home › Glossary
Scality Glossary
Clear, concise definitions of the object storage, S3, and scale-out infrastructure terms that matter when you evaluate enterprise storage. Browse the terms below.
A
Active-Active Disaster Recovery — Two or more sites that all serve live production traffic, so losing one requires no failover — the survivors are already running.
AI Data Architecture — The overall design for how AI data is collected, stored, processed and delivered — the blueprint a platform and its storage layer are built to satisfy.
AI Data Infrastructure — The storage, compute, networking and data architecture used to collect, process, protect and deliver the data behind AI and machine learning workloads.
AI Data Management — The practices and policies that keep AI datasets organized, governed, secured and retained across their lifecycle, from ingestion through archiving or deletion.
AI Data Pipeline — The stages data passes through on its way to and from a model — ingestion, preparation, storage, training, checkpointing, inference and the feedback loop back to the start.
AI Data Platform — An integrated environment that stores, governs and delivers the data AI applications run on, sitting between enterprise data sources and the compute that consumes them.
AI Data Readiness — Whether the data and infrastructure behind a given AI use case are accessible, usable, governed, secure and scalable enough to support it.
AI Storage — Storage infrastructure built for the capacity, throughput and access patterns of AI and machine learning workloads, from training datasets to checkpoints and RAG source content.
B
Backup Compression — Lossless algorithms applied to backup data to reduce its stored size, typically by 40–60% for ordinary enterprise data.
Breach Containment — The urgent work of cutting off an attacker’s access once a breach is detected — revoking credentials, isolating systems and removing persistence mechanisms.
C
Cloud Archive Storage — The lowest-cost cloud storage tier, for data that must be kept for years and is retrieved rarely, with retrieval times measured in hours rather than seconds.
Cloud Storage Replication — Automatically maintained copies of data across storage systems, regions or providers, so it survives the failure of any one of them.
Cold Site — A facility with space, power, cooling and connectivity but no computing equipment in it — the cheapest alternate site to hold and the slowest to bring up.
D
Data Fabric — An architecture that connects data across storage systems, applications, clouds and locations so it can be accessed, governed and moved consistently.
Data Gravity — The tendency for applications, services and computing resources to move closer to large concentrations of data as datasets grow.
Data Lake — A centralized repository that stores structured, semi-structured and unstructured data in its original format for analytics, data science and AI.
Data Locality — Keeping data physically or logically close to the compute and users that need it, so less of the work is spent moving bytes across the network.
Data Mobility — The ability to move data between storage systems, sites and clouds without losing its accessibility, integrity or the metadata applications depend on.
Data Replication — Creating and maintaining copies of data across systems, sites or clouds so it stays available when one of them goes down — synchronously or asynchronously.
Distributed File System — Files spread across many servers but presented as one ordinary file system, with the same paths, permissions and read-write semantics.
Distributed Storage — Storage that spreads data across many independent nodes, so capacity, throughput and fault tolerance grow with the number of machines.
E
Erasure Coding — A data protection method that splits data into fragments plus parity fragments and spreads them across drives, nodes or sites, so the original can be rebuilt from the survivors.
G
GPU Storage — Storage built to keep GPUs fed — enough throughput, low enough latency and enough concurrency that accelerators are computing rather than waiting on I/O.
H
Hot Site — A fully equipped secondary data center holding continuously replicated, current data — able to take over production workloads within minutes of the primary site becoming unavailable.
I
Incremental Backup — A backup that captures only what changed since the last backup of any kind, creating a chain that has to be restored in sequence.
M
Multi-Cloud — A method of leveraging multiple cloud computing platforms for independent or orchestrated tasks, rather than depending on one single cloud provider.
Multi-Tenant Storage — Multiple tenants share the same storage infrastructure while their data, permissions and resources remain logically isolated.
O
Object Storage vs. NAS — Two ways of holding unstructured data that differ in how applications reach it: a mounted filesystem that can be edited in place, or whole objects addressed by key over HTTP.
P
Petabyte Storage — Storing a thousand terabytes or more in one system, where drive failure becomes routine and protection overhead and rebuild time drive the design.
R
RAG Hallucination — When a retrieval-augmented generation system produces an answer its retrieved sources do not support — a faithfulness failure that usually originates in the pipeline rather than the model.
S
S3 Compatible Storage — Any storage system that can be read from and written to using the Amazon S3 API, without being Amazon S3 — inheriting the S3 application ecosystem while the data sits where you choose.
Sequential vs Random I/O — Reading data in the order it is physically stored versus jumping between scattered locations — the distinction that governs throughput, IOPS and hardware choice.
Storage Latency — The time one storage operation takes, dominated far more often by queueing and network round trips than by the media itself.
Storage Lifecycle Management — Policy-driven management of data and the storage beneath it across its whole lifecycle — placement, protection, retention and eventual deletion.
Storage Management — The ongoing work of running a storage system: provisioning, protection, monitoring, lifecycle and access control.
Storage Performance Tuning — Systematic optimisation of cache, controller, network and application layers to raise throughput and lower latency without buying more hardware.
T
Thin Provisioning — Allocating capacity to a volume as data is actually written rather than reserving it all up front, at the cost of having to watch real usage.
W
Warm Site — An alternate facility that is partially equipped and partially current — a range rather than a rung, whose recovery time is set mostly by how stale its data is.


















