Glossary
Storage pooling
Storage pooling combines the capacity of many physical drives, arrays or servers into one shared pool, from which volumes, file systems or buckets draw space as they need it. The pool tracks free space centrally, and allocations do not depend on which physical devices supply them.
Why storage pooling matters
Capacity bought for one application and locked to it is stranded the moment that application grows more slowly than forecast. Across an estate of many systems, stranded capacity and stranded performance add up to a large share of what was purchased. Pooling releases both: free space and drive performance become available to whichever workload needs them next.
Pooling also moves risk. When everything draws from one pool, the pool's free space, its contention and its failure modes are shared by every data set in it. Petabyte-scale pools make both the gain and the shared exposure larger.
How a pool is built
- Devices. Drives, or whole servers in a distributed system, are assigned to the pool.
- Protection groups. Devices are organised into RAID groups, replica sets or erasure-coding groups so that device failures are recoverable.
- Pool. Protected capacity is presented as one quantity of free space, tracked in allocation units (extents, chunks or slabs).
- Allocations. Volumes, file systems or buckets draw allocation units as data is written, without knowing which drives supply them.
That indirection is the same mechanism described under storage virtualization, applied to capacity.
Thick and thin allocation
A thick allocation reserves its full size at creation. A thin allocation draws units only as data arrives, allowing provisioned sizes to exceed physical capacity; see thin provisioning.
| Measure | Value |
|---|---|
| Physical pool capacity | 1 PB |
| Sum of provisioned sizes | 2.5 PB |
| Overcommit ratio | 2.5 ÷ 1 = 2.5:1 |
| Data written | 700 TB (70% utilisation) |
The pool can take 300 TB more, even though its volumes report 1.8 PB free between them. When a thin pool fills, writes fail on every allocation in it, including those far below their provisioned size.
Object storage pools behave much like thin pools by default. Buckets have no preset size and consume capacity only as objects arrive, so the controls that limit growth are quotas per bucket, account or tenant rather than reservations at creation.
Wide striping and shared contention
When an allocation's units are spread across every drive in the pool, a random workload against it is served by all of them. Taking about 100 random reads per second per 7,200 RPM drive, a volume confined to a 6-drive RAID group draws on roughly 600; the same volume striped across a 48-drive pool draws on about 4,800. The pool lends each workload performance that would otherwise sit idle elsewhere.
The counterpart is that every allocation competes for the same drives, so a burst on one raises latency for the others. Storage quality of service limits are the usual control. Failure scope follows the same logic: if every volume is striped across eight protection groups, losing any one group affects every volume in the pool.
What storage pooling means for capacity planning at scale
With pooling, the unit a storage team watches is the pool, and the number that matters is time to full against purchase lead time. A pool growing by 40 TB a week with an eight-week lead time for new hardware needs at least 40 × 8 = 320 TB of headroom at the moment an order is placed, and more if growth is uneven. Overcommitted thin pools add a second number: the gap between provisioned and physical capacity, which shows how badly a sudden fill by tenants could go.
For service providers and platform teams running multi-tenant storage, pooling is what makes per-tenant pricing efficient and also what makes one tenant's burst everyone's latency problem, so quotas, QoS and isolation become part of the service definition. Pool size is a balance: larger pools improve utilisation and wide-striping performance while enlarging the set of data any single fault or full-pool event can reach. Distributed systems change that balance by protecting data across servers, racks or sites, so a pool spanning hundreds of servers can lose whole nodes without losing data.
A single large pool also does not have to mean a single protection layout. Systems that apply protection per class let archive data use a wide erasure code for capacity efficiency and smaller, hotter data use replication for fast access, both drawing from the same drives. The pool then removes stranded capacity without forcing every data set onto the same cost and performance trade-off.
Storage pools in Scality RING
In Scality RING, the drives of a cluster of standard x86 servers form the pool. The RING product page describes multi-tenancy with hard isolation, with capacity, performance, tenants and sites scaling independently of one another. Erasure coding schemes are defined per storage class, so an archive class and a performance class can draw on the same RING with different protection layouts.














