Glossary

Scale-out storage

Scale-out storage grows by adding nodes to a cluster, with each node contributing drives, processors, memory and network bandwidth to one logical system. Capacity and performance rise together as nodes join, where scale-up storage adds drives behind a fixed pair of controllers.

Why scale-out matters at petabyte scale

A scale-up array has an end. When its controllers run out of processing, cache or drive slots, the next step is a new array and a migration. At tens of terabytes that is a project. At petabytes it is a programme: copying 5 PB at a sustained 5 GB/s takes 5,000,000 GB ÷ 5 GB/s = 1,000,000 seconds, about 11.6 days of uninterrupted transfer, before cutover, verification and the inevitable retries.

Scale-out designs remove that event. The system grows a node at a time and retires old nodes the same way, so the cluster itself has no end-of-life date. For organizations whose unstructured data grows every year, that changes storage from a series of replacements into one long-lived platform.

Scale-up and scale-out

PropertyScale-upScale-out
Unit of growthDrive shelfNode
Controller resourcesFixed, shared by all capacityAdded with every node
Performance per TB as capacity growsFallsRoughly constant
Upper boundController limitsCluster software and network
Hardware refreshWhole system replacedNodes added and retired individually

How data is placed and protected across nodes

A scale-out system computes where each object or file lives without a central lookup for every request, usually by hashing its identifier onto a key space whose ranges belong to nodes. The hashing method decides how much data moves when the cluster changes. With simple modulo placement, growing from 10 to 11 nodes reassigns about 91% of keys. With consistent hashing, the new node takes over only its share, about 1 ÷ 11 or 9%, and the rest stays put. This is the basis of distributed storage.

Protection is spread across nodes in the same way, by replication or erasure coding:

SchemeRaw capacity per usable TBLosses tolerated
3-way replication3.0 TB2 copies
Erasure coding 8+4(8 + 4) ÷ 8 = 1.5 TB4 fragments
Erasure coding 10+4(10 + 4) ÷ 10 = 1.4 TB4 fragments

At 10 PB usable, the gap between 3.0 and 1.5 raw per usable terabyte is 15 PB of drives. The placement rule matters as much as the scheme: fragments are spread so that no node, and in larger clusters no rack or site, holds more of an object than the scheme can lose.

Where scale-out designs reach their limits

Adding nodes increases aggregate capability, rarely in exact proportion. Metadata is the usual bottleneck: designs that kept names, versions and directory trees on one or two servers hit a ceiling as every create, list and delete passed through them, and later designs distribute metadata across nodes as well. The network is the other constraint, because protection traffic, rebalancing after expansion and rebuilds after failures all cross it alongside client traffic. Rebalancing and rebuild are throttled for that reason, which trades recovery speed for steady client performance.

Rebuild time is where the architecture shows most clearly. A failed 20 TB drive rebuilt from many drives in parallel at an aggregate 2 GB/s takes 20,000 GB ÷ 2 GB/s = 10,000 seconds, under three hours. The same drive rebuilt onto a single spare at 200 MB/s takes 100,000 seconds, more than a day, with reduced protection for the whole period.

What scale-out means for capacity planning and operations

For a storage team, purchasing shifts from sizing a system for five years to adding nodes in step with measured growth, and refresh becomes continuous: on a five-year server life, roughly a fifth of the nodes are replaced each year while the cluster keeps serving. Mixed hardware generations become normal, with placement weighting larger nodes to receive more data.

Failure handling also changes character. With thousands of drives, a drive failure is a routine weekly event absorbed by the protection scheme, and the operational question becomes how fast the cluster restores full protection and how much that work slows clients. Lean teams benefit most, because one system to manage stays one system as capacity grows. The trade-off is coupling: in designs where every node carries both compute and capacity, a capacity-heavy archive pays for processing it does not use, and a performance-heavy AI workload pays for drives it does not need. Architectures that let those dimensions scale separately address that.

Multi-site designs extend the same idea. A cluster can span data centres with fragments placed so that the loss of a whole site is survivable, at the price of inter-site bandwidth on every write and latency that depends on the distance between sites.

Scale-out architecture in Scality RING

RING is scale-out object and file storage on standard x86 servers, and a single RING holds up to 300 billion objects. Its MultiScale architecture separates the compute that serves data from the storage that holds it, and the RING product page describes capacity, performance, tenants and sites as scaling independently. When a drive fails, RING writes data across the remaining drives in that server and rebuilds only data that was written, which keeps rebuild work local to the affected server.