Glossary

Cloud scalability

Cloud scalability is the ability of a system to handle growth in data, workload or users by adding resources, while performance and cost per unit stay within acceptable bounds. Scaling up adds resources to an existing machine; scaling out adds more machines and spreads the work across them.

Compounding growth over a platform's life

A 10 PB estate growing 30 percent a year reaches 10 × 1.3⁵ ≈ 37 PB in five years without a single new workload. Storage platforms are not swapped lightly, since moving tens of petabytes off one takes months, so the design chosen today has to hold at that later size, along with the object counts, request rates and site counts that arrive with it. What gives out first at that size is seldom raw capacity. It is metadata, rebuild time, rebalancing and the staff hours needed to run the result. The capacity side of that growth is covered under petabyte storage.

Bigger controllers or more servers

Scale up (vertical)Scale out (horizontal)
How growth happensBigger controllers, more drives behind themMore servers joining a cluster
CeilingThe largest machine availableThe software's coordination limits
Throughput as capacity growsFixed by the controllerRises as each server adds CPU, network and drives
Failure impactA controller fault affects the whole systemA server fault removes a fraction of the cluster
End of lifeForklift replacement and migrationOld servers retired while new ones join

Public clouds scale out almost everywhere because no single machine reaches their size, and scale-out storage applies the same principle to data on owned hardware. Gains are never perfectly linear, since servers spend part of every addition agreeing on where data sits and keeping metadata consistent. Designs that route each request through a central component flatten once it saturates. Designs that spread data and metadata across every node, as many distributed storage systems do, stay closer to linear.

Object count as the hidden ceiling

Storage grows along several axes at once, and the axis that scales worst sets the real limit: bytes stored, number of objects or files, aggregate throughput and request rate, tenants and identities in the control plane, and the number of sites. Bytes and object count part ways quickly. 300 billion objects averaging 1 MB hold 300 PB; the same count averaging 100 KB holds 30 PB with an identical metadata load, so a system sized only in bytes can hit its object limit at a tenth of its rated capacity.

Two effects grow with size whatever the design. Drive failures become continuous, because a cluster of thousands of drives always has some failing, and bigger drives mean more data to reconstruct after each one. Expansion triggers rebalancing as data spreads onto new servers, and adding servers to a nearly full cluster forces a large redistribution while it is already under pressure; smaller, earlier expansions move less each time. Both repair and rebalancing take bandwidth from applications.

Cloud elasticity and bursts storage cannot shed

Cloud elasticity is the short-term counterpart of scalability: resources added and released automatically as demand rises and falls over minutes or hours. NIST SP 800-145 lists rapid elasticity among the five essential characteristics of cloud computing. Scalability sets the ceiling, and elasticity is how closely usage tracks demand beneath it. An autoscaler measures a load signal, compares it with a target, starts capacity and releases it after a stabilization delay that prevents oscillation. The Kubernetes horizontal pod autoscaler computes desired replicas as ceil(current replicas × current metric ÷ target metric), so 4 replicas running at twice the target become 8.

Stored data is the part that does not shrink when traffic falls. Capacity follows retention and is freed only by deletion, expiry or moving data elsewhere, while throughput demand spikes during a training run, a restore or a month-end batch and then drops. Elastic compute concentrates those spikes on storage: an analytics or GPU cluster growing from 100 to 1,000 workers sends ten times the read requests at one dataset. In public cloud, per-request and egress charges scale out with the job. On premises, elasticity equals installed headroom plus the speed at which servers can be added, with thin provisioning letting tenants see large allocations while physical capacity follows consumption. The biggest burst most estates ever face is recovery, when restore jobs read backup and archive capacity at full rate at the worst possible moment.

Scality RING and cloud scalability

RING scales out by adding standard x86 servers, to 300 billion objects in a single RING, and Scality reports approximately 6 exabytes under management across its customer base. Scality states that capacity, performance, tenants, sites and protocols scale independently, so an archive can add capacity without buying throughput and a training tier can add throughput without buying capacity. After a drive failure, RING writes data across the remaining drives in the server and rebuilds only data that was written. Owned servers do not hand capacity back when demand falls, so on RING elasticity means installed headroom plus servers added to the running cluster.