Glossary
Cloud elasticity
Cloud elasticity is the ability of a system to add and release resources automatically as demand rises and falls, so that the capacity in use follows the workload closely over time. It is judged by how quickly and how precisely allocation tracks demand, in both directions.
NIST SP 800-145 lists rapid elasticity as one of the five essential characteristics of cloud computing, describing capabilities that can be provisioned and released to scale outward and inward with demand.
Why elasticity matters at scale
Elasticity is the economic case for cloud computing: capacity is paid for while it is used and returned when it is not. That case is strong for compute, where a web tier or a batch job can shrink to almost nothing overnight. It is weak for stored data, which does not disappear when traffic falls. A dataset that was read heavily during the day still occupies the same petabytes at night.
For an architect, the useful question is which costs fall when demand falls. Compute, network throughput and request rates can be elastic. Stored capacity grows with retention, shrinks only when data is deleted or moved, and on owned infrastructure is bounded by the servers already installed.
Elasticity and scalability
| Elasticity | Scalability | |
|---|---|---|
| Time scale | Minutes to hours | Months to years |
| Direction | Up and down | Mostly up |
| Trigger | Load metrics, schedules | Growth in data, users, workloads |
| Mechanism | Automated scaling of instances or pods | Architecture that accepts more servers without redesign |
A system can only be elastic within the limits its architecture allows. Cloud scalability sets the ceiling; elasticity is how closely usage follows demand beneath it.
How the elasticity cycle works
- Measure: a monitoring system collects a load signal, such as CPU use, request rate or queue length.
- Decide: a scaling policy compares the signal with a target. The Kubernetes horizontal pod autoscaler uses desired replicas = ceil(current replicas × current metric ÷ target metric), so 4 replicas running at twice the target become 8.
- Provision: new instances or pods are started and join the service.
- Release: when the signal falls, capacity is removed, usually after a stabilization delay that stops the system from oscillating.
The weak point is lag. Starting new capacity takes seconds for containers and minutes for virtual machines, so demand that rises faster than provisioning briefly overloads whatever is already running. Every layer behind the elastic tier, including storage, carries that peak in the meantime.
Elasticity of storage
Storage has two dimensions that behave differently. Capacity follows retained data and grows almost monotonically: it is released only by deletion, expiry or moving data elsewhere. Performance demand, in throughput and requests per second, can be highly elastic, spiking during a training run, a restore or an end-of-month analytics job and falling back afterwards.
Public cloud object storage presents capacity as effectively unlimited and bills only what is stored, which looks elastic from the consumer's side. On owned infrastructure, the equivalent is a pool with headroom plus thin provisioning, which lets tenants see large allocations while physical capacity is added behind them as it is consumed.
What elasticity means for petabyte-scale storage
- Elastic compute concentrates load on storage. A GPU or analytics cluster that grows from 100 to 1,000 workers sends ten times the read requests at the same dataset. Storage has to be sized for the throughput of the largest burst, while its capacity is sized for what is retained.
- The storage bill does not shrink on its own. Scaling compute down cuts its cost the same day. Storage cost falls only when data is deleted, expired or moved to a cheaper tier, which makes lifecycle policy the storage equivalent of scaling in.
- On premises, headroom is the elasticity. Owned storage absorbs bursts only up to the capacity and throughput installed. The pool's spare margin, and how fast servers can be added to it, decide how much variation it absorbs.
- Recovery is the largest burst. After an outage or a ransomware incident, restore jobs read from backup and archive storage at the highest rate the environment ever demands, at the moment it matters most. Storage sized for steady ingest can become the bottleneck of the recovery itself.
- Request charges are elastic too. In public cloud object storage, a workload that scales out also scales its per-request and egress charges, so an elastic job reading a large dataset many times can cost more in requests than in storage.
Elasticity and Scality RING
On owned infrastructure, Scality RING provides the storage pool that elastic services draw buckets and capacity from. RING runs on standard x86 servers and grows by adding them, and Scality states that capacity, performance, tenants, sites and protocols scale independently rather than in lockstep. Throughput can therefore be expanded for bursty workloads without buying capacity they do not need, and capacity can be added for retention without overbuying performance.














