Glossary

Multi-region storage

Multi-region storage keeps data in two or more geographically separate regions, so that the loss of a data center, metropolitan area or cloud region leaves the data available elsewhere. A region here is a location far enough from the others that a single power, network or natural event is unlikely to affect more than one.

Why multi-region storage matters for distributed enterprises

Organizations with operations on several continents, regulated recovery obligations or service commitments to customers cannot depend on a single site. A region can be lost to a power failure, a fiber cut, a flood or an operator error at a provider. Multi-region storage determines how much data is lost when that happens, how long service takes to resume, and what it costs in capacity and network to stay ready.

The term covers two goals that are often conflated. Disaster recovery keeps a copy that can take over after a regional loss. Data locality keeps data close to the users and compute in each region so reads stay fast. A design built for one does not automatically serve the other: an asynchronous replica in a distant region protects against site loss but does nothing for users far from the primary, and read replicas placed for locality can lag too far behind to serve as a recovery point.

Multi-region architectures

ArchitectureWhat each region holdsWrite path
Asynchronous replicationA full copyAcknowledged locally, copied to other regions afterwards
Synchronous replication or stretched clusterA full copy, or a share of one clusterAcknowledged only once every region has it
Geo-distributed erasure codingA subset of fragmentsFragments spread across regions; any sufficient subset rebuilds the data
Active-active multi-writerA full writable copyWrites accepted anywhere and exchanged; conflicts resolved by rule

Distance, latency and recovery point

Light in fiber travels about 5 microseconds per kilometre, so 500 km adds roughly 5 ms to every round trip before switching or storage time. A synchronous write waits for that round trip every time, which confines synchronous designs to metropolitan distances. Asynchronous replication keeps local writes fast and lets the remote copy trail behind, and the lag at the moment of failure becomes the actual recovery point.

Lag depends on bandwidth. A site writing 2 GB/s over a 10 Gb/s link, which carries 1.25 GB/s, falls 0.75 GB further behind every second until the write rate drops.

Replication tools also differ in what they copy. Some replicate only objects written after a rule is created, leaving existing data to a separate batch copy, and deletes, versions and metadata changes may or may not travel with the data. A replica that looks complete by capacity can still differ from the source in ways that surface only during failover.

Erasure coding across regions

Erasure coding splits each object into data and parity fragments, and any set of fragments equal to the data count rebuilds it. Spread so that no region holds more fragments than the parity count, the scheme survives the loss of a whole region. With three regions and an 8+4 scheme, four fragments sit in each region; losing one region removes four, which the scheme tolerates. Stored overhead is 12 ÷ 8 = 1.5 times the data, against 3 times for full copies in three regions. The saving is paid in inter-region traffic, since every read and every rebuild draws fragments from at least two regions.

What multi-region storage means for multi-site operations

For an architect spreading petabytes across regions, the architecture sets three bills: capacity, network and recovery. Full replicas in three regions triple the stored footprint; geo-distributed erasure coding cuts it to about one and a half times at the cost of inter-region bandwidth on every read. Links sized for average write rates fall behind during bursts, which widens the recovery point at exactly the moment the most data is changing.

Writes in more than one region create conflicts. Multi-writer designs resolve them by rule, often last writer wins, which can silently discard an update. Single-writer designs avoid conflicts but put distant writers at long-distance latency. During a network partition, synchronous designs stop accepting writes on at least one side. Each choice amounts to a statement about which failure the business tolerates.

Failover involves clients as much as data. Promoting a secondary region means redirecting endpoints through DNS or a global load balancer, and writes taken after promotion exist only in the new primary until failback. Copies in several regions can also sit in several jurisdictions, so data residency rules limit which regions may hold which datasets, and replication rules are scoped to match.

Multi-region storage with Scality

RING provides a unified S3 namespace across sites and clouds and supports multi-geo and stretched-cluster deployments. A stretched RING, one cluster spanning sites synchronously, operates within an envelope of 10 Gb/s or greater bandwidth and less than 5 ms latency between sites, and inside that envelope a site loss does not lose acknowledged writes. Scality's ADI adds multi-site replication and policy-enforced data residency at the namespace level.