Glossary
Failover
Failover is the transfer of a workload from a failed or unreachable component to a standby, which takes over its role so that service continues.
It applies at every scale, from a disk path or server to a cluster or a whole site. The move back once the original is repaired is failback.
Why failover matters for large platforms
At scale, component failure is routine. A platform with thousands of drives and hundreds of servers loses some every week, and the services on top of it are expected to carry on without anyone noticing. Failover is the mechanism that makes that possible inside a site, and the same idea, applied between sites, decides what happens when a whole data centre goes dark. The difference between the two is mostly about how confident the system can be that the primary has really failed, and what a wrong decision costs.
Failure detection
Failover begins with a decision that the primary has failed. The usual mechanism is a heartbeat: the primary sends a short message at a fixed interval, and an observer declares failure after a set number are missed. Detection time is the interval times the threshold, so a 1-second heartbeat with a threshold of 3 detects failure in about 3 seconds. Shorter detection means less downtime and more false alarms, since a brief network delay can look like a failure. Health checks go further by testing function, issuing a query or request and checking the answer.
Automatic failover is common within a site, where failures are frequent and a wrong call is contained. Between sites, failover is often manual or approval-gated, because a mistaken switch moves many systems at once and can cause data to diverge. Planned failover, or switchover, is used for maintenance: the primary is quiesced and synchronized first, so nothing is lost.
Failover configurations and headroom
| Configuration | Arrangement | Effect of a failure |
|---|---|---|
| Active-passive | One node or site serves; a standby waits | Standby takes the full load |
| Active-active | All nodes or sites serve at once | Survivors absorb the failed share |
| N+1 / N+M | N active members share one or M spares | Spares replace up to that many failures |
In active-active designs capacity sets the limit. Two sites each running at 60% cannot absorb the loss of one, because the survivor would need 120%. Full service after a single loss requires each of two members to stay at or below 50%, or each of three at or below about 67%.
Quorum and split-brain
A network partition can leave two members running, each convinced the other has failed. If both act as primary, both accept writes and the data diverges, a condition called split-brain. Clusters prevent it with a quorum: a member continues as primary only if it can reach a majority of voters. With three voters, two form a majority and one failure is tolerated; with four, three are needed and still only one failure is tolerated, which is why voting groups usually have an odd size. A two-site design cannot form a majority when the link between them breaks, so a third vote is placed at a separate witness location.
What failover means for storage platforms
Failover only helps if clients reach the survivor. Applications find storage through virtual IP addresses, load balancers, DNS names or endpoint lists in their SDKs, and each has its own delay: a DNS record with a 300-second time to live can add up to five minutes before cached clients follow. Connections open at the moment of failure are dropped, and whether a retried write is safe depends on whether the application can repeat it without duplicating its effect.
Data state decides the outcome for anything stateful. With synchronous replication the survivor holds every acknowledged write; with asynchronous replication, whatever had not yet crossed is gone, which is the lag that defines the achievable recovery point objective. Shared-storage clusters avoid replication between servers by moving the single point of failure into the storage system, which makes storage high availability a prerequisite for everything above it.
Site-level headroom is the cost that is easiest to miss. A two-site active-active storage estate at 70% utilization on each side has no room to fail over, however well the software handles the switch. Multi-site resilience is paid for in idle capacity and in a third location for the quorum witness.
Failover within Scality RING
Scality RING spreads data across servers with replication or erasure coding, so a disk or server failure involves no promotion step: requests are served from the remaining copies or fragments. When a drive fails, RING writes data across the remaining drives in that server and rebuilds only data that had been written, keeping the rebuild local to the server. In a stretched deployment across two or three sites, writes are synchronous, and the loss of a whole site plus a further server or disk group leaves read-write access in place when the sites are connected within 10 Gb/s or greater bandwidth and under 5 ms latency.














