Glossary
Failback
Failback is the return of production workloads and data from a recovery site to the primary site, or to a replacement for it, after a failover. It reverses the direction of replication, brings the primary up to date with everything written while it was out of service, and then moves applications and clients back in a planned cutover.
Failover happens under pressure, often in minutes. Failback starts with both sites running, so it is a scheduled change, and at large data volumes it is usually the longer of the two.
Why failback matters for multi-site storage
Most recovery planning concentrates on getting service running somewhere else. Far less attention goes to the return trip, yet the recovery site is often the weaker one: fewer servers, less capacity headroom, slower media, or space rented from a provider by the month. Every week spent there carries cost and runs with reduced protection, because the site that normally receives the replica is the one that failed.
For an organization holding petabytes, failback is mostly a data movement problem. Applications can be redirected in minutes. Moving the data they wrote while away, or re-seeding a rebuilt primary from scratch, can take days or weeks over the links between sites. That duration, more than the cutover itself, decides how long the business stays exposed after an incident.
How a failback proceeds
- Restoration of the primary. The cause of the outage is resolved: hardware replaced, power or network restored, or infrastructure rebuilt after an attack.
- Reverse replication. A replication relationship is set up from the recovery site back to the primary.
- Resynchronization. The primary receives either the changes made since failover or a complete copy of the data.
- Catch-up. Replication continues while the recovery site still serves production, until the gap between the sites shrinks to normal replication lag.
- Cutover. Applications are quiesced, the last changes are sent, site roles are swapped, and DNS, load balancers or client endpoints are pointed back at the primary.
- Validation and re-protection. Applications are checked at the primary, and replication resumes in its original direction.
Only the cutover interrupts service, and its length depends on how small the remaining backlog of changes is when it starts.
Delta and full resynchronization
A delta resynchronization sends back only what changed at the recovery site since failover. It is possible when the primary's storage survived intact and its last replicated state is known. A full resynchronization copies everything, and it is required when the primary is new hardware or when its data can no longer be trusted.
The gap between the two grows with scale. Transfer time is data volume × 8 ÷ link speed in bits per second. For a 2 PB dataset over a 10 Gb/s link running at line rate:
| Scenario | Data to move | Time at 10 Gb/s |
|---|---|---|
| Delta after two weeks at 1% daily change | Up to 280 TB | About 2.6 days |
| Full resynchronization | 2 PB | About 18.5 days |
Links shared with production traffic rarely sustain line rate, so both figures stretch. A full re-seed at this scale is often done by shipping drives or servers, or by adding temporary bandwidth, which turns a networking question into a logistics and budget one.
Divergence between sites
With asynchronous data replication, the primary may have accepted writes in its last seconds that never reached the recovery site. Those writes fall inside the recovery point objective window. Once production has continued elsewhere they conflict with newer history, and failback either discards them by rolling the primary back to its last replicated point or hands them to the application to reconcile.
A worse case is split-brain, where a network partition leads each site to treat the other as failed and both accept writes. Clustered systems prevent it with a quorum or a witness at a third location that decides which side continues. Synchronous replication removes the first kind of divergence entirely, since a write is acknowledged only once both sites hold it.
What failback means for multi-site operations
For teams running replicated object or file storage across sites, the mechanics above have direct consequences.
- The return trip sets the exposure window. A failover measured in minutes can be followed by weeks on a single site with no second copy, if the primary has to be re-seeded in full.
- Cyber incidents change the arithmetic. After ransomware or another intrusion the original primary is treated as compromised and rebuilt, so a delta resynchronization becomes a full one. Data may also pass through an isolated environment first, as described under clean room recovery.
- Link sizing is a recovery decision. Inter-site bandwidth chosen for steady-state replication is often far below what a full re-seed needs, and the shortfall only shows up during failback.
- Applications return in dependency order. Storage comes back before the databases on it, and databases before the services that use them, the sequence that disaster recovery orchestration runbooks encode.
- Symmetric designs avoid the question. Where both sites serve production, as in active-active disaster recovery, there is no passive copy to return from.
In effect, failback time is decided at design time, by replication topology and inter-site capacity, long before any outage.
Failback and Scality RING
Scality RING supports two- and three-site stretched deployments in which writes are synchronous across sites. Within the published envelope of 10Gb/s or greater bandwidth and less than 5ms latency between sites, an entire site plus another server or disk group can fail while read-write access continues. That configuration has no passive replica to promote, so the active-passive failback sequence does not apply to it in the same way.
Separate RING clusters at greater distances replicate asynchronously. Moving production between them follows the familiar order of reverse replication, resynchronization and cutover.














