Glossary
Disaster recovery
Disaster recovery (DR) is the combination of second copies of data, alternate locations and documented procedures used to restore IT systems after an event that disables a whole site or destroys the data it holds.
It is the IT part of business continuity, which also covers people, premises and suppliers.
Why disaster recovery matters at scale
For a few terabytes, recovery can mean restoring from backup onto new hardware. At petabyte scale that approach stops working because of transfer time alone. Restoring 2 PB over a 10 Gb/s link, running flat out, takes 2 × 10¹⁵ × 8 ÷ 10¹⁰ = 1.6 million seconds, about 18.5 days. No business process tolerates that, so large environments decide their recovery capability in advance, by where the second copy already lives and how quickly it can serve applications.
The other pressure is the type of event. A flooded data centre and a ransomware attack both count as disasters, and they leave very different data behind.
Types of disaster and the copies that survive them
| Event | What happens to data | Copies that remain usable |
|---|---|---|
| Site loss (fire, flood, power, network) | Intact but unreachable, or destroyed with the site | Any copy at another site |
| Platform failure (array, hypervisor cluster, identity service) | Site survives, dependent services stop | Copies on independent platforms |
| Logical corruption (software fault, admin error) | Damage written to production and replicated within seconds | Point-in-time copies taken before the damage |
| Ransomware | Data and often backups encrypted or deleted deliberately | Copies the attacker could not alter or delete |
Replication handles the first two rows. Only versioned, retained copies handle the last two, because every replica faithfully receives the corruption.
Recovery approaches
DR designs are described by two targets: the recovery time objective (RTO), how long a system can be down, and the recovery point objective (RPO), how much recent data can be lost. Both are set per system, usually from a business impact analysis.
- Restore from backup: lowest standing cost, recovery point equal to the backup interval, recovery time set by hardware provisioning plus restore throughput.
- Standby with asynchronous replication: a replica at a second site trails production by seconds to minutes; recovery means promoting it and starting compute.
- Synchronous replication: every acknowledged write exists at both sites, at the cost of a network round trip per write, which limits distance.
- Active-active: both sites serve production, and a site loss is absorbed by the survivor, provided it has the capacity.
The approaches differ mainly in what stands ready before the event. Restore from backup keeps almost nothing running at the second site and pays for it in recovery time. Replication keeps a full copy of the protected data at the second site, roughly doubling raw capacity for that data, plus the network to keep it current. Synchronous and active-active designs add the requirement that sites sit close enough for every write to cross between them, which in practice means metropolitan distances.
Recovery sites range from a cold site with power and space to a hot site running current data.
What disaster recovery means for multi-site storage operations
At large scale, the storage architecture is the DR plan for data. A replicated or stretched platform turns site loss into a capacity and network question; a single-site platform with offsite backups turns it into weeks of restore. That choice is made when the platform is deployed, and changing it later means moving petabytes.
A second site also doubles exposure to the same mistakes. A deletion, a corrupt application write or an attacker with administrator rights propagates to every synchronous and asynchronous copy. Environments that cover both site loss and logical damage therefore combine geographic distribution with versioned copies that hold their retention independently of whoever controls production.
Location carries obligations of its own. Organizations under sovereignty or residency rules need the recovery copy inside the same jurisdiction, which rules out some cloud regions and shapes where a second or third site can go. For lean operations teams, every manual step in recovery is a step that depends on specific people being available on the day, so designs that need no data promotion or resynchronisation at the storage layer remove a large part of the work at the worst possible moment.
Disaster recovery with Scality RING
Scality RING, software-defined object and file storage on standard x86 servers, can be deployed as a stretched cluster across two or three sites. Writes are synchronous, and Scality documents that an entire site plus a further server or disk group can fail while read-write access continues with zero RPO and RTO, within an inter-site envelope of 10 Gb/s or greater bandwidth and under 5 ms latency. For corruption and ransomware, RING supports S3 Object Lock in governance and compliance modes on versioned buckets; compliance-mode versions cannot be deleted by any user, including the account root, until their retain-until date, whereas governance mode can be bypassed by an identity holding the bypass permission.














