Glossary
Disaster recovery plan
A disaster recovery plan (DRP) is the documented procedure an organization follows to restore its information systems at an alternate location after a disruption leaves the primary location unusable. It records who acts, in what order, on which systems and from which copies of data.
Why the disaster recovery plan matters
In a large environment a real recovery involves hundreds of systems, several teams and decisions taken under pressure, often by people who did not design the systems they are restoring. The plan is the only thing that coordinates them. It also carries weight outside IT: auditors, regulators and customers with contractual recovery commitments ask to see it, and the results of executing it.
A plan describes the environment at the moment it was written, while infrastructure keeps changing. The gap between the two is where most recovery failures start.
What a disaster recovery plan contains
- Scope: the systems, sites and types of event covered.
- Roles and authority: the recovery team, deputies, the person who declares a disaster, and vendor contacts, held somewhere that survives the event.
- Recovery objectives: the recovery time and recovery point objective for each system, taken from the business impact analysis.
- Data sources: which replica, snapshot or backup each system is recovered from, and where it is held.
- Alternate site: capacity, access arrangements and the steps that bring it into service.
- Procedures: step-by-step recovery instructions in execution order.
- Validation and return: the checks that confirm recovered systems are correct, and the procedure for failback.
In the framework of NIST SP 800-34, the DRP is the plan that relocates information systems, sitting alongside the business continuity plan, single-system contingency plans and the incident response plan. Because it is written to be executed by people who may not have written it, under time pressure and possibly without access to the systems it describes, current copies are kept outside the environment it covers, at the recovery site and offline.
Recovery order and dependencies
Systems cannot all start at once. An application that authenticates against a directory fails if the directory is not running, and the directory depends on DNS, network and storage. The plan therefore groups recovery into tiers: network and storage first, then core services such as DNS, directory and key management, then databases, then applications in business priority, and finally user access.
The shortest possible recovery time for an application is the sum of its dependency chain. If storage takes 30 minutes to bring into service, directory services 20, the database 45 and the application itself 25, the application cannot be available in under 30 + 20 + 45 + 25 = 120 minutes, whatever its own recovery time objective says. Shared services inherit the strictest objective among the applications that depend on them: if a payments application has a two-hour target and depends on the directory, the directory recovers in under two hours however low its own business profile. Plans that automate this ordering become workflows, the subject of disaster recovery orchestration.
What a disaster recovery plan means for large storage environments
Storage sits at the bottom of every dependency chain, yet it is often described in a plan with a single line. At scale that line hides the decisions that matter most. The plan has to state which copy serves which kind of event: a replica at the second site for site loss, and a retained point-in-time version for corruption or ransomware, since the replica carries the damage. If the plan names only the replica, a ransomware recovery starts from encrypted data.
Capacity and throughput at the recovery site are part of the plan whether written down or not. A plan that assumes a full restore from backup has implicitly accepted the transfer time for every petabyte involved, and a plan that assumes failover to a second site has accepted that the site can carry the full load. Credentials, management consoles and encryption keys for the storage platform also have to be reachable from the surviving site; a plan that stores them only in the lost data centre stalls at the first step.
Drift is the ongoing cost. Every new bucket, storage class, replication rule or AI pipeline changes what the plan describes, and results from disaster recovery testing are the main way that drift is found before a real event finds it.
Storage entries for Scality RING in a DR plan
For data on Scality RING, the data-source and alternate-site entries describe the RING topology. A stretched RING runs synchronously across two or three sites within 10 Gb/s or greater bandwidth and under 5 ms latency, so the plan contains no replica promotion or resynchronisation step for that data. Where recovery depends on point-in-time copies, the plan identifies the object versions to restore from and their S3 Object Lock mode: governance mode, which an identity holding s3:BypassGovernanceRetention can override, or compliance mode, which no user can override before the retain-until date.














