Glossary

Disaster recovery testing

Disaster recovery testing is the practice of exercising a disaster recovery plan, in discussion or by actually carrying out its steps, to find out whether systems and data can be restored within their recovery objectives. Its output is measured recovery times, achieved recovery points and a list of what went wrong.

Why disaster recovery testing matters

Until a plan has been executed, its recovery times are estimates written by the people who designed the systems. A test replaces those estimates with measurements, and in large environments the two rarely match. Steps turn out to depend on a credential nobody can find, a service that was added after the plan was written, or a restore rate the recovery site cannot sustain. In multi-site environments a test is also the only controlled moment to discover whether the second site can really carry production load, before an outage forces the question.

Testing is also how an organization shows auditors, regulators and customers that recovery commitments are real. Sector regulators in finance, healthcare and government frequently require evidence of it.

Types of disaster recovery test

TypeWhat happensSystems involved
Plan reviewOwners check the document against the current environmentNone
Tabletop exerciseStaff talk through their roles in a described scenarioNone
Component testOne element is recovered: a database restored, a replica promotedOne system, isolated
Parallel testSystems are fully recovered at the alternate site while production runsFull recovery site, isolated
Full interruptionProduction stops and operations move to the recovery siteProduction and recovery site

Each step down the table costs more and disrupts more, and each tells more about what would happen in a real event. Tests can also be announced or unannounced. Announced tests measure the procedure; unannounced tests also measure detection, escalation and the time taken to assemble the team, which in a real event often exceeds the technical work.

What a test measures

  • Measured recovery time: from the declared start of the scenario to confirmed service, per system and per phase, including the human stages. If the technical restore takes 90 minutes and locating a credential takes 40, the measured figure is 130 minutes.
  • Achieved recovery point: the gap between the scenario start and the newest consistent data recovered. A scenario starting at 14:00 that recovers a database consistent at 13:42 achieves 18 minutes, compared against the RPO.
  • Data integrity: record counts, checksums or application checks against the source.
  • Procedure accuracy: steps that were missing, out of order or blocked.

Scenarios matter as much as metrics. A site-loss test and a ransomware test exercise different copies: the first can use the replica, the second has to start from isolated point-in-time copies because the replica is assumed compromised.

What disaster recovery testing means for storage and platform teams

Scale distorts small tests. Restoring 10 TB in one hour does not mean 1 PB restores in 100 hours, because concurrency limits, network links, metadata operations and the recovery site's own capacity all start to bind at volumes a component test never reaches. A test plan that only ever recovers a sample measures the procedure and leaves the throughput question unanswered, and throughput is usually what decides the outcome at petabyte scale.

Parallel tests need somewhere to run. Recovered copies of production cannot share a network with production without conflicting hostnames and addresses, so they run in an isolated segment, fed by writable clones or snapshots of the replicated data. That consumes capacity at the recovery site for the duration of the test, and a recovery site sized only for steady-state replicas may not have it.

Frequency follows from cost. A full interruption test of a multi-petabyte platform is a major undertaking, so a common pattern combines frequent component and parallel tests with rarer full-scale exercises, and the smaller tests are only as informative as their data volumes and concurrency are realistic.

For ransomware scenarios, the meaningful test restores from the copies that would be used after an attack, typically retained, locked versions, and checks how long it takes to identify the last clean point before restoring from it. The difference between the plan's documented times and the measured ones, tracked across successive tests, is the clearest indicator of whether the plan still describes the environment. Automating recovery through disaster recovery orchestration makes those measurements repeatable.

Scality RING in recovery tests

With Scality RING the storage behaviour in a test depends on topology. A stretched RING, synchronous across sites within 10 Gb/s or greater bandwidth and under 5 ms latency, can be tested by isolating one site and confirming that read-write access continues from the others. Restore tests that read object versions held under S3 Object Lock in compliance mode leave those versions untouched, because no user, including the account root, can delete or overwrite them before the retain-until date, so the same protected copies remain available for the next test or a real event.