Cyber resilience

Ransomware recovery exercises: What to measure

A recovery exercise is only useful if it produces numbers. What to measure so the next one shows real improvement.

5 min read
Operations command room with blank glowing wall displays under violet lighting

Most organizations run some form of recovery exercise, and most produce a report that says it went well, with a list of observations and a note to repeat next year. The exercises that change anything produce numbers instead: how long each phase took, where the time went, and what the measured recovery capability is compared with the objective the business has been given.

The difference is not effort. It is deciding beforehand what will be measured, and then measuring it during an exercise rather than reconstructing it afterwards from memory and calendar entries.

Decide the scenario before the metrics

A ransomware exercise is not a disaster recovery exercise with a different label. The distinguishing assumptions are that the production environment is untrusted, that identity infrastructure may be compromised, that the most recent restore points may contain the intrusion, and that the recovery happens under investigation constraints rather than at full speed.

Those assumptions change which measurements matter. In a hardware failure scenario the interesting number is data movement rate. In a ransomware scenario the interesting numbers include how long it took to establish a clean restore point, how long identity recovery took, and how long systems waited for examination before being allowed onto the network.

State the scenario in a paragraph and have the participants agree it before starting. Exercises drift toward the comfortable version otherwise, which is the one where the backup repository is healthy and everyone can log in.

The measurements worth taking

Timestamps are the raw material. Record when each phase started and ended, and record what the team was waiting for whenever nothing was progressing.

MeasurementWhat it revealsFrequently surprising
Time to assemble the team and startEscalation and contact effectivenessOften hours, rarely counted
Time to retrieve credentials and keysWhether break-glass access worksA single custodian is a common blocker
Time to recover the backup catalogFront of the dependency chainAdds directly to everything downstream
Time to identify a clean restore pointForensic readiness and retention depthUsually the longest single phase
Time to first restored systemEnd of the setup phaseThe point where the plan becomes real
Aggregate restore throughputThe rate that sets the total durationWell below single-stream figures
Per system overheadProvisioning, verification and handoverDominates when the count is large
Time to first business service restoredWhat the business actually experiencesLater than the first restored system
Idle time and its causesWhere the plan stallsWaiting for a person or a decision

Idle time deserves its own tracking. In most exercises a substantial share of elapsed time is spent waiting for a decision, an approval or a person, and that share is invisible unless someone records it. It is also the part that improves fastest once it is named.

Scope the exercise so it can be completed

A full scale recovery is impractical for most organizations, and an exercise that is too ambitious gets abandoned halfway and produces no measurements at all. A narrower exercise that runs to completion is more useful than a broad one that does not.

Two shapes work well. The first is a depth exercise: recover a small number of systems completely, through every phase from credential retrieval to business validation, measuring each step. The second is a rate exercise: restore a larger number of systems concurrently purely to measure aggregate throughput and per system overhead, without the full procedural wrapper.

Alternating between the two across successive exercises gives both the procedural findings and the numbers needed to extrapolate a full scale recovery timeline without ever having to run one.

Include the parts people skip

Three phases are routinely omitted and are among the most valuable to include. Verification is the first: confirming that a restored system is functional and that its data is correct, which takes longer than restoring it and is usually done by application owners rather than by the infrastructure team. Examination is the second: the checks performed before a restored system is trusted, which is what a ransomware scenario requires and a hardware failure scenario does not.

The third is communication. Who was told what, when, and how, including the case where the usual communication tools are unavailable. Exercises that use a working email system to coordinate a scenario in which email is down are measuring the wrong thing.

Running the exercise in an isolated environment, along the lines of a clean room recovery, makes all three realistic without touching production.

Where Scality fits in an exercise

An exercise that restores from the real repository produces real numbers, which is preferable to restoring from a copy created for the occasion. Where backups sit on Scality RING or ARTESCA with S3 Object Lock, an exercise can restore from locked copies exactly as a real recovery would, without any risk of the exercise altering or consuming them.

Two figures are worth capturing from the repository side specifically. One is aggregate read throughput with many concurrent restore streams, which sets the duration of a large recovery. The other is the same figure with a node deliberately failed, since a real incident may not find the system in perfect health. Both belong in the recovery plan next to the objective they are measured against.

Compare against the last exercise, not against the plan

The report that matters is a comparison table: this exercise against the previous one, phase by phase, with the differences explained. Improvement shows up as reduced idle time, faster credential retrieval and a shorter path to the first restored system, and those improvements come from specific fixes made after the last exercise.

Two or three prioritized fixes carried out before the next exercise beat a long observation list that nobody owns. The measure of a good exercise programme is that the numbers move, which requires that they were numbers in the first place.

See Scality in action

Exabyte-scale object storage for AI data and cyber resilience. Talk to our team about what it can do for yours.

Request a demo