Glossary
Recovery point objective (RPO)
Recovery point objective (RPO) is the maximum amount of recent data, expressed as a span of time before a disruption, that an organization accepts losing when a system is recovered. An RPO of one hour means recovered data may be up to an hour older than the moment of failure.
Why RPO matters at scale
RPO decides which protection technology a dataset needs, and those technologies differ in cost by large factors. A nightly backup and a synchronous copy at a second site both protect data, yet one accepts a day of loss and the other accepts none, and the second requires a second site within short network distance. Across petabytes, setting RPO per dataset rather than one value for the whole estate is one of the main levers on what resilience costs.
Time is used as the unit because loss depends on when the last usable copy was made. Turning that time into data volume shows what is at stake: a system writing 30 GB of changes per hour with a four-hour copy interval has up to 30 × 4 = 120 GB at risk at any moment.
How the recovery point is set
With copies taken at a fixed interval, a failure loses everything written since the last copy completed. A failure just before the next copy loses almost a full interval, so the worst case equals the interval, and the RPO is met only when that worst case is within it. Daily backups at midnight give a worst case of 24 hours and an average of 12, and the recovery point actually achieved depends on which copy is used: seconds from a healthy replica, a full day from the backup if the replica is unusable.
| Protection method | Recovery point |
|---|---|
| Daily full or incremental backup | Up to 24 hours |
| Periodic snapshots | The snapshot interval, commonly minutes to hours |
| Asynchronous replication | Replication lag at the moment of failure, seconds to minutes |
| Continuous data protection | Seconds, to any point in a retained journal |
| Synchronous replication | Zero for acknowledged writes |
RPO describes data; the recovery time objective describes time out of service. The two are independent, and both are set per system in a business impact analysis.
Replication lag and distance
Asynchronous replication acknowledges writes at the primary and sends them afterwards, so the replica trails by whatever is still queued. Lag stays small while the link carries the change rate and grows when writes exceed it. A 100 Mb/s link carries 45 GB per hour; a batch job writing 60 GB per hour for two hours leaves 30 GB queued, which takes another 40 minutes to drain after the job ends. A failure at the peak loses much of the batch.
Synchronous replication removes that lag by acknowledging a write only once it is stored at both sites, which adds a network round trip to every write. Light in fibre covers roughly 200 km per millisecond, so 100 km of separation adds about 1 ms of round trip per write before equipment delay. That is why zero-RPO designs stay within metropolitan distances, and longer distances accept a non-zero RPO.
What RPO means for multi-site storage
An asynchronous RPO is a function of workload peaks as much as link size. Nightly ingest, model checkpoints and batch analytics produce bursts that a link sized on daily averages cannot keep up with, so the achieved recovery point is worst exactly when the most data is being written. Monitoring lag during peaks gives a more honest picture of the real RPO than the configured replication schedule.
Logical damage changes the effective recovery point entirely. Replication copies deletions, corrupted records and ransomware encryption as faithfully as good data. If corruption began six hours before detection, a system with a five-minute replication RPO recovers to a point at least six hours old, taken from whichever retained copy precedes the damage. The recovery point that matters for ransomware is set by point-in-time copies and how long they are kept, rather than by replication.
Different data in the same estate tolerates very different loss. Raw sensor or log data that can be re-ingested from source, training corpora that can be rebuilt, and transactional records that cannot be recreated do not need the same RPO, and treating them alike either overspends on the first two or underprotects the last.
Recovery points in Scality RING deployments
For site loss, Scality RING in a stretched configuration is synchronous across two or three sites: data is acknowledged only once committed to persistent media, and Scality documents zero RPO through the loss of an entire site plus a further server or disk group, within an envelope of 10 Gb/s or greater bandwidth and under 5 ms latency between sites. For logical damage, RING supports S3 Object Lock on versioned buckets; in compliance mode no user, including the account root, can remove a locked version before its retain-until date, which keeps a known-good recovery point available after the live data has been altered.














