Glossary

Secondary storage

Secondary storage holds copies of production data and data no longer in active use: backups, replicas, archives, snapshot copies and test or analytics clones. Primary storage serves live application reads and writes; secondary storage is sized for capacity, throughput and cost per terabyte.

In computer architecture the same phrase means any non-volatile storage reached through an I/O interface, as opposed to memory. In enterprise IT it describes a role, and the same media can serve either role.

Why secondary storage matters at enterprise scale

Secondary storage is usually larger than primary storage, often by a multiple, because it holds many historical versions of each data set plus archives that primary systems have shed. An enterprise with a few petabytes in production can easily hold several times that in backups, replicas and archives.

It is also where recovery comes from. When production is lost to a failure, a mistake or ransomware, the speed and integrity of the secondary copies decide how long the outage lasts and whether it ends without paying anyone. That makes secondary storage a resilience system that happens to be priced like a capacity tier.

What secondary storage holds

WorkloadData heldAccess pattern
BackupFull and incremental copies kept for set retention periodsLarge sequential writes; bulk reads during restore
Replication targetA copy of production at a second siteContinuous or scheduled writes; reads on failover
ArchiveData retained for compliance or referenceWrite once, rarely read
Snapshot exportPoint-in-time copies from primary arraysPeriodic writes; occasional reads
Test, dev and analytics copiesClones of production used outside itMixed reads and writes

Each workload has a different relationship with time: backups expire on a rolling schedule, archives accumulate for years, and replicas track production continuously. Backup is the largest of these in most organizations, and the term often means the backup target specifically.

How secondary storage is sized

Backup data is highly redundant. Thirty daily backups of a 100 TB data set with 2% daily change occupy 30 × 100 TB = 3 PB as full copies. With deduplication or incremental backup, the first copy stores 100 TB and each later one about 2 TB: 100 + (29 × 2) = 158 TB, a ratio near 19:1. Compression reduces the unique data further, depending on content.

Throughput matters more than latency, since backup jobs and restores move data in bulk. Cost per terabyte comes down through high-capacity drives, erasure coding in place of mirroring, and data reduction. Media vary by role: disk or object storage for recent, fast-restore copies, tape or cloud archive storage for long retention.

What secondary storage means for recovery

For a large estate, restore throughput sets the real recovery time. Restoring 500 TB at a sustained 10 GB/s takes 500,000 GB ÷ 10 GB/s = 50,000 seconds, about 14 hours. At 2 GB/s it takes about 69 hours, close to three days. Secondary storage sized only for cheap capacity and backup ingest can turn a recoverable incident into a multi-day outage, which is why restore performance under concurrent load is part of its design.

Integrity under attack is the second consequence. Attackers target backups before they encrypt production, so copies reachable with production credentials survive only as long as those credentials do. S3 Object Lock changes that only in compliance mode: governance-mode retention can be bypassed by any identity holding s3:BypassGovernanceRetention that sends the bypass header, while compliance mode resists all users, including the account root, until the retain-until date. Object Lock requires versioning, and retention expires.

A third shift is reuse. Large organizations increasingly consolidate backup, archive and replica data onto one object platform and read it for analytics, e-discovery and AI work, so secondary data is less cold than its name implies.

Placement completes the picture. A secondary copy in the same data centre, on the same network and under the same identity directory as production shares many of production's failure modes. Copies held at a second site, under a separate administrative domain, or both, are the ones that survive a site loss or a compromised directory. At enterprise scale, the data reduction arithmetic above is what makes keeping those extra copies affordable: the difference between 3 PB and 158 TB per protected data set decides how many copies, sites and retention periods the budget allows.

Scality and secondary storage

ARTESCA is Scality's software-defined S3 object storage for backup, deployed on standard servers or as a hardware appliance and validated to 8.5 PB. It supports S3 Object Lock in governance and compliance modes, retention periods and legal holds, applies S3 Lifecycle rules per bucket, and has published compatibility with backup applications including Veeam, Commvault, Cohesity, Rubrik, HYCU, Veritas NetBackup and Zerto. For larger estates, Scality RING supports the same Object Lock modes, retention periods and legal holds on software-defined object and file storage that scales to 300 billion objects.