Glossary

Cold storage

Cold storage holds data that is rarely read after it is written but is retained, often for years, for compliance, reference or later reuse. It trades access speed and access cost for the lowest price per stored terabyte.

Why cold storage matters for growing data estates

Data accumulates far faster than it is deleted. Completed projects, closed financial records, medical images past active care, surveillance footage, instrument output and older backups pile up year after year while only recent data stays in use. In most large estates cold data makes up the majority of stored capacity, so its price per terabyte, and the power and floor space it consumes, set the long-run cost of keeping data at all.

Cold describes access frequency, and age is a separate attribute. Some old data stays hot, and some new data, such as a compliance copy, is cold from the moment it lands.

Online and offline cold storage

  • Online cold storage keeps data on media that serve a request directly, typically high-capacity hard drives in object storage. Reads take milliseconds, slower than flash but with no restore step.
  • Offline or nearline cold storage keeps data on media that are loaded or staged before a read: tape cartridges, optical media, or drives powered down until needed. A read first requests a restore, waits for staging, then reads the staged copy.

The two differ in time to first byte by several orders of magnitude. Applications reading from offline tiers are built around an asynchronous request, wait and read sequence, which many analytics and AI tools were never designed for.

Media and archive classes

MediumAccess modelTime to first byteNotes
High-capacity hard drivesOnlineMillisecondsLowest cost per terabyte among online media
Spun-down hard drivesNearlineSeconds, to spin upLess power for rarely read data
TapeOfflineTens of seconds to minutesHigh streaming rate once positioned
Cloud archive classesEither, by classMilliseconds to as long as 48 hoursMinimum storage durations of 90 or 180 days are common

Power is part of the price. Capacity disk spinning around the clock draws power whether or not anything is read, which is why spun-down disk and tape remain in use for the deepest tiers despite slower access.

Long-term retention in public cloud archive classes is covered in cloud archive storage.

Cost arithmetic for cold data

Cold pricing has three parts: capacity per month, retrieval per gigabyte or request, and a minimum storage duration. The minimum changes the price of short-lived data. An object deleted after 30 days from a class with a 90-day minimum is charged for 90 days, three times the capacity it used.

A break-even follows from the two prices. If a cold class saves 3 cost units per terabyte per month over a warm class and charges 10 units per terabyte retrieved, data read in full about once every 3.3 months (10 ÷ 3) costs the same on both. Data read less often is cheaper cold; data read more often is cheaper warm. Retrieval cost scales with the amount read: for data read once a decade it is small next to the capacity saving, and for data read monthly it can exceed it.

What cold storage means for long-term retention at scale

For an organization keeping petabytes for a decade or more, the cost that surprises is usually access. Cold pricing assumes data stays put. A legal discovery request, an audit or an AI project that wants years of history reverses that assumption: in a cloud archive class, restoring 100 TB brings retrieval charges on all of it, staging time measured in hours, and egress fees if the data leaves the provider. Online cold tiers on capacity disk avoid the restore step and per-gigabyte charges, at a higher capacity price than tape or deep archive.

Retention periods outlive hardware. Data kept for twenty years spans several generations of drives and servers, so cold storage relies on fixity checks that detect silent corruption, repair from replicas or erasure-coded fragments, and migration to new media before old media leaves support. A platform that refreshes hardware underneath data in place avoids the bulk migrations that otherwise recur every few years.

Cold data is also a ransomware target, because it is often the last clean copy. S3 Object Lock in compliance mode stops any user, including the account root, from deleting or overwriting a locked version until its retain-until date. Governance mode can be bypassed by any identity holding s3:BypassGovernanceRetention that sends x-amz-bypass-governance-retention:true. Object Lock requires versioning, and protection ends when retention expires.

Scality RING as cold storage

Within Scality's ADI, cold data sits on capacity media under the same namespace and policy-driven lifecycle as hot and warm data, with its power draw visible alongside the rest of the system. The underlying RING stores object and file data on standard x86 servers and supports S3 Object Lock in governance and compliance modes, with retention periods and legal holds, so a cold dataset kept there for a decade carries a defined retention window for that decade.