Glossary
Cloud storage tiering
Cloud storage tiering is the movement of data between storage classes of different price and performance, based on how often it is read. It operates within a cloud service, between classes such as standard and archive, or between on-premises storage and the cloud.
Why cloud storage tiering matters for large datasets
Most data is written once, read heavily for a short time and then rarely touched. At petabyte scale, keeping all of it on the fastest and most expensive class wastes money on data nobody reads, and moving all of it to the cheapest class turns every read into a delay and a retrieval charge. Tiering is the mechanism that matches each object to the cheapest class whose access characteristics still fit how it is used. Because the price difference between top and bottom classes is large, getting the placement right is often the single biggest lever on a storage bill.
Storage classes and what each trades
| Class | Capacity price | Retrieval charge | Time to first byte | Minimum duration |
|---|---|---|---|---|
| Standard | Highest | None | Milliseconds | None |
| Infrequent access | Lower | Per GB read | Milliseconds | About a month |
| Archive, instant retrieval | Lower still | Higher per GB | Milliseconds | About three months |
| Archive, delayed retrieval | Low | Per GB, by speed | Minutes to hours, after a restore request | About three months |
| Deep archive | Lowest | Per GB | Hours | About six months |
The upper rows correspond to hot storage and the lower ones to cold storage and cloud archive storage. Each step down lowers the capacity price and raises some other cost.
Rule-based and access-based tiering
Lifecycle rules attach to a bucket or prefix and name a condition, usually age since creation, and an action: transition to a class or expire. A common pattern moves objects to infrequent access at 30 days, to archive at 180 days, and deletes them at the end of their retention period. Age-based rules cannot see reads, so an old object still in daily use moves with the rest. The wider discipline is storage lifecycle management.
Access-based tiering tracks the last read of each object and moves it down after a period without access, and back up when it is read again. Providers charge a per-object monitoring fee for this, and very small objects are usually excluded, which matters for datasets made of millions of small files.
Objects in delayed-retrieval archive classes stay listed, but a read returns an error until a restore request has completed, after which a temporary readable copy exists for a set number of days. Applications that read archived data are written to issue the restore and wait.
Tiering also crosses environments. An on-premises system can keep active data locally and transition ageing data to a cloud class, keeping the metadata needed to find it, one of the patterns of cloud-integrated storage. Reads of that data then add network transfer charges to the provider's retrieval charge.
What cloud storage tiering means for cost and recovery at scale
The saving depends on read volume. With hypothetical rates of $0.023 per GB-month for standard, $0.0125 for infrequent access and $0.01 per GB retrieved, moving 1 PB (1,000,000 GB) saves 1,000,000 × 0.0105 = $10,500 a month in capacity. Reading back 200 TB of it each month costs 200,000 × 0.01 = $2,000, leaving $8,500. The saving reaches zero only when the whole petabyte is read back slightly more than once a month, but archive classes reach their break-even far sooner, because retrieval and request prices are higher.
Minimum durations punish churn. Data deleted or overwritten before the minimum period is billed for the remainder, so workloads that rewrite data frequently gain little from cold classes and can end up paying more than if they had stayed in standard.
Object count matters as much as capacity. Per-request transition charges and per-object monitoring fees are trivial for large media files and significant for datasets of hundreds of millions of small objects, where the one-time cost of moving the data can outweigh a year of capacity savings.
Restore time becomes recovery time. When backup or disaster-recovery copies sit in a delayed-retrieval class, the hours needed to restore them add directly to the recovery time of anything that depends on them. Backup software handles this by keeping recent restore points on fast storage and aging older ones down, the arrangement described under backup storage tiers.
Cold data does not always stay cold. AI and analytics projects regularly reactivate years of archived content as training or retrieval corpora. A dataset tiered to deep archive on the assumption it would never be read again then has to be restored, at retrieval cost and over days, before a GPU cluster can use it.
Tiering in Scality storage
Scality ADI places "hot, warm, and cold data" on suitable media "under one namespace and one policy-driven lifecycle". Scality's Azure partner page describes tiering from on-premises Scality storage to Azure Blob, where tiered data remains accessible to Azure services. ARTESCA applies S3 Lifecycle rules per bucket for backup data.














