Glossary
Hot storage
Hot storage is the storage tier for data that applications read and write constantly, built on media and systems chosen for low latency and high operation rates. It sits at the top of the temperature scale used in storage tiering, above warm and cold storage.
Why hot storage matters for performance-sensitive workloads
Hot storage is the most expensive capacity in an estate and the capacity applications wait on. Transaction commits, virtual machine boots, index lookups and training batches all read from it, so its latency and throughput appear directly in application response times and, for AI infrastructure, in how long accelerators sit idle waiting for data. A hot tier that is too small slows the business; one that is too large spends flash budget on data nobody is reading.
Characteristics of hot data
Data is hot when it is read or written many times an hour, when applications wait on each operation, when access is random across a large address space, or when many clients reach it at once. The hot portion is usually a small share of the whole. A 2 PB estate in which 5% of the data is touched in a given week has a working set of about 100 TB, and sizing the hot tier to that working set, with the remainder on capacity media, is the basis of tiered design. Working sets also turn over: the 100 TB touched this week overlaps only partly with last month's, so a hot tier sized exactly to one week runs short the next.
Media and workload profiles
| Medium | Typical read latency | Role |
|---|---|---|
| DRAM | About 100 ns | Volatile cache in front of persistent media |
| NVMe flash | 50 to 100 µs device read | The common medium for hot tiers |
| SAS or SATA flash | Similar NAND latency, more interface overhead | Lower parallelism and bandwidth per device |
| Hard drives | 5 to 15 ms random read | Hot only for large sequential streams |
Flash dominates because a 7,200 RPM drive waits about 4.17 ms on average for the sector to rotate under the head, which limits it to fewer than two hundred random reads per second, while a single NVMe device completes hundreds of thousands. Hard drives still serve hot workloads that stream large objects sequentially, such as media delivery or analytics scans, where throughput governs performance; the distinction is covered in sequential vs random I/O and storage latency.
Hot workloads split into two profiles. Databases, virtual machines and search indexes are sensitive to small-block IOPS and to latency at the 99th percentile. AI training, content serving and analytics are sensitive to aggregate throughput across many parallel streams. A hot tier specified for one profile can underperform on the other, so hot storage is specified by workload as well as by capacity.
Cost structure of hot storage
| Cost element | Hot storage | Cold storage |
|---|---|---|
| Capacity per terabyte | Highest | Lowest |
| Per-request or retrieval charge | Low or none | Higher, sometimes per gigabyte read |
| Minimum storage duration | None | Often 90 or 180 days in cloud archive classes |
| Time to first byte | Microseconds to milliseconds | Milliseconds to hours |
The two ends invert the balance between capacity price and access price. Data read many times costs less on hot storage despite the higher price per terabyte, because the access charges and delays of a cold tier accumulate with every read. In public clouds, hot classes carry no retrieval charge and no minimum duration, which makes them the default for data whose access pattern is unknown.
What hot storage means for AI and enterprise infrastructure
For teams running GPU clusters, the hot tier is sized from the accelerators backward. The read throughput that keeps every GPU busy, plus the write bursts from checkpoints, sets its performance floor, and capacity follows from the portion of the dataset in active use. A tier that meets capacity targets and misses on throughput leaves accelerators idle, which costs more per hour than the storage itself. The wider design is covered under GPU storage.
Hot storage design also depends on how quickly data cools. Most records are hottest just after creation, but some data reheats: an archived project reopened for audit, or historical data pulled into a model training run. Policies that demote quickly keep the hot tier small and pay recall costs when data reheats; policies that demote slowly pay for flash holding idle data. The policy decides the bill more than the media does.
Protection works differently on data that changes continuously. Frequent snapshots, replication and backups taken from snapshots keep copies consistent while applications keep writing, and the replication link carries the full change rate of the hot tier, which sets its bandwidth.
Hot storage on Scality platforms
In Scality's ADI, hot data is one of three temperatures placed on matching media under one namespace and policy-driven lifecycle. For the hot end of that range, Scality has published measured RING latencies of 511 microseconds for a GET and 741 microseconds for a PUT on 4 KB objects with NVMe drives over a 100 GbE network. Those figures belong to that configuration, and other media and networks produce other numbers.














