Glossary
Flash storage latency
Flash storage latency is the elapsed time for a read or write on flash media, from the moment the host submits the command to the moment it receives the completion. It combines the time NAND cells take to be read or programmed with time spent in the drive controller, on the host interface and in queues.
Network and protocol contributions belong to the wider measurement of storage latency.
Why flash latency matters beyond the drive
Moving from disk to flash cut device time from milliseconds to tens of microseconds, and in doing so moved the bottleneck. In a networked storage system the drive is now often the smallest share of a request's total time. What still traces back to the drive is variation: flash is fast on average and occasionally much slower, and in large systems those occasional slow responses decide how applications behave under load.
NAND operations and their timing
| Operation | Unit | Typical duration |
|---|---|---|
| Read (sense a page) | Page, typically 16 KB | Tens of microseconds; longer on QLC |
| Program (write a page) | Page | Hundreds of microseconds to over a millisecond |
| Erase | Block of many pages | Several milliseconds |
A 4 KB read on an idle NVMe drive follows a short path: the host places a command in a submission queue, the controller looks up the physical location in its mapping table, the die senses the page, error correction decodes it, and the data is copied to host memory before a completion is posted. End to end this takes roughly 60 to 100 microseconds.
Reads, programs and erases share the same dies, so a die busy programming or erasing delays any read addressed to it. Enterprise drives can suspend a program or erase to let a read through, which shortens the wait without removing it.
Write latency and buffering
Enterprise drives acknowledge a write once the data is in DRAM protected by power-loss capacitors and program the NAND afterwards. A small write can therefore complete faster than a small read. The buffer is finite: under sustained writing it fills, and write latency rises to the pace at which NAND can be programmed. QLC drives add a second stage, an area of flash run in single-bit mode that absorbs writes quickly until it fills.
Tail latency: garbage collection, read retry and contention
- Garbage collection: the drive relocates valid pages and erases blocks to make room. A read queued behind an erase can wait milliseconds.
- Read retry: on worn or long-stored cells, error correction may fail on the first pass and the page is re-read with shifted reference voltages, adding a read time for each retry.
- Die contention: two requests for the same die are served one after the other.
- Host queueing: as queue depth rises, requests wait longer before the drive sees them.
The average stays low while the 99th and 99.9th percentiles climb, which is why flash latency is described by percentiles as well as a mean.
What flash latency means for distributed storage and AI pipelines
Drive time is a minority of what clients see. A network round trip inside a data centre, request handling and the software that locates data typically add more than the NAND read. Comparing systems by drive latency says little about their end-to-end response.
Fan-out multiplies the tail. A request that waits for several drives finishes with the slowest of them. If each drive answers slowly 1% of the time, a request touching ten drives meets at least one slow answer 1 − 0.9910 ≈ 9.6% of the time. Erasure-coded reads and distributed metadata lookups are fan-out patterns, so drive tail behaviour becomes cluster tail behaviour.
Writes disturb reads. Heavy ingest triggers garbage collection on the same drives that serve latency-sensitive reads; an AI cluster writing checkpoints while data loaders fetch training samples sees read tails widen during each checkpoint.
Media choice moves the tail more than the median. QLC's longer program times and more frequent read retry widen its distribution relative to TLC, so two flash tiers with similar median latency can differ substantially at the 99.9th percentile.
Latency caps single-stream throughput. A reader that waits for each 4 KB object before requesting the next, at 500 microseconds per request, completes 2,000 requests a second and moves about 8 MB/s, however fast the cluster is. Small-object pipelines reach high throughput only by keeping many requests in flight, which is why data loaders run many parallel workers.
The slowest fetch sets the pace. In synchronous AI training, a step waits for every worker's batch, so the time GPUs spend idle tracks the slow end of the latency distribution far more closely than its average.
Flash latency in Scality RING XP
Scality's measured figures for RING XP, its all-flash configuration on NVMe servers, are 511 microseconds per GET and 741 microseconds per PUT for 4 KB objects. These are complete object operations over the object API, and each one includes the network, request handling and the software that locates or places data on top of the drive time described above.














