Glossary
NVMe performance
NVMe performance describes how quickly an NVMe drive or NVMe-based system completes work, expressed as latency per operation, operations per second (IOPS) and data rate (throughput).
The flash media, the drive controller, the PCIe link and the host software set these figures together, and each published number holds only for the conditions under which it was measured.
Why NVMe performance matters for infrastructure teams
NVMe datasheets quote figures in the millions of IOPS and many gigabytes per second per drive. Those numbers are real, but each comes from a different test, and none describes a storage cluster serving a mixed production load. Teams sizing flash for databases, analytics or AI pipelines need to know which part of the stack actually limits the result, because buying faster drives only helps when the drive is the limit.
How latency, IOPS and throughput relate
Two relationships connect the three measures. By Little's law, operations in flight = IOPS × latency, so IOPS = queue depth ÷ latency. And throughput = IOPS × operation size. Storage latency, IOPS and storage throughput each cover one measure in depth.
A flash drive has many dies that work in parallel, so it only reaches its rated IOPS when the host keeps enough requests in flight. With illustrative values:
| Queue depth | Latency per operation | IOPS |
|---|---|---|
| 1 | 80 µs | 12,500 |
| 8 | 90 µs | ≈ 88,900 |
| 32 | 120 µs | ≈ 266,700 |
| 128 | 400 µs | 320,000 |
At low depth, extra requests are absorbed by idle dies and IOPS climbs almost in proportion. Past saturation, requests only wait: going from 32 to 128 adds about 20 percent more IOPS and more than triples latency. Headline figures quoted at a queue depth of 128 or 256 describe that saturated region, while many applications run at far lower depths.
Block size, writes and steady state
Small operations are limited by how many commands the controller and flash can process; large ones by the PCIe link. A PCIe 4.0 ×4 link carries about 7.88 GB/s, which caps 1 MiB operations at roughly 7,500 per second, while 4 KiB operations hit controller limits long before the link.
Writes behave differently from reads. Programming flash is slower than reading it, and under sustained writes the drive also runs garbage collection to free blocks. A new or freshly erased drive has no garbage to collect and posts figures it cannot sustain once full. Industry test specifications therefore precondition a drive until results reach a steady state before measuring. Mixed workloads add a further effect: a read that reaches a die busy with a program or erase waits for it, so the slowest 1 percent of reads can be ten times slower than the average.
The host software path
Once device latency falls below about 100 µs, host software becomes a visible share of every operation. The parts that matter most are:
- Queue placement: one queue pair per core avoids locking, if the application spreads I/O across cores.
- Completion handling: interrupts cost a context switch; polling saves it at the price of a busy core.
- File systems and caches: each layer adds copying, locking and metadata work.
- NUMA placement: I/O issued from the socket that does not own the drive crosses the inter-processor link.
What NVMe performance means for AI and analytics infrastructure
In a distributed storage system, the drive is rarely the slowest stage. A GET to an object store crosses the client's network stack, the network, the storage software that locates and assembles the object, and only then the drive. With NVMe media, the drive is often the smallest of those parts, so the network design and the efficiency of the storage software decide most of the end-to-end result.
For GPU clusters, tail latency matters more than averages. A training step that reads thousands of samples waits for the slowest one, so a system with good average latency and a poor 99th percentile can leave accelerators idle. Write-heavy phases such as checkpointing push drives into garbage collection, which is exactly when read tails lengthen.
Drive figures also do not add up to system figures. Data protection multiplies back-end writes, metadata lookups add reads, and the network caps what each server can export. Sizing that starts from steady-state, mixed-workload behaviour at realistic queue depths, measured at the protocol the applications use, predicts production far better than a sum of datasheet maxima. Storage benchmarking covers how such tests are built.
NVMe performance in Scality RING
Scality measured 511 µs per GET and 741 µs per PUT for 4 KB objects on its all-NVMe RING XP configuration with a 100GbE network (Solved by Scality). Those are object operation times at the S3 API, end to end, measured on small objects where fixed per-request costs weigh most.
If the drive's own read takes 50 to 100 µs, roughly 410 to 460 µs of each GET is spent in the network round trip, HTTP handling and the RING software path, which is where an all-flash object store earns or loses its latency.














