Glossary
Flash storage IOPS
Flash storage IOPS (input/output operations per second) is the number of read or write operations a flash drive or flash-based system completes each second. The figure is meaningful only with its test conditions: block size, queue depth, read/write mix and whether the drive was in steady state.
The general metric across all media is covered under IOPS.
Why flash IOPS matter for large-scale workloads
Many of the heaviest workloads in a large estate are counted in operations rather than in bytes: billions of small objects, metadata lookups, image datasets with files of a few hundred kilobytes, log ingestion, and the listing and HEAD requests that precede reads. For these, operations per second set how long a job takes and how many clients one system can serve. A single NVMe drive can be rated for hundreds of thousands to over a million random reads per second, a rate that once needed thousands of hard drives. At scale the open question is how much of that capability reaches the application.
Conditions behind an IOPS figure
| Condition | Effect on the figure |
|---|---|
| Block size | Usually 4 KB. Larger blocks give fewer operations and more bytes per second. |
| Queue depth | More requests in flight raise IOPS until the drive saturates, then raise only latency. |
| Read/write mix | Random writes rate well below random reads; mixed ratings such as 70/30 sit between. |
| Steady state | A fresh drive writes faster than one that has been filled and is garbage collecting. |
| Fill level | A nearly full drive has less free space for garbage collection, which lowers sustained write IOPS. |
| Random or sequential | Datasheet IOPS are random; sequential access is usually quoted as throughput. |
Parallelism and Little's law
An SSD contains many NAND dies spread across several channels, and each die handles one operation at a time. High IOPS come from keeping many dies busy at once, which requires many outstanding requests. NVMe was built for this, with deep, parallel command queues.
Little's law links the quantities: IOPS = requests in flight ÷ latency. At 100 microseconds per read, one request at a time yields 1 ÷ 0.0001 = 10,000 IOPS. Thirty-two requests in flight at the same latency yield 320,000. Past saturation, latency climbs as the queue grows and IOPS stop rising, so very high IOPS figures and very low latency figures on a datasheet usually come from different tests.
Write IOPS and IOPS per terabyte
Random write IOPS are lower than read IOPS because programming NAND takes longer than reading it and because garbage collection competes with host writes once a drive is full.
Small operations saturate on count before bandwidth. At 4 KB, 1,000,000 IOPS moves only about 4 GB/s, well below what an NVMe link carries, so small-object workloads exhaust a drive's operation rate long before its throughput.
Capacity has also grown faster than IOPS. Two hypothetical drives each rated at 1,000,000 random read IOPS deliver very different densities: a 7.68 TB drive gives about 130,000 IOPS per terabyte, a 61.44 TB drive about 16,300. A hard drive with around 200 random IOPS on 20 TB gives 10. The same dataset on fewer, larger drives has fewer operations available per terabyte of data.
What flash IOPS mean for small-object and AI workloads
Drives are rarely the limit. A server holding 24 NVMe drives has a theoretical aggregate in the tens of millions of IOPS, far beyond what its processors, network ports and storage software can handle. Sizing a system by adding up drive datasheets overstates what clients will see by a wide margin.
Object operations are heavier than block operations. One object GET involves a metadata lookup and one or more data reads, and with erasure coding a single read can touch drives in several servers. The number that describes a cluster is object operations per second at a stated object size, and it sits well below the sum of its drive IOPS.
Concurrency carries the load. A single client issuing one request at a time is bound by latency regardless of how fast the drives are. Scale-out systems reach high operation rates through many clients and many requests in flight, spread across many servers, which is how AI data loaders with many parallel workers are built.
Write-heavy small-object services meet the lower write figures. Log and telemetry ingest, thumbnail generation and event streams arrive as large numbers of small PUTs, and each PUT also writes metadata and protection fragments, so the operations reaching drives are a multiple of the operations clients send.
Density counts on capacity flash. A dataset accessed by many concurrent jobs places a request load per terabyte, and on very large QLC drives that load is shared by fewer devices.
Flash IOPS and Scality RING
For its all-flash RING XP configuration, Scality publishes latency figures of 511 microseconds per GET and 741 microseconds per PUT on 4 KB objects, measured over the object API. By Little's law, one client issuing one GET at a time completes about 1 ÷ 0.000511 ≈ 1,957 requests per second. Rates beyond that come from concurrency across clients and the servers of the cluster.














