Glossary
IOPS
IOPS (input/output operations per second) counts how many separate read and write operations a storage device or system completes each second.
An IOPS figure says nothing by itself about how much data moved. It only has meaning alongside the operation size, the read/write mix, the access pattern and the number of operations in flight.
Fewer operations per terabyte as hard drives grow
A hard drive completes roughly 150 random operations a second whether it holds 4 TB or 24 TB, because its heads still travel to every piece of data they read. As capacities climb, the ratio falls from about 37 operations per terabyte to about 6. A 10 PB raw tier built from 20 TB drives has 500 spindles and, at that rate, around 75,000 random operations a second for the whole tier, before any protection overhead.
That budget is ample for large sequential objects such as backup images or video. It is thin for workloads that generate many operations per terabyte: billions of small objects, metadata-heavy file trees, log analytics and AI datasets made of small files. A cluster sized purely on capacity can sit half empty and still be too slow.
Flash storage IOPS and the conditions behind a datasheet figure
Flash storage IOPS, the operation rate of an SSD or flash-based system, comes from parallelism. A drive holds many NAND dies spread over several channels, each handling one operation at a time, and its rated figure appears only when enough requests are outstanding to keep those dies busy. Little's law states the relationship: IOPS = requests in flight ÷ latency. At 100 µs per read, one outstanding request yields 1 ÷ 0.0001 = 10,000 IOPS, and thirty-two outstanding at the same latency yield 320,000. Beyond saturation, extra requests only queue, so the highest IOPS and the lowest latency printed on one datasheet usually come from different runs. Hard drives gain little from deeper queues, since a single set of heads can do no more than reorder its seeks.
A flash rating also hides its test conditions. It is normally taken with 4 KB random operations. Random writes rate well below random reads because programming NAND takes longer than reading it. A fresh drive writes faster than one that has been filled and is running garbage collection, and a nearly full drive has less spare area to collect into, which lowers sustained write IOPS again.
Density works against flash as well. Two drives each rated at 1,000,000 random reads a second give about 130,000 IOPS per terabyte at 7.68 TB and about 16,300 at 61.44 TB. A dataset read by many concurrent jobs spreads its request load over fewer devices when it sits on large QLC drives.
Front-end requests against back-end device operations
The operation count an application sees and the count its media performs are different numbers. In a RAID 5 group one small write becomes four device operations (read old data, read old parity, write new data, write new parity), so ten thousand small writes a second land as forty thousand on the drives. Mirroring doubles writes, and replicated or erasure-coded layouts multiply them across servers, with metadata reads and updates on top.
Object storage counts requests instead: GET, PUT, HEAD, LIST and DELETE applied to whole objects. One GET involves a metadata lookup and one or more data reads, and with erasure coding a single read can touch drives in several servers. Summing drive datasheets therefore overstates a cluster by a wide margin. A server holding 24 NVMe drives has a theoretical aggregate in the tens of millions of IOPS, far more than its processors, ports and storage software can turn into object requests.
Data rate is IOPS multiplied by operation size. At 4 KB, a million operations a second carries only about 4 GB/s, so small-object workloads exhaust an operation budget long before they fill a link, the split described under sequential vs random I/O.
Sizing an object tier on request rate
On a large object platform, request rate per bucket or per application belongs in the plan next to capacity. LIST and HEAD calls from analytics engines and backup catalogues consume operations without moving payload, and they often outnumber reads and writes. Rebuilds after a drive failure draw on the same budget as production traffic, so the headroom left during a rebuild is what the busiest hour actually receives. A lone client is bound by its latency however fast the drives are, which is why high request rates come from many clients with many requests in flight, the way AI data loaders with parallel workers behave. Hybrid designs answer the hard-drive shortfall by holding metadata on flash and keeping disks for capacity.
Scality RING and IOPS
Scality RING handles the per-terabyte shortfall through placement. A validation of just over 100 x86 RING nodes, stretched across three availability zones, kept data on hard drives and put flash only under metadata, and with 4 KB objects it recorded "173,456 S3 PUT/sec and 151,091 S3 GET/sec" (Solved by Scality). At that object size the PUTs carry about 694 MB/s of payload, so the test measured operation handling, most of it on the metadata flash.
Erasure coding schemes in RING are defined per storage class, so the number of back-end device writes behind each small PUT depends on the class the data lands in.














