Glossary
Storage performance testing
Storage performance testing is the controlled measurement of a storage system's latency, IOPS and throughput under defined workloads, carried out to characterise the system, confirm that it meets a requirement or detect change after an upgrade or reconfiguration.
Testing concerns one environment and the questions asked about it. Comparing different systems under common rules is the separate discipline of storage benchmarking.
Why performance testing matters for large deployments
A multi-petabyte platform is chosen on published figures, but it is accepted, expanded and upgraded on tests. Testing is how a team finds out whether a new cluster sustains the ingest rate a backup estate needs, whether a software release changed tail latency, or whether losing a node drops throughput below what an AI pipeline consumes. The stakes are high because a platform carrying a dozen applications cannot be swapped out easily after go-live.
Types of performance test
| Test type | What it establishes |
|---|---|
| Baseline | Reference figures for a known configuration and workload, used to judge every later result |
| Load | Behaviour at the expected production load and concurrency |
| Stress | The point where latency climbs sharply or throughput stops rising, and the behaviour beyond it |
| Soak | Stability over hours or days, exposing garbage collection, cache exhaustion and resource leaks |
| Failure mode | Performance with a drive, node or link down and while redundancy is rebuilt |
| Scalability | How results change as clients, nodes or capacity are added |
| Regression | Whether a software, firmware or configuration change moved results relative to the baseline |
Defining the workload
A result means something only in relation to the workload that produced it. A workload definition fixes the object or I/O size (or a size distribution), the read/write mix, the access pattern, concurrency (clients, threads and outstanding requests), the working set relative to cache, the data content and the duration, including any warm-up period left out of the measurement.
Two distortions recur. A working set that fits in cache measures the cache, so a test on a few terabytes says little about a platform that will hold several petabytes. Fresh flash has no garbage collection to perform, so short runs on new drives report figures that sustained use does not reproduce. Preconditioned media and runs long enough to reach steady state remove both effects.
Load generation and measurement
A closed-loop generator issues a request, waits for the reply, then issues the next, so its offered load drops whenever the system slows. An open-loop generator issues requests on a schedule, as independent users do. During a two-second stall at 1,000 requests per second, a closed-loop tool records a handful of slow requests, while 2,000 scheduled requests would have waited. Leaving them out, an effect known as coordinated omission, understates tail latency.
Results are distributions. Percentiles such as p50, p99 and p99.9 come from full latency histograms, merged across clients before percentiles are computed. Repeated runs show run-to-run variation, and a difference between two configurations smaller than that variation is indistinguishable from noise. The load generator can also be the bottleneck: a host with one 100 Gb/s link tops out near 12.5 GB/s, so measuring a cluster that sustains far more takes many clients working in parallel.
What performance testing means for platform teams
For teams running shared storage at scale, the value of testing lies less in a peak number than in how the platform behaves in the conditions it will actually meet. Tests built from the real estate expose the gap between datasheet figures and the second month of production, and they turn storage performance from a quoted figure into a measured property of that specific platform.
- Object size distribution decides the outcome. Backup data written as large objects and an AI corpus of small images produce operation rates that differ by orders of magnitude on identical hardware.
- Failure-mode results describe the platform's worst ordinary day: throughput and p99 latency with a node down and a rebuild running.
- Scalability results show whether performance grows with nodes. A cluster that scales evenly from 6 to 12 nodes keeps capacity and performance planning a single exercise; one that flattens creates a separate performance plan.
- Regression results protect consumers who never read the change log. A release that adds 20% to p99 on small PUTs shows up in their jobs first.
- Multi-site results include the inter-site link, because synchronous writes pay its round trip on every operation.
Scality and performance measurement
Scality publishes measured latencies of 511 microseconds for a GET and 741 microseconds for a PUT on Scality RING. Those figures describe the configuration on which they were taken; a test against a specific estate's object sizes and concurrency answers a different question.
For stretched RING deployments, Scality specifies 10 Gb/s or greater bandwidth and less than 5 ms latency between sites, the network conditions under which a multi-site test reflects a supported layout. Storage performance tuning covers the configuration changes that testing is used to evaluate.














