Glossary

Storage benchmarking

Storage benchmarking is the practice of running a defined, repeatable workload against storage systems under fixed rules so that their results can be compared on a common basis. The workload, the run rules and the reporting format stay constant, and the system under test is the only variable.

Run rules, disclosure and what inflates a result

A benchmark fixes five things: the workload specification (operation mix, sizes, access patterns, dataset), run rules (warm-up, duration, dataset size relative to the system, permitted tuning, repetitions), the metrics reported, full disclosure of hardware and software, and in some suites the system price for price-performance figures. Holding those constant lets results from different vendors share one table. That is also why benchmarks shape shortlists: published results are often the first numbers an architecture team lines up, well before any proof of concept.

The frame still leaves room for distortion. A dataset that fits in controller memory measures cache. Highly compressible or duplicate data inflates systems with inline reduction. Disclosed configurations often carry more drives, more flash and lighter protection than typical deployments, and short runs on fresh flash dodge the garbage collection of sustained use. A maximum quoted without response time hides the load curve, which is why some suites report throughput only under a stated latency ceiling. Vendor tests outside any suite lack shared rules, so dataset size, client count or protection level can explain much of a reported gap between two systems.

SPC, SPECstorage, IO500 and MLPerf Storage

Load reaches the system from one of three sources: a synthetic generator with set size, mix and depth that isolates one variable at a time; a real application, whose results depend on how it is tuned; or replayed production traces, which reflect only the environment they were recorded in. The recognised suites each model a narrow slice. SPC-1 represents random, transaction-style block I/O and SPC-2 large sequential transfers. SPECstorage Solution runs application-based workloads such as genomics, electronic design and AI image processing on file systems. IO500 ranks high-performance computing file systems on bandwidth and metadata phases. MLPerf Storage measures whether storage keeps emulated AI accelerators busy during training, and a result counts only above a utilisation threshold set per workload.

MLPerf Storage also covers checkpointing, where training jobs periodically write full model state and read it back after a failure. Checkpoint time is state size divided by achieved write bandwidth: 1 TB at 20 GB/s takes 50 seconds, during which accelerators in a synchronous checkpoint sit idle. AI storage and GPU storage describe the requirements these workloads model. Object storage at petabyte scale has no dominant standard suite, so comparisons there rest mostly on vendor-published figures and on tests run with the buyer's own object-size mix.

Storage performance testing on the platform being bought

Storage performance testing measures one environment against the questions asked of it, giving up comparability in return. A multi-petabyte platform is chosen on published figures but accepted, expanded and upgraded on tests, and once a dozen applications depend on it, replacing it is slow and expensive. Tests fall into recognisable kinds:

  • Baseline: reference figures for a known configuration, against which later results are judged.
  • Load: behaviour at expected production load and concurrency.
  • Stress: the point where latency climbs sharply or throughput stops rising.
  • Soak: stability over hours or days, exposing garbage collection, cache exhaustion and leaks.
  • Failure mode: performance with a drive, node or link down and during the rebuild.
  • Scalability: how results change as clients, nodes or capacity are added.
  • Regression: whether a software, firmware or configuration change moved results.

A result means something only relative to its workload definition: object or I/O size distribution, read/write mix, access pattern, client and thread counts, outstanding requests, working set relative to cache, data content and duration. For platform teams the size distribution often settles the outcome, since backup data written as large objects and an AI corpus of small images produce operation rates orders of magnitude apart on identical hardware. Failure-mode runs describe the worst ordinary day. A cluster that scales evenly from 6 to 12 nodes keeps capacity and performance planning as one exercise, while one that flattens forces a second plan. A release adding 20 percent to p99 on small PUTs surfaces first in the jobs of consumers who never read the change log.

Coordinated omission and load-generator limits

A closed-loop generator sends a request, waits for the reply and sends the next, so its offered load drops whenever the system slows. An open-loop generator sends on a schedule, as independent users do. During a two-second stall at 1,000 requests a second, a closed-loop tool records a handful of slow requests while 2,000 scheduled ones would have waited. Leaving those out, an effect called coordinated omission, understates tail latency.

Percentiles such as p99 and p99.9 come from full latency histograms merged across clients before percentiles are computed, and a gap between two configurations smaller than run-to-run variation cannot be told apart from noise. The generator can be the bottleneck too: a host with one 100 Gb/s link tops out near 12.5 GB/s, so measuring a cluster that sustains hundreds of gigabytes per second takes many clients working in parallel. Storage performance tuning covers the configuration changes such tests are used to evaluate.

Scality RING and storage benchmarking

Scality publishes RING measurements tied to described configurations. Its hard-drive validation, just over 100 nodes across three availability zones with flash for metadata, reported reads and writes sustained over a two-hour window with 10 MB objects, the kind of soak result a short peak run cannot show (Solved by Scality).

The all-NVMe RING XP build reports per-request time instead, 511 µs for a GET and 741 µs for a PUT on 4 KB objects (Solved by Scality). Read against a planned estate, the two answer separate questions: aggregate sustained rate on large objects, and single-request latency on small ones.