Glossary

Storage quality of service (QoS)

Storage quality of service (QoS) is the set of controls a storage system, host or hypervisor applies to divide performance among workloads, by capping, reserving or weighting the IOPS, bandwidth or request rate each one receives.

Workloads that share drives, nodes or network links otherwise compete freely, and the most aggressive consumer sets the latency the others experience.

Why QoS matters on shared storage platforms

Consolidation keeps storage costs under control for large organisations and service providers: one platform, many applications or tenants. It also creates the noisy-neighbour problem. Shared resources have shared queues, and queueing delay scales roughly as 1 ÷ (1 − utilisation). A device kept 40% busy by tenant A gives A a response time of about 1.67 times service time. When tenant B adds another 45%, the device runs at 85% and A's response time becomes about 6.7 times service time, four times worse, although A changed nothing.

QoS puts bounds on that interaction. The effect is strongest in multi-tenant storage, where the competing workloads belong to different organisations or business units with their own expectations.

Types of QoS control

ControlEffectBehaviour when the system is idle
Limit (cap)Maximum IOPS, bandwidth or request rate for a workloadThe cap still applies
Reservation (floor)Minimum performance held for a workloadUnused reservation is lent to others in most designs
Shares (weights)Relative proportion of resources under contentionNo effect until contention occurs
Priority classesOrdering of requests between classesNo effect until contention occurs
Latency targetThrottling of other workloads when a protected workload's latency risesNo effect until the target is exceeded

How QoS is enforced

Rate limits are commonly enforced with a token bucket. Tokens accumulate at the permitted rate up to a maximum depth, each operation spends one, and an operation that finds the bucket empty waits. The depth sets the burst allowance: at 5,000 IOPS with a depth of 50,000 tokens, a workload that has been idle can run at 15,000 IOPS for 50,000 ÷ 10,000 = 5 seconds before falling back to its limit.

Shares divide contended capacity by weight. Weights of 300 and 100 give two busy workloads 75% and 25% of the resource; when one goes idle, the other can use everything, a property called work-conserving. Limits lack that property, since a capped workload stays at its cap on an otherwise idle system.

Enforcement can sit in several places. A host or hypervisor limits per process, container or virtual disk, but sees only its own load. A storage system applies policies per volume, share, bucket or tenant with a view of every client. Object storage commonly limits requests per second per bucket or account and answers excess requests with an HTTP 503 so clients back off and retry.

Inside a storage system the scheduler sits ahead of the device queues, because once a request reaches a drive it is served in the drive's order regardless of whose it is. Latency-target schemes work indirectly: rate controls act on requests admitted, while latency is an outcome of everything sharing the system, so these schemes throttle lower-priority traffic until the protected workload returns within target. Storage performance and storage performance tuning cover the factors behind that outcome.

What QoS means for multi-tenant and service-provider storage

For a service provider or a central platform team, QoS turns shared capacity into a service with a predictable level. Without it, the performance any tenant receives depends on whatever every other tenant happens to be doing.

  • Caps protect neighbours at the price of idle headroom. A capped tenant cannot use spare performance at night, so strict limits leave capacity unused that weights would hand out.
  • Shares protect relative position only. If total demand doubles on fixed hardware, every tenant's absolute performance halves while its share stays the same, which makes shares a contention policy more than a service level.
  • Background work behaves like a tenant. Rebuilds, replication and lifecycle jobs need a place in the scheduling policy or they either starve or starve everyone else.
  • Tying performance to capacity simplifies chargeback. At 10 IOPS per gigabyte, a 2 TB volume receives 20,000 IOPS and a 500 GB volume 5,000, so performance density stays constant as tenants grow.
  • At object scale, request rate is often the more useful lever. Many small requests exhaust metadata capacity long before they exhaust disk bandwidth, so a bandwidth cap alone leaves the platform exposed.

Multi-tenancy in Scality RING

Scality RING is software-defined object and file storage on standard x86 servers. Scality describes RING as offering high-concurrency multi-tenancy with hard isolation, and states that each dimension of the system can be scaled without paying for the others.

Isolation and performance allocation operate at different layers of a shared platform. Headroom added by scaling reduces how often allocation controls have to arbitrate at all.