Glossary

Storage high availability

Storage high availability (HA) is the ability of a storage system to keep data readable and writable through component failures, maintenance and upgrades. It is achieved by removing single points of failure and recovering from faults automatically, and it is measured as the share of time the system can serve requests.

Why high availability matters at petabyte scale

At small scale, a failed drive is an event. At petabyte scale it is routine. If each drive has a 1% chance of failing in a year, a system with 5,000 drives loses about 50 of them a year, close to one a week, before counting servers, power supplies, network ports and software upgrades. Failure stops being an exception and becomes a normal operating condition the system absorbs.

Availability also compounds upward. Backup jobs, analytics, AI training runs and customer-facing applications often sit on the same storage, so an hour without access stops all of them at once. For a service provider, that hour is also an SLA breach multiplied across every tenant.

Availability, durability and disaster recovery

PropertyWhat it describesUsual measure
AvailabilityWhether data can be read and written right nowPercentage of time in service, or downtime per year
DurabilityWhether stored data survives over timeProbability of losing data in a year
Disaster recoveryWhether service resumes after losing a siteRecovery time and recovery point objectives

The three overlap but do not substitute for one another. Data in an offsite tape vault is durable and slow to reach. A system that stays online while silently returning corrupted data is available and still losing data.

Availability is usually written as MTBF ÷ (MTBF + MTTR): mean time between failures over total time. Shortening repair time improves it as much as lengthening time between failures, which is why automatic failover, cutting effective repair time to seconds, sits at the center of HA design. At 99.99% availability, a year of 8,760 hours allows about 53 minutes of downtime.

Redundancy and failure domains

Redundant parts protect each other as long as they fail independently. Two servers each available 99% of the time give 1 − 0.01² = 99.99% when either can serve a request. Components in series work the opposite way: a request that passes through three components at 99.9% each succeeds 0.999³ ≈ 99.7% of the time, lower than any one of them.

The independence assumption is where many designs fall short. A shared power feed, one top-of-rack switch, a firmware defect on every node or an operator error pushed to the whole cluster removes all the redundant copies together. HA engineering therefore groups hardware into failure domains (drive, server, rack, room, site) and places replicas or erasure-coded fragments so that losing any one domain still leaves enough data to answer every request.

A cluster also needs a rule for who continues when members lose contact. Without one, both halves accept writes and diverge, a condition called split-brain. A majority quorum prevents it: three voting members tolerate one failure, five tolerate two. Two sites alone cannot form a majority when the link between them drops, which is why two-site designs add a witness at a third location.

Degraded operation, rebuilds and upgrades

After a failure the system keeps serving with reduced protection until the lost redundancy is rebuilt. A second failure inside that window can exceed what the protection scheme tolerates, so rebuild time matters as much as the scheme itself. Reconstructing 20 TB at 200 MB/s takes 20 × 10¹² ÷ (200 × 10⁶) = 100,000 seconds, about 28 hours, with rebuild traffic competing against application I/O throughout. Distributed storage designs shorten the window by rebuilding in parallel and by reconstructing only data that was actually written.

Planned work counts as well. Capacity expansion, hardware refresh and software upgrades touch every node over a system's life. Scale-out designs that support rolling upgrades and online expansion keep running through them; designs that do not turn each one into a maintenance window.

What high availability means for large-scale storage teams

  • Failure handling is daily work. With thousands of drives, replacements are routine, and behavior while degraded matters as much as behavior when healthy.
  • Performance during rebuild is part of availability. A system that stays online but slows enough to stall ingest pipelines or backup windows delivers less usable availability than its uptime figure suggests.
  • Protection is only as wide as its domains. Fragments spread across servers in a single rack still go offline together when that rack loses power, so placement tracks how hardware is racked, powered and cabled.
  • Site loss is a separate tier. Surviving a whole site takes either a stretched cluster within tight latency limits or replication to another site with a failover process, and the two recover very differently.
  • Lean teams rely on automation. When a handful of engineers run many petabytes, availability depends on the system detecting faults, failing over and starting rebuilds without waiting for someone to act.

High availability in Scality RING

Scality RING is software-defined object and file storage running on standard x86 servers, with data distributed across the servers in the cluster and erasure coding schemes defined per storage class. When a drive fails, RING writes data across the remaining drives in the server and rebuilds only data that was written, within that server rather than across the whole cluster.

For site-level availability, RING supports two- and three-site stretched deployments with synchronous writes. Given 10Gb/s or greater bandwidth and less than 5ms latency between sites, an entire site plus another server or disk group can fail while read-write access continues.