Glossary

Storage throughput

Storage throughput measures the amount of data a storage device or system actually moves per second, summed across every operation in progress and quoted in MB/s or GB/s.

It is achieved performance, and it always sits at or below the bandwidth of the narrowest component in the data path.

From gigabytes per second to hours of work

For data-heavy work, throughput converts straight into elapsed time. Reading 1 PB at 10 GB/s takes 100,000 seconds, nearly 28 hours; at 100 GB/s it takes under three. That one division sets how long an AI training epoch spends reading, how long a large restore runs, how soon an analytics scan finishes and how long a cluster stays exposed while it rebuilds after a failure.

Capacity growth erodes the figure quietly. Refreshing a cluster with larger drives in the same number of servers raises terabytes per node without raising data rate per node, so the time to read the whole dataset lengthens with every refresh unless servers and ports grow alongside capacity.

Storage bandwidth ceilings from drive link to switch uplink

Storage bandwidth is the fixed maximum rate one link or component can carry, set by its signalling rate and width; throughput is what a workload achieves beneath it. Data crosses a chain of such links (drive interface, server CPU and storage software, network ports, switch fabric, client), and sustained throughput cannot exceed the narrowest. Adding drives to a server whose ports are already full therefore adds capacity and nothing else.

Units cause frequent misreadings. Networks are quoted in bits and storage in bytes, so 100 Gb/s is 12.5 GB/s before overhead and 10 Gb/s is 1.25 GB/s. Encoding and headers take a further share. SATA III signals at 6 Gb/s yet carries 600 MB/s after 8b/10b encoding, and a standard 1,500-byte Ethernet frame delivers about 94.9 percent of the wire rate as TCP payload, which leaves a 100 Gb/s port near 11.9 GB/s. Jumbo frames lift that to about 99 percent.

Shared links are usually provisioned below the sum of what feeds them. A leaf switch with 48 server ports at 25 Gb/s (1,200 Gb/s) and four 100 Gb/s uplinks (400 Gb/s) is oversubscribed 3:1, so the bandwidth any one server gets depends on what its neighbours are doing at that moment.

Fixed per-request cost with small objects

Every object request pays for connection handling, authentication, metadata lookup and a response round trip before payload moves. With an illustrative 500 µs of fixed cost on a 12.5 GB/s path, a single stream of 4 KB objects moves about 8 MB/s, while a stream of 10 MB objects moves about 7.7 GB/s. Small-object workloads are governed by request rate, measured as IOPS, and large-object workloads by transfer rate.

One stream rarely fills a system. Aggregate throughput grows with concurrent requests until some stage saturates, and scale-out storage lifts that ceiling with servers that each bring drives, processors and ports. On AI infrastructure the target comes from the accelerators: the read rate that keeps a GPU cluster busy, multiplied across concurrent jobs, reachable only by spreading load over enough servers. Checkpoints push bursts the other way, and their size and frequency decide how much write rate the platform absorbs without stalling training. Caches in drives, controllers and clients soak up short bursts at high speed and then drain at media rate, so a few minutes of testing can flatter what a multi-hour job will see.

Inter-site circuits, replication backlog and recovery time

Between sites, bandwidth turns into days. A 10 Gb/s circuit running flat out moves 1.25 GB/s × 86,400 s = 108 TB a day, or 54 TB at a realistic 50 percent average. A site that changes more than that each day builds a replication backlog, and the gap keeps widening until the change rate falls. Seeding a remote copy of 1 PB over the same fully used link takes about 800,000 seconds, more than nine days, before overhead, and resynchronising after a long outage or restoring from the remote copy follows identical arithmetic. Pulling data back from public cloud adds egress fees to the transfer time.

Distance imposes a second limit. A link only runs full when enough data is in flight to cover its round trip, an amount equal to bandwidth × round-trip time: 6.25 MB at 10 Gb/s and 5 ms, 62.5 MB at 100 Gb/s and 5 ms. When TCP windows or request concurrency fall short of that, achieved throughput stays below the circuit rate.

Inside each site, rebuilds and replication compete with production reads for the same drives and links, and protection multiplies the bytes behind every client write, as set out under read/write performance. Circuits between sites come in fixed increments and often take months to provision, so headroom for growth in daily change rate is decided well before replication lag becomes visible.

Scality RING and storage throughput

Scality tested sustained data rate on a RING cluster of just over 100 nodes in three availability zones, built on hard drives with flash holding metadata. Across a two-hour window with 10 MB objects it held "approximately 420 GB per second" of S3 reads and "approximately 250 GB per second" of writes (Solved by Scality).

For RING stretched between sites, Scality publishes an envelope of "10Gb/s or greater bandwidth and <5ms latency" (Solved by Scality); at that floor, about 6.25 MB in flight keeps the circuit full. A stretched layout leaves the daily replication arithmetic untouched, since backlog still follows circuit size against change rate.