Glossary

Solid-state drive (SSD)

A solid-state drive (SSD) is a storage device that keeps data in NAND flash chips managed by an on-board controller, with no moving parts.

The controller's firmware makes flash, which is written in small pages and erased only in much larger blocks, look to the server like an ordinary disk made of numbered blocks.

Controller, DRAM map and flash translation layer

Five parts make up a data-centre SSD. NAND flash packages hold the data, organised into dies, blocks and pages. The controller runs firmware that schedules work across many dies in parallel, corrects bit errors and manages wear. DRAM, roughly 1 GB per TB of flash on most data-centre drives, holds the map from logical blocks to physical flash locations. Capacitors on drives with power-loss protection supply enough energy to finish in-flight writes and map updates if power fails. The host interface is SATA, SAS, or PCI Express carrying the NVMe protocol.

Because flash pages cannot be rewritten in place, the flash translation layer sends every update to a fresh page, records the new location and marks the old page invalid. When free pages run low, the controller picks a block with many invalid pages, copies the valid ones elsewhere and erases it. That copying is write the host never requested, and it grows as the drive fills and as writes get smaller and more random. Spare flash the host cannot address (overprovisioning) and host hints about blocks that no longer hold data (TRIM on SATA, Deallocate on NVMe) keep it in check.

Enterprise and client SSD ratings

Enterprise flash storage is the class of SSDs and flash systems built for continuous data-centre duty. The JEDEC JESD218 standard sets different rating conditions for client and enterprise drives, reflecting very different lives:

ConditionClient classEnterprise class
Active use8 hours a day at 40°C24 hours a day at 55°C
Powered-off data retention1 year at 30°C3 months at 40°C
Uncorrectable bit error rate1 in 1015 bits or better1 in 1016 bits or better

The shorter retention figure fits drives expected to stay powered and to be written hard at higher temperature. Beyond the rating conditions, enterprise drives add features that client drives usually lack:

  • Power-loss protection, so an acknowledged write survives an outage and the drive can safely acknowledge from DRAM, which also lowers write latency.
  • End-to-end data protection, with checksums carried from interface to NAND and back so corruption inside the drive is caught.
  • Endurance grades, typically read-intensive models around one drive write per day over five years and mixed-use models around three, with the arithmetic set out under flash storage endurance.
  • Self-encryption, with keys that can be erased to retire a drive.

Performance is rated differently too. Client drives are usually measured fresh, when every block is empty. Enterprise drives are measured in steady state, after being filled and overwritten until garbage collection runs continuously, the condition a drive in a busy cluster actually lives in, and their specifications state latency at high percentiles such as the 99.99th.

Interfaces and form factors in current servers

InterfaceLink ceilingTypical place in large estates
SATA III600 MB/s, one command queueBoot drives, older servers
SASSet by SAS generation; dual-portedShared-shelf storage arrays
NVMe over PCIeSet by PCIe generation and lanesCurrent server and all-flash designs

Physical shapes include the 2.5-inch U.2 drive, M.2 boot modules and the EDSFF family (E1.S, E3.S and related sizes) designed for dense data-centre servers.

Correlated wear, tail latency and power events across a fleet

SSDs have no mechanical parts to fail, but their flash wears with writing, and each drive reports how much rated endurance it has used. Across thousands of drives that makes part of the replacement plan a forecast. It also creates correlated risk, since drives installed together, from the same batch, under the same load, approach their limits together.

Drive grade becomes a decision per role. Large read-intensive drives suit capacity tiers holding data written once, while mixed-use drives suit metadata, journals and write-heavy services. The same NAND is often sold as a 3.84 TB read-intensive drive and a 3.2 TB mixed-use drive, the difference held back as spare area, so choosing mixed-use across a petabyte gives up roughly a sixth of raw flash for endurance and steadier writes.

Tail latency compounds in distributed systems. A request in an erasure-coded cluster that touches several drives completes when the slowest one answers, so one drive model with erratic garbage collection slows requests on every server it is fitted to. Consistency at high percentiles is a stronger selection criterion than peak IOPS, and one drive model and firmware level across hundreds of nodes keeps that behaviour predictable. A failed power feed drops hundreds of drives at once, and power-loss protection is what keeps acknowledged writes and metadata intact through it. Dual porting, by contrast, exists so two array controllers can share a drive and goes unused in shared-nothing designs that protect data across servers.

Scality RING and SSDs

SSDs play two roles in RING deployments. In hard-drive-based clusters they hold metadata while disks hold the data; a Scality validation of just over 100 such nodes sustained about 420 GB/s of S3 reads on 10 MB objects (Solved by Scality). In the all-flash RING XP configuration they hold both, and the published RING XP test used 3.2 TB PCIe NVMe mixed-use drives in Dell PowerEdge R7615 servers.