Glossary

All-flash array

An all-flash array (AFA) is a shared storage system in which every persistent capacity device is flash, with no hard disk drives in the data path. Controller software presents block, file or object storage to servers over the network and decides how data is placed, protected and reduced on the flash.

Why all-flash arrays matter at scale

All-flash arrays took over primary storage for databases and virtual machines because they removed the mechanical delay of disk from every random read. For workloads measured in tens or hundreds of terabytes, a higher price per terabyte was offset by fewer devices, less tuning and response times that stayed predictable under load.

At petabyte scale the same design is judged on different grounds. The speed of flash is rarely in doubt. What decides the fit is how the system grows past its first footprint, how much the stored data really reduces, and what it costs to put data on flash that is read once a quarter.

How an all-flash array is built

A conventional all-flash array has four parts:

  • Controllers: servers running the array software, usually in pairs so that one can take over if the other fails.
  • Protected write buffer: memory backed by batteries or capacitors and mirrored between controllers. A write is acknowledged once it sits in both buffers.
  • Flash media: NVMe solid-state drives or flash modules built by the array maker.
  • Front-end ports: Fibre Channel, iSCSI or Ethernet, and on newer systems NVMe over Fabrics.

Flash cannot overwrite a page in place, so most arrays gather incoming writes into large stripes and write them to fresh space, reclaiming old space in the background. This log-structured layout keeps write amplification low and preserves drive endurance.

In a scale-up array the controller pair is fixed. Adding shelves adds capacity but no processing, so performance per terabyte falls as the array fills. Scale-out designs replace the pair with a cluster of nodes, each contributing processing, network bandwidth and flash, so capacity and performance grow together.

Data reduction and effective capacity

Most all-flash arrays apply inline deduplication and compression and quote an effective capacity after reduction. The chain runs from raw capacity, to usable capacity after parity, to effective capacity after reduction:

StepCalculationResult
Raw capacity24 drives × 7.68 TB184.32 TB
Usable after dual parity184.32 TB × 22 ÷ 24168.96 TB
Effective at 3:1 reduction168.96 TB × 3506.88 TB
Effective at 1.1:1 reduction168.96 TB × 1.1185.86 TB

The ratio belongs to the data. Virtual machine images that share operating system blocks reduce heavily. Video, images, encrypted files, compressed backups and most AI training data reduce very little, and their effective capacity stays close to usable capacity.

Types of all-flash systems

TypeAccess methodCommon workloads
Block all-flash arrayVolumes over Fibre Channel, iSCSI or NVMe-oFDatabases, virtual machines
File all-flash arrayNFS, SMBShared file systems, analytics, media
All-flash object storageS3 and other object APIsAI data pipelines, small-object workloads
Software-defined all-flashAny of the above, on standard servers with NVMe drivesDepends on the software

What all-flash arrays mean for large-scale storage teams

The first consequence is economic. A purchase sized in effective terabytes at an assumed 3:1 reduction delivers roughly a third of the expected capacity when the data turns out to be media, logs, backups or training sets. In the table above that is the gap between 506.88 TB and 185.86 TB, and it surfaces after the data has landed, as an expansion that arrives years early.

The second is the scaling unit. A dual-controller array has a ceiling for both capacity and performance, and growth beyond it means another array, another namespace and another migration at every hardware refresh. Estates that grow this way accumulate arrays with separate headroom, support contracts and end-of-life dates, the pattern described as storage sprawl.

The third is placement. In most multi-petabyte estates a minority of data drives most requests. Putting everything on flash pays flash prices for the cold majority, which is why large environments commonly keep flash for small objects, metadata, active AI datasets and databases and hold bulk capacity on hybrid or disk-based tiers through storage tiering.

Software-defined all-flash on standard servers changes the lifecycle question. Nodes are added, retired and replaced individually, so a refresh becomes a rolling change inside one system instead of a migration between two.

All-flash configurations of Scality RING

Scality RING is software-defined object and file storage on standard x86 servers. RING XP is its all-flash configuration, running on AMD EPYC-based NVMe servers from Lenovo, Supermicro, Dell and HPE, with published measurements of 511 microseconds per GET and 741 microseconds per PUT on 4 KB objects. RING defines erasure coding schemes per storage class, so the protection overhead on flash is set in the class configuration, and the same software runs on servers that mix flash and disk where the economics of a tier call for it.