Glossary
NVMe
NVMe, short for NVM Express, is an open protocol for talking to non-volatile storage: it defines the commands, the queues that carry them and how a host and a drive exchange work.
It was written for flash on PCI Express, with a first specification published in 2011, and has since been extended across networks and, from version 2.0, to hard drives.
Per-core queue pairs in place of one command queue
The interfaces flash inherited were built for spinning disks. SATA with AHCI offers one queue of 32 commands, and SAS with SCSI a deeper single queue per device behind a host adapter or RAID controller. A flash drive has dozens of dies able to work at once and a modern server has dozens of cores issuing I/O, so a single shared queue becomes the choke point.
NVMe moves work through pairs of queues held in host memory. The host places a command in a submission queue and rings a doorbell register on the drive; the drive fetches the command, executes it and posts the result to the paired completion queue. The specification allows up to 64K I/O queues of up to 64K commands each, far beyond what any drive uses. What counts is that each core can own a queue pair and work without waiting on a lock held by another core. Completions arrive by interrupt to the issuing core, or the host polls for them, spending a core to save the interrupt cost, a trade storage software at very high operation rates often makes.
A drive presents its capacity as one or more namespaces, each a separate block device. Since NVMe 2.0 the specification has been a family of command sets: conventional block, Zoned Namespaces (ZNS), whose zones are written sequentially, and Key Value. The queue model is defined independently of the wire, which is how NVMe over Fabrics carries it across networks.
PCIe flash, lane budgets and generations
PCIe flash is flash wired directly to a server's PCI Express lanes, as a drive in a PCIe bay or as an add-in card, with no SATA or SAS controller between the processor and the SSD. Almost all of it speaks NVMe. Each lane is a pair of signal paths that send and receive at the same time; U.2 and EDSFF drives typically take four lanes, add-in cards eight or sixteen, and each generation roughly doubles the rate per lane.
| Generation | Rate per lane | Usable per lane, one direction | ×4 drive link |
|---|---|---|---|
| PCIe 3.0 | 8 GT/s | ≈ 0.985 GB/s | ≈ 3.94 GB/s |
| PCIe 4.0 | 16 GT/s | ≈ 1.97 GB/s | ≈ 7.88 GB/s |
| PCIe 5.0 | 32 GT/s | ≈ 3.94 GB/s | ≈ 15.75 GB/s |
From PCIe 3.0 onwards the 128b/130b encoding costs under 2 percent, so usable rates sit close to the transfer rate divided by eight. Form factor governs cooling, power and density: U.2 is the long-standing 2.5-inch front-bay drive, and EDSFF sizes such as E1.S and E3.S suit dense chassis.
Lanes come from a fixed supply on the processors, shared with network cards and accelerators. A chassis with 24 NVMe bays at four lanes each needs 96 lanes for direct attachment. When the processors cannot supply them, a PCIe switch fans a smaller upstream link out to many drives that then share it, and 24 drives behind one ×16 upstream port are oversubscribed about 6:1. In two-socket servers each drive belongs to one processor, and I/O from the other socket crosses the inter-processor link.
NVMe performance at realistic queue depths
NVMe performance, meaning the latency, IOPS and throughput an NVMe drive or system delivers, depends heavily on how many requests the host keeps in flight. With illustrative values, a drive at a queue depth of 1 and 80 µs per operation completes 12,500 IOPS. At 8 outstanding and 90 µs it reaches about 88,900, and at 32 and 120 µs about 266,700. At 128 outstanding, latency rises to 400 µs for 320,000 IOPS: roughly 20 percent more operations for more than three times the latency. Headline figures quoted at depths of 128 or 256 describe that saturated region, while many applications run far lower.
Small operations hit controller and command-processing limits, large ones the link, so a PCIe 4.0 ×4 drive caps 1 MiB operations at roughly 7,500 a second. Industry test specifications precondition a drive until results reach steady state, because a fresh drive has no garbage to collect and posts write figures it cannot hold once full. Mixed loads add a tail effect: a read that reaches a die busy with a program or erase waits for it, so the slowest 1 percent of reads can take ten times the average. Below about 100 µs of device latency, host software becomes a visible share of each operation, from file system and cache layers to I/O issued by the socket that does not own the drive.
Storage servers whose drives outrun the network port
For platform teams the main effect of NVMe is that the drive stops being the slow part. One PCIe 5.0 ×4 drive moves about 15.75 GB/s, more than the 12.5 GB/s of a 100 Gb/s Ethernet port. A server with 24 such drives and two 100 Gb/s ports has around 378 GB/s of drive bandwidth behind 25 GB/s of network, and its processors give out before the drives do once storage software, checksums and erasure coding run at line rate. Balanced designs are therefore sized on cores, memory bandwidth and ports per drive. Throughput targets for AI training are met with more servers and more ports, while capacity targets favour denser servers that put more data inside each failure.
Generation upgrades interact with that balance. Moving from PCIe 4.0 to 5.0 halves the lanes needed for the same drive bandwidth, often the cheapest route to more per-node throughput at a refresh, provided network ports rise in step. ZNS drives need less spare flash and less garbage collection under software that already writes sequentially, as log-structured and object stores do, at the price of explicit support in that software.
Scality RING and NVMe
Scality RING uses NVMe between its software and local drives; clients use S3 and file protocols. The all-NVMe RING XP build paired 3.2 TB PCIe NVMe drives with one 100GbE port per server and measured 4 KB GETs at 511 µs and PUTs at 741 µs at the S3 API (Solved by Scality).
Two of its PCIe 4.0 ×4 drives together outrun the port. With a drive read of 50 to 100 µs, roughly 410 to 460 µs of each GET belongs to the network, HTTP handling and RING software.














