Glossary
PCIe flash
PCIe flash refers to flash storage connected straight to a server's PCI Express bus, as a drive in a PCIe-wired bay or as an add-in card, with no SATA or SAS controller in between.
Almost all PCIe flash today uses the NVMe protocol, and the bandwidth each device gets is set by the PCIe generation and the number of lanes it is given.
Why PCIe flash matters for storage servers
Putting flash on the PCIe bus removes the host bus adapter and the 600 MB/s SATA link that used to sit between the processor and an SSD. A single current drive can now move more data than the server's network port. That reverses the old balance of a storage server: the drives stop being the slow part, and the design questions move to how many lanes the processors offer, how many drives the chassis holds, and how much network bandwidth leaves the box.
Lanes, generations and bandwidth
A PCIe link is built from lanes, each a pair of signal paths that send and receive at the same time. Drives in U.2 and EDSFF form factors typically use four lanes; add-in cards can use eight or sixteen. Each generation roughly doubles the rate per lane.
| Generation | Rate per lane | Usable bandwidth per lane, one direction | ×4 drive link |
|---|---|---|---|
| PCIe 3.0 | 8 GT/s | ≈ 0.985 GB/s | ≈ 3.94 GB/s |
| PCIe 4.0 | 16 GT/s | ≈ 1.97 GB/s | ≈ 7.88 GB/s |
| PCIe 5.0 | 32 GT/s | ≈ 3.94 GB/s | ≈ 15.75 GB/s |
From PCIe 3.0 onwards, the line encoding (128b/130b) costs under 2 percent of the raw rate, so these figures sit close to the transfer rate divided by eight. PCIe 6.0 doubles the rate again with a different signalling scheme. Links carry the same rate in both directions at once. Almost every drive on these links speaks NVMe, and NVMe over Fabrics extends the same protocol beyond the server.
Form factors
- U.2: a hot-pluggable 2.5-inch drive in a front bay, the long-standing server format.
- EDSFF (E1.S, E3.S and related sizes): drive shapes designed for dense data-centre chassis, with better airflow and higher power budgets than U.2.
- Add-in card: a card in a standard slot, able to use eight or sixteen lanes.
- M.2: a small module, used in servers mainly for boot devices.
The electrical link is PCIe in every case. Form factor decides cooling, power per drive, hot-swap behaviour and how many drives fit in a rack unit.
Lane budgets and PCIe switches
Every lane given to a drive comes from a fixed supply on the server's processors, shared with network cards and accelerators. A chassis with 24 NVMe bays at four lanes each needs 24 × 4 = 96 lanes for direct attachment. When the processors cannot supply them, a PCIe switch fans a smaller upstream link out to many drives, and those drives then share the upstream bandwidth. In two-socket servers, each drive also belongs to one processor, and I/O issued from the other socket crosses the link between them, adding latency.
What PCIe flash means for server and cluster design
For a scale-out storage cluster, the consequence is that the network usually sets the ceiling. One PCIe 5.0 ×4 drive can move about 15.75 GB/s; one 100 Gb/s Ethernet port carries 12.5 GB/s. A server with 24 such drives and two 100 Gb/s ports has around 378 GB/s of drive bandwidth behind 25 GB/s of network. Adding drives to that server adds capacity and operation rate, but no deliverable throughput. Storage bandwidth sets out these ceilings link by link.
That changes how teams size flash clusters. Throughput targets for AI training or analytics are met by spreading drives across more servers and more network ports, while capacity targets favour denser servers. Dense servers also concentrate risk: a node with hundreds of terabytes of flash takes that much data offline when it fails, and the cluster has to rebuild or serve around it. Power and cooling follow drive count too, which is why EDSFF formats appear in dense flash designs.
The lane budget is a constraint worth reading off the server specification before anything else. Drives behind a heavily oversubscribed switch, or attached to the far socket, deliver less than their datasheet figures, and the gap only shows under load.
Generation upgrades interact with this budget. Moving from PCIe 4.0 to PCIe 5.0 doubles the bandwidth per lane, so the same drive bandwidth needs half the lanes, or the same lanes carry twice as much. For a storage platform that refreshes servers every few years, that is often the cheapest way to raise per-node throughput, provided the network ports are upgraded in step. Otherwise the extra drive bandwidth sits idle behind the same ports.
PCIe flash in Scality RING
RING runs on standard x86 servers and uses PCIe flash as a server component. Scality's published RING XP configuration ran on Dell PowerEdge R7615 servers with "3.2TB PCI/NVMe mixed-use drives" and one 100GbE connection per server, measuring 511 µs per GET and 741 µs per PUT on 4 KB objects (Solved by Scality).
In that build the network port is the narrower pipe: 100 Gb/s is 12.5 GB/s, while two PCIe 4.0 ×4 drives already offer about 15.8 GB/s of link bandwidth between them.














