Glossary
NVMe over Fabrics (NVMe-oF)
NVMe over Fabrics (NVMe-oF) extends the NVMe protocol across a network, so a server can use a remote flash device much as if it were plugged into its own PCIe bus.
It runs over RDMA networks, Fibre Channel or ordinary TCP/IP, and keeps NVMe's many parallel queues from host to target.
Why NVMe-oF matters for large infrastructures
Local NVMe ties flash to the server it sits in. If one server needs more flash and its neighbour has spare, the capacity cannot move. NVMe-oF breaks that tie. Flash can sit in dedicated enclosures and be allocated to compute servers over the network, so compute and storage grow on separate schedules. It also gives storage area networks a protocol designed for flash, replacing SCSI-based access to all-flash arrays. For platform teams running large shared clusters, it is one of the main routes to disaggregated block storage.
How NVMe-oF works
On a local bus, an NVMe drive reads commands straight from queues in host memory. Across a network there is no shared memory, so NVMe-oF wraps each command and its response in a message called a capsule. A small write can carry its data inside the command capsule, saving a separate exchange.
The queue model survives the move. Each NVMe queue pair maps to its own network connection, so a host with many cores still opens many independent queues to the remote target. A discovery service tells hosts which storage subsystems exist on the fabric and how to reach them. To the operating system, a remote namespace then appears as one more NVMe block device, and multipathing across several target controllers keeps it reachable if one path fails.
Transport options
| Transport | Network | Trade-offs |
|---|---|---|
| RDMA | RoCE on Ethernet, InfiniBand, iWARP | Lowest latency and CPU cost; RoCE needs carefully configured, low-loss Ethernet |
| Fibre Channel (FC-NVMe) | Existing Fibre Channel SANs | Runs beside SCSI traffic on fabrics many enterprises already operate |
| TCP (NVMe/TCP) | Any IP network | No special adapters; somewhat higher CPU use and latency than RDMA |
RDMA and Fibre Channel were part of the original specification in 2016. NVMe/TCP followed in 2018 and has become the common choice where operational simplicity outweighs the last few microseconds.
NVMe-oF compared with SCSI SANs and object storage
A traditional SAN carries SCSI commands over Fibre Channel or iSCSI. When the array's media is NVMe flash, it translates SCSI into NVMe internally, and the host-side SCSI stack offers far fewer parallel queues. NVMe-oF removes the translation and extends per-core queues end to end.
NVMe-oF remains a block protocol. It moves drives, or block volumes, across a network; it does not provide files, objects, data protection across sites or a namespace shared by many applications. Those come from the software above it. Object storage sits at a different layer, with the network crossing placed above the storage software, through S3.
What NVMe-oF means for disaggregated and multi-site designs
Inside a data centre, NVMe-oF adds only a few microseconds to a flash read of 50 to 100 µs, which keeps remote flash close to local performance. Distance changes that. Light in fibre adds about 1 ms of round trip per 100 km, ten times a flash read, so NVMe-oF is a local-fabric technology. Multi-site resilience still needs replication at a higher layer.
Disaggregation also moves the bandwidth limit to the network. An enclosure holding 24 PCIe 4.0 ×4 drives has about 189 GB/s of drive bandwidth; two 100 Gb/s ports carry 25 GB/s. Enclosure designs balance drive count against port count, and the number of drives a pool can usefully hold is set by its ports.
The failure domain changes shape too. A flash enclosure shared by dozens of hosts takes all of their volumes with it if it fails, unless multipathing, redundant controllers or host-side mirroring cover that case. And the transport choice is an operations decision as much as a performance one: RoCE fabrics reward specialist network skills, while NVMe/TCP runs on the IP networks lean teams already manage.
Cost follows from the same choices. Disaggregated flash raises utilisation, because capacity stranded in one server can be lent to another, but it adds enclosures, switch ports and fabric design to the bill. The saving is largest in environments with many servers whose flash needs vary widely, and smallest where every node already uses its local drives fully.
NVMe-oF and Scality RING
RING works at a different layer from NVMe-oF. Its servers use local NVMe or hard drives, and applications reach the data through S3 and file protocols over the network. Protection across servers comes from replication or erasure coding, with schemes defined per storage class.
Across sites, a stretched RING operates at the object layer within a published envelope of "10Gb/s or greater bandwidth and <5ms latency" between sites (Solved by Scality), distances at which block-level NVMe-oF would lose most of its advantage.














