Glossary
Storage virtualization
Storage virtualization is the abstraction of physical storage devices into logical resources, so that hosts address volumes, files or objects while a mapping layer tracks where the data physically sits. SNIA defines it broadly as the application of virtualization to storage services or devices.
Why storage virtualization matters in large estates
A large storage estate is never one generation of hardware. Arrays of different ages and vendors, new drive densities, and data that has outgrown the system it started on all coexist. Each time data has to move between physical devices, the applications using it are at risk of downtime unless something hides the move from them.
Virtualization is that something. Because hosts address logical locations, the layer underneath can relocate data, rebalance it across new devices, or retire old hardware without the host noticing. At petabyte scale, where migration is one of the largest recurring operational costs, that property is the main reason the layer exists.
How the logical-to-physical map works
Every virtualization layer keeps a map. A host asks for block 1,000,000 of volume 7, or the file /data/a.csv; the layer looks up which device and offset hold it and redirects the request. The size of that map depends on its granularity:
| 1 PB mapped in | Map entries | Map size at 16 bytes per entry |
|---|---|---|
| 1 GiB extents | 10¹⁵ ÷ 1,073,741,824 ≈ 931,323 | About 15 MB |
| 4 KiB blocks | 10¹⁵ ÷ 4,096 = 244,140,625 | About 3.9 GB |
Fine granularity enables thin provisioning and snapshots at small block sizes; coarse granularity keeps the map small enough to hold in memory. Across tens of petabytes, the map is itself a large, critical data structure, and losing it is equivalent to losing the data it describes.
Where virtualization runs
| Location | Mechanism | Scope |
|---|---|---|
| Host | Logical volume manager or file system spanning devices | One server |
| Array | Controller firmware maps LUNs onto RAID groups and drives | One array |
| Network or appliance | A virtualizing controller sits between hosts and several arrays | Several arrays, possibly from different vendors |
| Distributed software | Cluster software maps volumes, files or objects onto drives across all nodes | The whole cluster |
Network-based virtualization pools heterogeneous arrays behind one point of control and has long been a tool for storage consolidation. Distributed software virtualization is the basis of software-defined storage.
In-band and out-of-band designs
An in-band layer sits in the data path: every read and write passes through it. Caching and data services are easy to add, and the layer's throughput becomes a ceiling for all traffic behind it, so in-band appliances run in redundant pairs or clusters with mirrored write cache. An out-of-band layer holds only the map. Hosts fetch mapping information from a metadata service and then send data directly to devices, which removes the throughput ceiling at the cost of client software on every host. Parallel file systems commonly use this model.
Distributed object and file platforms combine elements of both: placement is computed in software spread across every node, so there is no single appliance in the path and no separate metadata server that every client has to consult for each request.
What storage virtualization means for refresh and migration
For teams planning a refresh, virtualization decides whether moving data is an outage or a background task. Placing a virtualization appliance in front of existing arrays lets their contents move to new hardware while hosts stay connected, and can extend the life of older arrays by pooling them, though the pooled arrays keep their own failure characteristics and support end dates.
The layer also becomes part of every data path it covers. An in-band appliance in front of several petabytes is a throughput limit and a shared failure point for all of it, and its upgrades are change windows for every application behind it. Distributed software designs avoid the separate appliance by building the map into the storage platform, so new servers join, data rebalances onto them and old servers drain out with the same mechanism. At large scale, how the map is protected, replicated and rebuilt matters as much as how fast it is.
The map is also what makes placement policy possible. Because data can move under a logical address, cold extents or objects can shift to denser, cheaper media while hot data stays on flash, which is the mechanism behind storage tiering. For an estate where most capacity is rarely read, that placement freedom has a direct effect on cost per terabyte.
Virtualization in Scality RING
In Scality RING, buckets and file systems map onto erasure-coded or replicated data spread across a cluster of standard x86 servers, and protection is set per storage class, so a data set's physical layout follows the class it belongs to. When a drive fails, RING writes data across the remaining drives in the server and rebuilds only data that was written. The MultiScale architecture lists S3, NFS and SMB as the access protocols over this layer.














