Glossary
Hyper-converged infrastructure (HCI)
Hyper-converged infrastructure (HCI) combines compute, storage and virtualization on one cluster of standard servers, with the local drives in each node pooled by software into shared storage for the virtual machines running on those nodes. There is no separate storage array or storage network.
HCI has become a common base for virtual machine estates. Storage architects meet its limits when unstructured data, backups and AI data sets grow far beyond what compute nodes hold economically.
Why HCI matters to storage architects
HCI simplifies the virtualization tier: one cluster of identical building blocks, one management interface, growth by adding a node. That simplicity comes from a fixed coupling. Every node brings compute, memory, hypervisor licensing and drives in a set proportion, so capacity and compute grow together whether or not both are needed.
For VM workloads whose storage grows roughly in line with compute, that coupling is efficient. For data that grows on its own curve, such as file archives, backup repositories, media libraries and training data, it is expensive, and it puts the copies meant to protect the cluster inside the same cluster.
Three-tier, converged and hyper-converged designs
| Design | Compute | Storage |
|---|---|---|
| Three-tier | Rack or blade servers | Separate SAN or NAS array, integrated by the customer |
| Converged (CI) | Servers | Separate array, pre-integrated and validated by the vendor |
| Hyper-converged (HCI) | Same nodes as storage | Local drives in each node, pooled by software |
HCI replaces the array and the storage area network with software-defined storage running on the compute nodes, either as a controller VM on each node or inside the hypervisor kernel.
How HCI stores and protects data
When a VM writes to its virtual disk, the storage software on the local node accepts the write and sends copies to one or two other nodes before acknowledging it. The replication factor (RF) sets the number of copies: RF2 survives one node loss, RF3 survives two. Some platforms offer erasure coding for capacity-oriented data at a higher write cost. Reads are served preferentially from a local copy, a property called data locality; after a VM migrates, its reads cross the network until data follows it.
When a node fails, its VMs restart elsewhere and the surviving nodes recreate the lost copies in parallel. During that rebuild, RF2 data has a single remaining copy, and rebuild traffic competes with the VMs for the same drives and network. The larger each node's drives, the longer that window lasts, so dense capacity nodes lengthen the period of reduced protection after every node failure.
Capacity arithmetic
For an 8-node cluster with 20 TB of raw capacity per node:
- Raw capacity: 8 × 20 TB = 160 TB.
- Reserve one node's worth for rebuild: 160 − 20 = 140 TB.
- At RF2: 140 ÷ 2 = 70 TB usable. At RF3: 140 ÷ 3 ≈ 46.7 TB usable.
Scaled to unstructured data, the same arithmetic gets heavy. Holding 1 PB usable at RF2 needs about 2 PB raw plus rebuild reserve, which at 20 TB per node is more than 100 nodes, each carrying processors, memory and hypervisor licences the data does not use. Compression and deduplication improve the ratio by an amount that depends on the data.
What HCI means for large-scale storage
In enterprise estates HCI usually ends up as the compute and VM tier, with a separate storage platform beside it for data that grows independently. Three pressures drive that split. Cost: buying compute nodes to add capacity becomes the most expensive way to store petabytes. Failure domain: backups or replicas held in the same cluster as the VMs they protect share its software, credentials and upgrade cycle, so an incident affecting the cluster affects both. Workload fit: backup ingest, large sequential reads for analytics, and AI data loading generate traffic that competes with latency-sensitive VMs for the same drives and network.
The resulting pattern pairs HCI for virtual machines with external object storage reached over S3 for backup targets, archives and data lakes, scaling capacity on its own curve. Lifecycle reinforces the split. An HCI refresh replaces compute and drives together on the server vendor's cycle, while data held for seven or ten years outlasts several generations of compute nodes and has to be migrated each time if it lives inside the cluster. Licensing that counts cores or nodes adds to the cost of every terabyte held there.
HCI remains one implementation of software-defined infrastructure; the external platform is another part of the same software-defined estate.
Scality RING and the HCI model
Scality RING sits beside HCI. It is software-defined object and file storage on standard x86 servers, consumed over S3, NFS and SMB, and its MultiScale architecture separates the compute that serves data from the storage that holds it. The RING product page describes capacity, performance, tenants and sites as scaling independently, which is the opposite arrangement to HCI's fixed-proportion nodes.














