Glossary
Hybrid cloud
Hybrid cloud is an architecture that combines two or more distinct cloud infrastructures, typically a private environment and one or more public clouds, joined by technology that lets data and applications move between them. Each environment stays separate; what makes it hybrid is the connection and a shared way of operating across both.
The definition follows NIST SP 800-145. In practice most large organizations are hybrid by default, with some workloads on owned infrastructure and others in a public cloud, whether or not the connection between them was designed.
Compute moves in minutes, petabytes in days
With small datasets, hybrid cloud is a question of where applications run. At petabyte scale it turns into a question of where data lives. Compute starts in any environment within minutes, whereas a single petabyte needs over a week on a fully used 10 Gb/s link and, leaving a public cloud, a per-gigabyte egress charge on top. That pull of large datasets on everything around them is data gravity, and it makes hybrid cloud an architecture decision instead of a deployment option.
Three pressures usually meet in that decision: sovereignty and regulatory rules that keep some data on owned or in-country infrastructure; AI training and analytics that want public cloud services or rented GPU capacity; and the cost of holding large, frequently read datasets in a public region for years.
Placement patterns priced in data movement
| Pattern | What runs where | Consequence for data |
|---|---|---|
| Data on premises, compute bursts to cloud | Primary data stays owned; extra compute rented at peaks | Every burst reads across the link, or works from a copy kept current |
| Tiering to cloud | Hot data on premises; cold data in public cloud object storage | Cheap to hold, slow and costly to bring back; see cloud archive storage |
| Cloud as recovery site | Production on premises; replicas in a public cloud | A second full copy, plus egress if recovery ever brings it home |
| Sovereign split | Regulated data on owned infrastructure; the rest in public cloud | Classification decides placement, so data is labeled before it lands |
| Cloud-born data repatriated | Data created in cloud services moved to owned storage as it grows | A one-time egress cost traded against recurring storage charges |
Cloud-integrated storage on the private side
Cloud-integrated storage is on-premises storage that treats public cloud capacity as an extension of itself, tiering, replicating or archiving data to a cloud endpoint under its own policies while applications keep addressing the local system. A rule such as "replicate every new object" or "transition objects older than 180 days" is evaluated against metadata; matching items are queued, sent with retries and bandwidth limits, and recorded. Tiered items release their local copy, and the metadata keeps their cloud location. A cloud storage gateway turns the arrangement around, with the cloud holding everything and the local device acting as a cache.
The form of the cloud copy decides how hybrid the result really is. In native format each object lands as one recognizable cloud object, readable by cloud analytics and AI services and usable without the on-premises system. In a system-specific format data is packed into containers only the originating system can reassemble, which allows deduplication across objects but makes the copy opaque to everything else and turns any exit into a full recall. File systems that tier often leave stubs in place of moved files, and a backup job, virus scan or indexing crawl that walks the tree recalls everything it touches.
Recall costs grow with volume: retrieval charges, egress and, from archive classes, a delay measured in hours. At a hypothetical $0.09 per GB, bringing back 50 TB (51,200 GB) costs 51,200 × 0.09 = $4,608 in transfer alone, so rules that tier out data still read every month lose more on recall than they save on capacity. Losing the on-premises metadata that records where tiered objects went can strand cloud data that is itself intact.
One S3 interface on both sides of the link
A hybrid design is only as portable as its interfaces are alike. When applications use the same storage API in both environments, moving a workload means a new endpoint and new credentials; when the APIs differ, every move is a code change. Identity and policy follow the same rule, since keys, bucket policies, encryption and retention defined in one environment need to mean the same in the other, or a dataset locked in one place arrives unprotected in the next. Distance adds its own tax: an application in a public region reading from on-premises storage pays a wide-area round trip per request, so workloads issuing many small reads slow down far more than streaming ones.
Sovereignty and recovery across the boundary
A replica or cache in another jurisdiction is data held in that jurisdiction, which makes replication and tiering rules compliance rules. A lifecycle rule that expires data on premises and has no counterpart on the cloud copy leaves the organization keeping data it believes is gone. Without a primary location and a list of permitted copies for each dataset, copies multiply and capacity is bought several times over.
Direction carries the price. Moving data into a public cloud is usually free and moving it out is charged, so designs that read cloud-resident data back on premises pay on every read. Keeping the large dataset owned and sending compute or extracts outward reverses that flow, the same logic behind most repatriation of cloud-born data once it grows.
Recovery crosses the same boundary. With the recovery copy in a public cloud, bringing a large dataset home is limited by the link and billed as egress, and recovering into the cloud avoids that only if the applications are ready to run there. Replicated buckets serve as recovery copies; tiered data has no second copy unless one is made. Two sets of consoles, keys, monitoring and lifecycle rules also weigh on a lean team, so a common storage interface is what lets one group operate both sides.
Scality RING and hybrid cloud
RING presents an S3 interface on standard x86 servers in an organization's own data centers. Scality states that RING offers full S3 fidelity and mobilizes data across on-premises and the major clouds without re-architecting applications, with a unified S3 namespace across sites and clouds. Lifecycle tiering and one-to-many replication to public clouds date from RING8 in 2019, and data tiered to Azure "stays accessible to Azure Machine Learning, Power BI, Video Indexer and other Azure services", as Scality's Azure partner page puts it, which is the native-format case. Stretching one RING across two sites requires a link of 10Gb/s or greater with latency under 5 ms between them.














