Glossary
Cloud-integrated storage
Cloud-integrated storage is on-premises storage that uses public cloud capacity as an extension of itself, moving, copying or tiering data to the cloud under its own policies. Applications keep addressing the local system, which tracks where each item lives.
Why cloud-integrated storage matters for hybrid estates
Few large organizations put all their data in one place. Active datasets stay close to the applications and GPUs that use them, while older data, disaster-recovery copies and burst workloads can make use of public cloud capacity. Running these as separate silos means separate tools, separate copies and manual movement between them. Integration makes the on-premises platform responsible for that movement, driven by policy, so the decision of where data lives becomes a rule on a bucket or path instead of a migration project.
Integration patterns
| Pattern | What goes to the cloud | Authoritative copy |
|---|---|---|
| Tiering | Data matching an age or access rule, released from local media | Cloud copy, reached through the local system |
| Replication | Copies of selected buckets, kept in step with the source | Local; cloud copy for recovery |
| Backup or archive target | Point-in-time copies kept for a retention period | Local |
| Cloud bursting | Working data copied for processing on cloud compute | Local |
A cloud storage gateway inverts the arrangement: the cloud holds everything and the local device is a cache. Cloud-integrated storage keeps a full storage system on premises, with its own capacity and protection, and uses the cloud for part of the data or for extra copies. Cloud storage replication and cloud storage tiering are covered in their own entries.
How data moves and in what form
The system evaluates policies, such as "replicate every new object" or "transition objects older than 180 days", against its metadata, queues the matching items, sends them to the cloud endpoint with retries and bandwidth limits, then updates its own records. For tiered data, the local copy is released and the metadata records the cloud location, so later reads are resolved by fetching the data back or redirecting the client.
The form of the cloud copy matters. In native format, each object lands as one recognizable cloud object with its data unchanged, readable by cloud analytics and AI services and usable without the on-premises system. In a system-specific format, data is packed into chunks or containers that only the originating system can reassemble, which allows deduplication and compression across objects but makes the cloud copy opaque to everything else.
File systems that tier often leave a stub, a small placeholder carrying the original name and attributes, so directory listings still show the file. Opening a stub triggers a recall. That is invisible to users until a backup job, an antivirus scan or an indexing crawler walks the whole tree and recalls everything it touches.
What cloud-integrated storage means for hybrid architectures
Exit depends on format. Native-format data can be read in place if the on-premises platform is retired or replaced. Data in a proprietary format has to be recalled through the originating system first, which at petabyte scale is a long and expensive operation.
Bandwidth sets the pace. A 10 Gb/s link carries at most 1.25 GB/s, so seeding 1 PB (1,000,000 GB) takes at least 1,000,000 ÷ 1.25 = 800,000 seconds, about 9.3 days, at full utilization with nothing else on the link. The steady state is easier to plan: it is the daily volume of new or changed data the policies send out. Large initial transfers are often staged over weeks or shipped on physical devices, a practical face of data gravity.
Recall has a cost that grows with volume. Bringing tiered data back incurs retrieval charges, egress fees and latency, and from archive classes a restore delay measured in hours. At a hypothetical $0.09 per GB, recalling 50 TB (51,200 GB) costs 51,200 × 0.09 = $4,608 in transfer alone. Tiering policies that send out data still read every month lose more to recall than they save on capacity.
Governance follows every copy. A replica or tier in a public cloud region is subject to that region's location, retention and access rules. Residency and retention policies apply to the cloud side as much as to the primary site, and a lifecycle rule that expires data locally has to be matched in the cloud copy.
Disaster recovery depends on which copy is authoritative. A replicated bucket in the cloud can serve as a recovery copy if the primary site is lost. Tiered data has no second copy unless one is created, and losing the on-premises metadata that records where tiered objects live can strand cloud data that is itself intact.
Cloud integration in Scality RING
Cloud data management arrived in RING with eXtended Data Management in RING8, which Scality announced in 2019 with lifecycle tiering and one-to-many replication across RING and public clouds. Data tiered from Scality storage to Azure "stays accessible to Azure Machine Learning, Power BI, Video Indexer and other Azure services", as Scality's Azure partner page puts it, which is the native-format case described above. The RING product page now presents this as a "unified S3 namespace across sites and clouds", through which applications keep addressing one system while the data sits in several places.














