Glossary

Cloud migration

Cloud migration is the process of moving applications, data and workloads from one environment to another: from an on-premises data center to a public cloud, between clouds, or from a public cloud back to owned infrastructure. A migration covers both the transfer of data and the move of the applications that use it, and ends with a cutover to the new environment.

Why migration matters at petabyte scale

Moving an application is mostly a matter of configuration. Moving its data is a matter of physics and money. The network sets a floor on transfer time: at a sustained 10 Gb/s, 1 PB takes 8 × 10¹⁵ bits ÷ 10¹⁰ bits per second = 800,000 seconds, about 9.3 days, and that assumes the link is saturated around the clock. A 10 PB estate at the same rate needs about three months of transfer, during which the source keeps changing.

Cost follows direction. Moving data into a public cloud is usually free; moving it out is charged per gigabyte as egress fees. A migration into a cloud is therefore also a decision about what leaving that cloud will cost later. Once a large dataset has moved, applications and services tend to follow it, the effect known as data gravity, which makes the first placement of data the one that sticks.

Migration directions

DirectionTypical driverMain data concern
On premises to public cloudData center exit, access to cloud servicesTransfer time; long-term storage and egress costs
Cloud to cloudPricing, services, consolidation after acquisitionsEgress from the source; API differences between providers
Public cloud to owned infrastructureCost of large steady datasets, sovereignty, performanceOne-time egress; rebuilding operations on owned hardware
Partial moves into a hybrid cloudSome workloads move, others stayApplications split across sites; latency and egress between the halves
Platform to platform on premisesHardware end of life, vendor changeDuration of dual running; preserving metadata and retention

Migration strategies

Applications are usually sorted into a few strategies before any data moves.

  • Rehost: move the application as it is onto equivalent infrastructure.
  • Replatform: move with targeted changes, such as swapping a file share for an object store.
  • Refactor: rewrite to use the target environment's services, the most expensive option and the hardest to reverse.
  • Retain or retire: leave the application where it is, or switch it off and archive its data.

For storage, the strategy largely depends on whether the application's storage interface exists at the destination. Applications that already use the S3 API can usually be rehosted against a different S3 endpoint; those that depend on a specific file or block system often need replatforming.

Moving object data

Large object migrations usually run in two phases. A bulk copy moves the existing data while the source stays live, then incremental passes copy what changed until the gap is small enough for a short cutover window, when writes switch to the destination.

Object count matters as much as volume. Each object costs at least one listing entry, one read and one write, so a petabyte of small objects can take far longer than a petabyte of large ones on the same link. The metadata has to travel too: object tags, versions, access policies and, for data kept under S3 Object Lock, the retention settings. Data copied without its retain-until dates arrives at the destination unprotected.

What cloud migration means at petabyte scale

  • Migrations are long projects. Months of transfer, incremental sync and dual running are normal for multi-petabyte estates, and both environments are paid for throughout.
  • Capacity doubles temporarily. During dual running, the source and destination both hold the full dataset, which has to fit in budgets and in the destination's growth plan.
  • Exit cost is set at entry. Every petabyte placed in a public cloud carries a future egress bill if it ever leaves. Repatriation projects are often sized by that bill.
  • Retention obligations travel with the data. Regulatory holds and lock periods have to be reproduced at the destination before the source copy is deleted, or the organization holds data it can no longer prove was protected.
  • Rollback lasts as long as the source. Until the source copy is decommissioned, a failed cutover can be reversed by pointing applications back. Once it is gone, the destination is the only copy, so the timing of decommissioning is a risk decision as much as a cost one.
  • Interface compatibility decides the scope. If source and destination share an S3 interface, applications follow the data by changing an endpoint; otherwise the migration becomes an application project as well. That data mobility is what lets the next migration be smaller.

Scality RING at either end of a migration

Scality RING can be the source or the destination when object data moves on or off owned infrastructure, presenting an S3 interface on standard x86 servers. Scality states that RING offers full S3 fidelity and mobilizes data across on-premises and the major clouds without re-architecting applications. RING supports S3 Object Lock in governance and compliance modes with retention periods and legal holds, the settings a migration of retained data has to carry across intact.