Glossary

Cloud orchestration

Cloud orchestration is the coordinated automation of many provisioning and configuration tasks into one workflow, run in the correct order across compute, network and storage, so that a complete environment or service is delivered and kept in its intended state. Automation performs individual tasks; orchestration decides their order, handles their dependencies and reacts when something fails.

Why orchestration matters for infrastructure teams

A single service in a modern platform depends on dozens of resources: networks, load balancers, compute, identities, secrets, volumes and buckets with their policies. Built by hand, each environment comes out slightly different, and rebuilding one after a failure depends on someone remembering how it was done. Orchestration replaces that with a definition that produces the same environment every time, at any site.

For teams running infrastructure across several sites with few people, that repeatability is what keeps growth from translating into headcount. It also changes the risk profile: the same machinery that builds an environment in minutes can remove one just as quickly.

Orchestration is also the layer where storage stops being configured by storage administrators. Buckets, policies and keys are created by pipelines written by platform and application teams, so the storage team's influence moves into the templates, modules and API permissions those pipelines use.

How orchestration works

  1. Desired state: a definition describes the intended environment, usually in declarative files kept in version control.
  2. Dependency graph: the orchestrator works out which resources depend on which. A bucket policy needs its bucket; an application needs its network and credentials.
  3. Execution: independent tasks run in parallel and dependent ones in order, each through the target system's API.
  4. Reconciliation: the orchestrator compares what exists with the definition and corrects any difference, once per run or continuously in a control loop, as Kubernetes controllers do.
  5. Failure handling: failed steps are retried, rolled back or left for an operator, depending on how the workflow was written.

Steps that can be repeated safely, producing the same result whether run once or several times, are what make retries and reconciliation reliable.

Types of cloud orchestration

TypeWhat it coordinatesTypical tools
Infrastructure orchestrationNetworks, machines, storage, identitiesTerraform and similar infrastructure-as-code tools
Container orchestrationPlacement, scaling and recovery of containersKubernetes; see cloud native
Workflow and data pipeline orchestrationOrdered jobs that move and transform dataWorkflow schedulers for ETL and AI data pipelines
Recovery orchestrationFailover and restart of services at another site in orderSee disaster recovery orchestration

Declarative orchestration states the end result and lets the tool work out the steps. Imperative orchestration lists the steps explicitly. Declarative definitions are easier to review and reconcile; imperative ones give finer control over order and timing, which matters in recovery runbooks.

What orchestration means for storage operations

  • Storage becomes an API target. Orchestrators create buckets, set policies and issue keys through the storage API. Any setting that cannot be reached through the API stays manual and drifts.
  • Data is the slowest part to rebuild. An orchestrator recreates compute and networks in minutes from a definition. It cannot recreate data; repopulating petabytes takes days. Environments can be disposable only because their data lives outside them in durable storage.
  • Automated mistakes run at machine speed. A definition that drops a bucket, or a reconciliation loop that reads an empty definition as intent, can delete storage across every site it manages. Scoped credentials for the orchestrator, and retention that the storage itself enforces, limit what a bad run can destroy.
  • Pipelines move data between tiers. Workflow orchestration in analytics and AI decides when datasets are staged onto fast media, read, written back and moved to cheaper capacity, which shifts storage tiering from a static plan to a function of the workflow.
  • Applied is not finished. Changing replication or lifecycle rules on a bucket holding petabytes takes effect over days as data is copied or moved. The orchestrator reports the change as complete the moment the API call succeeds, while the storage system is still working through it.
  • The orchestrator holds the keys. Whatever system runs the workflows holds credentials for every platform it manages, which makes it one of the most valuable accounts in the environment to an attacker.
  • Consistency across sites becomes checkable. When every site is built from the same definition, a difference between sites is visible as a difference in code.

Orchestration and Scality storage

Orchestrated environments reach Scality RING through its S3 API, so buckets, policies and keys are created by the same calls used against any S3 endpoint, and RING itself expands by adding standard x86 servers. For data under S3 Object Lock in compliance mode, retention holds against all users, including the account root, until the retain-until date, so an orchestration run with storage credentials cannot shorten it. At the data layer, Scality describes its Autonomous Data Infrastructure (ADI) as aligning storage media, performance and protection to each stage of the data lifecycle through policy-governed operations.