Glossary

Cloud automation

Cloud automation is the use of software to create, configure, scale and remove infrastructure without manual steps. The intended setup is written down as code or policy, and a program applies it on a schedule, in a pipeline or in response to an event.

Buckets, keys and retention as API calls

On an S3-compatible platform nearly every storage setting is an API operation: creating a bucket, attaching a policy, setting a quota, enabling versioning, writing a lifecycle rule, configuring replication to a second site, applying a default Object Lock period. Once those operations are scripted, the storage team stops being the people who click through a console and becomes the owner of the templates every pipeline calls. At a few hundred buckets the change is a convenience. At tens of thousands of buckets spread across tenants and sites, run by a team that has not grown with capacity, it is the only way configuration stays consistent. Inside the platform, placement, repair and rebalancing follow the same logic under the name storage automation.

Cloud provisioning is the front edge of this work: the request that turns a name and a few parameters into a ready bucket or volume with credentials attached. The request is authenticated and checked against quota, placement picks the site and storage class, the resource is created with its encryption, lifecycle and lock settings, and metering starts. Placement has the longest tail. The site and jurisdiction chosen at creation decide where the data lives, and changing them later means migrating data, so sovereignty and residency rules end up encoded in provisioning templates instead of being checked by hand after the fact.

Declarative plans, drift and repeatable runs

Automation tools split by what their instructions describe.

ApproachInstructions describeTypical formRun a second time
ImperativeSteps to perform, in orderScripts, CLI commands, SDK callsRepeats the steps unless the script checks state first
DeclarativeThe desired end stateInfrastructure-as-code definitionsChanges only the difference between desired and actual state

A script that creates three buckets has created six after its second run. A declarative definition of three buckets compares what exists with what is written and leaves an existing three alone, a property called idempotency that makes repeated runs safe against a live estate. The tool works in two stages: a plan that lists every create, update and delete it intends, then an apply once someone approves it. A state record maps each definition to the real resource it manages. A bucket edited by hand drifts from its definition, and the next plan shows the edit as a change to reverse.

Allocation is decided at the same moment. A pool using thin provisioning promises tenants more than it physically holds: 50 tenants allocated 200 TB each have been promised 10 PB, which on 4 PB of disks is a 2.5 to 1 overcommitment that holds only while consumption stays under 4 PB. Automated provisioning makes those promises in seconds, so thin pools are tracked against physical consumption and growth rate, never against allocations.

Cloud orchestration across sites and pipelines

Cloud orchestration is the layer above single tasks. It sequences many automated steps into one workflow across compute, network, identity and storage, using a dependency graph to decide what runs in parallel and what waits: a bucket policy needs its bucket, and an application needs its credentials and network before it starts. Failed steps are retried, rolled back or left for an operator. Kubernetes controllers extend the idea into a continuous control loop that keeps comparing actual state with desired state.

In a storage estate, orchestration shows up in four places. Infrastructure orchestration builds sites from code. Container orchestration places and restarts the services that read and write objects. Workflow orchestration in analytics and AI decides when a dataset is staged onto fast media, processed and returned to capacity storage, as in an AI data pipeline. Recovery orchestration brings services up at another site in a set order, covered under disaster recovery orchestration.

One gap is particular to data. An orchestrator reports a change to replication or lifecycle rules as done when the API call returns, while the platform may spend days copying or moving the petabytes behind that rule. Compute rebuilt from code is back in minutes. A bucket holding several petabytes is not, which is why environments can be treated as disposable only when their data lives outside them.

When a template deletes data

Errors travel at pipeline speed. A wrong retention value or an over-broad bucket policy in a shared module lands on every bucket built from it, in every tenant, and replication carries it to the second site. Deletion is the sharpest case. When a declarative tool no longer finds a resource in its definitions, its plan proposes removing it; on compute that rebuilds a server, on storage it removes data. Large estates usually keep data-bearing buckets out of automatic deletion paths and lean on versioning plus S3 Object Lock, so a faulty run can strip configuration without destroying stored objects.

The pipeline account becomes one of the most powerful identities in the estate. A service account that creates buckets and policies for every tenant holds keys that reach every tenant, so compromising the automation system is equivalent to compromising a storage administrator. Lock mode decides how much that account can undo. In governance mode, an identity granted s3:BypassGovernanceRetention can override retention by sending the x-amz-bypass-governance-retention:true header, and a broad pipeline role may well hold that grant. In compliance mode no user can shorten or remove the lock before its retain-until date, not even the account root.

The return on all this is portability. When on-premises object storage exposes the same S3 API as public cloud, one set of definitions and pipelines manages both, during repatriation or in a hybrid cloud estate, instead of two codebases that slowly diverge.

Scality RING and cloud automation

RING exposes an S3 API, so the infrastructure-as-code providers, modules and pipelines written for public cloud buckets can also create RING buckets, bucket policies, lifecycle rules and Object Lock settings on owned servers. Object Lock on RING covers both governance and compliance modes, plus retention periods and legal holds, so versions written under compliance mode stay out of reach of even the broadest pipeline role until they expire. RING applies a template as written: a wrong policy pushed through the API is a wrong policy on RING too.

Within the platform, Scality ADI includes Guardian, which Scality describes as autonomous AI operations that "reduce manual tasks, anticipate risks, and simplify maintenance".