Glossary

Multi-protocol storage

Multi-protocol storage exposes the same data through more than one access protocol, so a file written over NFS can be opened over SMB or read as an object over S3 without a second copy being made. The system keeps one copy and translates each protocol's names, permissions and locking rules onto it.

Why multi-protocol access matters for large data sets

Data at scale rarely stays inside one protocol for its whole life. Instruments, cameras, sequencers and render farms write files over NFS or SMB. Analytics engines, AI training jobs and cloud-native applications read over S3. Windows users browse the same results over SMB. Each time data crosses one of those boundaries on single-protocol systems, it becomes a copy job.

A copy at petabyte scale is expensive in two ways. It doubles the capacity held, and it takes time: moving a 2 PB data set at a sustained 10 GB/s takes 2,000,000 GB ÷ 10 GB/s = 200,000 seconds, about 2.3 days, before the second copy is usable. It also creates two sets of permissions to keep aligned and a window in which the copies disagree. Multi-protocol storage removes the copy by letting each tool reach the original.

How one copy is presented through several protocols

ProtocolData modelIdentityConcurrency
NFSFiles and directoriesUNIX user and group IDs; NFSv4 ACLsAdvisory byte-range locks
SMBFiles and directoriesWindows security identifiers, NTFS-style ACLsShare modes, byte-range locks, leases
S3Objects by key in a bucketAccess keys, users, roles, bucket policiesNone per object; each PUT replaces the whole object

The file /projects/alpha/report.pdf becomes the key projects/alpha/report.pdf in a bucket. Directories have no independent existence in S3; a listing by prefix and delimiter reconstructs them. An identity mapping, usually backed by Active Directory or LDAP, lets the system recognise one person behind a UNIX ID, a Windows SID and an S3 access key.

Where the protocol models differ

  • Renames. Renaming a directory of 10,000 files is one metadata update in a file system. Where the directory exists only as a key prefix, the same rename becomes a copy and a delete per object: 2 × 10,000 = 20,000 requests.
  • Names. SMB is case-insensitive and reserves characters such as : and *; NFS and S3 are case-sensitive. Report.pdf and report.pdf coexist in a bucket and collide over SMB.
  • Permissions. A UNIX permission mode such as 750 fits in nine bits; an NTFS ACL can hold dozens of entries with inheritance. Converting the richer model to the simpler one loses detail.
  • Locks. A file locked by an SMB client can still be replaced by an S3 PUT unless the system checks file locks on object writes.
  • Metadata. Files carry separate access, modification and change times; an object has one last-modified time and free-form user metadata.

Implementation approaches

ApproachMechanismTrade-off
Protocol gatewayA gateway translates file operations into object requests against a separate storeExtra hop; rename and partial-write costs land on the object store
Shared namespaceEach protocol service reads and writes one internal namespace directlyConsistent view; the platform carries the full reconciliation work
Copy and syncData is copied between a file system and an object store on a scheduleSimple per protocol; two copies with lag between them

What multi-protocol access means for data pipelines at scale

For a team running ingest-to-analysis pipelines, the gain is measured in copies avoided: capacity that is never bought twice, transfer windows that disappear, and one retention and protection policy per data set. Research groups, media houses and AI teams feel this most, because their files are produced over file protocols and consumed by S3-native tools.

The semantic gaps become operational facts. A permissions model chosen at deployment decides whether an S3 bucket policy and an NTFS ACL can disagree about who reads a folder, which matters to the security review of any shared data set. Mass renames or restructuring of large trees behave very differently depending on which protocol performs them. Concurrent writes to the same file from different protocols are where edge cases surface, which is why stable large deployments tend to have one protocol writing a given data set and the others reading it.

Protection and retention follow the single copy too. Where file and object views share one copy, a versioning, lifecycle or retention policy set on that data covers whatever was written over NFS or SMB, which closes the gap that appears when a file copy and an object copy are protected by different rules on different systems. For AI teams, it means training jobs reading over S3 see the same files that labelling or curation teams edit over SMB, without waiting for a sync cycle. The distinction from unified storage, where several protocols share capacity but not necessarily data, and the wider object storage and NAS comparison both shape which approach fits.

Protocol access in Scality RING

Scality RING, software-defined object and file storage on standard x86 servers, serves S3 with what its product page calls full S3 fidelity, and the MultiScale architecture lists S3, NFS and SMB as its protocols. Erasure coding is defined per storage class beneath whichever protocol a workload uses, so protection and capacity planning are handled once for the cluster.