Glossary
Multi-protocol storage
Multi-protocol storage exposes the same data through more than one access protocol, so a file written over NFS can be opened over SMB or read as an object over S3 without a second copy being made. The system keeps one copy and translates each protocol's names, permissions and locking rules onto it.
Why multi-protocol access matters for large data sets
Data at scale rarely stays inside one protocol for its whole life. Instruments, cameras, sequencers and render farms write files over NFS or SMB. Analytics engines, AI training jobs and cloud-native applications read over S3. Windows users browse the same results over SMB. Each time data crosses one of those boundaries on single-protocol systems, it becomes a copy job.
A copy at petabyte scale is expensive in two ways. It doubles the capacity held, and it takes time: moving a 2 PB data set at a sustained 10 GB/s takes 2,000,000 GB ÷ 10 GB/s = 200,000 seconds, about 2.3 days, before the second copy is usable. It also creates two sets of permissions to keep aligned and a window in which the copies disagree. Multi-protocol storage removes the copy by letting each tool reach the original.
How one copy is presented through several protocols
| Protocol | Data model | Identity | Concurrency |
|---|---|---|---|
| NFS | Files and directories | UNIX user and group IDs; NFSv4 ACLs | Advisory byte-range locks |
| SMB | Files and directories | Windows security identifiers, NTFS-style ACLs | Share modes, byte-range locks, leases |
| S3 | Objects by key in a bucket | Access keys, users, roles, bucket policies | None per object; each PUT replaces the whole object |
The file /projects/alpha/report.pdf becomes the key projects/alpha/report.pdf in a bucket. Directories have no independent existence in S3; a listing by prefix and delimiter reconstructs them. An identity mapping, usually backed by Active Directory or LDAP, lets the system recognise one person behind a UNIX ID, a Windows SID and an S3 access key.
Where the protocol models differ
- Renames. Renaming a directory of 10,000 files is one metadata update in a file system. Where the directory exists only as a key prefix, the same rename becomes a copy and a delete per object: 2 × 10,000 = 20,000 requests.
- Names. SMB is case-insensitive and reserves characters such as
:and*; NFS and S3 are case-sensitive.Report.pdfandreport.pdfcoexist in a bucket and collide over SMB. - Permissions. A UNIX permission mode such as 750 fits in nine bits; an NTFS ACL can hold dozens of entries with inheritance. Converting the richer model to the simpler one loses detail.
- Locks. A file locked by an SMB client can still be replaced by an S3 PUT unless the system checks file locks on object writes.
- Metadata. Files carry separate access, modification and change times; an object has one last-modified time and free-form user metadata.
Implementation approaches
| Approach | Mechanism | Trade-off |
|---|---|---|
| Protocol gateway | A gateway translates file operations into object requests against a separate store | Extra hop; rename and partial-write costs land on the object store |
| Shared namespace | Each protocol service reads and writes one internal namespace directly | Consistent view; the platform carries the full reconciliation work |
| Copy and sync | Data is copied between a file system and an object store on a schedule | Simple per protocol; two copies with lag between them |
What multi-protocol access means for data pipelines at scale
For a team running ingest-to-analysis pipelines, the gain is measured in copies avoided: capacity that is never bought twice, transfer windows that disappear, and one retention and protection policy per data set. Research groups, media houses and AI teams feel this most, because their files are produced over file protocols and consumed by S3-native tools.
The semantic gaps become operational facts. A permissions model chosen at deployment decides whether an S3 bucket policy and an NTFS ACL can disagree about who reads a folder, which matters to the security review of any shared data set. Mass renames or restructuring of large trees behave very differently depending on which protocol performs them. Concurrent writes to the same file from different protocols are where edge cases surface, which is why stable large deployments tend to have one protocol writing a given data set and the others reading it.
Protection and retention follow the single copy too. Where file and object views share one copy, a versioning, lifecycle or retention policy set on that data covers whatever was written over NFS or SMB, which closes the gap that appears when a file copy and an object copy are protected by different rules on different systems. For AI teams, it means training jobs reading over S3 see the same files that labelling or curation teams edit over SMB, without waiting for a sync cycle. The distinction from unified storage, where several protocols share capacity but not necessarily data, and the wider object storage and NAS comparison both shape which approach fits.
Protocol access in Scality RING
Scality RING, software-defined object and file storage on standard x86 servers, serves S3 with what its product page calls full S3 fidelity, and the MultiScale architecture lists S3, NFS and SMB as its protocols. Erasure coding is defined per storage class beneath whichever protocol a workload uses, so protection and capacity planning are handled once for the cluster.














