Glossary
Data loss prevention (DLP)
Data loss prevention (DLP) is a category of security controls that find sensitive information, track where it is stored and how it moves, and block or flag uses and transfers that break a defined policy. Despite the name, it deals with unauthorized disclosure; loss through deletion, corruption or hardware failure is handled by backup, replication and immutable storage.
The NIST glossary, reproducing the CNSS definition, describes DLP as protecting data in use, data in motion and data at rest.
Why DLP matters for data held at scale
DLP began at the edges of the network: email gateways, web proxies and laptop agents watching for a spreadsheet of customer records on its way out. The largest stores of sensitive data in most enterprises now sit somewhere else, in object stores, data lakes, backup repositories and the datasets feeding analytics and AI pipelines. These are written by applications, read by service accounts and shared across teams, often without anyone having catalogued what they contain.
That shift moves the hard part of DLP from catching a file in transit to knowing what is sitting in storage at all. A policy cannot protect data nobody has identified as sensitive, and at petabyte scale the unidentified portion is large.
Data in use, in motion and at rest
| State | Where it is inspected | Typical controls |
|---|---|---|
| In use | Endpoints and virtual desktops | Blocking copies to removable media, printing, clipboard and browser uploads |
| In motion | Email gateways, web proxies, network egress | Inspecting outbound traffic and attachments; blocking or encrypting matches |
| At rest | File shares, databases, object stores, SaaS applications | Scanning stored content, labeling it, restricting or quarantining it |
Endpoint and network DLP act at the moment of transfer. At-rest discovery runs in the background and feeds the other two: labels applied during a scan travel with the data and tell gateways and agents what they are handling.
How DLP recognizes sensitive content
- Patterns and checksums. Regular expressions find structured identifiers such as card or national ID numbers, and check-digit tests discard strings that only resemble them.
- Exact data match. Hashes of real records from a source database flag content that matches actual customer or employee data.
- Document fingerprinting. Hashes of sections of known sensitive documents detect copies and excerpts.
- Labels and metadata. Classification tags applied by people or tools are read directly.
- Classifiers. Trained models recognize kinds of document, such as contracts or medical forms, without fixed patterns.
Every method trades coverage against noise. Broad rules catch more real cases and raise more false alarms; tight rules cut alerts and miss more. Policies commonly start in monitor-only mode, logging matches without blocking, so the alert volume is known before anything is enforced.
Limits of content inspection
DLP acts only on content it can read. Traffic encrypted end to end is opaque unless a proxy decrypts it, and files encrypted or password-protected before they move cannot be inspected. Data also changes shape inside modern pipelines: records are compressed into columnar files, split into chunks, converted into embeddings for search or folded into training sets, and pattern-based rules lose track of it along the way.
Legitimate access is the other blind spot. Users and service accounts allowed to read data can move it through approved channels, in volumes each individual policy permits. That is the territory of insider threat programs, and the reason attackers staging data exfiltration favor encrypted channels and ordinary cloud services. Detection then leans on behavior (unusual volumes, destinations or hours), and those signals often appear first in storage access logs.
What DLP means for storage and platform teams
Storage teams rarely run the DLP product, but they own most of the surface it covers, and several consequences land on them directly.
- Discovery scans are storage workloads. Scanning content means reading it. A scanner reading 5 PB at a sustained 2 GB/s needs 5 × 10¹⁵ ÷ (2 × 10⁹) = 2.5 million seconds, about 29 days, and that read load lands on systems serving production.
- Copies multiply the scope. Each replica, backup, test copy and analytics extract holds the same sensitive records, often under different access rules.
- Access logs become evidence. Where content inspection is blind, a record of which identity read which objects, and how much, is often the only trace of a disclosure.
- Layout carries policy. Object tags, metadata and separate buckets or namespaces for sensitive datasets give DLP and access controls something to act on without reading every byte.
DLP and storage protection also answer different failures. DLP limits who sees data, while immutability and replication limit what an attacker can destroy, and regulated data at scale is exposed to both kinds of loss. The related controls are covered under cloud storage security.
Scality storage and data loss
Scality RING and ARTESCA are storage systems, and inspection of content for disclosure happens in DLP tools above the storage layer. Both address the other sense of data loss, deletion and overwrite, through S3 Object Lock in governance and compliance modes, with retention periods and legal holds, on versioned buckets. Governance-mode retention can be bypassed by any identity holding s3:BypassGovernanceRetention that sends x-amz-bypass-governance-retention:true. Compliance mode resists all users, including the account root, until the retain-until date, after which retention expires.
ARTESCA audit logs can be forwarded to Splunk, Graylog, Elasticsearch or syslog, where storage access events sit beside the alerts that DLP and other security tools send to the same platforms.














