Glossary

Personally identifiable information (PII)

Personally identifiable information (PII) is information that can distinguish or trace an individual's identity, on its own or when combined with other information linked or linkable to that person. The term comes from US federal policy; European law uses the related and broader concept of personal data.

Why PII matters for large-scale storage

PII is defined in privacy law, but most of it physically lives in storage: customer databases, log archives, call recordings, scanned documents, medical images, and the backups and replicas of all of them. Obligations attached to a person's data apply to every copy, wherever it sits, for as long as it is kept.

At enterprise scale the copies are the difficulty. One customer record created in one application ends up in logs, analytics extracts, a data lake, test environments, a second site and years of backups. Storage and platform teams inherit questions about all of them: where they are, who can read them, how long they stay and whether they can be erased on request.

PII and personal data

NIST SP 800-122 adopts a definition covering information that can distinguish or trace an individual's identity, and other information linked or linkable to that individual. Linked information is already associated with a person; linkable information identifies no one alone but can be joined to data that does. The EU GDPR, in Article 4(1), defines personal data as any information relating to an identified or identifiable natural person.

AspectPII (US federal usage)Personal data (GDPR)
TestDistinguishes or traces an individual, alone or combinedRelates to an identified or identifiable natural person
Online identifiersDepends on context and linkabilityNamed explicitly
Higher-risk dataAssessed by confidentiality impact levelSpecial categories listed in Article 9
Pseudonymized dataTreated by re-identification riskStill personal data

The GDPR definition captures more, so a dataset that falls outside PII in US usage can still be personal data in Europe.

Direct and indirect identifiers

Direct identifiers name or uniquely number a person: full name, passport number, email address, account number. Indirect identifiers, also called quasi-identifiers, point to no one alone but narrow the field when combined: postcode, date of birth, gender, job title, employer.

The narrowing is fast. About 365 × 80 = 29,200 birth dates cover people aged 0 to 80, and two recorded genders double that to 58,400 combinations. A postal area of 20,000 residents spread evenly across them averages 20,000 ÷ 58,400 ≈ 0.34 people per combination, so most residents are the only person with theirs. Removing names from a dataset rarely removes identity from it.

Machine-generated data behaves the same way. Device identifiers, IP addresses, location traces and behavioral logs identify people once joined to almost anything else, so log archives and telemetry stores often hold more personal data than their owners assume.

Pseudonymization and anonymization

Pseudonymization replaces direct identifiers with tokens while keeping a way to reverse the mapping. Anonymization removes the ability to re-identify altogether. Under the GDPR the difference is decisive: pseudonymized data remains personal data, while anonymous data falls outside the regulation. In storage terms, a pseudonymized dataset and the key table that reverses it are both sensitive, and the tokens only protect anyone while the two are held apart under separate access.

What PII means for storage at scale

  • Every copy is in scope. Daily backups kept for 30 days and replicated to a second site hold 30 × 2 = 60 copies of a record at any moment, before counting snapshots, logs and extracts.
  • Erasure requests meet retention. A record deleted in production remains in backups until each reaches the end of its retention period, and data under S3 Object Lock compliance mode stays until its retain-until date regardless of who asks. Retention periods become a privacy decision as much as a protection one.
  • Location is a legal attribute. Replicating a bucket to another country can move personal data across a border, which brings data residency and data sovereignty rules into replication topology.
  • AI pipelines create new copies. Training sets and the document stores behind retrieval-augmented generation absorb PII from their sources, and indexes built from them can surface it in generated answers.
  • Separation makes control practical. Sensitive datasets kept in their own buckets, namespaces or storage classes let access policy, encryption, retention and placement be set once for the data that needs them, and give data loss prevention scanners a defined scope.

Scality storage and personal data

Identifying and classifying PII happens in applications and data governance tools. The storage layer contributes control over where data sits, how long it is kept and who can delete it. Scality RING and ARTESCA support retention through S3 Object Lock in governance and compliance modes, with retention periods and legal holds, on versioned buckets. Governance-mode retention can be bypassed by any identity holding s3:BypassGovernanceRetention that sends x-amz-bypass-governance-retention:true. Compliance mode resists all users, including the account root, until the retain-until date, after which retention expires.

Scality ADI lists policy-enforced data residency at the namespace level, along with object-level immutability and retention enforcement, so placement and retention of personal data follow policies set on the namespace.