Glossary

Storage compression

Storage compression reduces the number of bytes needed to hold data by encoding it with a lossless algorithm, so the original is reconstructed exactly when read. It is applied by applications, file systems, storage systems or drives, and is expressed as a ratio of original size to stored size.

Compression of backup streams is covered separately in backup compression.

Why storage compression matters for capacity at scale

Capacity at petabyte scale is bought in racks, powered in kilowatts and refreshed every few years. Fewer stored bytes cut all three, which makes compression one of the few levers that lowers cost without touching the workload. Its value depends entirely on the data: some datasets shrink to a fraction of their size and others not at all, and a capacity plan built on an average ratio that does not match the estate runs short.

How lossless compression works

Lossless compressors remove two kinds of redundancy. A dictionary stage finds byte sequences that appeared earlier and replaces each repeat with a short reference to the earlier occurrence. An entropy stage then gives shorter codes to common symbols and longer codes to rare ones. DEFLATE, Zstandard and LZ4 all follow this pattern with different balances: LZ4 favours speed at moderate ratios, while Zstandard and DEFLATE offer selectable levels that trade speed for ratio. Decompression is much faster than compression for all of them, which suits storage, where data is read more often than written. Data kept compressed also costs less to move, because replication links and migrations carry fewer bytes.

Ratios, savings and data types

A ratio of N:1 stores 1/N of the original, a saving of 1 − 1/N. Savings grow more slowly than ratios: 2:1 saves 50%, 4:1 saves 75%, and 8:1 adds only another 12.5 points.

DataTypical compressibilityReason
Text, logs, CSV, JSON, telemetryHighRepeated words, field names and structure
Database filesModerateRepeated values and empty page space among dense data
Virtual machine imagesModerate to highZeroed and repeated blocks
JPEG, video, ZIP, compressed ParquetNear 1:1Already compressed by the format
Encrypted dataNear 1:1Ciphertext has no detectable patterns

Compressing dense data spends processor time for no saving, so storage systems commonly test a sample of each block and store it uncompressed when the gain falls below a threshold.

Placement and order of operations

Inline compression runs in the write path, reducing bytes written to media at the cost of processor time on every write. Post-process compression writes data as is and compresses it later, keeping write latency unchanged but holding the full-size data until the background pass reaches it and writing it twice.

When methods are combined, the order is fixed. Deduplication runs first, because identical blocks are recognizable only before compression changes their encoding. Compression runs next. Encryption runs last, because ciphertext does not compress.

Systems compress in independent units so a read does not decompress everything before it. Larger units compress better; smaller units waste less on reads. Reading 4 KB from the middle of a 64 KB compressed unit means reading and decompressing all 64 KB, sixteen times the data requested.

What storage compression means for large-scale storage

For architects planning capacity, the composition of the data matters more than the algorithm. Much of the fastest-growing unstructured data in large estates arrives already compressed: video, images, compressed columnar files used in analytics and AI pipelines, and encrypted backups. On that data, storage-level compression returns close to nothing, while logs, text and uncompressed exports shrink well. A multi-petabyte platform sized from a blended vendor ratio, instead of from samples of the actual data, is a common source of capacity shortfalls a year after deployment.

Encryption placement decides whether compression works at all. Data encrypted by the application or client before it reaches storage is incompressible to the storage system. Estates moving to client-side encryption for sovereignty or security reasons lose storage-side savings on that data, and the capacity plan changes with it.

Compression and protection overhead multiply. Data compressed 2:1 and stored under a scheme with 1.5 times overhead occupies 0.5 × 1.5 = 0.75 of its original size on media, so a change in either factor changes the raw capacity required. Processor cost scales too: inline compression across hundreds of gigabytes per second of ingest consumes CPU that would otherwise serve requests, which is why fast algorithms dominate primary storage and higher ratios are kept for colder data.

Compression and protection overhead in Scality RING

In RING, protection overhead comes from the erasure coding scheme, which is defined per storage class. Different classes in one RING therefore carry different overheads, and any reduction applied before data reaches RING, by the application or the backup software, combines with the class overhead to set the raw capacity a dataset consumes. Planning the two together gives the physical footprint per class.