Glossary

Flash cache

A flash cache is a layer of flash that holds copies of frequently or recently used data in front of a slower, larger tier, usually hard disks, so that repeated requests are served at flash speed. The slower tier keeps the authoritative copy of the data.

Why flash caching matters for large storage systems

A few percent of capacity in flash, placed in front of disk, can serve the busiest data at flash speed while the system keeps the cost profile of disk. Everything depends on reuse. Data read once gains nothing from passing through a cache, and data read repeatedly within a short period gains the most. Whether a flash cache pays off in a large estate is therefore a question about access patterns, more than about the cache itself.

Where flash caches sit

  • Host-side cache: flash in the application server, saving a network trip, but private to that server.
  • Array or controller cache: flash inside a hybrid flash array, shared by all hosts.
  • Storage-node cache: flash in each server of a scale-out system, caching data and especially metadata held on that server's disks.
  • Edge or gateway cache: flash in a cloud storage gateway or remote-site appliance, holding local copies of data that lives in a distant or cloud tier.

Read and write caching modes

ModeWhen a write is acknowledgedTrade-off
Read cacheWrites do not use the cacheAccelerates rereads only
Write-throughAfter the slower tier has the dataSafe; no gain for write latency
Write-backOnce the data is on flashFast writes; flash holds the only copy until destaged, so it needs mirroring or power-loss protection
Write-back, mirroredOnce the data is on flash in two nodesSurvives the loss of one server; adds a network hop to each write
Write-aroundAfter the slower tier has the data, bypassing flashKeeps write-once data from pushing hot data out

Eviction policy decides what leaves when the cache is full, commonly least recently used or a frequency-weighted variant. Admission policy decides what enters at all; many caches skip large sequential reads, which disks stream well and which would otherwise flush the cache.

Hit ratio, working set and warm-up

The hit ratio is the share of reads the cache serves. It is high when the working set, the data actually touched within a period, fits in the cache, and it falls quickly once the working set outgrows it. A 5 PB system whose weekly working set is 3% of its data needs about 150 TB of cache to hold it.

A cache also has to be filled. Warming 100 TB of flash from disk at an aggregate 2 GB/s takes 100 × 1012 ÷ (2 × 109) = 50,000 seconds, roughly 14 hours, during which requests run largely at disk speed.

What flash caching means for AI and data lake workloads

Repeat-read workloads benefit most. Metadata lookups, dashboards over recent data, model files loaded by many inference servers, and the frequently retrieved documents in a RAG corpus all reread a compact set of data, and a cache serves them well.

Scan and shuffle workloads defeat it. An AI training job reads its whole dataset in shuffled order every epoch; when the dataset is larger than the cache, each item is evicted before its next read and the hit ratio approaches zero. Analytic full scans, large restores and migrations behave alike. For those jobs, performance comes from the aggregate throughput of the capacity tier, or from holding the dataset on flash outright as all-flash storage.

Events reset the cache. After a node replacement, failover or restart, a cold cache spends hours warming, and the performance seen in that window is the performance of the disks behind it.

Write-back adds a dependency. Until destaged, recent writes exist only in the cache, so the design of the cache's own protection determines what a failure costs.

Large caches approach the cost of flash capacity. Once the working set is a large share of the total, the cache needed to hold it costs close to what an all-flash tier for the same data would cost, while still paying the disk penalty on every miss. At that point a separate flash tier with tiering or a dedicated flash class often fits the data better than a bigger cache.

Working sets drift. A cache sized for last year's projects can be undersized after a new pipeline starts reading historical data, and the change shows up as rising latency with no change in capacity used.

Flash caching and Scality RING

Scality RING runs on standard x86 servers, and the flash in each server is chosen per deployment. The large RING customers Scality describes run hybrid servers mixing flash and hard drives. For small-object, latency-sensitive workloads, RING XP is an all-flash configuration on NVMe servers in which the capacity layer is itself flash, so there is no slower tier behind it to miss to.