RAMindex

RAM caching for mechanical drives: how to set it up

By Harry Saarinen · Updated

The cache you already have

You already have one, and you did not switch it on. Every mainstream operating system keeps a read cache for storage in main memory, by default, with no configuration, and it is already using every byte you are not. On Linux that is the page cache, described by the kernel in one sentence: “Whenever a file is read, the data is put into the page cache to avoid expensive disk access on the subsequent reads” (Documentation/admin-guide/mm/concepts.rst). On Windows the Cache Manager does the same job by mapping views of files into ordinary physical pages, which then sit in the system working set, the standby list or the modified page list. Neither system asked permission, and neither has an off switch worth using.

That answers the question as it is normally asked, and it turns out to be the wrong question. Nobody decides to cache storage in RAM. The kernel decided, at boot, for every file on the machine. What is left is narrower and harder: is the cache big enough, and is it holding the right bytes? Those are two different questions with two different measurements, and only one of them is usually solved by tuning.

Free memory is memory doing nothing

The number most people quote is the wrong column. free(1) defines free as “Unused memory (MemFree and SwapFree in /proc/meminfo)” and available as an “Estimation of how much memory is available for starting new applications, without swapping” - the latter counts reclaimable page cache, the former does not. Current procps-ng also documents used as total minus available, not as total minus free minus buffers minus cache, which is why old blog arithmetic no longer reconciles. A healthy Linux box shows a small free and a large buff/cache. That is the cache working.

Windows says the same thing in different words. Microsoft’s own walkthrough of the Task Manager memory view gives Cached as the sum of the system working set, the standby list and the modified page list, and Available as the sum of the standby pages, the free pages and the zero page lists. The standby pages appear in both totals, deliberately: they hold cached file data and they are the first memory handed to a new allocation. “Cached 12 GB” is not 12 GB spent.

One caveat before anyone treats Cached as a spare-capacity figure: on Linux it includes tmpfs and shmem, which cannot be dropped, only swapped. Subtract Shmem from Cached before calling the remainder free.

The value of all this is measurable in two commands, and measuring it is the honest way to start:

# cold: discard clean cache, then read
sync; echo 3 | sudo tee /proc/sys/vm/drop_caches > /dev/null
time cat bigfile > /dev/null      # served by the device
time cat bigfile > /dev/null      # served by RAM

Do that once, to see the size of the effect on your own hardware, and then stop touching drop_caches. The kernel’s own documentation is blunt: “This file is not a means to control the growth of the various kernel caches (inodes, dentries, pagecache, etc…) These objects are automatically reclaimed by the kernel when memory is needed elsewhere on the system.” It also logs the culprit by command and PID, so dmesg | grep drop_caches will tell you whether a cron job on your machine is throwing away a warm cache every hour to make a graph look tidy.

Three different things get called RAM caching

Keep them apart, because they fail in completely different ways and the safeguards for one are useless for the others.

A read cache of clean pages. A clean file page is a duplicate of bytes that exist on the device. Reclaim unlinks it from the LRU and reuses the frame with no I/O at all. Losing the whole cache costs exactly one re-read per page and nothing else. There is no durability question here, no data to protect and no safeguard to configure. This is what most people mean, and it is the only one of the three that is free to get wrong.

A write cache of dirty pages. A successful write(2) guarantees only that the kernel has taken custody of your bytes. Until writeback, the page in RAM is the authoritative copy. The kernel defaults, read from the initialisers in mm/page-writeback.c, are dirty_expire_interval = 30 * 100 centiseconds and dirty_writeback_interval = 5 * 100: the flusher threads wake every 5 seconds, and the comment above the expire interval in that file calls it “the longest time for which data is allowed to remain dirty”. It is a ceiling on the age of a dirty page under periodic writeback rather than a minimum delay: once dirty memory crosses dirty_background_ratio the flushers start writing out regardless of age, and fsync or sync shortcuts it entirely. On a lightly loaded machine the ceiling is what you get, so an acknowledged write can sit in RAM for around half a minute before anything reaches the platter. That window is tunable, it is bounded by fsync, and it is the subject of a later section. It is not avoidable.

An explicit in-memory store. tmpfs, /dev/shm, a brd RAM disk, Redis with persistence turned off. Here there is no shadow copy anywhere: the data exists only in volatile memory until something deliberately writes it down. These also fail differently under pressure. A tmpfs instance stops at its size= and refuses further writes; ramfs takes no size parameter at all, its pages are unevictable, and its only limit is how much memory the machine has. The kernel’s warning about sizing is blunt: “if you oversize your tmpfs instances the machine will deadlock since the OOM handler will not be able to free that memory.”

Read cache, write cache, RAM disk. One costs a re-read, one costs seconds of acknowledged writes, one costs everything.

The uncomfortable part

For most readers the entire optimisation is: buy more memory. The page cache needs no tuning to be effective, it already holds everything it can, and if the hot data does not fit then no sysctl will make it fit. The cheapest way to raise a hit ratio is almost always more capacity, and capacity has a published price per gigabyte you can weigh against the tuning time you were about to spend - every generation, ranked, is at RAM prices per gigabyte. Everything after this section is worth doing only once that question is settled.

Capacity above a desktop platform’s ceiling is a different purchase, not just a larger one: the big modules are registered or load-reduced, and they need a memory controller and a socket that accept them. Which modules a given machine takes is recorded per server model on the compatibility pages, and the trade-offs between the module types are set out in registered vs unbuffered.

There is also a floor under all of this. If the hot working set runs to tens of terabytes, memory is not the answer at any price, and the next question is what sits under the disks rather than above them - a flash tier in front of spinning media. The two questions are the two terms of one equation: this one raises the hit ratio, that one lowers the cost of a miss.

What a RAM cache is worth

About a thousandth of the latency of the flash underneath it and a hundred-thousandth of the latency of a disk, and neither figure decides anything. The value of a cache is set by its miss rate, not by its hit speed, and the miss rate is set by one thing: whether the working set fits.

That is why “RAM is fast” is a useless sentence. Fast is the constant in the equation. The variable is a fraction, and the fraction moves by orders of magnitude for reasons that have nothing to do with the memory.

The ladder, with sources

Tier Figure Where it comes from
DRAM random load 81 ns Measured, DDR4 on Cascade Lake, arXiv:1903.05714 §3.1
Datacentre NVMe, 4 KB 130 us at four nines Solidigm D7-P5520 product brief endnotes
Previous-generation NVMe 230 us at four nines Solidigm D7-P5510, same brief
SATA SSD, 4 KB 180 to 240 us at five nines Kingston DC600M datasheet, by capacity
7200 rpm nearline, rotation only 4.16 ms Seagate Exos 7E10 datasheet and SAS product manual §4.1

One gap in that table is more informative than the entries, and the SATA row is not it: Kingston publishes both a typical and a five-nines latency for the DC600M, while the Solidigm D3-S4520 product brief publishes IOPS and endurance and no latency at all, so the class is inconsistent. Measure your own with fio at --rw=randread --bs=4k --iodepth=1 --direct=1 and publish the flags beside the number.

The gap that is left is worse. Seagate specifies 4.16 ms of average rotational latency for the Exos 7E10 and specifies no average seek time anywhere in the product manual. Rotational latency is half a revolution at 7200 rpm; a random access is rotation plus seek. The 8.5 ms seek figure that circulates for nearline drives is no longer in any current datasheet, and this guide will not claim one. Use 10 ms as a round working number for a random 4 KB read and know that it is a modelling assumption, not a specification.

The arithmetic that decides

T_eff = h * T_hit + (1 - h) * T_miss

T_hit = 0.1 us   (the DRAM access; a page-cache hit through read(2) also
                  costs a syscall and a copy, so treat this as a floor)

Backed by NVMe, T_miss = 100 us
  h = 0.90      0.090 + 10.000  = 10.09 us
  h = 0.95      0.095 +  5.000  =  5.10 us
  h = 0.99      0.099 +  1.000  =  1.10 us
  h = 0.999     0.100 +  0.100  =  0.20 us
  h = 0.9999    0.100 +  0.010  =  0.11 us

Backed by 7200 rpm, T_miss = 10000 us
  h = 0.95      0.095 + 500.0   = 500.10 us
  h = 0.99      0.099 + 100.0   = 100.10 us
  h = 0.999     0.100 +  10.0   =  10.10 us
  h = 0.99999   0.100 +   0.1   =   0.20 us

The miss term stops dominating when (1 - h) < T_hit / T_miss:
  NVMe    at 100 us   ->  h > 99.9 per cent
  7200rpm at 10 ms    ->  h > 99.999 per cent

Read the shape before the numbers. Every row is the miss term with a rounding error attached. At 90 per cent on NVMe the cache contributes under one per cent of the answer; the drive contributes the rest.

The claim you will see made from a table like this is that the hit ratio matters more than the class of storage underneath. Check it. Going from 95 to 99 per cent on NVMe buys 4.6 times. Going from a 7200 rpm disk to that NVMe at a fixed 95 per cent buys 98 times. On that comparison the storage wins, and the margin is a factor of twenty. Anyone who tells you otherwise has not done the multiplication.

Three things are true instead, and they are more useful. Within one class of device, hit ratio dominates: four percentage points is worth 4.6 times, while the best case for swapping a SATA SSD for an NVMe on the misses is a small multiple that the datasheets will not let you pin down. Second, the returns accelerate rather than flatten, because each decade of miss rate removed buys a whole decade of effective latency, right up to the crossover. Third, below the crossover the two levers are the same lever - and that is precisely why a cache in front of spinning disks is a different question from this one. If the misses are going to a disk array, lowering T_miss with flash is worth more than any plausible amount of DRAM, and the companion piece on arrays with NVMe drives is the article you want first.

Bandwidth or latency

These are different ceilings and a workload only ever hits one of them.

Per 64-bit DDR5 channel:  GB/s = MT/s * 8 bytes
  DDR5-4800   4800 * 8 =  38.4 GB/s
  DDR5-5600   5600 * 8 =  44.8 GB/s
  DDR5-6400   6400 * 8 =  51.2 GB/s

  8-channel socket at DDR5-4800    8 * 38.4 = 307.2 GB/s
 12-channel socket at DDR5-6400   12 * 51.2 = 614.4 GB/s

PCIe 4.0 x4 = 16 GT/s * 4 * 128/130 / 8 =  7.88 GB/s
PCIe 5.0 x4 = 32 GT/s * 4 * 128/130 / 8 = 15.75 GB/s

AMD publishes “up to 614 GB/s” for EPYC 9005 over twelve DDR5 channels, which is the same number the formula gives to three significant figures, so the method is sound. Nameplate is a ceiling; published streaming runs land well short, commonly between 60 and 85 per cent, and it should be quoted as a ceiling rather than as a throughput.

The population arithmetic matters more than the grade. One module of DDR5-6400 gives 51.2 GB/s. Two of DDR5-4800 give 76.8 GB/s. The slower kit in both slots beats the faster kit in one by half. On an eight- or twelve-channel socket with four modules fitted you are running at a third to a half of the ceiling you paid for, and the fix is more DDR5 modules of the grade you already have rather than a faster grade.

Which ceiling binds is measurable. Divide bytes moved by elapsed time. Land near the nameplate fraction above and you are bandwidth-bound, and more channels help. Land at a few per cent of it with cores stalled on dependent loads, and you are latency-bound; more channels do nothing and the hit ratio is the only lever.

The working set, and the cliff

A cache smaller than the working set does not degrade gently. Take the textbook case: a loop that touches N+1 blocks in strict order, against an LRU cache that holds N. Every block is evicted exactly one access before it is needed again, and the hit ratio is zero. One block of shortfall, total loss. Real workloads smooth this, but the shape survives, and it is why “another 25 per cent of RAM changed nothing” is such a common report.

Measure before buying, because the measurement exists on every platform.

  • Linux, the direct answer: the cachestat(2) syscall, real since 6.5. Its nr_recently_evicted field counts pages in the range whose eviction was recent enough that their return would indicate memory pressure - the shortfall, in pages.
  • Linux, the pressure answer: /proc/pressure/memory. A full avg60 above zero means tasks are genuinely stalling on memory. Flat means they are not.
  • Linux, the reclaim answer: sar -B for majflt/s and %vmeff. Low efficiency under pressure means the kernel is scanning hard and finding little.
  • ZFS: arcstat’s mrug and mfug columns, the MRU and MFU ghost-list hits. Those are requests that would have hit had the ARC been larger: the same shortfall cachestat reports, priced in hits rather than in pages.
  • Windows: \Memory\Long-Term Average Standby Cache Lifetime (s). Microsoft’s own threshold is 1800 seconds; below it the cache is being churned.

If the ghost hits are flat and pressure is zero, more memory buys nothing and the money belongs downstream. If they are not, size the purchase to the measured shortfall and check the platform will take it: a working set in the hundreds of gigabytes means registered or load-reduced modules and a socket that accepts them, which is a different machine from the one on your desk.

The server memory capacity limits guide has the per-platform ceilings, the compatibility pages list what each server model takes, and price per gigabyte has what that capacity costs today.

Where it is transformative, and where it is not

Transformative wherever accesses have reuse. Metadata-heavy trees are the clearest case: the dentry and inode caches turn a path lookup into a hash lookup, and mail spools, rsync runs and many-small-file build trees live entirely in that behaviour. Databases with index locality, compile trees read over and over, and VM hosts where dozens of guests share the same base images all qualify. So does anything whose hot fraction is small and stable, which is most transactional work.

Nearly worthless where there is no second touch. A single-pass sequential read gets exactly one hit per page and pays for the frame. Media streaming, backup targets, large restores and bulk ingest are all read-once by construction, and the cache is not neutral there - it is negative, because each read-once page evicts a page that had reuse. Every serious cache has a defence against precisely this and each picked a different one: the Linux inactive file list, ARC’s ghost lists, InnoDB’s 3/8 midpoint insertion, PostgreSQL’s 256 kB ring buffer for sequential scans. For a backup target the honest answer is that memory is not the purchase at all, and capacity at the lowest price per terabyte is: see every drive ranked by price per terabyte.

The Linux page cache in detail

Clean. Almost everything Linux reports as cache is clean file data, and a clean page is dropped by unlinking it from an LRU list, with no disk I/O at all. It is free memory that happens to still be useful.

That answers the question people ask when free frightens them, and it is the wrong question. The number that decides whether the machine is safe is not how much is cached but how much of the cache is dirty, because dirty pages are the only ones that cost latency to reclaim and the only ones whose contents exist nowhere else yet.

Clean, dirty, and what the meminfo fields really count

Documentation/admin-guide/mm/concepts.rst puts the mechanism in two sentences: file data read from disk is put into the page cache “to avoid expensive disk access on the subsequent reads”, and written pages “are marked as dirty and when Linux decides to reuse them for other purposes, it makes sure to synchronize the file contents on the device with the updated data”. Reclaiming a clean page is bookkeeping. Reclaiming a dirty one is a write, and the allocation waits for it.

Every /proc/meminfo field that carries a unit is labelled kB and means kibibytes, 1024 bytes, whatever the label says; the HugePages_* entries are bare page counts with no unit at all. The seven that matter here, in the kernel’s own words from Documentation/filesystems/proc.rst, except the two file LRU fields, which that file does not describe at all:

Field Kernel’s description
Cached “In-memory cache for files read from the disk (the pagecache) as well as tmpfs & shmem. Doesn’t include SwapCached.”
Buffers “Relatively temporary storage for raw disk blocks shouldn’t get tremendously large (20MB or so)”
Dirty “Memory which is waiting to get written back to the disk”
Writeback “Memory which is actively being written back to the disk”
SReclaimable “Part of Slab, that might be reclaimed, such as caches”
Active(file) not described; proc_meminfo(5) says “[To be documented.]”
Inactive(file) not described; proc_meminfo(5) says “[To be documented.]”

Those last two are genuinely undocumented: proc.rst defines Active and Inactive but never the file split, and the man page marks both of the split fields as pending. What they count has to be taken from behaviour: the file LRU is split in two, a page lands on the inactive list on first touch and is promoted to the active list on a second reference, and reclaim scans the inactive list first. A kernel running MGLRU replaces that two-list scheme with multiple generations, so the counters keep accounting while no longer describing the algorithm in use. The kernel’s MGLRU document gives no version number for its arrival, so read /sys/kernel/mm/lru_gen/enabled on the machine in front of you rather than trusting a number in an article.

A constructed reading for a 64 GB server, illustrative rather than captured from a machine, trimmed to the relevant lines:

MemTotal:       65806652 kB      Active(file):   12884992 kB
MemFree:         1042184 kB      Inactive(file): 32551040 kB
MemAvailable:   48231920 kB      Dirty:            624880 kB
Buffers:          812344 kB      Writeback:         18432 kB
Cached:         49238112 kB      Slab:            3128880 kB
Shmem:           4194304 kB      SReclaimable:    2471616 kB

Cached is 47.0 GiB, but it includes tmpfs and shmem, and shmem is not reclaimable by dropping it - it needs swap. Subtract the 4 GiB of Shmem and the real file cache is about 43 GiB. Dirty is 610 MiB, Writeback 18 MiB: around 1.4 per cent of the cache is at risk and the rest is free for the taking. Inactive(file) runs two and a half times Active(file), which says most of this cache is single-touch streaming data that reclaim will discard cheaply rather than a hot working set. A cache with that shape is not a working set waiting for more DIMMs: if the data streams past whatever memory is fitted, the question is how fast the device underneath can serve it rather than how much memory is fitted above it. The LRU counters and Cached do not tie out to an exact identity, so do not treat them as one. SReclaimable is 2.4 GiB of reclaimable slab, mostly dentry and inode caches, which is where a find over a large tree goes, not into Cached. And MemFree at 1 GiB is not a problem: free(1) now derives used as total minus available, and MemAvailable at 46 GiB is the honest figure.

The writeback machinery, and the window it leaves open

Defaults come from the initialisers in mm/page-writeback.c, not from vm.rst, which describes these knobs without stating their values.

sysctl Unit Default
vm.dirty_background_ratio per cent of dirtyable memory 10
vm.dirty_ratio per cent of dirtyable memory 20
vm.dirty_background_bytes bytes 0, disabled
vm.dirty_bytes bytes 0, disabled
vm.dirty_expire_centisecs centiseconds 3000, so 30 s
vm.dirty_writeback_centisecs centiseconds 500, so 5 s

Three things about that table are routinely misreported. The percentages are not of installed RAM: vm.rst defines them against “total available memory that contains free pages and reclaimable pages”, which moves as the workload moves. The ratio and bytes forms are mutually exclusive - writing one zeroes the other on read. And the two thresholds differ in kind. Crossing the background threshold wakes flusher threads and the writer runs on unthrottled. Past the midpoint of the two thresholds - dirty_freerun_ceiling() in mm/page-writeback.c returns (thresh + bg_thresh) / 2, so about 15 per cent on stock settings - balance_dirty_pages() begins pausing the writing process, and the pauses lengthen as dirty climbs towards dirty_ratio. Either way the delay lands in the application, at the worst possible moment.

The arithmetic that matters on a large machine, with the dirtyable figure and the drain rate both assumed rather than read:

dirtyable memory, as a round figure        ~50 GiB
dirty_background_ratio = 10  ->  flushing starts at   ~5 GiB dirty
dirty_ratio            = 20  ->  writers throttled at ~10 GiB dirty

10 GiB = 10,737,418,240 bytes
drained at an assumed 200 MB/s: 10,737,418,240 / 200,000,000 = 53.7 s

On stock settings, a 64 GB machine will happily hold several gigabytes of acknowledged writes that have never reached a disk. The time exposure is separate from the volume and smaller: a page must age dirty_expire_centisecs before it is eligible, and flusher threads look every dirty_writeback_centisecs, so the loss on a power cut is up to 30 s plus one writeback interval of writes the application was told had succeeded. Setting vm.dirty_writeback_centisecs to 0 disables periodic writeback altogether and makes that window unbounded until a threshold or an explicit sync is hit. The exposure scales with installed memory, which means the machines built to hold the largest caches - the ones taking registered or load-reduced modules rather than unbuffered ones, for the reasons set out in registered vs unbuffered - carry the largest window by default.

The standard fix is to replace the ratios with absolute limits sized to one to a few seconds of the device’s write bandwidth, then verify that the kernel agrees:

# background flush at 512 MiB, hard throttle at 2 GiB
sysctl -w vm.dirty_background_bytes=536870912
sysctl -w vm.dirty_bytes=2147483648

# read the computed thresholds back, in pages
grep -E 'nr_dirty_(background_)?threshold' /proc/vmstat
# nr_dirty_threshold 524288   ->   524288 x 4096 = 2 GiB

Persist it in /etc/sysctl.d/. Reading the threshold back from /proc/vmstat is the step that proves the setting actually took, and the one most easily skipped.

What fsync, fdatasync and O_DIRECT actually promise

A successful write(2) guarantees that the kernel has taken custody of the bytes. It guarantees nothing about durability. fsync(2) “transfers (‘flushes’) all modified in-core data” for the file “to the disk device”, flushes the file’s metadata, includes “writing through or flushing a disk cache if present”, and “blocks until the device reports that the transfer has completed”. fdatasync(2) skips metadata “unless that metadata is needed in order to allow a subsequent data retrieval to be correctly handled” - so extending a file still forces a metadata flush, and fdatasync only saves an I/O on overwrites in place. That is exactly why databases pre-allocate their files.

Two traps. Calling fsync() on a file “does not necessarily ensure that the entry in the directory containing the file has also reached disk. For that an explicit fsync() on a file descriptor for the directory is also needed” - the write-to-temp-then-rename pattern is not crash-safe without it. And since Linux 4.13, writeback errors are reported “to all file descriptors that might have written the data which triggered the error”; before that, one descriptor could consume the error and leave every other one believing the write had landed.

O_DIRECT is orthogonal to all of this. It exists to “minimize cache effects”, does its I/O “directly to/from user-space buffers”, and the man page is explicit that it does not provide the guarantees of O_SYNC, which must be combined with it for synchronised I/O. It bypasses the page cache; it does not bypass the drive’s volatile write cache. It imposes filesystem- and device-dependent alignment restrictions on buffer address, length and file offset, and mixing it with buffered I/O on overlapping regions of the same file is documented as something applications should avoid. O_SYNC gives file integrity, O_DSYNC data integrity, matching fsync and fdatasync respectively. O_DIRECT says “do not cache this, I cache it better myself”. fsync says “make it durable now”. Conflating the two is a common and expensive error.

One consequence worth carrying into the rest of this guide: for the seconds it sits dirty, the page in RAM is the authoritative copy of that data, and a bit flipped in it is written to storage as truth. That is an argument for ECC, not an argument about caching.

drop_caches is a diagnostic, not a tuning step

vm.rst says so flatly, and it is quotable verbatim: “This file is not a means to control the growth of the various kernel caches (inodes, dentries, pagecache, etc…) These objects are automatically reclaimed by the kernel when memory is needed elsewhere on the system.” Writing 1 drops pagecache, 2 drops reclaimable slab, 3 drops both; it is “a non-destructive operation and will not free any dirty objects”, which is why the incantation begins with sync. The kernel also logs who did it, by command and PID: cat (1234): drop_caches: 3. A dmesg | grep drop_caches that comes back full is the signature of a cron job destroying a warm cache on a schedule to make a monitoring graph look better.

The one legitimate use is a cold baseline for measurement: sync; echo 3 > /proc/sys/vm/drop_caches, then the timed run, then the same run warm. To evict a single file rather than everything, posix_fadvise(2) with POSIX_FADV_DONTNEED after an fsync is the precise tool, and POSIX_FADV_WILLNEED is the honest way to warm one.

The other knobs, and the symptom that justifies each

Leave them alone. Every tunable below ships with a default that is right for the machine you have not yet measured, and each has exactly one honest reason to change it: a reading that says it is wrong.

That answers the question as it is usually asked, and it is the wrong question. The interesting part is never the value. It is the observation that licenses the value. A knob with no symptom attached is a knob nobody should turn, so each one here comes with its default, what it actually weighs, and the reading that justifies moving it.

vm.swappiness

Default 60. Range 0 to 200.

The range is the whole story, and it is why most advice about this knob is stale. Documentation/admin-guide/sysctl/vm.rst calls swappiness “the rough relative IO cost of swapping and filesystem paging”: at 100 the VM “assumes equal IO cost and will thus apply memory pressure to the page cache and swap-backed pages equally; lower values signify more expensive swap IO, higher values indicates cheaper”. It is a price ratio between two evictions, not a dial marked eagerness. Dropping a clean file page costs a re-read. Swapping an anonymous page costs a write now and a read later. Swappiness says which of those you would rather pay.

Values above 100 are legitimate, and the kernel document derives one:

swap random IO is 2x faster than filesystem IO
  x + 2x = 200   ->   2x = 133.33
  vm.swappiness = 133

That case is real whenever swap is zram or zswap, where the swap write never leaves DRAM. The old instruction to set swappiness to 0 was written when the range stopped at 100 and zero read as the floor. It is not an off switch: the doc’s own wording is conditional - “At 0, the kernel will not initiate swap until the amount of free and file-backed pages is less than the high watermark in a zone”. Reports circulate of zero producing OOM kills under cgroup v2 where an older kernel would have swapped; this guide will not attach a version number to that, because the published 0 to 200 range is already sufficient evidence that the advice predates the kernel you are running.

Symptom for lowering it: vmstat 1 shows sustained si and so next to a large cache column, meaning anonymous pages are leaving while file cache stays. Symptom for raising it past 100: your swap is compressed and in memory, or, in the kernel document’s other case, on a device faster than the filesystem it competes with. Nothing else.

Adjacent, and usually forgotten: vm.page-cluster, default 3, logarithmic, so eight pages are read per swap-in fault. Against zram there is no seek to amortise and 0 is the reasoned setting, though the kernel document defines the knob without making that recommendation.

vm.vfs_cache_pressure

Default 100. The dentry and inode caches are the cheapest cache in the system and the one bad advice most often throws away.

The dentry cache turns a path lookup into a hash lookup; the inode cache holds what a stat() would otherwise fetch from the filesystem. Both land in SReclaimable, not Cached, which is why a find / across a large tree grows Slab while Cached stays flat, and why anyone reading free alone concludes the metadata cache is not there. vm.rst: at the default the kernel reclaims dentries and inodes “at a ‘fair’ rate with respect to pagecache and swapcache reclaim”; lower values make it “prefer to retain dentry and inode caches”, and raising it well past the denominator “may have negative performance impact”.

Zero is not “keep everything” in any useful sense. The document is blunt about it: at zero the kernel “will never reclaim dentries and inodes due to memory pressure and this can easily lead to out-of-memory conditions”.

Recent kernels add vm.vfs_cache_pressure_denom, default 100, which is also its minimum. The effective pressure is the fraction vfs_cache_pressure / vfs_cache_pressure_denom, so a larger denominator buys resolution below one per cent. Run ls /proc/sys/vm/ before writing to it; older kernels do not have it.

Symptom: a metadata-heavy workload whose second pass is barely faster than its first - rsync, a backup walk, a many-small-file build, a mail spool. Time a cold find over the tree, repeat it, and compare. If the warm run is not dominated by dentry hits, 50 is defensible. The mirror image, a database with one large hot data set, argues for leaving it at 100, because there the page cache should win the contest.

Readahead

VM_READAHEAD_PAGES is 128 KiB (32 pages at a 4 KiB page size), but that is a floor, not the number you will read back. blk_apply_bdi_limits() in block/blk-settings.c takes the largest of that floor, twice the device’s optimal IO size and whatever is already set, and a rotating disk that advertises no optimal IO size is given twice its maximum request size instead, so mechanical disks commonly come up well above 128 KiB. Read the device before assuming a number. Two interfaces, two units, and the units are the trap:

/sys/block/<dev>/queue/read_ahead_kb    KiB
blockdev --getra / --setra              512-byte sectors
128 KiB = 256 sectors                   same setting, different number

Raising it pays when reads are large, sequential and served by a device that is slow to get going: a restore stream, a media scan, a sequential backup off rotating disks. It wastes memory when reads are small and random, because every fault pulls the whole window and those pages evict working set. You lose twice, once in bandwidth and once in the cache entries that were dropped to make room.

It is per device and it does not propagate. A device-mapper or MD stack carries its own read_ahead_kb on the top-level device; setting it on the members changes nothing the filesystem will ever see. That failure is common and silent.

Where the application can speak for itself, posix_fadvise(2) beats tuning the device: on Linux POSIX_FADV_SEQUENTIAL doubles the readahead window for one file and POSIX_FADV_RANDOM “disables file readahead entirely”, per file, without imposing that choice on every other reader of the same disk.

Symptom for raising it: sequential throughput below the device’s rating at queue depth 1 with the CPU idle. Symptom for lowering it: a random-read workload whose hit ratio will not climb while nr_file_pages churns.

vm.min_free_kbytes and the watermarks

This is the reserve the allocator keeps so reclaim has somewhere to stand. Raising it takes memory away from the cache permanently, and vm.rst is blunt in both directions: “Setting this too high will OOM your machine instantly”, and below 1024 KB the system risks deadlock.

Two neighbours price the same thing more precisely. vm.watermark_scale_factor, default 10, expressed in fractions of 10,000, so the gap between watermarks is 0.1 per cent of memory; maximum 3000. vm.watermark_boost_factor, default 15000, so up to 150 per cent of the high watermark is reclaimed when fragmentation is detected.

Symptom: page allocation failure in dmesg, or higher-order allocation failures under bursty network or storage load, on a machine with plenty of reclaimable cache. That is kswapd waking too late. Raise watermark_scale_factor first and in small steps: it widens the band kswapd works in without pinning a fixed slab of RAM out of the cache forever.

Overcommit

vm.overcommit_memory defaults to 0, “Heuristic overcommit handling. Obvious overcommits of address space are refused.” Mode 1 always overcommits. Mode 2 never does, and caps the total at

CommitLimit = swap + (vm.overcommit_ratio% of RAM)   # ratio default 50

with vm.overcommit_kbytes as the absolute alternative. /proc/meminfo reports CommitLimit and Committed_AS; that pair, not free, is the reading. Mode 2 ignores MAP_NORESERVE, and two reserves are held out of what is left: vm.admin_reserve_kbytes, default min(3% of free pages, 8MB), so root can still log in and kill something, and vm.user_reserve_kbytes, default min(3% of the current process size, 128MB), which applies in mode 2 only.

Symptom for mode 1: a process that forks a copy of itself to write a snapshot and fails on memory it was never going to touch. Redis states the requirement outright, “Add vm.overcommit_memory = 1 to /etc/sysctl.conf”, because an RDB save or AOF rewrite is fork() plus copy-on-write and the heuristic refuses a reservation the child will not spend. Symptom for mode 2: a machine whose whole job is to hold a cache and which must fail an allocation rather than let the OOM killer pick a victim. The PostgreSQL manual asks for exactly that, sysctl -w vm.overcommit_memory=2, on the grounds that, while it will not prevent the OOM killer from being invoked altogether, it lowers the chances significantly. Expect refusals well before RAM runs out unless you raise the ratio, and expect anything built on sparse mappings to break.

Every knob above redistributes pressure. Not one of them adds a page. If the measurement that sent you here was a demand-miss rate that will not fall, or ghost-list hits that keep arriving, the fix is modules rather than sysctls - and past what an unbuffered board will hold, which board vendors now put at 256 GB across four 64 GB modules on current consumer DDR5 platforms and 128 GB on most DDR4 desktops, that means a platform which accepts buffered memory, registered modules or, at the top capacities, load-reduced ones, because the electrical load of the ranks is what caps an unbuffered board long before the address space does. Registered vs unbuffered covers which of the two your board will actually take, and the server compatibility pages list what each model accepts per channel. And when the working set will never fit at any module count the board accepts, the tier below memory is what you are buying rather than another sysctl.

Explicit RAM storage and compressed memory

Mount a tmpfs and copy the files onto it. That is the answer, and it is the one technique in this guide that is not caching. A cache holds a second copy of data that exists elsewhere; a RAM disk holds the only copy. Putting anything on one that you cannot regenerate is a durability decision wearing the costume of an optimisation.

tmpfs against ramfs

fs/ramfs/inode.c defines exactly one mount parameter: mode. There is no size, no nr_inodes, no quota, no NUMA policy. ramfs_get_inode() calls mapping_set_unevictable(), so its pages are never on an LRU, never reclaimed, never swapped. The consequence is the one that matters: a ramfs has no ENOSPC. It grows until the machine is out of memory, and the OOM killer cannot free what it finds. The kernel’s own words: “The size limit of a ramfs filesystem is how much memory you have available, and so care must be taken if used so to not run out of memory.”

tmpfs is the same idea with accounting attached. Its pages live in the page cache, count as Shmem in /proc/meminfo and Shared in free(1), and can be written to swap. When it fills, mm/shmem.c reports the accounting failure as -ENOSPC, an ordinary filesystem error the writer can handle. With neither size nor nr_blocks given the default is size=50% of physical RAM, with nr_inodes defaulting to half the physical RAM pages. It can be resized on remount but never shrunk below current usage, noswap exists since Linux 6.4, huge= defaults to never, and there is no direct IO path at all, so O_DIRECT on tmpfs fails.

ramfs tmpfs
Mount options mode only size, nr_inodes, noswap, huge, mpol, quotas
Pages unevictable page cache, reclaim via swap
When full grows until OOM -ENOSPC to the writer
Resize no on remount, not below usage
O_DIRECT no no

The safeguard sentence is the kernel’s: “If you oversize your tmpfs instances the machine will deadlock since the OOM handler will not be able to free that memory.” Always set size=. For a sense of what a sane default looks like, systemd’s own tmp.mount sets size=50% and nr_inodes=1m on top of mode=1777,strictatime,nosuid,nodev, and its per-service credential mounts are 1 MiB each with noswap.

brd, when a block device is the point

modprobe brd rd_nr=1 rd_size=8388608 max_part=1 gives one 8 GiB /dev/ram0. rd_size is in kbytes, and all three parameters are mode 0444, so they are set at modprobe time and not afterwards. Pages are allocated lazily on first write, one at a time, into an xarray keyed by sector, so rd_size is a ceiling rather than a commitment. Discard genuinely frees memory (discard_granularity is PAGE_SIZE, and brd_do_discard() erases the entry and puts the page), so mount with -o discard or run fstrim, or the device only ever grows. brd pages cannot be swapped and the device cannot be resized after creation.

brd beats tmpfs in exactly one circumstance: you need a block device. O_DIRECT, a device-mapper or md target stacked on top, mkfs and filesystem work, block-layer benchmarking without storage noise. For a fast scratch directory tmpfs wins on every axis, and the kernel says so: “Contrary to brd ramdisks, tmpfs has its own filesystem, it does not rely on the block layer at all.” Ignore ramdisk.rst where it describes brd using the buffer cache; current brd.c owns its pages outright.

zram: compression behind a block device

zram is /dev/zramN backed by compressed pages in zsmalloc, and its disksize is virtual, an advertised capacity rather than a cost. Because the device sets BLK_FEAT_SYNCHRONOUS, mm/swapfile.c marks it SWP_SYNCHRONOUS_IO. Until Linux 7.0 that let the swap-in path skip the swap cache for a page mapped only once; phase II of the swap table work removed the bypass, routed every swap-in through the cache, and turned readahead off for SWP_SYNCHRONOUS_IO devices instead. On 7.0 the kernel therefore reaches the vm.page-cluster=0 conclusion for you, and on an older kernel you set it yourself: readahead of 2^page-cluster pages amortises a seek, and there is no seek to amortise, only wasted decompression.

The default algorithm upstream and on Ubuntu is lzo-rle (CONFIG_ZRAM_DEF_COMP), the zram default since Linux 5.1: LZO plus run-length encoding for zero runs, on the observation that page data is full of zeros. The patch series measured 4178 4 KiB pages captured from a swapping Chromebook and found the set bimodal, “44% of pages in this dataset contain 5% or fewer zeros, and 44% contain over 90% zeros”, with the run-length step reported to add 50 to 300 MB/s on the zero-heavy half of that distribution. zstd buys ratio for CPU; lz4’s level is an acceleration level, where higher means a worse ratio. Select the algorithm before writing algorithm_params, set every parameter in one write because changing the algorithm resets that priority’s parameters, and expect invalid parameters to be reported by the later disksize write rather than by the write that caused them.

zram fixed cost: 16 bytes of table per 4 KiB of disksize (64-bit)
  16 / 4096 = 0.39 % of disksize, allocated the moment you set it
  8 GiB disksize -> 32 MiB       32 GiB disksize -> 128 MiB
(the doc's "about 0.1% of the size of the disk" is stale by roughly 4x)

r = orig_data_size / mem_used_total       (mm_stat columns 1 and 3)
net memory gained by holding S bytes compressed = S x (1 - 1/r)
  r=2 -> 50 %    r=3 -> 66.7 %    r=4 -> 75 %
worst case, device full = D/r + 0.0039 x D
to bound that at fraction f of RAM:  D <= f x M / (1/r + 0.0039) ~ f x M x r
32 GiB box, D = 32 GiB, r = 3: 10.7 GiB of pages + 128 MiB of table,
                               holding 32 GiB of anon pages
CPU:  cores_busy = (pages_per_second x 4096) / bytes_per_second_per_core

Two readings fall out. Returns diminish steeply, since r = 3 to r = 4 buys eight further percentage points of S, which is the argument against an expensive primary compressor. And decompression rate matters more than compression rate for interactive feel, because compression happens on the reclaim path while decompression happens on the synchronous page-fault path. The worst case is memory you must physically own, and a machine sized to hold both a working set and a compressed pool runs out of DIMM slots long before it runs out of address space. Compression buys a multiple of the memory you already own; it does not buy a slot.

Use mem_used_total, not compr_data_size, for r: the first includes allocator fragmentation and metadata and is what the device actually costs. Watch huge_pages, the count of pages that would not compress below zsmalloc’s largest size class and got a whole page anyway. A high count means the data is already compressed and the CPU is being spent for nothing. The ratio is a property of your data, not of the algorithm - the same build returns a high ratio on repetitive text and close to 1.00 on a file that has already been compressed once.

zram as swap is safer than zram carrying a filesystem, because swap pages are reclaimable by design while filesystem pages pin memory. mem_limit is a backstop, not a tuning knob: once hit, writes fail, which on a swap device means a failed swap-out, reclaim stall and eventually OOM. None of this moves the bound above, either: compression multiplies the memory you already own, so a working set that was never going to fit at any DIMM count is a storage question rather than a larger pool.

Writeback, and where compression stops paying

zram’s writeback sends chosen pages to a real block device, set through backing_dev before disksize and required to be a block device rather than a file. Targets are selected by class: huge, idle, huge_idle, incompressible, or explicit page indexes; Linux 6.16 reworked the interface to uniform key=value and added page_indexes=1-100 ranges. Because that device is usually flash, budget the wear:

MB_SHIFT=20; 4K_SHIFT=12
echo $((400<<MB_SHIFT>>4K_SHIFT)) > /sys/block/zram0/writeback_limit  # 400 MiB
echo 1 > /sys/block/zram0/writeback_limit_enable

Reading writeback_limit returns the remaining budget, in 4 KiB units, and the counter resets on device reset or reboot, so carrying it across reboots is userspace’s job. This is the seam where the argument leaves memory entirely: the pages being written out are the ones compression could not pay for, and what sits underneath decides what they cost to fetch back, which is a question about the device rather than about the DIMMs.

zswap, and how it differs

zswap is not a device. It is a compressed cache in front of a real swap device: CONFIG_ZSWAP depends on SWAP, and when the pool hits its limit zswap evicts on an LRU basis to the backing device. That single structural fact settles most of the choice. zram swap has no overflow path unless you configure one; zswap always has, so when compression stops paying the system degrades smoothly rather than failing. zswap sees only anonymous pages headed for swap, and it is memcg-aware (memory.zswap.max, memory.zswap.writeback); zram is a block device that can also carry a filesystem, and is not.

Current parameters live in /sys/module/zswap/parameters/: enabled, compressor, max_pool_percent (default 20), accept_threshold_percent (default 90, a hysteresis band expressed as a percentage of the max pool size, so the default pool fills at 20 per cent of RAM and reopens at 18), and shrinker_enabled, which is off unless CONFIG_ZSWAP_SHRINKER_DEFAULT_ON is selected. Two removals matter, because stale advice is everywhere: zbud and z3fold went in Linux 6.15, leaving zsmalloc as the only allocator, and the zpool parameter that selected them has gone with them, so zswap.zpool=z3fold on a current kernel does nothing at all.

Choose zswap where a swap device already exists and the working set is unpredictable, since it needs no capacity decision and its failure mode is “slower” rather than OOM. Choose zram where there is no swap device, where swap is undesirable, or where you need a block device in compressed RAM. Running both means compressing a page and then handing it to be compressed again; the kernel docs do not forbid it, but the design leaves no reason to.

What actually ships, and when

Ubuntu 26.04.1 LTS, kernel 7.0.0-34-generic, as it shipped in September 2026: CONFIG_ZSWAP=y with CONFIG_ZSWAP_DEFAULT_ON unset, so zswap is compiled in and off, with compressor=lzo, max_pool_percent=20 and shrinker_enabled=Y, that last differing from upstream’s Kconfig default of n. zram is a module with lzo-rle and writeback compiled in, CONFIG_ZRAM_MULTI_COMP unset (so recompress does not exist), and nothing configures a zram device at all: swap is a plain file.

Fedora took the other road, and dated: the F33 change page (2020-10-13) enabled swap-on-zram by default through zram-generator-defaults, at 50 per cent of RAM capped at 4096 MiB, with no swap partition created; the F34 change page (last updated 2021-01-27) raised that to the full memory size capped at 8192 MiB. The upstream zram-generator default, where a distribution ships the tool without its own configuration, is zram-size = min(ram / 2, 4096) MiB with swap-priority 100 and options=discard. Android’s mmd defaults mmd.zram.size to 50 per cent of device RAM with a fast primary algorithm and zstd for recompression, driving idle, writeback and recompress as a three-tier hierarchy. Arch packages zram-generator and enables nothing. For Debian, openSUSE and ChromeOS defaults this guide will not claim one.

Before writing any tuning command, read what your kernel shipped: ls /sys/module/zswap/parameters/, ls /sys/block/zram0/, and grep -E 'ZSWAP|ZRAM' /boot/config-$(uname -r).

The ZFS ARC

More than you would ever have given it, and it took that much by default. On OpenZFS 2.3 and later the default ARC maximum on Linux is the larger of installed memory minus 1 GiB and five eighths of installed memory, which on any machine with more than about 2.7 GiB of RAM means memory minus 1 GiB. On a 64 GiB server the ARC will grow to 63 GiB unless told otherwise.

That is the answer, and “how do I stop it” is the wrong question. The ARC gives memory back under pressure, it is accounted separately from the page cache, and it is the only cache in this article whose replacement policy is worth the memory it holds. The question is what it should hold, not how little.

The formula is in module/os/linux/zfs/arc_os.c and arc_default_max(), and it is identical to the FreeBSD one since commit 6a629f3 (PR openzfs/zfs#15437, October 2023, “arc_default_max on Linux should match FreeBSD”). Before that, Linux used half of memory, which is why so much published advice is wrong:

arc_default_max(allmem) = MAX(allmem * 5 / 8, allmem - 1 GiB)

allmem - 1 GiB  >  allmem * 5/8   when   allmem * 3/8 > 1 GiB
                                  i.e.   allmem > 8/3 GiB = 2.67 GiB

64 GiB installed:  MAX(40 GiB, 63 GiB) = 63 GiB

Ghost lists, and why this is not an LRU

ARC is the Adaptive Replacement Cache of Megiddo and Modha (FAST ’03, and the readable version in IEEE Computer, April 2004). It keeps two LRU lists: one of pages seen once recently, one of pages seen at least twice recently. ZFS calls them MRU and MFU. Each has a ghost list behind it, holding the metadata of evicted pages and none of their data, so the cache directory covers twice the pages the cache holds.

The adaptation is one number. A hit in the MRU ghost list says recency is being under-served and grows the MRU target; a hit in the MFU ghost list shrinks it. Nothing else tunes it. The paper’s claim for scan resistance is mechanical: a block not in either list enters at the top of MRU and must be requested a second time before it reaches MFU, so “a long sequence of one-time-only reads will pass through L1 without flushing out possibly important pages in L2”. A backup, a find /, a full table scan: the paper’s guarantee is that such a scan flushes only L1, the recency list, so blocks read more than once survive it while blocks read a single time do not.

The cost of that is 0.75 per cent of cache size in the real implementation the paper reports, with per-request time independent of cache size. The benefit, at identical cache size, from Table 4 of the 2004 paper:

Workload LRU hit ratio ARC hit ratio
P3 3.57 % 17.12 %
SPC1-like, 4 GB cache 9.19 % 20.00 %
Merge(S), 4 GB cache 27.62 % 40.44 %

ZFS does not implement that algorithm exactly, and the header comment of module/zfs/arc.c says so, listing three departures: some buffers are pinned by an outstanding reference and cannot be evicted at all; the cache size is a target that moves with memory pressure rather than a constant; and blocks are variable in size, so the ARC budgets in bytes, not pages, and eviction has to weigh size as well as list position.

Setting the bounds, and making them stick

Both zfs_arc_max and zfs_arc_min are in bytes and both default to 0, meaning derived. zfs_arc_min derives to the larger of 32 MiB and memory divided by 32.

cat /sys/module/zfs/parameters/zfs_arc_max          # read
echo 17179869184 > /sys/module/zfs/parameters/zfs_arc_max   # 16 GiB

The new cap applies at once, but man 4 zfs warns that lowering it below the current ARC size will not shrink the ARC without memory pressure to induce shrinking, and that it cannot be set back to 0 while the module is loaded.

Persistently on Linux, in /etc/modprobe.d/zfs.conf:

options zfs zfs_arc_max=17179869184

On root-on-ZFS the module is loaded from the initramfs, so the file has to be rebuilt (update-initramfs -u on Debian and Ubuntu) or the boot-time ARC takes the derived default rather than your value. Defaults move between releases; run man 4 zfs on the machine in front of you rather than trusting any web page, including this one. OpenZFS says exactly that about its own documentation, the man page shipped with your version being the authoritative one, and its Workload Tuning page still states the stale “half of system memory on Linux” figure as proof of the hazard.

Metadata against data

The tunable everyone still recommends no longer exists. zfs_arc_meta_limit, zfs_arc_meta_limit_percent (75 %), zfs_arc_meta_min, zfs_arc_meta_prune and zfs_arc_meta_strategy were all removed in OpenZFS 2.2 and replaced by one knob:

zfs_arc_meta_balance - “Balance between metadata and data on ghost hits. Values above 100 increase metadata caching by proportionally reducing effect of ghost data hits on target data/metadata rate.” Default 500.

The split is now earned through the ghost lists, the same mechanism the paper uses for recency against frequency, applied a second time. The fix for a badly behaved tunable was to delete it. zfs_arc_dnode_limit_percent, default 10 %, survives.

Per dataset, primarycache and secondarycache take all (the default), metadata or none, for the ARC and the L2ARC respectively. They control what is admitted from then on, not what is already resident. Two documented cases for primarycache=metadata: a ZFS swap zvol, where the OpenZFS FAQ says to set it “to avoid keeping swap data in RAM via the ARC”, and InnoDB, where the Workload Tuning page says to set it on the data and log datasets “to prefer InnoDB’s caching”. That is the double-caching problem in one property: on ZFS the page cache for file data is the ARC, so a buffer pool and an ARC hold the same block against the same RAM.

secondarycache is where the argument leaves memory. If the working set will never fit, the tier below the ARC is flash, and l2arc_write_max defaults to 33554432 bytes, 32 MiB per second per device, so an L2ARC warms slowly and is a different purchase from more DIMMs.

Reading the ratio, and which fields are real

arcstat and arc_summary were renamed zarcstat and zarcsummary in OpenZFS 2.4.0 to stop colliding with an unrelated package; both names appear in the wild. They read /proc/spl/kstat/zfs/arcstats.

Modern ARC has three outcomes, not two. From zarcstat’s own arithmetic:

v["read"]  = v["hits"] + v["iohs"] + v["miss"]
v["hit%"]  = 100 * v["hits"] / v["read"]
v["dread"] = v["dhit"] + v["dioh"] + v["dmis"]
v["dh%"]   = 100 * v["dhit"] / v["dread"]

The headline hit% is polluted by prefetch. A pool doing large sequential reads shows a superb overall ratio that says nothing about whether the working set is resident. The fields that answer “is the ARC big enough” are dh% and dm%, the demand hit and miss percentages, and above all mrug and mfug, the ghost list hits per second. A ghost hit is a request that would have hit had the ARC been larger. It is the only directly measured argument for buying memory that this article can offer.

arcstat -f time,read,dread,dh%,dm%,mrug,mfug,arcsz,c 1

Use the interval form. arc_summary prints counters accumulated since boot, which on a long-lived box is archaeology, not measurement.

The 1 GB per TB rule, honestly

It is in no OpenZFS document. The FAQ’s only figure is “8GB+ of memory for the best performance”, with no per-terabyte term at all. The TrueNAS CORE hardware guide, the page usually cited as the source, says 8 GB for up to eight drives and “add 1 GB for each drive added after eight” - per drive, not per terabyte - and its only per-terabyte number is for deduplication. The rule appears to be early FreeNAS guidance, written when the practical problem was 4 GB machines, later generalised by people who were not reading it carefully.

Replace it with the mechanism: ARC demand is set by the working set and the access pattern, not by pool capacity. An 80 TB archive read sequentially is content with 32 GB; an 8 TB pool of random 8 K database reads is not. dm% and mrug settle it, and arithmetic on pool size does not.

Deduplication is the trap

Dedup is the one feature that does turn capacity into a RAM bill, because the deduplication table must be resident or every miss becomes a random read. The OpenZFS figure is “slightly more than 320 bytes of memory” per cached entry, and entries scale with block count, not bytes:

10 TiB unique / 128 KiB recordsize =  83,886,080 blocks -> ~26.8 GB of DDT
10 TiB unique /  16 KiB recordsize = 671,088,640 blocks -> ~215  GB of DDT

The recordsize chosen for an InnoDB dataset multiplies the dedup bill eightfold on the same data. TrueNAS quotes 1 GB to 3 GB per TB for pools with a ratio of 3x or better, and 5 GB per TB as a planning figure. OpenZFS 2.3’s fast dedup adds a flat entry format and a dedup quota that stops deduplicating new blocks rather than letting the table consume RAM; check with zpool status -D and estimate beforehand with zdb -S.

An ARC of this size is a platform decision before it is a software one. Memory minus 1 GiB is only a large number on a socket with enough channels and enough slots, and enough slots means modules that buffer the command and address lines rather than loading the controller once per rank, which is what registered DDR4 buys, and, where even that runs out of capacity, load-reduced modules, which buffer the data lines as well. It should also be ECC. Every block ZFS writes is checksummed from whatever is in the ARC at that moment, so a bit flipped before the checksum is computed produces a corrupt block with a perfectly valid checksum, and scrub will defend it forever. That is the ECC argument in its sharpest form.

Database buffer pools

As large as practical. MySQL’s manual says “up to 80% of physical memory is often assigned to the buffer pool” on a dedicated server; PostgreSQL’s says 25 % of RAM and that more than 40 % is unlikely to work better than a smaller amount. Same machine, same question, a factor of three between the answers.

That gap is not a disagreement about how much memory a database deserves. It is a disagreement about who owns the copy. InnoDB keeps the only cached copy itself and, since 8.4, opens its data files with O_DIRECT by default, so there is little left in the kernel to leave room for. PostgreSQL reads and writes buffered on purpose and expects the operating system’s page cache to hold most of the working set. Both numbers are correct, and applying either one to the other engine wastes memory.

InnoDB: one cache, sized to the machine

innodb_buffer_pool_size is a global, dynamically resizable integer whose default is 134217728 bytes (128 MB) with a minimum of 5242880. That default is a shipping compromise and is wrong on any server doing real work. It caches table and index pages, and, because MySQL 8.4 turned innodb_adaptive_hash_index from ON to OFF and innodb_change_buffering from all to none, a modern pool holds rather less besides those pages than an 8.0 pool did.

Sizing has an arithmetic constraint that usually surfaces only after the server has quietly changed your number:

innodb_buffer_pool_size must equal, or be a multiple of,
    innodb_buffer_pool_chunk_size * innodb_buffer_pool_instances

innodb_buffer_pool_chunk_size default = 134217728 B (128 MB), min 1048576
chunks = innodb_buffer_pool_size / innodb_buffer_pool_chunk_size
manual's ceiling: chunks should not exceed 1000

Set something that does not conform and “buffer pool size is automatically adjusted” upward, which is why SHOW VARIABLES sometimes disagrees with the config file. The 8.4 system-variable table gives the default for innodb_buffer_pool_instances as “see description”; this guide does not print a number the table does not give.

The cleanest answer is to stop doing the arithmetic. --innodb-dedicated-server derives the pool from detected memory: under 1 GB it keeps the 128 MB default, 1 GB to 4 GB gives detected memory times 0.5, and above 4 GB detected memory times 0.75.

InnoDB is not a plain LRU, and the difference is the same scan-resistance problem every cache has. The list is split, with 3/8 of the pool in the old sublist; new pages enter at that midpoint and only reach the head on a second access. innodb_old_blocks_pct defaults to 37 (range 5 to 95) and innodb_old_blocks_time to 1000 milliseconds. The manual’s own advice for a table scan that cannot fit is to drop innodb_old_blocks_pct to 5, confining read-once data to 5 % of the pool.

Restart behaviour is worth configuring rather than rediscovering: innodb_buffer_pool_dump_at_shutdown and innodb_buffer_pool_load_at_startup are both enabled by default, innodb_buffer_pool_dump_pct defaults to 25, and what lands in ib_buffer_pool is only tablespace and page IDs, not data. A cold 512 GB pool that has to re-read its working set from storage is usually the longest stall a restart produces.

Eighty per cent of physical memory is a buying instruction as much as a configuration one, and the platform decides it before MySQL does. A pool worth the name needs modules that a desktop board cannot electrically carry, which means registered or load-reduced DIMMs on a socket that accepts them; registered vs unbuffered explains why the wall is loading rather than address space, and the server compatibility pages list what a given machine takes.

PostgreSQL: two caches, one of them the kernel’s

shared_buffers “sets the amount of memory the database server uses for shared memory buffers”, defaults to 128 MB unless the kernel will not support that, has a floor of 128 kB, and can only be set at server start. A unit trap: “if this value is specified without units, it is taken as blocks, that is BLCKSZ bytes, typically 8kB”, so shared_buffers = 4096 is 32 MB and not 4 GB.

The 25 % figure is architecture, not modesty. The manual states the reason in the same paragraph: “because PostgreSQL also relies on the operating system cache”. Its replacement policy is a clock sweep over per-buffer usage counters, and its scan resistance is a ring buffer rather than a midpoint: the buffer manager’s own notes give 256 kB for sequential scans, 16 MB but not more than 1/8th of shared_buffers for bulk writes, and vacuum_buffer_usage_limit for VACUUM. A SELECT over a huge table therefore evicts a ring of that order rather than your working set, which is a large part of why a small shared_buffers survives contact with reporting queries.

effective_cache_size allocates nothing

This one is misread constantly, and the manual disposes of it in one sentence: “This parameter has no effect on the size of shared memory allocated by PostgreSQL, nor does it reserve kernel disk cache; it is used only for estimation purposes.” It is a number the planner reads when costing an index scan, defaulting to 4 GB, and taken as 8 kB blocks if written without a unit.

Two failures follow. Treating it as an allocation produces “I set effective_cache_size to 48 GB and PostgreSQL still only uses 128 MB” - nothing was allocated, and nothing was going to be. Leaving it at 4 GB on a 128 GB machine is the expensive one: the planner believes the cache is tiny, systematically over-costs index scans, and drifts to sequential scans and merge joins on queries that should be index nested loops. That costs real query time with no memory misconfigured at all. The manual’s sizing instruction is to consider shared buffers plus the portion of the kernel’s disk cache that will be used for PostgreSQL data files, and to allow for the expected number of concurrent queries on different tables, since they have to share that space.

The same page, twice

A buffered read leaves one copy in the kernel’s page cache and one in the buffer pool, both charged to the same RAM. O_DIRECT is the answer where it exists, and as of MySQL 8.4 innodb_flush_method defaults to O_DIRECT if supported, otherwise fsync on Linux, where 8.0 defaulted to fsync. PostgreSQL ships the opposite verdict: debug_io_direct exists, takes data, wal and wal_init, and carries the line “Currently this feature reduces performance, and is intended for developer testing only” - because direct I/O leaves the database to do its own readahead, and that machinery has only recently arrived: PostgreSQL 18 added asynchronous I/O, an io_method that can dispatch reads to worker processes or io_uring, and read streams for sequential scans, bitmap scans and VACUUM.

Two cautions. O_DIRECT is not durability: the man page notes it does not provide the guarantees of O_SYNC, and it bypasses the page cache, not the drive’s volatile write cache. And on ZFS the “OS cache” is the ARC, so OpenZFS’s workload guidance for InnoDB is primarycache=metadata on the data and log datasets with recordsize=16K for data - telling the filesystem to keep only metadata rather than bypassing it. ZFS gained direct I/O only in OpenZFS 2.3.0, through the per-dataset direct property; from ZFS on Linux 0.8.0 onward the flag was accepted but the data was cached anyway, so the ARC held a second copy unless primarycache said otherwise, and older versions rejected it outright.

The measurement that settles the argument

SHOW ENGINE INNODB STATUS prints Buffer pool hit rate 1000 / 1000, per 1000 accesses over the window since the last monitor output. It is a windowed figure: 1000 / 1000 means no misses in this window, not a perfect cache. PostgreSQL’s pg_stat_database.blks_hit is narrower still and the docs say so - it counts “only hits in the PostgreSQL buffer cache, not the operating system’s file system cache”, so a block served from the page cache in microseconds counts as a read exactly like one fetched from a disk. Chasing that ratio to 99 % by inflating shared_buffers can easily make things worse; pg_buffercache shows what is actually resident.

What a good ratio looks like depends entirely on what a miss costs:

round figures, chosen to show the crossover, not vendor specifications

T_eff = h*T_hit + (1-h)*T_miss        T_hit = 100 ns = 0.0001 ms

miss to datacentre NVMe, T_miss = 0.1 ms (100 us)
  h = 0.99    -> 0.000099 + 0.001000 = 0.001099 ms = 1.10 us
  h = 0.999   -> 0.000100 + 0.000100 = 0.000200 ms = 0.200 us
  h = 0.9999  -> 0.000100 + 0.000010 = 0.000110 ms = 0.110 us

The miss term stops dominating only when (1 - h) falls below T_hit / T_miss: a hit rate of about 99.9 % against NVMe, and about 99.998 % against a 7200 rpm disk whose average rotational latency alone is 4.17 ms. So a 94 % buffer pool hit rate is a catastrophe on spinning media and merely mediocre on flash. Where the working set will never fit in RAM at any sane price, the lever is the other term: cutting T_miss by a factor of a hundred multiplies the miss rate you can afford by the same hundred, which is the companion question and belongs to the NVMe drives rather than to the DIMMs.

Caching in the application rather than the kernel

Because the kernel cannot see what your bytes mean. The page cache stores file offsets and can only ever save you a read; an application cache stores the result of work, and saves you the work.

That is the whole difference, and speed is the wrong reason to reach for Redis. The page cache already holds the bytes your web root and many of your databases read, at no cost, with no configuration, and with no failure mode worse than a miss. The exception is a database opening its files with direct I/O: InnoDB’s innodb_flush_method defaults to O_DIRECT from MySQL 8.4 onward, and those reads never touch the page cache at all. An application cache is a second copy taken out of the same DIMMs, which you then have to size, bound, evict, monitor and keep alive across restarts. The question is not which cache is faster. It is which copies you can afford to lose.

Redis is not a cache until you configure it as one

maxmemory defaults to 0 on 64-bit builds, meaning no limit (32-bit builds carry an implicit 3 GB ceiling), and maxmemory-policy defaults to noeviction. Out of the box Redis is a data-structure server that keeps allocating as it sees fit, which its own documentation warns can gradually eat up all your free memory. The out-of-memory error on writes only arrives once you have set a limit for noeviction to enforce. The policies are allkeys-lru, allkeys-lfu, allkeys-random, volatile-lru, volatile-lfu, volatile-random, volatile-ttl and noeviction, joined since Redis 8.6 by allkeys-lrm and volatile-lrm, which key on least recently modified rather than least recently used, and the volatile-* family behaves exactly like noeviction when no key carries a TTL - a production trap that presents as a cache which has simply stopped accepting writes.

Both LRU and LFU are approximations. maxmemory-samples defaults to 5: Redis samples that many keys and evicts the best candidate among them, and raising it to 10 approaches true LRU at some CPU cost. LFU counts with an 8-bit Morris counter, range 0 to 255, tuned by lfu-log-factor (default 10) and lfu-decay-time (default 1 minute). Exact LRU over tens of millions of keys would cost a pointer pair per key, and cost in memory is the only reason Redis gives for not implementing it; the approximation is the price of the cache being cheap.

Persistence is where an application cache stops being free. RDB snapshots (save 60 1000) and AOF rewrites both work by fork() and copy-on-write, so the host must supply whatever the parent dirties while the child runs. Redis documents that a write-heavy instance can use up to twice its normal memory during a save, requires vm.overcommit_memory = 1, and asks for transparent huge pages to be turned off:

Copy-on-write cost of one byte written to a shared page after fork:
  4 KiB base page  ->    4 KiB copied
  2 MiB huge page  -> 2048 KiB copied        (512x amplification)

Page tables for a 24 GiB instance (Redis's own arithmetic):
  24 GiB / 4 KiB x 8 B per PTE = 48 MiB of leaf page tables

appendfsync defaults to everysec, bounding loss at roughly one second; always is durable and slow; no leaves it to the kernel, which Redis describes as flushing about every 30 seconds - the same number as vm.dirty_expire_centisecs, because it is the same mechanism. For a pure cache, turn persistence off. Redis lists “no persistence” as a first-class configuration, and a cache that needs a durable write path is not a cache, it is a database you have not budgeted for.

Sizing follows from the dataset, not from the traffic: an instance holding a few hundred gigabytes of objects is a capacity purchase, and capacity on a single socket means registered or load-reduced modules on a platform that accepts them, for the electrical reasons set out in registered vs unbuffered. Leave headroom beyond maxmemory: replication and AOF buffers are not counted against it, and INFO memory exposes the excluded total as mem_not_counted_for_evict.

Memcached: the cache that admits it is one

Defaults are -m 64 (megabytes), -I 1m maximum item size, -c 1024 connections, -t 4 threads, -f 1.25 slab growth factor, -n 48 minimum chunk, port 11211, UDP off. Memory is cut into 1 MB pages, each assigned to a slab class and chunked; the documentation is blunt that once a page is assigned to a class it is never moved, and each class runs its own LRU. That produces slab calcification: a workload whose item sizes shift can evict from one class while another sits half empty, which the page cache does not do to you, because it has no per-size classes to strand memory in. Slab reassignment and the automatic page mover exist to fight it, and have been on by default since 1.5.0 rather than opt-in. Overhead beyond -m, for the hash table and connection buffers, is documented as a few per cent.

Since 1.5.18 memcached can survive a restart with -e /path/memory_file, where the path must be a RAM disk or a DAX mount; you stop it with SIGUSR1, and it refuses to reuse the file if -m, the maximum item size, the chunk sizes or the CAS setting changed, or if the clock jumped while it was down. That is a RAM disk used not for speed but as a place to leave memory behind. When the working set will never fit in RAM at any sane module price, memcached’s answer is extstore (-o ext_path=...), which tiers cold values onto flash and keeps only keys and metadata resident.

Varnish stores the answer, not the file

Note the naming first: the open-source project renamed itself Vinyl Cache in March 2026, while Varnish Software continues to ship a product called Varnish Cache as a downstream distribution under its own governance, so both names are current and they are not the same tree.

Varnish keeps assembled HTTP responses, keyed by request, with TTL, grace and VCL-driven invalidation. A page cache cannot do any of that: it has no concept of a Vary header, cannot purge by URL, and caches the template rather than the rendered page. Storage is -s default[,size] (an alias for umem where it exists and malloc otherwise), -s malloc[,size], -s file,path[,size...], or the persistent store the users guide now spells -s deprecated_persistent,path,size. The file backend is not a persistent cache - it is memory backed by an unlinked file, deliberately handed to the kernel to page out. Documentation disagrees about the default malloc size (the manual page says that with no -s option the daemon applies -s default,100m, while the users guide calls the default size unlimited, and an upstream issue has been filed about exactly this), so this guide will not claim one: size it explicitly. Size Transient too, since its default is an unbounded malloc store and therefore a real path to OOM.

The architectural argument is Poul-Henning Kamp’s, and it is the best statement of the whole problem: Squid wrote objects to “RAM”, the kernel paged them to swap without Squid knowing, and Squid then decided they were cold and wrote them to disk itself, so one object existed three times. “So what happens with squids elaborate memory management is that it gets into fights with the kernels elaborate memory management, and like any civil war, that never gets anything done.”

nginx proxy caching is on disk, and that is the right answer

proxy_cache_path takes a keys_zone=name:size, and that shared memory holds keys and cache metadata only, at roughly 8,000 keys per megabyte in the open-source build; a commercial subscription stores more per key and needs a larger zone. Response bodies are files on disk, named from the MD5 of the cache key. inactive defaults to 10 minutes; proxy_cache_valid has no default, so nothing is cached until you say so or the upstream says so in an X-Accel-Expires, Expires or Cache-Control header, which take priority over the directive; and proxy_cache_key is documented as $scheme$proxy_host$request_uri, which the documentation itself describes as close to, but not the same as, $scheme$proxy_host$uri$is_args$args.

10,000,000 objects / 8,000 keys per MB    = 1,250 MB of keys_zone in RAM
10,000,000 bodies x 40 KB assumed mean    =   400 GB of files on disk
resident hot subset                       = whatever the page cache holds

The 40 KB is an assumption, not a measurement: substitute your own mean body size and the disk figure moves with it, while the keys_zone figure does not. nginx’s own guidance is against putting that directory on tmpfs: a tmpfs cache is capped by available RAM, the same memory would serve a larger on-disk cache better through the page cache, and tmpfs is swapped out when memory is constrained anyway. An nginx cache on disk is a RAM cache for the hot set, for free, with no size cliff.

The judgement

A kernel cache is free and knows nothing about your data. It sizes itself, shares between processes, survives every process restart, and is reclaimed instantly when something needs the frames. An application cache costs configuration, an eviction policy, monitoring and headroom, and in exchange it knows everything: that these four rows are one rendered page, that this object graph took 400 ms of joins, that this response may be served to anyone without a cookie.

The decision is not about speed, at least not against a disk. Both live in the same DIMMs, but only the page cache hit is a memory access: Redis documents a round trip of about 30 us over a Unix domain socket and about 200 us over a 1 Gbit/s network. The decision is how much of your working set is reproducible. Anything the kernel caches for reading is reproducible by definition - the authoritative copy is on the device, and a lost clean page costs one read. Anything an application cache holds is reproducible only if you can recompute it, and the moment it is not, you have bought a durability problem: fork storms, appendfsync, warm-restart files, a second copy of your data to operate. Cache derived work in the application, cache bytes in the kernel, and be suspicious of any cache you cannot throw away.

Windows: the standby list, SysMain, and reading memory correctly

Nothing. Windows has been caching your storage in RAM since boot, the cache has no size setting, and there is no Windows equivalent of a vm.* sysctl to tune it. The Cache Manager fills the system cache by mapping views of files, and those cached pages are ordinary physical pages on the same lists as process memory.

That answers the question, and it is the wrong question, because the thing that actually goes wrong on Windows is not the cache. It is the reader. Task Manager uses words that mean something different from the words free(1) uses, and a person who reads one the way they read the other will conclude that Windows is either wasting memory or has run out of it, and will then go looking for something to switch off.

Cached, Available, Free: three numbers that overlap

Microsoft’s own definitions, from the Task Manager performance-tab write-up: Cached is the sum of the system working set, the standby list and the modified page list. Available is the sum of the standby pages, the free pages and the zero page list. Free is the free and zero lists alone.

Read those twice. Standby pages are counted in both Cached and Available. A machine reporting 12 GB cached and 14 GB available has not spent 12 GB; it has a warm file cache that the memory manager will repurpose without a stall the moment an allocation needs it. The Linux reader’s habit transfers cleanly if you map the terms rather than the layout: MemAvailable is \Memory\Available MBytes, and in both systems a small free figure is health, not scarcity. The pages that genuinely cost something to reclaim are the dirty ones, counted by \Memory\Modified Page List Bytes, which is the nearest counterpart to Dirty in /proc/meminfo, counting the dirty pages that have already left a working set.

Windows page states, per Windows Internals: Active (valid), Transition, Standby, Modified, Modified no-write (NTFS uses it for filesystem metadata so transaction log entries can be flushed first), Free, Zeroed, ROM and Bad. The standby list is segmented by priority 0 to 7, and the memory manager repurposes the low priorities first.

Two limits worth knowing before anyone tries to cap the cache. In Microsoft’s memory-limits table the 64-bit system cache virtual address space is “always 1 TB regardless of physical RAM”, 16 TB for Windows 8.1 and Server 2012 R2, and its physical size is limited only by physical memory rather than by address space. And the only supported way to bound it is the pair GetSystemFileCacheSize and SetSystemFileCacheSize in memoryapi.h, which take minimum and maximum working-set sizes in bytes and require SE_INCREASE_QUOTA_NAME. There is no registry value and no GUI for this.

Microsoft does publish one threshold, and it is the closest thing to a cache-sizing signal the platform offers:

Get-Counter '\Memory\Long-Term Average Standby Cache Lifetime (s)',
            '\Memory\Available MBytes',
            '\Memory\System Cache Resident Bytes' -Continuous

Under 1800 seconds of average standby lifetime is the first counter Microsoft lists for a system cache that has grown until memory is nearly gone: pages are arriving and being evicted faster than they are being reused, which is the measured form of “this machine wants more RAM”. Read it with the other two, because Microsoft lists all three counters together - short lifetime, low available memory, and a system cache holding a significant part of physical memory - and the prescribed next step is RAMMap. Hours of lifetime with flat available memory means more RAM buys nothing.

SysMain, and the advice to disable it

SysMain, formerly SuperFetch, runs in svchost.exe from sysmain.dll. It profiles page usage over time and asks the memory manager to preload file pages into the standby list at appropriate memory priorities. It does not reserve memory: everything it loads is reclaimable immediately, which is why “SysMain is eating my RAM” is a misreading of the Cached figure described above.

The documented control surface is the MMAgent PowerShell module, not the registry:

Get-MMAgent
Disable-MMAgent -ApplicationLaunchPrefetching   # alias -alp
Enable-MMAgent  -MemoryCompression              # alias -mc

with -ApplicationPreLaunch (apl), -OperationAPI (oa) and -PageCombining (pc) alongside. Memory compression is on by default on Windows 10 1511 and later and shows as “In use (Compressed)”; page combining periodically merges physical pages with identical content.

Microsoft publishes no general recommendation to disable SysMain. The one Microsoft support article on SysMain CPU use is narrow: it applies to Windows 7 SP1, describes a one to two minute CPU spike when a 64-bit application is built /LARGEADDRESSAWARE:NO, and the prescribed fix is to rebuild the application. The widespread advice to disable SysMain on an SSD has no Microsoft document behind it, and this guide will not supply one. Leave it on. Disable it as a diagnostic when the SysMain-hosted svchost is demonstrably and persistently burning CPU or disk, and expect slower cold application launches in exchange. Server SKUs are often described as shipping with prefetching not fully enabled; that could not be confirmed from a current Microsoft page, so check Get-MMAgent on the box rather than trusting a blog.

RAMMap, because no counter will tell you which file

RAMMap (Sysinternals, Vista and Server 2008 upward) is the tool Microsoft’s cache troubleshooting guidance points to for “what is in the cache”. Its tabs are Use Counts (usage by type against paging list), Processes, Priority Summary (the prioritised standby sizes), Physical Pages, Physical Ranges, File Summary and File Details. Columns are the paging lists: Active, Standby, Modified, Modified No Write, Transition, Zeroed, Free, Bad. File Summary is the one that earns the download, because it names the files resident in RAM.

Microsoft diagnoses exactly two cache-growth pathologies with it. High active Metafile pages on a server touching millions of files, formerly mitigated by the DynCache tool, “should no longer exist” after the Server 2012 redesign. And high active Mapped File pages caused by applications opening files with FILE_FLAG_RANDOM_ACCESS, which hints the Cache Manager to keep mapped views resident; since Server 2016 the flag is ignored for trimming but still disables prefetching, and Microsoft’s wording is that it is “highly recommended” the flag “not be used by applications”.

RAMMap’s Empty menu (working sets, system working set, modified page list, standby list, priority 0 standby) is a diagnostic gesture, not a tuning step. Emptying the standby list frees nothing that was in use; it discards a warm cache that Windows must now re-read from disk. Scheduling EmptyStandbyList.exe every few minutes, as several forum recipes suggest, is a cache-destruction cron job.

Pressure, working sets and the dirty window

Windows accounts allocations as commit charge against a commit limit of roughly RAM plus pagefile. Past the limit, allocations fail and applications report out-of-memory; there is no OOM killer picking a victim. That is the failure mode difference worth carrying: Linux kills a process, Windows refuses an allocation. A working set is the pageable physical memory currently resident for a process, trimmed by the balance set manager under pressure. Note the definitional trap for large pages: Microsoft states they are part of process private bytes but not part of the working set, “because the working set by definition contains only pageable memory”.

Applications can cooperate through CreateMemoryResourceNotification, and Microsoft warns there is a band where neither the low nor the high event is signalled, in which “applications should attempt to keep the memory use constant”.

One tunable that does exist, for file servers: the remote-file dirty page threshold, default 5 GB per file since Server 2016.

HKLM\SYSTEM\CurrentControlSet\Control\Session Manager\Memory Management
  RemoteFileDirtyPageThreshold  REG_DWORD, units are 4096-byte PAGES

10 GiB target:  10,737,418,240 / 4096 = 2,621,440 (decimal)

Microsoft’s guidance: raise it in 256 MB increments until performance is satisfactory, keep the value between 128 MB and 50 per cent of available memory, -1 disables it and is not recommended, reboot required.

Explicit RAM caches: all of them are third party

Windows ships no supported general-purpose RAM disk. The inbox ramdisk.sys driver is there to boot Windows PE from a WIM or SDI image, not to give you a formattable volume, and the general-purpose RAMDisk driver in the WDK is sample code. Where Linux hands you tmpfs in the kernel, Windows makes you install somebody’s kernel driver, and that asymmetry should inform how much you trust the result.

ImDisk is the usual free answer, and its own upstream now describes the design as legacy and not recommended for recent versions of Windows, with maintenance fixes but no new features; AWEAlloc and DevIoDrv have moved out of the project and feature work is focused on Arsenal Image Mounter. SoftPerfect RAM Disk is commercial and can save contents to an image file on shutdown, which means a clean shutdown; a crash still loses everything.

PrimoCache is the one with a real write cache, and therefore the one that needs stating precisely. Its Level-1 cache is physical memory, optionally including “OS Invisible Memory” the running Windows cannot address; Level-2 is flash and is persistent across restarts. Cache block size runs 4 KB to 512 KB, and the vendor advises a value at or below the filesystem cluster size. Defer-Write acknowledges the write to Windows before the disk has it, and the window is up to the Latency value you type, the interval in seconds at which deferred data is flushed, with 0 disabling deferral. The vendor’s own words: on “a sudden power failure, system crash or freeze, the data will have no chance to be written to the disk, resulting in loss”, the deferred data “may include file system metadata”, and in extreme cases the filesystem “may even be damaged”. The risk is identical for L1 and L2, because the cache index lives in memory either way, and an SSD tier does not make a deferred write crash-safe. A ten-second Latency is a ten-second promise you are making to yourself.

That L2 tier is also the honest exit from this section. When the working set will not fit in RAM at any price the platform supports, the next tier is flash rather than more DIMMs, and the sizing question moves to the drive. Before it does, check the two Windows ceilings people hit first: client editions cap at 128 GB on Windows 11 Home and 2 TB on Pro, and the standby list can only be as large as the board’s DIMMs allow. Past roughly a quarter of a terabyte the answer is usually a server platform and buffered memory, which is a different purchase from desktop kit: see registered vs unbuffered, the registered and load-reduced listings, and the per-model compatibility pages for the boards that take them.

Huge pages and the TLB

Nothing, if the machine in front of you is a desktop with 32 GB. Huge pages pay in proportion to how much memory one process addresses at random, and the documented failure cases are concentrated in exactly the software most people run. That is the honest answer, and it is the wrong question, because the question underneath it is not “should I enable transparent huge pages” but “how much of my cache can the CPU translate without walking a page table”.

The arithmetic

x86-64 uses 4 KiB base pages and four-level paging. Every virtual address the CPU cannot resolve from the TLB costs a page-table walk of up to four dependent memory references, five under LA57, each of which can itself miss cache. TLB reach is entries multiplied by page size, and the entry count is fixed in silicon: Intel’s Golden Cove is documented with a 2048-entry L2 STLB, AMD’s Zen 4 with a 3072-entry L2 DTLB and Zen 5 with 4096. Those second-level TLBs are shared between page sizes, and the share is not always the whole array: Intel’s optimisation manual states that on Golden Cove 4 KiB pages can use all 2048 entries while 2 MiB and 4 MiB pages can use only 1024 of them.

Golden Cove L2 STLB, 2048 entries (2 MiB pages may use 1024 of them):
  4 KiB pages:  2048 x 4 KiB  =    8 MiB of reach
  2 MiB pages:  1024 x 2 MiB  =    2 GiB of reach       256x

A 64 GiB in-RAM cache, addressed at random:
  64 GiB / 4 KiB              = 16,777,216 translations
  8 MiB / 64 GiB              = 0.012 per cent TLB-covered
  2 GiB / 64 GiB              = 3.125 per cent TLB-covered

Leaf page tables, 8-byte PTE per 4 KiB page:
  64 GiB / 4 KiB x 8 B        = 128 MiB of page tables
  same region at 2 MiB        = 256 KiB                 512x

The page-table figure is the one people quote, and it is the less interesting one. Redis’s own documentation works the same sum: “a large 24 GB Redis instance requires a page table of 24 GB / 4 kB * 8 = 48 MB.” Reach is what matters for a cache, because a cache is by construction accessed all over.

What a miss costs: the walk is cached at several levels, so the common case is cheap. The bound is not. Four dependent references that all miss to DRAM, at the roughly 80 ns a random local DRAM load is measured to cost, come to several hundred nanoseconds spent translating an address whose data then takes 80 ns to fetch. That multiplication is arithmetic rather than a published page-walk latency, but it is the shape of the worst case: the translation can cost more than the access. That is the whole case for huge pages, and it is a case about reach, not about 128 MiB of page tables.

Transparent huge pages

THP promotes anonymous mappings to 2 MiB pages without the application asking. Three settings, read and written at /sys/kernel/mm/transparent_hugepage/enabled:

  • always - promote wherever the kernel can
  • madvise - promote only in regions that called madvise(MADV_HUGEPAGE)
  • never - no automatic promotion

A second file, defrag, governs how hard the kernel works to find a contiguous 2 MiB block, and it is the one that decides whether a fault waits. At always an application requesting THP “will stall on allocation failure and directly reclaim pages and compact memory”; defer wakes kswapd and kcompactd in the background instead, and defer+madvise does the same everywhere except in regions that called madvise(MADV_HUGEPAGE), which still enter direct reclaim and compaction. Per-size controls exist as well: PMD-sized THP defaults to inherit, every other size to never. Distribution defaults differ and have changed between releases, so this guide will not claim one: run cat /sys/kernel/mm/transparent_hugepage/enabled and read the bracketed value on your own box.

Two counters tell you whether the setting is working. grep thp_fault /proc/vmstat gives thp_fault_alloc and thp_fault_fallback. A rising fallback count against a flat alloc count is a fragmented machine paying compaction cost and getting 4 KiB pages anyway.

None of these knobs governs the Linux page cache the way it governs anonymous memory. THP for tmpfs and shmem is a separate control, historically defaulting to never. If your cache is the kernel’s page cache, this section is not your optimisation: the next question is what that cache sits in front of, and replacing a spinning disk moves the number further than any page size will.

Where the vendors say no

The mechanism is usually one of three things: synchronous compaction on the allocation path, khugepaged collapsing sparsely-touched regions and inflating RSS, or copy-on-write at 2 MiB granularity.

Redis documents the third precisely: fork creates two processes sharing huge pages, “a few event loops runs will cause commands to target a few thousand of pages, causing the copy on write of almost the whole process memory”, and “this will result in big latency and big memory usage”. A one-byte write copies 4 KiB normally and 2 MiB under THP - 512 times the write amplification, on the exact code path a background save uses. Redis’s instruction is echo never > /sys/kernel/mm/transparent_hugepage/enabled.

PostgreSQL’s manual says THP “has been known to cause performance degradation with PostgreSQL for some users on some Linux versions, so its use is currently discouraged (unlike explicit use of huge_pages)”. Couchbase: “THP must be disabled in order for Couchbase Server to function correctly on Linux”. Oracle’s public documentation carries a “Disabling Transparent HugePages” page while continuing to recommend standard HugePages.

MongoDB reversed. Through 7.0 it said disable. From 8.0 it says enable, before mongod starts, with enabled=always, defrag=defer+madvise, khugepaged/max_ptes_none=0 and vm.overcommit_memory=1, crediting an upgraded TCMalloc with per-CPU caches. That combination is the modern compromise: promotion that defers reclaim and compaction to the background outside madvised regions, and no collapsing of regions that are mostly untouched.

Explicit hugetlbfs

hugetlbfs pages are reserved at a fixed pool size, pinned, never swapped, and unavailable to the ordinary page allocator. You trade memory the rest of the system cannot borrow back for translation behaviour that cannot degrade. grep Huge /proc/meminfo reports HugePages_Total, _Free, _Rsvd, _Surp and Hugepagesize; vm.nr_hugepages sizes the pool, and the per-node files under /sys/devices/system/node/ place it.

PostgreSQL’s recipe needs no arithmetic from you:

postgres -C shared_memory_size_in_huge_pages -D $PGDATA   # prints a count
sysctl -w vm.nr_hugepages=<that count>                    # then persist it
# postgresql.conf: huge_pages = on   (start fails if they cannot be got)
# after restart:   SHOW huge_pages_status;

MySQL wants large_pages=ON, off by default, with the pool sized as innodb_buffer_pool_size / Hugepagesize; on failure InnoDB falls back silently apart from the log line Warning: Using conventional memory pool. Oracle divides the SGA by the page size and rounds up, and warns that Automatic Memory Management and HugePages are incompatible, because AMM allocates the SGA in /dev/shm.

Windows has no THP at all. Large pages are explicit only: acquire SeLockMemoryPrivilege, call GetLargePageMinimum(), then VirtualAlloc with MEM_LARGE_PAGES. Microsoft’s own caveat is the one to internalise - allocations “may be difficult to obtain after the system has been running for a long time because the physical space for each large page must be contiguous”, so “allocate all large pages one time, at startup”.

Who this is for

A process whose resident working set is measured in tens or hundreds of gigabytes, accessed without locality, where the buffer pool or dataset is one allocation the application owns. A database. A key-value store. Not a file server, whose cache is the kernel’s. Not a workstation, where the whole of RAM sits within a small multiple of a 2 MiB TLB’s reach anyway and the compaction stalls are pure cost.

The machine that reaches that size is a machine with slots: eight or twelve channels populated, often two DIMMs per channel, and modules the memory controller can drive at that electrical load, which means registered or load-reduced parts rather than unbuffered desktop memory. Registered vs unbuffered explains the buffering, and the compatibility pages list what a given server accepts.

NUMA and memory bandwidth

More RAM stops helping the moment the memory controller holding it is not the one your threads are attached to. On a multi-socket box, capacity and locality are separate purchases, and only one of them costs money.

That is the answer, and it is the wrong question for anyone sizing a cache. Placement is a tuning problem with a tuning fix. Underneath it sits a ceiling that no amount of capacity lifts: a socket moves only so many bytes per second, that figure is set almost entirely by how many channels you populated, and a cache that is bandwidth-bound does not get faster when you add a DIMM to a channel that was already full.

What the tools show

numactl --hardware prints, in the man page’s words, an “inventory of available nodes on the system”, and with it a node distances matrix read from the ACPI SLIT. SLIT normalises self-distance to 10 and scales everything against it, so 12 means roughly 1.2 times local and 32 means roughly 3.2 times local. Two-socket EPYC systems split into several NUMA nodes per socket, which means NPS2 or NPS4, are commonly reported as showing 12 between nodes in the same socket and 32 across sockets; under the NPS1 default, with one node per socket, the matrix reads 10 and 32. Treat those as firmware’s declaration rather than a measurement: they are a hint to the scheduler, vendors populate them approximately, and the way to get a real number is mlc --latency_matrix and mlc --bandwidth_matrix from Intel’s Memory Latency Checker.

The policy flags are exact and worth quoting, because three of them fail in different ways. --membind means “Only allocate memory from nodes. Allocation will fail when there is not enough memory available on these nodes.” --preferred will “preferably allocate memory on node, but if memory cannot be allocated there fall back to other nodes.” --interleave allocates “using round robin on nodes”. --cpunodebind pins execution rather than memory, and pinning one without the other is how a process ends up running on node 1 against an allocation on node 0.

numastat tells you whether any of it worked. Its four load-bearing counters, per node: numa_hit, “memory successfully allocated on this node as intended”; numa_miss, “memory allocated on this node despite the process preferring some different node”; numa_foreign, “memory intended for this node, but actually allocated on some different node”; and other_node, “memory allocated on this node while a process was running on some other node”. A rising other_node on a box you believed was partitioned is the whole diagnosis. numastat -p <pid> narrows it to one process, and /proc/<pid>/numa_maps narrows it to one mapping.

The failure this section exists to prevent

The folk version of this warning says a large cache on the wrong node is slower than a small one placed right. Taken literally that is not true, and the arithmetic elsewhere in this guide says why: a remote DRAM hit costs something on the order of 1.5 to 3 times a local one, while an NVMe hit costs roughly a thousand times a local one. A misplaced RAM cache still beats the drive. That changes only when the working set will never fit in RAM at a price you would pay, and at that point the argument leaves memory altogether: the question stops being which node holds the cache and becomes which drive sits underneath it.

Three defensible versions, which are bad enough:

  • A badly placed cache can be slower than a correctly placed cache a third its size, and you paid three times as much RAM to be slower.
  • For bandwidth-bound scans the latency ratio understates the damage. Remote traffic consumes interconnect bandwidth that is a fraction of local memory bandwidth and is shared with every other cross-socket transaction on the machine, so the misplaced cache degrades unrelated work as well.
  • Under --membind, a misplaced allocation does not get slower, it fails. The failure mode is an out-of-memory error on a machine with free RAM.

And one case where the folk version is literally true: vm.zone_reclaim_mode. The kernel documents it as disabled by default, with OR-able bits 1 (reclaim on), 2 (write dirty pages) and 4 (swap). Enabled, it makes the kernel reclaim memory from the local node rather than allocate from a neighbour, which means it throws away local page cache to avoid a cheap remote access. That is a one-line sysctl that makes a cache worse than no cache. Check it before you tune anything else.

Explicit huge-page pools are placed, not merely sized. The per-node file is /sys/devices/system/node/node<N>/hugepages/hugepages-<size>kB/nr_hugepages, and vm.nr_hugepages_mempolicy respects the caller’s NUMA policy where plain vm.nr_hugepages does not. Reserving a large pinned pool that lands entirely on node 0 while the process runs on node 1 is self-inflicted and silent.

Automatic NUMA balancing, and when to trust it

kernel.numa_balancing “enables/disables and configures automatic page fault based NUMA memory balancing. Memory is moved automatically to nodes that access it often.” Values: 0 off, 1 normal, 2 memory-tiering. It works by taking sampling faults and migrating pages towards whoever touches them, so it charges you faults and copies to buy locality back.

Trust it where access is unpartitioned and drifts: mixed general-purpose loads, VMs whose vCPUs float, anything you were never going to pin by hand. Do not trust it to fix a shared read cache, where every node touches the same pages and there is no correct node to migrate to. Interleaving is the honest policy for that case, and it is why NPS1 - one NUMA node per socket with memory interleaved across all the socket’s channels - is the setting AMD’s tuning guides treat as the default and recommend for most workloads, while AMD points HPC and other highly parallel workloads, the ones that do their own placement, at NPS4. What a given board ships with is the OEM’s choice, so read the BIOS rather than assume it. In tiering mode (2), also read kernel.numa_balancing_promote_rate_limit_MBps: the documentation warns that too much promotion and demotion traffic hurts application latency and suggests keeping it under a tenth of the slower tier’s write bandwidth.

Bandwidth is the real ceiling

Per channel:  GB/s = MT/s x 8 bytes per transfer   (64-bit channel)

  DDR5-4800    4800 x 8  =  38.4 GB/s
  DDR5-5600    5600 x 8  =  44.8 GB/s
  DDR5-6400    6400 x 8  =  51.2 GB/s

Per socket:
  desktop, 2 channels of DDR5-5600      2 x 44.8 =  89.6 GB/s
  8-channel server, DDR5-4800           8 x 38.4 = 307.2 GB/s
  12-channel server, DDR5-6400         12 x 51.2 = 614.4 GB/s

Same money, dual-channel board:
  one DIMM  of DDR5-6400                1 x 51.2 =  51.2 GB/s
  two DIMMs of DDR5-4800                2 x 38.4 =  76.8 GB/s  (+50%)

The 12-channel line is not a guess: AMD publishes “up to 614 GB/s over twelve DDR5 memory channels” for EPYC 9005, which the arithmetic reproduces to three significant figures. Quote such numbers as ceilings and never as throughput. Published STREAM Triad results on real servers fall short of the nameplate figure by a margin that depends on the platform and on how carefully the run is tuned - NVIDIA’s benchmarking guide expects a well tuned Triad to reach 80 to 95 per cent of theoretical peak - and a single core never approaches it at all: saturation is a property of the whole machine, which is what mlc --loaded_latency exists to show. The latency-against-bandwidth curve has a knee, and past it small bandwidth gains cost enormous latency.

Channel population matters more than rated speed. The last block is the entire buying argument: the slower kit in both slots beats the faster kit in one by half again, because a half-populated board discards 50 per cent of its ceiling while a 33 per cent speed bump recovers 33 per cent. On an 8-channel or 12-channel socket with four DIMMs fitted you are running at a half or a third of what you paid for. Populate every channel first, then argue about MT/s - and when you do argue about it, MHz is not MT/s and the number on the label is the one to read carefully. Choosing the modules to fill those channels with is a capacity-per-channel question before it is a speed question, and filling eight or twelve of them at the capacity a large cache needs means registered or load-reduced modules and a platform rated to take them. The DDR5 and DDR4 listings rank by what you are actually short of.

Two caveats. A DDR5 DIMM presents a 64-bit channel split into two independent 32-bit sub-channels, so consumer marketing sometimes calls a two-DIMM desktop “four-channel”. It is two channels; compare sockets, not sub-channels. And server platforms generally derate the supported data rate at two DIMMs per channel, trading bandwidth for capacity. That derating is specific to the CPU generation and the module rank, the vendor population tables are the only authority on it, and this guide will not print a number for it. What this site can tell you is which modules a given machine is listed as taking, on the compatibility pages.

ECC, and why a large RAM cache changes the argument

Yes, if the cache holds dirty pages. A cache is data in transit, and a bit flipped in a dirty page is written to storage as truth. The filesystem checksum, the backup and the replica downstream all then record the corrupted value faithfully, because by the time any of them sees it, it is the only version there is.

That answers it, and “do I need ECC” is still the wrong question, because ECC does not stop bits flipping. It tells you one did.

What the code corrects, and what it only detects

The commodity server baseline is SECDED: single error correct, double error detect. A single flipped bit in a word is corrected in hardware and logged, and the machine carries on. Two flipped bits are detected and not corrected, which normally halts the machine rather than returning the word. That asymmetry is the whole product: the common case is repaired silently, the rarer case becomes a crash instead of a lie. Chipkill and its equivalents extend the code far enough to correct a whole failed 4-bit-wide device, so a dead DRAM chip becomes a service ticket rather than an outage.

The overhead is one extra device’s worth of width per channel, and DDR5 doubled it:

DDR4 DIMM   64 data + 8 check = 72 bits     8/64  = 12.5 %
DDR5 RDIMM  two sub-channels of 32 + 8      16/64 = 25 %
            = 40 bits each, 80 per DIMM

A second mechanism matters specifically for caches: memory scrubbing, usually exposed in server firmware as a patrol scrub option, where the memory controller walks memory while idle and rewrites corrected words so that single-bit soft errors do not accumulate into uncorrectable pairs. In the fleet Schroeder, Pinheiro and Weber studied, three of the six platforms scrubbed, at a typical rate of 1 GB every 45 minutes; the rest found errors only on access. Memory you never read is memory whose errors you never find. A 512 GB cache holding a cold tail is exactly that memory.

One correction worth making before anyone buys on it: DDR5’s mandatory on-die ECC is not system ECC. It corrects errors inside the DRAM array before data leaves the chip, it reports nothing to the host, and it does not protect the transfer across the bus to the memory controller. A non-ECC DDR5 module is a non-ECC module. See ECC vs non-ECC for what the two words mean on a listing.

The field data, and where it runs out

Two large studies, both unusually specific. Schroeder et al. (SIGMETRICS ’09) covered Google’s fleet over about two and a half years and reported “25,000 to 70,000 errors per billion device hours per Mbit and more than 8% of DIMMs affected by errors per year”, with 1.3 per cent of machines seeing an uncorrectable error annually and errors “dominated by hard errors, rather than soft errors”. Meza et al. (DSN 2015), across fourteen months of Facebook’s fleet, found around 9.62 per cent of servers experienced correctable errors, and that failures follow a Pareto distribution in which “the average exceeds the median amount by around 55x”.

Take the distribution seriously before taking the averages. In the Google data the top 20 per cent of DIMMs with errors produced over 94 per cent of all observed errors, and among affected DIMMs the median annual count was 42 to 167 against means of 20,000 to 140,000. The failure mode is a small number of bad parts producing enormous counts, not a uniform drizzle of cosmic rays. That also makes correctable errors predictive: a DIMM with a correctable error was 27 to more than 400 times more likely to see an uncorrectable one in the same month, depending on platform, though the absolute probability stayed at 1.7 to 2.3 per cent. A rising count is a reason to schedule a replacement, not to evacuate.

Both fleets predate DDR4. There is no comparable public field study of DDR4 or DDR5 error rates at that scale, so this guide will not claim one. What the Facebook data does say about direction is that density made things worse rather than better: 4 Gb chips showed 1.8 times the failure rate of 2 Gb chips.

Why the cache size is the argument

A read cache holding clean pages has a bounded blast radius. A flipped bit returns one wrong value; the authoritative copy on disk is intact, and a re-read after eviction is correct. A write-back cache has no such floor. The dirty page is the authoritative copy until writeback completes, and Linux’s defaults leave it that way for a while: vm.dirty_expire_centisecs defaults to 3000 and vm.dirty_writeback_centisecs to 500, set as dirty_expire_interval and dirty_writeback_interval in mm/page-writeback.c, so periodic writeback leaves a page alone until it has been dirty for 30 seconds, and the flusher threads wake every five. Background writeback, memory pressure, a journal commit and an explicit fsync can all move it sooner.

Volume scales with the cache, not with that timer. vm.dirty_ratio defaults to 20, taken as a percentage of dirtyable memory rather than of installed RAM, so a 256 GB machine can legitimately hold tens of gigabytes of unflushed writes at the instant a bit flips. Per-DIMM error rates then do the rest of the arithmetic: sixteen DIMMs carry roughly sixteen times the annual exposure of one.

ZFS does not close this, and the belief that it does is a common mistake in this area. A ZFS checksum is computed from whatever is in RAM at the moment the block is written. A flip that happens before checksumming produces a block whose checksum is perfectly valid and whose contents are wrong, and every subsequent scrub will confirm it, for years. Checksums protect the medium and the transport. They do not protect the buffer that fed them. The same holds for the ARC on the read side: a flip in a cached buffer is served to the application without any checksum being consulted again.

Which platforms take it, and what buffering is for

ECC support is a platform property, not a module property. On the desktop side Intel gates ECC UDIMM support behind its workstation chipsets, W680 on LGA1700 and W880 on LGA1851, while Intel’s support note for the larger W790 platform says Xeon W-2400 and W-3400 take ECC RDIMMs and 3DS RDIMMs only; on AM5, Ryzen desktop parts will often run ECC UDIMMs but support is unofficial and board-dependent, while AMD does list ECC support on its Ryzen PRO desktop parts; the least ambiguous route on that socket is EPYC 4004 or 4005 on a server-class board. Everything above that is server memory by construction, which is where the ECC DDR4 modules worth buying live.

Buffering is a separate axis, and it is the one that decides how much cache you can install at all. A registered module buffers address, command and clock in a register so the controller drives one load instead of many; a load-reduced module adds buffers on the data lines too, presenting a single load regardless of how many ranks are inside, which is how eight-rank modules exist. Neither makes memory more reliable. They make capacity electrically possible, and a cache argument is a capacity argument. If the plan is hundreds of gigabytes of page cache or ARC, the module type is decided before the price is: registered vs unbuffered covers which of the two a given board will actually post with, and the largest modules sit in the load-reduced DDR4 listings.

Monitoring costs nothing once the hardware is right. On Linux, rasdaemon and ras-mc-ctl --error-count report corrected and uncorrected counts per DIMM from EDAC, and ras-mc-ctl --register-labels maps them to the silkscreened slot, which is the question you need answered at three in the morning. On Windows the equivalent record is the WHEA-Logger source in Event Viewer. A machine whose job is to hold a large cache should be alarming on that counter, because on this evidence the errors arrive concentrated, repeatedly, and from one bad part.

How much to buy

Enough to hold the working set. Nobody can tell you that number from here, and every ratio you have read - one gigabyte of RAM per terabyte of pool, a quarter of the database, twice the size of the index - is either the answer to a different question or a guess that acquired a decimal point on the way round.

The per-terabyte rule is worth dismantling first, because it is the one people quote with most confidence. It appears in no OpenZFS document; the OpenZFS FAQ gives a flat “8GB+ of memory for the best performance” with no capacity term at all. The TrueNAS hardware guide, usually cited as the source, says 8 GB for up to eight drives and “Add 1 GB for each drive added after eight” - per drive, not per terabyte - and reserves its only per-terabyte figure, 5 GB per TB, for deduplication, where the table really does scale with stored blocks. An 80 TB archive of large sequential reads is content with 32 GB. An 8 TB pool of random 8 KiB database reads is not. Capacity is not the variable.

Measuring a working set on Linux

Read the available column, subtract Shmem from Cached before calling the rest reclaimable, and then stop, because none of that is a working set. Three mechanisms measure one properly.

# 1. Per-file shortfall. cachestat(2) is a real syscall since Linux 6.5.
#    nr_recently_evicted counts pages that were cached and were evicted
#    recently enough that reading them again would count as a refault:
#    the size of the shortfall you are paying for, per file.

# 2. Per-process referenced set over a window.
echo 1 > /proc/<pid>/clear_refs        # reset ACCESSED/YOUNG bits
sleep 300
grep -E '^(Referenced|Pss)' /proc/<pid>/smaps_rollup

# 3. Whole-machine, region-sampled, bounded overhead: DAMON (Linux 5.15+;
#    the sysfs interface below since Linux 5.18),
#    /sys/kernel/mm/damon/admin, driven by the 'damo' userspace tool.

# Is the shortfall costing anything?
cat /proc/pressure/memory              # alarm on full avg60 > 0
sar -B 1                               # majflt/s, and %vmeff = pgsteal/pgscan

The kernel also ships idle page tracking (/sys/kernel/mm/page_idle/bitmap, with tools/mm/page-types), documented explicitly for working-set estimation: mark pages idle, wait, count the bits that were cleared.

The Active(file) and Inactive(file) counters are a useful free reading and a trap. The proc_meminfo(5) manual page still says “[To be documented.]” against both, and the kernel’s own proc.rst defines only the undivided Active and Inactive, so describe them by behaviour: a page enters the inactive file list on first touch and is promoted on a second reference, so a large inactive share means most of your cache is single-touch streaming data that reclaim will discard for nothing. On a kernel running MGLRU the two-list scheme is not the reclaim algorithm any more, only the accounting.

On ZFS the measurement is already built. arcstat (renamed zarcstat in OpenZFS 2.4.0) exposes mrug and mfug, the MRU and MFU ghost-list hits per second: requests that missed but would have hit had the ARC been larger. That is the most direct argument for buying memory that any of these systems gives you, naming misses a bigger cache would have caught. Watch dh% and dm% rather than the headline hit%, which prefetch inflates.

Measuring a working set on Windows

Windows has no vm.* knobs and a different vocabulary, but it publishes a better single indicator than Linux does.

Get-Counter -Continuous `
  '\Memory\Long-Term Average Standby Cache Lifetime (s)',
  '\Memory\Available MBytes',
  '\Memory\System Cache Resident Bytes',
  '\Memory\Modified Page List Bytes'

Microsoft’s own troubleshooting guidance names the first counter and the threshold: a long-term average standby cache lifetime under 1800 seconds is its indicator of memory pressure, which is the Windows way of saying the working set does not fit. Available MBytes is “the sum of memory assigned to the standby (cached), free, and zero page lists”, so standby pages are counted as available: “all my RAM is used” is a misreading of Task Manager, not a diagnosis. When Available MBytes is low while System Cache Resident Bytes is large, Microsoft’s instruction is to open RAMMap, whose File Summary tab answers the question no counter does: which files are resident.

Four starting points, each to be measured rather than obeyed

Desktop: 32 GB. The working set is a browser, an editor and one toolchain; everything else is page cache earning its keep on the second build, not the first. If full avg60 in /proc/pressure/memory sits at zero for a working week and vmstat’s si and so stay at zero, more memory buys nothing and the money belongs on the drive. A starting point, not a rule.

Home server or NAS: 32 GB to 64 GB. The driver here is usually metadata, not data - directory trees and inodes, accounted in SReclaimable rather than Cached. On ZFS, note what you are buying by default: since OpenZFS 2.3 the default ARC maximum is the larger of all memory minus 1 GiB and 5/8 of memory, where the Linux default had been half of memory, so a 64 GB box hands ARC about 63 GB unless you bound zfs_arc_max. Size by ghost hits. A starting point, not a rule.

Database host: enough for the indexes and the hot tables, plus a copy. The “plus a copy” is architectural. PostgreSQL’s manual puts a reasonable starting value for shared_buffers at 25 % of system memory and says above 40 % is unlikely to help, explicitly because “PostgreSQL also relies on the operating system cache” - two copies, by design. MySQL 8.4 defaults innodb_flush_method to O_DIRECT on Linux and --innodb-dedicated-server takes 75 % of detected memory above 4 GB, because InnoDB owns the only copy. Decide per layer who owns the copy, then buy for that. A starting point, not a rule.

Virtualisation host: the sum of the guests’ measured working sets, not the sum of their configured RAM. Guests size their caches to their apparent memory, so ten 32 GB guests will between them fill most of 320 GB with page cache that looks touched from the host. Reclaim then hits every guest at once and ten cache-miss storms land on the same array together - overcommit converts independent caches into a correlated failure. Use cache=none on guest disks so a block is not held in the host’s page cache and the guest’s at once. A starting point, not a rule.

The platform ceiling, which is where this becomes a purchase

Past a certain capacity the answer stops being a bigger kit and becomes a different machine. A consumer board gives four DIMM slots across two 64-bit channels and takes unbuffered modules only, because the memory controller drives address, command and clock into every device directly and electrical load, not address space, is the wall.

Consumer board, 4 slots, UDIMM only
  4 x 48 GB              = 192 GB   (board- and BIOS-dependent support)
Server socket, 12 channels, 1 DIMM per channel
  12 x 64 GB RDIMM       = 768 GB
  12 x 128 GB LRDIMM/3DS = 1.5 TB

Both columns are arithmetic on parts that exist rather than vendor population tables: 48 GB unbuffered modules are sold, but board and BIOS support for them is uneven, and the twelve-channel rows assume a twelve-channel socket, which is what EPYC 9004 and 9005 and Xeon 6900P give you.

A registered module buffers address, command and clock in an RCD. A load-reduced module adds buffers on the data lines too, so the whole module presents a single electrical load however many ranks are inside, which is how eight-rank parts exist at all; the trade is drawn out in registered vs unbuffered. The step from one column to the other changes the price per gigabyte, the slot count, the supported ranks and the ECC story at once, and that last part moves with each platform generation: Intel has gated ECC UDIMM support behind workstation chipsets such as W680 and W790, ECC on AM5 is board-dependent and unofficial, and officially supported ECC on that socket means a Ryzen PRO part, which AMD lists as supporting ECC DDR5, or an EPYC 4004 or 4005. Read the board’s QVL rather than the family name. Check the edition ceiling too - Microsoft’s specifications stop Windows 11 Home at 128 GB and Pro at 2 TB.

So the sequence is: measure, then price 64GB DDR4 modules against the slots you have; if the number needs more slots or more ranks than a desktop board can carry, you are shopping for LRDIMMs and a board that takes them, and the per-model maxima are in server memory capacity limits and the compatibility pages.

And when the measurement comes back in terabytes - a media archive, a backup target, anything read once - no memory platform holds it, the hit ratio will never get high enough to matter, and the money belongs one tier down, in flash in front of the bulk disks.

Measuring: a cache nobody checks is an expensive assumption

Run free -m, and read the available column. Every other number on that line describes where memory currently sits, not whether the cache is doing any work.

That answers the question as it is usually asked, and it is the wrong question. Occupancy is not effectiveness. A box with 200 GB in buff/cache and a 40 per cent hit rate is in worse shape than one with 20 GB and 99.9 per cent, and free cannot tell the two apart. Hits, misses and evictions that came back are what you are actually after.

free, and the one column that is not an artefact

The free(1) man page defines available as an “Estimation of how much memory is available for starting new applications, without swapping”, read straight from MemAvailable. Two things follow. First, used in current procps-ng is derived as total - available, not the old total - free - buffers - cache, so blog arithmetic from a decade ago no longer reconciles. Second, free’s cache column is Cached + SReclaimable, and Cached includes tmpfs and shmem, which are not droppable at all - they need swap. Subtract Shmem before treating that figure as reclaimable.

/proc/meminfo, field by field

Values are in kB (kibibytes, despite the label), bar the HugePages_* counts. The fields that carry information, drawing on the kernel’s own descriptions in Documentation/filesystems/proc.rst:

Field What it is
MemAvailable estimate of memory available without swapping
Cached pagecache for files, plus tmpfs and shmem, excluding SwapCached
Buffers raw block-device pagecache; the doc’s “20MB or so” is stale
Dirty “waiting to get written back to the disk”
Writeback “actively being written back”
Shmem tmpfs and shared memory, counted inside Cached
Mapped “files which have been mmapped, such as libraries”
SReclaimable the dentry and inode caches, reclaimable slab

Active(file) and Inactive(file) are literally “[To be documented.]” in proc_meminfo(5), and proc.rst does not describe them at all, so read them behaviourally: a file page enters the inactive list on first touch and is promoted on a second reference. A large inactive share means the cache is mostly single-touch streaming data that reclaim will drop cheaply; a large active share means a genuine hot set is resident. On a kernel with MGLRU enabled the two counters no longer describe the reclaim algorithm in use, only the accounting.

/proc/vmstat is the companion nobody reads. nr_dirty_threshold and nr_dirty_background_threshold report the kernel’s computed writeback limits in pages, which is how you verify a dirty_bytes setting instead of trusting your own multiplication: 524288 pages at 4096 bytes is exactly 2 GiB.

vmstat, sar and the pressure question

vmstat 1        # si/so: any sustained non-zero is real swapping
                # bi/bo: blocks in/out, i.e. the misses that reached disk
                # wa:    CPU time blocked on I/O
sar -B 1        # majflt/s, pgscank/s, pgscand/s, pgsteal/s, %vmeff
sar -r          # MemAvailable over time, from the sysstat archive

sar -B’s %vmeff is pgsteal/pgscan: reclaim efficiency. A low value under load means the kernel is scanning hard and finding little, which is short of memory rather than merely full of cache. majflt/s is the honest “we went to the device for memory” counter.

The safeguard metric is PSI. /proc/pressure/memory reports some (“at least some tasks are stalled”) and full (“all non-idle tasks are stalled”), each averaged over 10, 60 and 300 seconds. Alarm on full avg60, never on free memory. A healthy cache-heavy machine looks 95 per cent used and has a flat PSI line.

Hit ratio, measured rather than inferred

cachestat(2) has been a real syscall since Linux 6.5. It takes a file descriptor, faults nothing in, and returns nr_cache, nr_dirty, nr_writeback, nr_evicted and nr_recently_evicted for a byte range. That last field is the one to build on: a page that was cached, was evicted, and came back is a page you were short of.

The older bcc and bpftrace tool of the same name gives a per-second hit ratio across the whole system, and cachetop splits it per process. Be aware of what it is: it traces kernel page-cache functions by name, so it has had to be patched as those functions were renamed (filemap_add_folio, folio_mark_accessed, the writeback_dirty_folio tracepoint). Its own man page puts overhead as high as 34 per cent at extreme event rates. The syscall does not have either problem.

On ZFS, arcstat 1 is the instrument, renamed zarcstat in OpenZFS 2.4.0 along with arc_summary becoming zarcsummary. Ignore the headline hit%: prefetch inflates it, so a pool doing large sequential reads shows a superb ratio that says nothing about the working set. Read dh% and dm%, the demand hit and miss percentages. arc_summary reports counters cumulative since boot, which on a long-lived box is archaeology; arcstat N differences them over the interval and is the one to trust after a change.

perf answers a different question. perf stat -e major-faults,minor-faults,page-faults COMMAND attributes faults to a process, and perf record -e block:block_rq_issue shows which I/O actually reached the device. perf mem and perf c2c work at cache-line granularity inside the CPU: the right tools for a memory-bandwidth problem, the wrong layer entirely for a page-cache one.

Windows

\Memory\Available MBytes is Microsoft’s MemAvailable, and it explicitly includes the standby list, which kills “Windows is using all my RAM” on sight. The counter that answers the cache question is \Memory\Long-Term Average Standby Cache Lifetime (s): Microsoft’s own troubleshooting guidance treats a value under 1800 seconds as cache churn. RAMMap’s File Summary tab then says which files are resident, which no counter will. Its Empty menu is a diagnostic, not a tuning step - emptying the standby list discards a warm cache that Windows must re-read.

What a good ratio is, and why

T_eff = h·T_hit + (1 - h)·T_miss        T_hit = 0.1 µs (DRAM)

  h        miss = NVMe 100 µs    miss = 7200 rpm HDD 10 ms
  90%           10.09 µs                1000.1 µs
  99%            1.10 µs                 100.1 µs
  99.9%          0.200 µs                 10.1 µs
  99.999%        0.101 µs                  0.200 µs

The miss term dominates until (1 - h) < T_hit/T_miss. Both miss figures are round illustrative numbers rather than datasheet values; substituting your own device changes the arithmetic, not the shape. The hit ratio you need is a property of what is underneath you, not a universal target: about 99.9 per cent over NVMe, about 99.999 per cent over a spinning disk. Anyone quoting “95 per cent is good” has not said good for what.

Warm-up, and the benchmark that proves nothing

Every first measurement is a lie, in both directions. A cold cache makes the system look terrible; a run repeated three times makes it look perfect. Get a deliberate baseline with sync; echo 3 > /proc/sys/vm/drop_caches, then the warm run, and quote both. The kernel logs who dropped the caches, by command and PID, so dmesg | grep drop_caches is also how you catch a cron job destroying the cache nightly “to free memory”.

The second trap is structural. A synthetic sequential fio read shows the cache contributing nothing, because a streaming read is single-touch by construction: it lands on the inactive list and is evicted before it is ever requested twice, and readahead was already covering the latency. That is the cache working as designed on data with no reuse, not a fault. Benchmark the access pattern you have, or benchmark nothing.

The one measurement that settles the purchase

Ghost hits. ZFS exposes them directly as mrug and mfug in arcstat: requests that missed but would have hit had the cache been larger. Linux’s equivalent is nr_recently_evicted from cachestat(2); Windows’ proxy is the standby cache lifetime above. Sustained ghost hits are the most direct measured argument for buying memory to cache with. A flat ghost-hit line means more RAM buys the cache nothing, whatever free shows.

If the number is high and stays high, size the purchase to it rather than to the dataset, and note what a cache that large implies: capacity beyond a couple of DIMMs needs registered or load-reduced modules and a board whose controller will take them, which is the mechanism covered in registered vs unbuffered; what a given machine actually accepts is on the server compatibility pages. And if ghost hits stay high after every channel is populated with the largest module the platform supports, the working set will never be resident, and the next question is a faster tier underneath rather than a bigger one above: that is the point where the answer stops being memory at all.

Safeguards: bound the window, because you cannot close it

Every cache in this article is volatile, and no setting changes that. The only question a safeguard answers is narrower: at any instant, how many of your bytes exist in RAM alone, and for how many seconds?

The opening split the cache in three, and that split is the whole of this section. A clean page is a second copy of bytes the device already holds; losing it costs a re-read and nothing else, which is why the kernel can drop clean pages without any I/O at all. A dirty page is the authoritative copy until writeback completes, for a window you configured, knowingly or not. An explicit RAM store - tmpfs, ramfs, brd, a Windows RAM disk, a Redis instance with persistence off - is dirty by construction: nothing underneath it has ever seen the data, and there is no window, only a promise you have not made.

The dirty window, in numbers

The defaults are in the initialisers in mm/page-writeback.c, not in Documentation/admin-guide/sysctl/vm.rst, which describes the knobs without stating their values. Cite the source file; a great deal of published advice quotes dirty_ratio = 40, which stopped being the default in 2.6.22, in 2007.

static int dirty_background_ratio = 10;          /* % : flushers start      */
static int vm_dirty_ratio = 20;                  /* % : writer is throttled */
unsigned int dirty_writeback_interval = 5 * 100; /* centiseconds -> 5 s     */
unsigned int dirty_expire_interval = 30 * 100;   /* centiseconds -> 30 s    */

Both percentages are of dirtyable memory - free plus reclaimable pages - not of installed RAM; the kernel documentation says plainly that “the total available memory is not equal to total system memory”, and the figure moves as the workload moves. On a server with 256 GiB mostly holding reclaimable page cache, where the two numbers are at least close, the arithmetic is unpleasant:

256 GiB x 0.10  = 25.6 GiB dirty before background writeback is forced by volume
256 GiB x 0.20  = 51.2 GiB dirty before the writing process is throttled
25.6 GiB        = 27.5 GB
27.5 GB / 200 MB/s = ~137 s to drain, if the device sustains 200 MB/s

Two separate bounds, and you need both. Volume is bounded by the ratios above. Age is bounded by dirty_expire_centisecs, 30 s, plus up to one dirty_writeback_centisecs wakeup at 5 s, plus however long the writeback itself takes: up to 35 s of acknowledged writes before the drain even begins, and tens of gigabytes of them, on stock settings. Setting dirty_writeback_centisecs=0 disables periodic writeback entirely, which makes the age bound infinite.

The fix is absolute limits sized to a few seconds of the device’s write bandwidth, then a read-back to confirm the kernel agrees:

sysctl -w vm.dirty_background_bytes=536870912    # 512 MiB
sysctl -w vm.dirty_bytes=2147483648              # 2 GiB
grep -E 'nr_dirty_(background_)?threshold' /proc/vmstat   # in PAGES
# 524288 pages x 4096 = 2147483648 bytes, so the knob took effect

Only one form applies at a time: when one is written the kernel evaluates the limits from it, and the other reads back as 0. Windows exposes the same quantity as Memory\Modified Page List Bytes and offers no sysctl equivalent; its one documented threshold, RemoteFileDirtyPageThreshold, is per file, applies to remote writes, defaults to 5 GB on Windows Server 2016 and later, and is expressed in 4096-byte pages.

What fsync guarantees, and who below it can still lie

fsync(2) “transfers (‘flushes’) all modified in-core data … to the disk device”, flushes the file’s metadata, includes “writing through or flushing a disk cache if present”, and blocks until the device reports completion. fdatasync(2) omits metadata not needed to read the data back, which is why databases pre-allocate files: an overwrite in place needs no size update.

What it does not cover, in order of how often each bites:

  • The parent directory. The man page is explicit: fsync() “does not necessarily ensure that the entry in the directory containing the file has also reached disk”. Write-temp-then-rename is not crash-safe without an fsync on the directory descriptor.
  • The error you already consumed. Since Linux 4.13 writeback errors reach all descriptors that might have written the data; the page is then marked clean, so a retried fsync() can return success over data that is gone.
  • O_DIRECT. It bypasses the page cache. It is not a durability primitive and does not bypass the drive’s own volatile write cache; O_SYNC is the flag that asks for that.
  • The block layer and the device. cat /sys/block/sda/queue/write_cache returns write back or write through; nvme get-feature /dev/nvme0 -f 6 reads the Volatile Write Cache feature. Mounting ext4 with nobarrier turns every fsync() above it into a lie; the kernel documentation calls disabling barriers safe only “if your disks are battery-backed in one way or another”.
  • The filesystem’s own interval. ext4 defaults to commit=5, and the kernel documentation says plainly that “if you lose your power, you will lose as much as the latest 5 seconds of metadata changes”.

The NVMe specification contains the cleanest definition of the hardware that ends this chain: if a controller “is able to guarantee that data present in a write cache is written to non-volatile media on loss of power, then that write cache is considered non-volatile and this feature does not apply to that write cache”. That is power-loss protection, and it is why a capacitor-backed datacentre drive is fast at fsync() while a consumer one is either slow or dishonest. If the durable copy has to land somewhere, it lands on a device with that property, and that is a different purchase from one without it.

What a UPS buys

Time, and an ordered shutdown during which the OS flushes, the database checkpoints and the drive empties its cache. That is genuinely valuable and it is all it buys. It does nothing for a kernel panic, a hypervisor crash, a watchdog reset, an OOM kill, a PSU failure, a tripped internal breaker, a pulled C13, or a PDU fault downstream of the battery. It does nothing at all if no daemon is wired to it, in which case it is a device that delays the crash by ten minutes, and nothing if the runtime is shorter than the drain time calculated above. VRLA batteries age out in a few years and fail quietly. PostgreSQL’s manual puts the general case better than a safeguards checklist can: “High quality hardware alone is not a sufficient justification for turning off fsync.”

Battery-backed and flash-backed write caches exist precisely because a UPS cannot cover the last few hundred milliseconds. They carry their own operational tax: a MegaRAID BBU’s write cache “remains disabled while a battery learn cycle is in process”, so the array silently drops to write-through on a schedule. Supercapacitor CacheVault modules dump cache to NAND on power loss instead.

Crash consistency, stack by stack

Stack On an unclean stop
Page cache over ext4/XFS Journal replays; up to 30 s of data gone. On ext4, commit=5 bounds metadata loss at 5 s and the default data=ordered orders data before the metadata that names it; XFS has neither option and relies on its own log
tmpfs, ramfs, brd Everything, always. tmpfs pages may sit in swap, which is not persistence
zram swap, zswap Nothing to lose: swap contents are meaningless across a boot. zram carrying a filesystem is a RAM disk and loses everything
ZFS ARC The ARC is a cache; durability lives in the intent log and transaction groups. OpenZFS tells PostgreSQL users to set full_page_writes = off “as ZFS will never commit a partial write”
InnoDB buffer pool innodb_flush_log_at_trx_commit=1 is full ACID; 0 and 2 mean roughly a second, and the manual warns that once-per-second flushing “is not 100% guaranteed”
PostgreSQL shared buffers fsync and full_page_writes on by default, the latter against a page write “only partially completed”, leaving “a mix of old and new data”
Redis, memcached appendfsync everysec is the default once the append-only file is on, and loses a second; RDB snapshots lose minutes; memcached loses all of it
Windows write caching Memory\Modified Page List Bytes is the exposure. PrimoCache’s Defer-Write documentation is the honest version: the data “will have no chance to be written to the disk, resulting in loss”, possibly including filesystem metadata

Why this cannot be bought away

A power-loss-protected SSD is safe because a capacitor finishes the write. Ordinary DRAM has no equivalent. NVDIMM-N modules, which pair DRAM with flash and a supercapacitor that saves the contents on power loss, are a niche server part and not what a cache is built from. ECC does not change that; it changes the other failure, and a write-back cache is exactly where that failure is worst. A dirty page is the authoritative copy, so a bit flipped between the application’s write and the flush is written down as truth, and every checksum, replica and backup downstream records it faithfully. Filesystem checksums protect the medium and the transport, never the buffer they are computed from. Google’s fleet study found more than 8 per cent of DIMMs hit by correctable errors in a year, and in a month with one an uncorrectable error is 27 to over 400 times more likely. A large cache means many modules, so the sensible buy is ECC and, past a couple of hundred gigabytes, registered or load-reduced parts on a platform that takes them. The argument for ECC is set out in ECC vs non-ECC.

So the mitigation is not safety, it is a smaller window: absolute dirty_bytes sized to seconds rather than gigabytes, an application-level fsync policy you have actually read, a journalling or copy-on-write filesystem that recovers cleanly, and a device that does not lie about its cache.

And none of it is a backup. A cache is a copy of something else, right up until it is the only copy, and a snapshot of a corrupted buffer is a faithful snapshot. Back up the store, verify a restore, and treat the cache as what it is: speed you are renting from a volatile device.

The enterprise and SMB tier

Buy modules until the ghost lists stop hitting, and buy them in the shape the memory controller asks for. Every technique in this guide is free; the only thing that costs money is how many channels the socket exposes and how many modules the board will take.

That answers it, and it is the wrong question for the person signing the order. At this tier a cache is no longer one machine’s page cache. It is thirty guests’ page caches on one host, a buffer pool that must never be swapped, and a platform that either accepts twenty-four modules or does not. Those are purchasing decisions, and they are made once.

Overcommit, ballooning, and the host that swaps

Broadcom documents ESXi’s reclamation order plainly: page sharing first, then ballooning, then memory compression, then host-level swapping. The balloon driver vmmemctl “collaborates with the server to reclaim pages that are considered least valuable by the guest operating system”, and per-VM you can cap it with sched.mem.maxmemctl, in megabytes. Compression exists because “decompression latency is much smaller than the swap-in latency”. Host swapping is the floor, and the guidance is to stay off it: “swapped pages might be active, which can cause virtual machine performance to degrade significantly”.

KVM has the same ladder with different parts. virtio-balloon gained free page reporting, which is a weaker hold than a standard balloon: the guest may reuse a reported page as soon as the host acknowledges the report, so the saving lasts only as long as the guest leaves that page alone. libvirt’s autodeflate attribute turns on the QEMU balloon’s ability to give memory back at the last moment, before the guest’s out-of-memory killer takes a process. KSM is opt-in and its run default is 0; the kernel documentation says its pages_to_scan of 100 and sleep_millisecs of 20 were “chosen for demonstration purposes”, which is a polite way of saying nobody running stock KSM has tuned it.

Now the failure this section exists for. A guest sizes its cache to its apparent RAM, because that is the only number it can see.

10 guests x 32 GiB allocated             = 320 GiB guest-visible
host physical                            = 256 GiB
each guest settles at ~20 GiB page cache = 200 GiB the guests call "in use"
clean pages the host can take with no I/O =  0 GiB

Every one of those 200 GiB is clean, reclaimable file cache that the guest would drop instantly. To the host it is touched, resident memory inside a guest’s address space. The balloon cannot see it as cold, so it inflates, and ten guests reclaim their caches in the same second and send ten cache-miss storms at one array. Overcommit converts independent caches into a correlated failure. Two mitigations, both structural: size guests against the sum of their working sets rather than the sum of their allocations, and open guest disks with cache=none so the host page cache and the guest page cache do not hold the same blocks twice. Double caching is a common and avoidable waste of RAM in a virtualised estate.

What an in-memory database demands

It demands more than it stores. Redis’s maxmemory bounds the dataset; persistence bounds the machine. RDB snapshots and AOF rewrites fork(), and every page the parent dirties during the child’s run must be copied, so Redis asks for vm.overcommit_memory = 1, warning that without it “a background save or replication may fail under low memory condition”, and asks for transparent huge pages off, because a copy-on-write fault at 2 MiB instead of 4 KiB is a 512-fold write amplification. Replication and AOF buffers are not counted against maxmemory at all. The widely repeated rule that maxmemory should sit at half of RAM has no authoritative source, and this guide will not claim one; the defensible statement is that the headroom must cover a full fork under your write rate, and that mem_not_counted_for_evict in INFO memory is where you watch it.

The buffer-pool engines differ by design, not by taste. InnoDB manages its own cache and is commonly configured with O_DIRECT so that the same page is not cached twice; MySQL’s --innodb-dedicated-server sizes the buffer pool at 75 % of detected memory once that memory is above 4 GB. PostgreSQL shares deliberately, so 25 % is the documented starting point and more than 40 % is “unlikely to work better”. SQL Server states the tiering outright: Buffer Pool Extension is capped at 32 times physical memory on Enterprise and 4 times on Standard on a physical machine, and Microsoft’s advice is to keep “the ratio between the size of physical memory and the size of the buffer pool extension” at “1:16 or less”, with “1:4 to 1:8” described as possibly optimal. A vendor whose product sells the flash tier still tells you to buy the RAM first.

Persistent memory in Memory Mode, and why it is not an answer

Memory Mode was two-level memory: DRAM as a transparent, hardware-managed cache in front of Optane DC Persistent Memory (DCPMM), with the operating system seeing only the far capacity. It was volatile. Intel’s own support note is titled “Why Is the Intel Optane Persistent Memory in Memory Mode Not Persistent?”, because the volatile key is regenerated each power cycle.

The measured cost of the far tier, in a published characterisation of a Cascade Lake platform: 305 ns for a random load against 81 ns for DRAM, 169 ns sequential, and with the DRAM cache disabled, memcached and Redis lost 20.1 % and 23.0 % respectively. That is the honest shape of it. A hit in near memory is DRAM; a miss is roughly four times DRAM latency.

Optane was on the open market from April 2017. Intel announced the wind-down in its Q2 2022 earnings, reported 28 July 2022, taking a 559 million dollar inventory impairment and citing the industry’s move to CXL. The 300 series was cancelled and the business was wound down over the years that followed. Intel’s support page says only that it “plans to cease future development” and that the five-year warranty terms from date of sale are unchanged, so on a module sold in 2020 that clock has already run out.

Second-hand today, DCPMM buys terabyte-class capacity per socket for very little money. It does not buy a supported platform: it needs a matching Cascade Lake or Ice Lake Xeon, a BIOS that still exposes the mode, firmware nobody will update again, and a media-wear counter you cannot reason about from a listing photograph. For a read-mostly cache on a machine you already own it is a defensible experiment. As the memory plan for a business, it is a dead end with an expired warranty.

CXL, stated as a direction

CXL 2.0 (November 2020) added switching and memory pooling, 3.0 (August 2022) added fabrics, and the specification has kept moving through 3.1, 3.2 and 4.0 (November 2025). The number worth carrying is from Microsoft Research’s Pond: small pools of 8 to 16 sockets add 70 to 90 ns over NUMA-local DRAM, and rack-scale pooling adds more than 180 ns. With either 64 ns or 140 ns of additional latency emulated, 43 % and 37 % of 158 representative workloads stayed within 5 % of the performance of local DRAM. Read that the way it is written: roughly two workloads in five tolerate pooled memory, and the majority do not. CXL is a capacity and cost technology for organisations whose DRAM bill is a line item on an earnings call. There is nothing here for a small business to buy this year.

What a small business should actually buy

Measure first, for a week. On Linux, full avg60 in /proc/pressure/memory plus majflt/s and %vmeff from sar -B. On Windows, Microsoft’s own threshold: \Memory\Long-Term Average Standby Cache Lifetime (s) under 1800 seconds means the cache is being churned. If pressure is flat and standby lifetime runs to hours, more RAM buys nothing, the working set already fits, and the money belongs one tier down on storage instead.

If it does not fit, buy shape before speed.

DDR5-4800, 8 bytes per transfer: 4800 x 8 = 38,400 MB/s per channel
1 x 128 GB module, one channel populated:     38.4 GB/s
4 x  32 GB modules, four channels populated: 153.6 GB/s

Same capacity, same generation, four times the bandwidth. Two modules per channel derates the supported rate on server platforms, by an amount specific to the CPU generation and module rank; check the vendor population table rather than a rule of thumb.

Two things at this tier are not upgrades. ECC is not an upgrade. Google’s fleet study found more than 8 % of DIMMs affected by errors per year and 1.3 % of machines hit by an uncorrectable error per year, with errors dominated by hard faults rather than cosmic rays. A read cache that flips a bit returns one wrong value; a write-back cache that flips a bit writes the lie to disk as truth, and every checksum downstream records it faithfully. DDR5’s mandatory on-die ECC does not close this: it corrects inside the die, it does not protect data in transit to the CPU, and it does not make a standard module a server module. ECC vs non-ECC has the rest of the argument.

A platform that takes enough modules is not an upgrade either. A desktop board runs out of channels, out of electrical budget for unbuffered modules, out of module capacity and out of ECC at roughly the same point. A cache measured in hundreds of gigabytes needs buffered memory and a socket that accepts it, which means registered DDR4 or load-reduced modules on a server board. Check the population rules for the chassis you actually have before ordering: the per-model pages for Dell servers list what each generation takes, and how to check what RAM fits covers reading it off a running machine. One last ceiling to check before the invoice: Windows 11 Home caps at 128 GB, Pro at 2 TB.

Four setups, end to end

Each recipe below ends with the measurement that proves it worked. Run the verification before the change as well as after it, or you have a story rather than a result. All commands are run as root unless marked otherwise.

1. A Linux file server over a metadata-heavy tree

For: a mail spool, an rsync target, a build tree, a home-directory export - millions of small files where stat() outnumbers read(). Not for: a database host (recipe 3), and not for a media server whose files are streamed once and never re-read, where the cache holds nothing worth holding.

Start by finding out whether the cache is full or churning. They look identical in free, and only the second one is a reason to spend money.

# is the file cache single-touch streaming data, or a resident hot set?
grep -E 'nr_file_pages|nr_active_file|nr_inactive_file' /proc/vmstat
cat /proc/pressure/memory          # full avg60 above 0 means real stalls
sar -B 1 5                         # majflt/s, and %vmeff = pgsteal/pgscan

Then tighten the dirty window. The percentage knobs are a fraction of dirtyable memory, not of installed RAM, and on a large box 20 % is a multi-gigabyte wall that a slow array needs minutes to drain. Absolute limits sized to a few seconds of the device’s write bandwidth are the standard fix. Writing the bytes variant zeroes the ratio variant, which is why the ratios read back as 0.

sysctl -w vm.dirty_background_bytes=536870912   # flush from 512 MiB
sysctl -w vm.dirty_bytes=2147483648             # throttle at 2 GiB
grep -E 'nr_dirty_(background_)?threshold' /proc/vmstat
# 524288 pages x 4096 B = 2147483648 B. The kernel agrees with you.

Persist that in /etc/sysctl.d/99-writeback.conf. Then bias reclaim towards dentries and inodes, and prove the bias with a timed walk rather than a feeling. vm.vfs_cache_pressure defaults to 100; 50 keeps metadata at the expense of file data, which is the right trade here and the wrong one in recipe 3.

sysctl -w vm.vfs_cache_pressure=50
sync; echo 3 > /proc/sys/vm/drop_caches     # benchmarking only
time find /srv/export -type f | wc -l       # cold
time find /srv/export -type f | wc -l       # warm
grep -E 'SReclaimable|^Cached' /proc/meminfo

The cold-to-warm ratio is the value of the cache, stated in seconds. Note that the metadata lands in SReclaimable, not in Cached - that is the tell that you are measuring the dentry and inode caches and not the page cache. drop_caches belongs in that baseline and nowhere else; the kernel’s own documentation says it “is not a means to control the growth of the various kernel caches”, and it logs the command and PID that used it, so dmesg | grep drop_caches will find the cron job someone wrote to “free memory”.

If the cold number is what your users get all day, the answer is capacity, not a sysctl. Holding a tree that large resident means a platform with enough slots and enough electrical budget, which in practice means registered or load-reduced modules on a board that takes them - the ceilings are set out in server memory capacity limits, and the cost per gigabyte across generations is ranked in RAM prices.

2. A ZFS host

For: a pool whose demand reads repeat. Not for: an archive that is written once and read never; the ARC cannot cache what nobody asks for twice.

Set zfs_arc_max because you decided, not because the default suited you. From OpenZFS 2.3 the default is the larger of all_system_memory - 1 GiB and 5/8 x all_system_memory, so a 64 GiB box defaults to 63 GiB of ARC; on 2.2 and earlier the Linux default was half of RAM. Read man 4 zfs on the machine in front of you rather than trusting either number.

cat /sys/module/zfs/parameters/zfs_arc_max
echo 68719476736 > /sys/module/zfs/parameters/zfs_arc_max   # 64 GiB, now
printf 'options zfs zfs_arc_max=68719476736\n' > /etc/modprobe.d/zfs.conf
# root-on-ZFS: the module loads from the initrd, so rebuild it the way your
# distribution does - on Debian and Ubuntu that is update-initramfs -u

Choose primarycache per dataset. It defaults to all; metadata is right where something else already caches the data, which is OpenZFS’s own advice for InnoDB datasets and for a swap zvol.

zfs set primarycache=metadata tank/mysql
zfs get primarycache tank/mysql tank/media

Prove it with zarcstat, which was called arcstat before OpenZFS 2.4. The headline hit% is polluted by prefetch and flatters a pool doing sequential reads. The demand columns answer the question, and the ghost-list columns answer the next one.

zarcstat -f time,read,dread,dh%,dm%,mrug,mfug,arcsz,c 1   # arcstat before 2.4

mrug and mfug count requests that would have hit had the ARC been larger. Sustained ghost hits are the closest thing to a direct measurement of what more memory would buy. Sample with an interval rather than reading zarcsummary once: the kstats are cumulative since boot, so a one-shot ratio on a long-lived box is archaeology. If dm% stays high with the ARC already as large as the board will allow, the working set does not fit in memory at any price you will pay, and the next tier is flash rather than DRAM - l2arc_write_max is documented as bytes per feed interval, 32 MiB by default against an interval of one second, so an L2ARC warms slowly.

3. A database host

For: a dedicated engine. Not for: a box also serving files, where both caches will fight for the same pages.

Decide, explicitly, who owns the copy. MySQL 8.4 defaults innodb_flush_method to O_DIRECT on Unix-like systems where the platform supports it, so the buffer pool is the only copy and the manual’s “up to 80 % of physical memory” follows from that. PostgreSQL buffers its reads deliberately and documents 25 % as a starting point “because PostgreSQL also relies on the operating system cache”; its debug_io_direct is labelled “intended for developer testing only”. Paying for both caches is the default outcome and the commonest waste.

# MySQL: let it size itself against detected memory
mysqld --innodb-dedicated-server   # 1 to 4 GB: x0.5; above 4 GB: x0.75
# size must be a multiple of chunk size x instances, or it is rounded up

For PostgreSQL, huge pages need no arithmetic on your part. The server computes the pool for you, and with huge_pages set to on it refuses to start if the reservation is missing.

postgres -C shared_memory_size_in_huge_pages -D "$PGDATA"   # prints N
sysctl -w vm.nr_hugepages=N        # substitute the printed number
# set huge_pages = on in postgresql.conf, then restart
psql -c 'SHOW huge_pages_status;'
grep -E 'HugePages_(Total|Free|Rsvd)' /proc/meminfo

HugePages_Rsvd jumping to the printed number when the server starts, and HugePages_Free falling as the pool is touched, is the proof. On ZFS, add the double-caching decision: primarycache=metadata, recordsize=16K on the data files and logbias=throughput, which hands the data cache to the buffer pool. Verify with SHOW ENGINE INNODB STATUS, whose “Buffer pool hit rate 1000 / 1000” is a windowed figure, not a lifetime one. On PostgreSQL, blks_hit counts only shared_buffers, never the kernel’s cache, so a block served from RAM in microseconds is recorded as a miss - chasing that ratio towards 99 % by inflating shared_buffers is how you end up past the 40 % the manual warns against.

4. A Windows workstation

For: deciding between a RAM disk and more memory. Not for: tuning; there is no vm.* equivalent, and the supported cache control is one API pair, GetSystemFileCacheSize and SetSystemFileCacheSize.

Read the numbers honestly first. Task Manager’s “Cached” is the system working set plus the standby and modified lists, and “Available” is standby plus free plus zeroed, so the two overlap: a large Cached figure is not memory spent. Microsoft publishes one threshold worth alarming on.

Get-Counter '\Memory\Available MBytes',
  '\Memory\Long-Term Average Standby Cache Lifetime (s)',
  '\Memory\System Cache Resident Bytes' -Continuous

Under 1800 seconds of average standby lifetime is the churn signal Microsoft names, and it counts alongside a low Available figure rather than on its own. Then open RAMMap and read File Summary, which lists file data in RAM by file; File Details breaks the same set down to individual pages. If lifetime runs to hours and Available is healthy, more memory buys nothing and a RAM disk buys less: Windows ships no supported general-purpose RAM disk, so every option is a third-party kernel driver, and a deferred-write cache layered over it holds writes in volatile memory, so an ungraceful shutdown loses whatever has not yet reached the disk. If lifetime is under the threshold and Available is low, buy memory, and check two ceilings before you order. The edition first: Windows 11 Home stops at 128 GB, Pro at 2 TB, Pro for Workstations and Enterprise at 6 TB. Then the board, which is what how to check what RAM fits is for. The verification is the same counter, left running for a working day after the modules go in.

When more RAM is the wrong purchase

When the measurement says it will not be spent. A cache pays only at the margin where it changes the hit ratio, and there are five common cases where the next module changes nothing at all. This is the section a shop that sells memory has the least incentive to write, which is why it is here.

The working set is larger than any amount you can buy

Effective latency is a weighted average, and the weight matters more than either term:

T_eff = h·T_hit + (1 - h)·T_miss

T_hit  =      0.1 us  (DRAM, ~100 ns)
T_miss =    100 us    (datacentre NVMe, 4 KiB random read)
T_miss = 10,000 us    (7200 rpm, 4.16 ms rotational latency plus seek)

h = 0.99, NVMe backing:  0.99·0.1 + 0.01·100    =   1.10 us
h = 0.99, HDD  backing:  0.99·0.1 + 0.01·10,000 = 100.10 us

The miss term dominates until (1 - h) < T_hit / T_miss. Against NVMe that crossover is a hit ratio of about 99.9 per cent; against a 7200 rpm disk it is about 99.999 per cent. A 40 TB working set of random reads with 512 GB of RAM is nowhere near either figure, and doubling the RAM moves h from about 1 per cent to about 2.5 per cent, nowhere near the crossover, while the miss term stays exactly where it was.

The arithmetic also says what to buy instead. Adding memory raises h. Putting flash under the disks lowers T_miss by roughly a hundredfold, which raises the miss ratio you can afford by roughly the same factor. When the hot data will never fit in any module count the board will take, the money buys flash for the hot set and leaves the cold remainder on spinning disks, where capacity is cheap and latency no longer decides anything.

The workload reads each block once

A backup, a media transcode, a log ship, a nightly extract: data enters the cache, is read once, and is never asked for again. The kernel already handles this correctly. A clean file page is dropped by unlinking it from the LRU with no I/O, and a single-touch page stays on the inactive file list where reclaim scans it first. The tell is in the counters: a large Inactive(file) against a small Active(file), nr_recently_evicted from cachestat(2) near zero, and a hit ratio that never improves no matter how much of the file is resident.

Nothing about that is a memory shortage, and a larger cache does not compress a single-pass stream. The knobs that do move it are the readahead window, /sys/block/<dev>/queue/read_ahead_kb, posix_fadvise with POSIX_FADV_SEQUENTIAL (which the man page says doubles the default readahead), and the sequential bandwidth of the device itself. Readahead is also per device: setting it on the members of a device-mapper or MD stack does not reach the top-level device, which is a common reason tuning appears to do nothing.

The platform has run out of slots, channels and electrical budget

Consumer sockets expose two DDR5 memory channels, take unbuffered modules only, and commonly drop their supported data rate when the second slot per channel is filled. There is no register to buffer address and command lines and no data buffers to hide rank loading, so the wall is electrical rather than architectural. A server socket takes registered or load-reduced modules, and an LRDIMM presents the controller a single load regardless of how many ranks are inside, which is how eight-rank modules exist at all. That is the platform question, not a module question, and it is set out in server memory capacity limits and registered vs unbuffered.

Past the ceiling the economics invert. The last supported module for an old generation frequently costs more than a whole newer machine, and the older board often has 2, 4 or 6 memory channels where a current one has 8 or 12, so the bandwidth ceiling was already the lower one before capacity ran out. Check the generation your board takes against every other one at price per gigabyte before committing: if the per-gigabyte figure for the old socket is above the current one, the honest purchase is a different platform or a second machine.

The bottleneck was never storage

Memory bought for a CPU-bound or network-bound workload sits idle and compresses nothing. Measure first, on the machine, over a week rather than a minute:

# Linux: is the machine genuinely short of memory?
grep -E '^(some|full)' /proc/pressure/memory   # full avg60 > 0 means stalls
sar -B 1 5                                     # majflt/s, pgscan, %vmeff
vmstat 1 5                                     # si/so should be 0
Get-Counter '\Memory\Long-Term Average Standby Cache Lifetime (s)',
            '\Memory\Available MBytes' -Continuous

Microsoft lists a standby cache lifetime under 1800 seconds among the counters to check when a server is sluggish and available memory is nearly gone; a short lifetime means cached pages are churned out before they can be reused. If PSI full avg60 is flat, major faults are near zero and the standby cache lives for hours, the storage path is not what is limiting the machine and more RAM will change a number nobody is waiting on.

The application’s own cache was never configured

Defaults ship for laptops. PostgreSQL’s shared_buffers default is 128 MB. InnoDB’s innodb_buffer_pool_size default is 134217728 bytes, also 128 MB. memcached defaults to -m 64. Redis ships maxmemory 0 with maxmemory-policy noeviction, so it has no limit at all until one is set, and once one is set it refuses writes rather than evicting. The varnishd reference gives default,100m as the storage an omitted -s leaves you with (Varnish Cache is now a distribution of the open source Vinyl Cache project). A 256 GB server running any of those at its default has hundreds of gigabytes of unused memory and an application cache sized for a workstation.

Two knobs in particular save query time without costing a byte: PostgreSQL’s effective_cache_size, default 4 GB, which allocates nothing and only tells the planner how much cache to assume when costing an index scan, and MySQL’s innodb_dedicated_server, which sizes the buffer pool at 0.75 of detected memory on a machine with more than 4 GB, in one line. Both are free. Neither requires a purchase order.

What the cache was, all along

Every one of these cases lands back where the article started. The cache is not something to install. It has been running since the machine booted, in the page cache on Linux, in the standby list on Windows, in the ARC on ZFS, and a healthy box has always shown a small free and a large buff/cache because that is what correct looks like. The only questions ever on the table were how large that cache is and what it is permitted to hold.

Capacity answers the first, and once the measurement says capacity is the binding constraint, the modules are usually second-hand server parts with their own checks to make, covered in buying used RAM on eBay. Configuration answers the second: dirty limits in bytes rather than ratios, an ARC or buffer pool sized so only one layer owns each copy, readahead matched to the access pattern, ECC underneath any of it that holds dirty data. And when the answer to both is that the working set will never fit, the next question is not about memory at all, but about what sits between the cache and the platters, which is a cache in front of the drives and a separate argument with its own arithmetic.

Related guides

Affiliate disclosure:We are a member of the eBay Partner Network and earn a commission from qualifying purchases made through links to eBay on this site. Prices and availability are captured periodically and may have changed - the live price is always the one shown on eBay.