Why your server won't take its maximum RAM
By Harry Saarinen · Updated
Because the advertised figure is the best processor, the largest qualified module and a fully populated chassis multiplied together, and a second-hand machine is almost never all three at once. HPE prints the formula in its own QuickSpecs boilerplate: “The maximum memory capacity is a function of the number of DIMM slots on the platform, the largest DIMM capacity qualified on the platform, and the number and model of installed processors qualified on the platform.” Every term in that sentence is a variable, and the marketing page quotes the best value of each.
Buy against the headline and the memory arrives perfectly good and perfectly useless. It seats, and then the machine either refuses to POST or posts at a capacity nobody promised, and nothing in the box tells you which of the eleven separate ceilings underneath the headline you have just run into.
This article takes them one at a time: the processor’s addressable maximum and its per-socket limit, the channels and the slots wired to them, the firmware table that decides which modules the machine has ever heard of, the chip-select budget that caps ranks rather than modules, the module type that has to be chosen once for the whole machine, the population rules that decide whether the capacity you installed is interleaved or stranded, the reliability modes that spend capacity on redundancy, the address space under four gigabytes that is not memory at all, and the operating system that may refuse the last terabytes for reasons of licence. Then the procedure for working the real figure out of three documents and the running machine, four machines worked end to end, and the consequence on the second-hand market, which is that the cheap route to a capacity is usually more modules rather than bigger ones. It is long because the subject has four layers - the silicon, the board, the firmware and the standards - and the advice that goes wrong is almost always advice that covers one layer and assumes the rest.
Three housekeeping notes, because they decide how to read everything below. The command output shown is the shape of the output, carrying one consistent example machine through the article, and is not a capture from one particular server. That machine is a two-socket Xeon Gold 6130 with 24 DIMM slots and twelve 32 GB dual-rank registered modules fitted, and it is used throughout so that the numbers in one section can be checked against the numbers in another. Where a figure is a vendor’s, the vendor and the document are named beside it. Where a figure is worked out here, the arithmetic is written out so that you can check it, and the text says that it is a derivation rather than a disclosure. And where the evidence is this site’s own catalogue, prices and inventory move week to week, so the claim is phrased as something you can go and re-check rather than frozen into a number that will be wrong by the time you read it.
Eleven ceilings, and only four of them move
Every one of these is a real limit that has stopped somebody’s memory upgrade. They are enforced by different things, they bind in different orders on different machines, and the binding one is rarely the one in the marketing copy.
| Ceiling | Enforced by | Can you buy around it |
|---|---|---|
| Addressable physical memory | The processor’s physical address width | Not on 64-bit hardware, and it never binds there |
| Per-socket maximum | The memory controller, sometimes a fuse and a price list | Yes - fit a different processor |
| Sockets populated | The chassis, and whether the second processor was ordered | Yes - fit the second processor |
| Channels per socket | The processor | No |
| Slots per channel | The board’s routing | No |
| Largest die the firmware knows | The memory reference code in the system firmware | Sometimes, with a firmware update |
| Ranks per channel | Chip-select pins on the memory controller | Partly - stacked modules dodge the count |
| Module type | Board, processor and firmware together, machine-wide | Yes, but only by rebuilding the whole machine on one type |
| Qualified configuration | The vendor’s test matrix | Little worth risking |
| RAS mode | A firmware setting you chose | Yes - turn it off, and accept what it was for |
| Operating system or licence | The kernel, or the edition you paid for | Yes - a different edition |
Two things to take from the shape of that table before any of the detail. These are not alternatives. They all apply at once, and the usable capacity is the minimum of all eleven rather than the one the vendor chose to print. And the four that move are worth knowing precisely, because each is a purchase rather than a fact about the machine: a processor, a second processor, a module type, and a firmware setting you can change for nothing.
The addressable maximum, which never binds now and once was everything
Start at the bottom, with the ceiling people reach for first and the one that has not mattered on server hardware for fifteen years.
A processor can only address as much physical memory as it has physical address bits. That number is architectural, the chip reports it, and no board, firmware or module changes it. On Linux it is one line:
$ lscpu | grep -E "Model name|Socket|NUMA node|Address sizes"
Model name: Intel(R) Xeon(R) Gold 6130 CPU @ 2.10GHz
Socket(s): 2
NUMA node(s): 2
Address sizes: 46 bits physical, 48 bits virtual
The arithmetic is a power of two, and it is worth tabulating once because the exponents are unintuitive:
2^32 = 4,294,967,296 bytes = 4 GiB
2^36 = 68,719,476,736 bytes = 64 GiB
2^40 = 1,099,511,627,776 bytes = 1 TiB
2^44 = 17,592,186,044,416 bytes = 16 TiB
2^46 = 70,368,744,177,664 bytes = 64 TiB
2^48 = 281,474,976,710,656 bytes = 256 TiB
2^52 = 4,503,599,627,370,496 bytes = 4 PiB
Forty-six bits is 64 tebibytes, which is more than forty times the largest DRAM configuration the example machine can reach. Newer parts widened it rather than narrowing it: Intel moved to 52 physical address bits alongside five-level paging, and AMD did the same on its fourth-generation EPYC. Four pebibytes is a number with no product behind it. The address width is a limit in the sense that the speed of light is a speed limit.
It was not always decorative, and the fossil is still in circulation as advice. The 32-bit x86 architecture addressed 4 GiB; Physical Address Extension widened the physical address to 36 bits and 64 GiB while leaving each process inside 4 GiB; and a generation of chipsets implemented a 36-bit or 40-bit decode regardless of what the processor could do. That is the origin of every “the chipset caps the memory” statement you will still read, and on hardware from 2009 onward it is wrong twice over. The memory controller has lived on the processor since AMD’s K8 in 2003 and Intel’s Nehalem in 2008, so the platform controller hub has nothing to do with capacity, and the addressable space stopped being the constraint at the same time.
What is still live from that era is the software half, and it has its own section below, because a 32-bit operating system on a machine with 128 GB installed is something people still buy into and it does not announce itself.
The per-socket maximum is a property of the chip, and sometimes of the price list
This is the first ceiling that genuinely binds, and the one people are most surprised by, because the socket, the board and the manual do not change when it does.
The memory controller is on the processor die. The capacity ceiling therefore moves when the processor moves, in the same chassis, with the same slots and the same documentation. A Dell PowerEdge R730 reaches 3,072 GB with E5-2600 v4 processors and 1,536 GB with v3, and nothing differs except what is under the heatsinks. Supermicro’s X11DPL-i prints both tiers in one table: up to 2 TB of 3DS RDIMM or LRDIMM at DDR4-2933 with 82xx and 62xx processors, and up to 1 TB at DDR4-2666 with 81xx, 61xx, 51xx, 41xx and 31xx. A Skylake-SP board and a Cascade Lake board are the same board.
Intel went further on Xeon Scalable and sold the ceiling as a product. HPE’s DL580 Gen10 documentation notes that “Intel memory processors (with suffix M or suffix L) are needed for supporting more than 1 TB memory per socket” on second-generation parts, with L parts reaching 4.5 TB. Two chips with the same core count, the same clock and the same thermal envelope, differing in one fused capability and a great deal of money. On the used market those suffixes are worth searching for and are frequently absent from the listing title.
Run it on the example machine and the size of the effect is plain. The Gold 6130 is a standard first-generation Scalable part, rated by Intel at 768 GB per socket:
board arithmetic 24 slots × 128 GB LRDIMM = 3,072 GB
processor ceiling 2 sockets × 768 GB per socket = 1,536 GB
usable ceiling = min(3,072, 1,536) = 1,536 GB
The modules would allow three terabytes and the processors allow half of that. Fit a pair of M-suffix second-generation parts and the same board, the same slots and the same modules reach 3,072 GB, because the term that was binding moved. That comparison is mine; the 768 GB per socket is Intel’s figure for the SKU, and 128 GB is the largest DDR4 load-reduced module in general circulation.
The same effect runs both ways across a product’s life. A PowerEdge R660 reaches 8 TB on a fourth-generation Xeon Scalable and stops at 4 TB on a fifth, because the octal-rank 256 GB module Dell lists against the 4800 MT/s parts has no row against the 5600 MT/s ones. The R660xs does the opposite: 1,024 GB on fourth-generation parts and 1,536 GB on fifth, because the 96 GB module Dell lists at 5600 MT/s does not exist at 4800. Same slots, same channels, same chassis. The ceiling moved because the parts list moved, and Dell states neither conditionality next to the headline figure.
AMD’s ceilings behave the same way and are stated more cleanly. AMD publishes 4 TB per socket for EPYC 7002 and 7003, and separately describes a two-socket EPYC 7001 server reaching 4 TB in total, which is 2 TB per socket. Naples is half of Rome on the same SP3 socket and the same boards. On SP5, AMD publishes 6 TB per socket as “256GB DIMMs * 2 DPC * 12 channels” and also publishes 9 TB with 384 GB modules that are not yet something you can put in a basket. This site records the purchasable figure, for the reason set out on the methodology page: a ceiling that needs a part nobody sells is not a ceiling a buyer can use.
One trap specific to reading the processor column of a modern technical guide. Dell’s R660 guide shows a memory capacity against every processor SKU, and the Xeon Max parts - the 9480, 9470, 9460 and 9462 - read 64 GB in that column where every other SKU reads 4 TB. That 64 GB is the on-package high-bandwidth memory, not a DIMM ceiling. A column that mixes two quantities is a column worth reading the footnote of.
Every processor family page on this site states the generations, form factor, buffering, ECC rule and speed ceiling it was researched against, and links the machines built on it. The processor index is the fastest way to find which of these numbers applies to the chip you actually own, and the Xeon E5-2600 v4 and EPYC 7002 pages are the two that come up most often on second-hand hardware.
Channels, sub-channels, and what a vendor means by “channel”
The per-socket ceiling is divided into channels, and the word does not mean the same thing in every document, which matters once you start multiplying.
A channel is one independent path from the memory controller to a set of modules. The count per socket is fixed by the processor: two on a desktop part, four on the older single-socket server lines, six on first- and second-generation Xeon Scalable, eight on EPYC 7001 through 7003 and on fourth-generation Xeon Scalable, twelve on EPYC 9004 and 9005. That count sets the bandwidth and, with the slots per channel, the slot count.
The bandwidth arithmetic is one multiplication, and it is worth doing because it confirms the channel count against a figure the vendor publishes separately:
per channel MT/s × 8 bytes per transfer
per socket per channel × channels
DDR4-3200, 8 channels 3200 × 8 = 25,600 MB/s × 8 = 204.8 GB/s
DDR4-2667, 4 channels 2667 × 8 = 21,336 MB/s × 4 = 85.3 GB/s
DDR5-4800, 12 channels 4800 × 8 = 38,400 MB/s × 12 = 460.8 GB/s
The middle row is the interesting one. AMD footnotes four EPYC 7002 parts - the 7282, 7272, 7252 and 7232P - as “performance optimized for 4 channels with DDR4-2667 DIMMS” and gives them 85.3 GB/s rather than 204.8 GB/s, while leaving their maximum data rate at 3200. The published 85.3 GB/s decodes exactly as four channels of DDR4-2667, which is what the footnote means without AMD having to spell it out: those SKUs are tuned around half the populated channels of the rest of the family, and filling all sixteen slots per socket on one of them does not buy the bandwidth the family headline promises. The multiplication is mine; both endpoints are AMD’s.
DDR5 complicates the vocabulary rather than the arithmetic. JEDEC’s DDR5 standard, JESD79-5, splits a module into two independent 32-bit sub-channels, each with its own command and address bus, and 40 bits wide once registered ECC is included. Intel counts eight channels on a fourth-generation Xeon Scalable, each of which is one DIMM slot group carrying two sub-channels. AMD counts twelve on SP5 the same way. A sub-channel is a different way of counting the same interface, not extra capacity, and a document saying “sixteen channels” about an eight-DIMM-per-socket machine is counting sub-channels. What that split costs and buys in rank terms is worked through in ranks and 3DS modules.
One consequence catches people moving from DDR4 to DDR5. Because each sub-channel carries its own ECC lanes, a DDR5 ECC module is 80 bits wide rather than 72, and that holds for the unbuffered parts as well as the registered ones, and JESD79-5 also requires on-die ECC inside every DDR5 device. The on-die correction is internal to the die, invisible to the memory controller, and not subtracted from the advertised density: a 32 GB DDR5 module is 32 GB of addressable capacity, and the on-die parity is extra silicon the vendor pays for. It is not system ECC and does not replace it, which is the distinction ECC vs non-ECC exists to make.
Slots: the ones on the board, the ones on a card, and the ones the empty socket owns
The capacity identity is four terms:
maximum capacity = sockets × channels per socket × slots per channel × module capacity
Every headline figure in this article is that product evaluated at its best value in every term. Evaluate it at the values of the machine in front of you and you have the real number, subject to everything else here. Worked on the example machine:
2 sockets × 6 channels × 2 slots per channel = 24 slots
24 × 32 GB, if it were fully populated = 768 GB
12 × 32 GB, as actually fitted = 384 GB at one module per channel
Three of those four terms are where the surprises live.
Slots that belong to a processor that is not there. Intel states the general rule for its own boards: “The memory slots associated with a given processor are unavailable if the corresponding processor socket is not populated.” That is wiring rather than policy, because those slots have no controller at the other end. HPE footnotes the current DL365 Gen11 with “24 DIMM slots require selection of 2 processors”, and its DL160 Gen9 QuickSpecs states that “if only one processor is installed in a two processor system, only half of the DIMM slots are available”, which takes that machine’s stated 1,024 GB down to 512 GB with one chip fitted. A PowerEdge R610 splits its twelve sockets into “two sets of six sockets, one set per each processor”. An ML350p Gen8 has 24 slots and twelve of them are dead in a single-processor build.
On the second-hand market this is the commonest reason a machine arrives with half the memory the seller described, and it is also the cheapest of the eleven ceilings to lift: a low-bin second processor for a decade-old platform costs less than four modules. That trade is worked out at the end of this article.
Slots that arrive on a card you have to order separately. The entry model of the 48-slot DL580 Gen10 ships two of its four processors, and HPE’s note reads “24 DIMM slots available with Entry Model; 2 more processor slots and 24 more DIMMs available via optional HPE DL5x0 Gen10 CPU Version 2 Mezzanine Board Kit (P07991-B21)”. Half the advertised slots are a part number. The DL580 G5 is the same story a generation earlier and more extreme: every one ships with 16 DIMM sockets on the system board, and the memory expansion boards option (452179-B21) adds four cartridges of four sockets each to reach the 32 that go with the headline. HP prints both figures, and the one people quote is the larger. Its own memory configuration table reads “Maximum (Without Memory Expansion Boards Option) 128 GB” against a headline of 256 GB. A quarter of a terabyte and an eighth of a terabyte are the same machine with a different bill of materials.
Slots per channel, which is a board decision rather than a processor one. Two is the common number on current servers and three was common on DDR3-era platforms. Intel’s own Server Board S2600CP routes two slots per channel, which is why its documented per-module ceilings multiply out below the socket’s stated maximum; the 768 GB figure for that generation assumes a three-slot board. AMD publishes the SP3 rule as eight channels at up to two DIMMs each, “a maximum of sixteen DIMMs per socket”, and then notes that SP3 boards ship with 8, 12 or 16 slots. The socket permits sixteen. The board in front of you has whatever it has, and a Supermicro X11DPL-i with eight slots reaches its 2 TB ceiling only because 256 GB modules exist, where a sixteen-slot board of the same generation gets there with 128 GB parts at a fraction of the price per gigabyte.
That last point matters more than any other term in the identity for a used buyer, because the slot count is what decides how small the modules are allowed to be, and small modules are where the used market is cheap.
The chassis can rule a module type out entirely
Two ceilings in the table above are enforced by sheet metal and airflow rather than by silicon, and neither appears in the memory section of any datasheet.
A PowerEdge R720xd with 3.5-inch drives does not support load-reduced memory at all, at any capacity, for thermal reasons. The same machine with 2.5-inch drives does. Nothing about the DIMM slots differs; the drive cage changes the air path across them, and a load-reduced module with a buffer on every data line dissipates more than a registered one. That single line in a Dell manual is the difference between a 384 GB machine and a 768 GB one.
The DL380 Gen11 does it with slots rather than module types. HPE’s footnote reads “32 DIMMs only with 8SFF or 16SFF, 16 DIMMs maximum with 24SFF”, so ordering the dense drive cage halves the slot count and therefore the ceiling, from 8 TB to 4 TB, before you have chosen a module at all. Less obviously, it also keeps the machine at one module per channel across its sixteen surviving slots, which is full 1DPC speed. The dense chassis costs capacity and buys clock; the dense population costs clock and buys capacity. They are the same trade written twice, and a used machine rarely lets you choose the drive cage.
The firmware table is not the silicon, and a QVL maximum is usually a firmware maximum
This is the ceiling people argue with hardest and understand least, and it is worth separating into three things that get conflated.
The memory reference code has to know the die. Before any of this is about capacity it is about training: at every power-on the firmware reads each module’s serial presence detect, looks up a timing and voltage recipe for that DRAM density and organisation, and trains the channel. JEDEC standardises what is in the SPD - JESD21-C for the older generations, and a DDR5-specific SPD document for the current one - so the firmware can always read the module. What JEDEC does not do is oblige a firmware to have a recipe for every die that will ever exist. If the reference code has no entry for the die on the module, the outcome is not a graceful downgrade. The machine posts at the wrong capacity, posts at a fraction of the speed, or does not post.
The clearest published case is at the small end. Intel documents DRAM technologies of 1 Gb, 2 Gb and 4 Gb for LGA1155, and nothing else. A 16 GB DDR3 unbuffered module is built on 8 Gb dies, which appear nowhere in any of the four datasheets for that socket, so it is outside Intel’s specification even though it seats, even though the socket’s stated 32 GB total has room for it, and even though plenty of them work on later Ivy Bridge boards. The pattern repeats at every generation boundary. DDR4 16 Gb dies needed firmware updates on boards that had shipped assuming 8 Gb. The 24 Gb DDR5 dies behind the 48 GB and 96 GB modules need a thirteenth- or fourteenth-generation Core processor on LGA1700, where a twelfth-generation part caps at 32 GB modules and 128 GB in total, on the same socket and frequently on the same board.
So a qualified vendor list is a snapshot of the reference code, not a statement about the memory controller. It is revised upward whenever a bigger part is validated, which is why an old PDF understates a machine and a current one may quote a module that was never sold in volume. It is also why a firmware update sometimes raises a ceiling and sometimes does not: if the constraint was a missing die entry, an update can fix it; if the constraint was routing or rank loading, no firmware will.
The validated-configuration limit is a document, not a mechanism. A PowerEdge T640 has 24 slots and takes a 64 GB registered module, which multiplies out to 1,536 GB, and Dell states 768 GB. Nothing stops the arithmetic working; the larger configuration simply was not qualified. This is the one to argue with least, because a configuration outside the test matrix is one the firmware team never trained against, and the failure mode of untrained memory is silent corruption rather than a refusal to boot.
Vendors also contradict themselves inside one document, which is the best evidence available that these tables are maintained by hand. Dell’s R660 technical guide lists 256 GB modules in its supported memory matrix and stops at 128 GB in its supported DIMMs table, which is the difference between an 8 TB machine and a 4 TB one depending on which table you happen to open. That is Dell’s inconsistency rather than a reading error, and it is recorded as an open question against that machine on this site rather than resolved by picking the flattering number.
What the firmware thinks, read off the running machine
The firmware’s own opinion is readable, and it is the fastest way to find out what your particular BIOS believes about your particular board.
$ sudo dmidecode -t 16
Physical Memory Array
Location: System Board Or Motherboard
Use: System Memory
Error Correction Type: Multi-bit ECC
Maximum Capacity: 3 TB
Number Of Devices: 24
Number Of Devices is the honest slot count and is worth trusting, because it
comes from the board’s own device table rather than from a marketing document,
and it includes the slots that belong to an empty socket. Cross-check it
against what is actually fitted:
$ sudo dmidecode -t 17 | grep -c "^Memory Device"
24
$ sudo dmidecode -t 17 | grep -c "No Module Installed"
12
Maximum Capacity is the firmware’s opinion and nothing more. On the example
machine it reads 3 TB, which is 24 slots multiplied by the largest module the
firmware has an entry for, and the processors fitted cap the machine at half of
that. Treat it as the board term of the identity, never as the answer.
There is also a reporting artefact in that field worth knowing before you misread one. The SMBIOS Physical Memory Array structure stores Maximum Capacity as a 32-bit value in kilobytes, so the largest capacity it can express is:
2^31 - 1 kilobytes = 2,147,483,647 KB
2^31 kilobytes = 2,147,483,648 × 1024 = 2,199,023,255,552 bytes = 2 TiB
When the real capacity exceeds that, the specification requires the field to carry the sentinel value 0x80000000 and the true figure to appear in a separate 64-bit Extended Maximum Capacity field, added in SMBIOS 2.7. A machine reporting exactly 2 TB there may be reporting a sentinel rather than a capacity, and a tool that does not read the extended field will print 2 TB for an 8 TB server. The arithmetic is mine; the field widths are the specification’s.
The SPD is how the machine finds out, and it is a very small file
Everything the firmware knows about a module before it has trained the channel comes from one serial EEPROM on the module itself, read over a slow side-channel bus while the DRAM is still dark. JEDEC standardises its contents, in JESD21-C for the earlier generations and in a DDR5-specific serial presence detect document for the current one, which is what lets any firmware read any compliant module.
What is in it decides most of this article. The density and organisation, so the controller knows how many ranks and how wide the devices are. The module type, so it knows whether there is a register in the address path. The JEDEC speed grades and the timings for each. The manufacturer, part number and serial number. On DDR5, the module also carries its own power management circuit and a hub device in front of the presence-detect memory, which is a change worth knowing because it moves voltage regulation from the board onto the module.
On Linux the module’s own copy is readable where the bus is exposed, which on servers it frequently is not:
$ sudo modprobe eeprom
$ sudo decode-dimms | grep -E "Part Number|Module Type|Size|Ranks|Speed"
Failing that, dmidecode -t 17 is a second-hand view of the same data, because
the firmware builds the SMBIOS memory device structures from what it read out
of the SPD. That is worth knowing in both directions. It means dmidecode is
trustworthy about the module’s identity, and it means dmidecode is reporting
what the module claims rather than what the machine is doing with it, which is
why the Configured Memory Speed field exists separately from Speed.
Three consequences for capacity.
A module whose presence detect cannot be read is not slow, it is absent. The firmware has no recipe to train it with and no way to size it, so it maps the slot out. A machine that counts eleven modules out of twelve usually has a seating or presence-detect problem rather than a DRAM problem, and swapping the module into a different slot separates the two in one reboot.
Overclocking profiles are extra SPD entries and server firmware ignores
them. A desktop kit sold as DDR4-3600 carries JEDEC base grades in the
standard part of its presence detect and the 3600 grade in an XMP profile above
it, EXPO being the equivalent AMD arrangement on DDR5. A server that reads only
the JEDEC section trains the kit at its base grade, commonly 2133 or 2400, and
reports exactly that in Configured Memory Speed. Nothing is broken and
nothing will fix it. The
mechanism and the unit confusion around it are covered in
RAM speed, MHz and MT/s.
The rank field reports package ranks, not logical ones. Firmware fills the SMBIOS rank value from the presence detect’s package rank byte, so a 3DS module reports 2 where its own label says 4Rx4 or 8Rx4. That is the correct number for the chip-select budget and the wrong number for the capacity tables, and it is the single most common reason two people reading the same machine disagree about how many ranks are in it.
The rank budget, not the slot count
A memory channel does not have a module limit. It has a rank limit, because each rank needs its own chip-select line and the controller has a fixed number of them. Intel’s population rules for its own server boards state it as a hard number: “A maximum of 8 logical ranks can be used on any one channel, as well as a maximum of 10 physical ranks loaded on a channel.” Its current desktop controller states the same kind of budget smaller: the Core Ultra 200S support matrix gives a maximum of two ranks per channel at one DIMM per channel and four at two.
Eight is reached long before the slots run out:
server DDR4 channel, two slots - 8 selectable ranks in total
2 × 64GB 4Rx4 LRDIMM 4 + 4 = 8 ranks budget full, 128GB
2 × 64GB 2Rx4 RDIMM 2 + 2 = 4 ranks budget half, 128GB
2 × 32GB 2Rx4 RDIMM 2 + 2 = 4 ranks budget half, 64GB
2 × 256GB 3DS LRDIMM 2 + 2 = 4 ranks budget half, 512GB
(eight ranks each, four dies deep behind
each of two chip selects)
The first two rows hold identical capacity and load the channel completely differently. The fourth holds four times as much on the same budget, because a 3DS stack hangs several dies behind one chip select and selects between them with chip-ID pins that the controller’s budget never sees. That mechanism is worked through in ranks and 3DS modules; what matters here is the consequence. The largest module on any vendor’s list is nearly always a stacked part, and stacked parts carry population rules of their own.
| Largest module of its class | Construction | What it commits you to |
|---|---|---|
| 64GB DDR4 RDIMM | 2Rx4, monolithic 16Gb |
Fewest restrictions of anything here |
| 64GB DDR4 LRDIMM | 4Rx4, non-3DS |
Whole machine must be LRDIMM |
| 128GB DDR4 LRDIMM | 4Rx4 non-3DS, or 2-high 3DS |
LRDIMM, and the two builds are not interchangeable |
| 256GB DDR4 LRDIMM | 8Rx4, 4-high 3DS |
LRDIMM and 3DS throughout |
| 96GB DDR5 RDIMM | 2Rx4 on 24Gb dies |
A processor generation whose tables list 24Gb dies |
| 256GB DDR5 RDIMM | 8Rx4, 4-high 3DS |
3DS throughout; DDR5 server tables list no LRDIMM at all |
| 512GB DDR5 RDIMM | 2S8Rx4, 8-high 3DS |
Listed by AMD for EPYC 9004 as pending ecosystem enablement |
Those commitments are machine-wide rather than per-channel, and they move the ceiling a long way. An HPE ProLiant DL380 Gen9 takes 768 GB of registered memory and 3,072 GB of load-reduced out of the same 24 slots, a four-fold difference decided once for the whole server because the two types cannot be mixed. A PowerEdge R440 states “up to 512 GB RDIMM and 1 TB LRDIMM” out of sixteen slots. A Supermicro X10DRi-T states 512 GB with registered memory and 2,048 GB with 128 GB 3DS load-reduced parts out of the same sixteen. Every one of those pairs is the same board, the same channels and the same firmware, and the ratio between them is the module type.
Supermicro’s X11 rules extend the prohibition to stacking, forbidding non-3DS and 3DS LRDIMMs together “in the same channel, across different channels, and across different sockets”, and AMD’s EPYC 9004 guide says “Do not mix 3DS and non-3DS memory modules in a 2DPC system.” Read those as what they are. The decision is taken once, for the whole machine, on the day you buy the first module.
The physical rank limit is a second number, and it is the one that bites on old platforms
Intel’s rule quotes two budgets, eight logical and ten physical, and the gap between them is exactly the 3DS mechanism. A stacked module presents fewer chip selects than it has logical ranks and fewer electrical loads than it has dies, because the master die buffers the slaves. The physical budget counts loads on the bus. The logical budget counts addressable ranks.
On DDR3-era platforms the physical number is the one that stops you, and it is why quad-rank registered modules were so restricted. A PowerEdge R620 takes “up to two quad-rank RDIMMs per channel” and Dell’s manual prints the price: those modules run 1333 MT/s at one per channel and 1066 at two, and the third slot of a three-slot channel is closed to them entirely, with only single- and dual-rank parts able to fill it and the whole system capped at 1333 if one does. A quad-rank module can spend a whole desktop channel’s budget by itself, which is why a board refuses a second one rather than derating it.
What the second module per channel costs
Filling slot two doubles the ranks on the channel and the controller answers by slowing down. How much depends on the generation and, the part most people miss, on the processor’s market segment. Cisco publishes the clearest table for current Intel parts:
| CPU tier | 4th Gen 1DPC | 4th Gen 2DPC | 5th Gen 1DPC | 5th Gen 2DPC |
|---|---|---|---|---|
| Platinum 8 | 4800 | 4400 | 5600 | 4400 |
| Gold 6 | 4800 | 4400 | 5200 | 4400 |
| Gold 5 | 4400 | 4400 | 4800 | 4400 |
| Silver 4 | 4000 | 4000 | 4400 | 4400 |
| Bronze 3 | 4000 | 4000 | 4400 | 4400 |
Supermicro’s X13 guide gives the same 4800/4400 and 5600/4400 ceilings, and HPE states its ML350 Gen11 maximum as “32 x 256 GB RDIMM @ 4400MT/s at 2 DPC” with both processor generations. At two modules per channel no current Xeon exceeds 4400 MT/s, and Silver and Bronze parts sit at 4000 either way, so DDR5-5600 modules in a fully populated machine run about a fifth below their grade:
(5600 - 4400) ÷ 5600 = 0.214 = 21.4 per cent off the data rate
Convert that into what it costs an eight-channel socket and the size is clearer, because bandwidth scales directly with the data rate:
8 channels × 5600 MT/s × 8 B = 358.4 GB/s 1DPC, half the slots empty
8 channels × 4400 MT/s × 8 B = 281.6 GB/s 2DPC, every slot full
Doubling the capacity costs 21 per cent of the bandwidth, which is a trade most people would take and almost nobody is told they are making. The derivation is mine; both data rates are Cisco’s and Supermicro’s.
DDR4 is messier, and the vendors disagree. HPE’s DL580 Gen10 table derates registered memory from 2933 at 1DPC to 2666 at 2DPC while leaving load-reduced memory at 2933 at 2DPC, because the buffering buys back what the extra ranks cost. Supermicro’s X11 guide, covering the same processors, lists 2933 at both 1DPC and 2DPC for every module type including quad-rank RDIMMs. Assume the pessimistic figure when shopping and trust the table for the machine actually in front of you.
AMD’s DDR4 platforms derate consistently:
| Processor | 1DPC | 2DPC | Stated by |
|---|---|---|---|
| EPYC 7001 (Naples) | 2666 | 2133 | AMD |
| EPYC 7002 (Rome) | 2933 | 2666 | HPE, DL385 Gen10 |
| EPYC 7003 (Milan) | 3200 | 2933 | HPE, DL385 Gen10 Plus v2 |
The Naples row is the one to notice: filling every slot on a first-generation EPYC machine costs a fifth of the memory clock. It is also a row this site records as uncertain rather than confirmed, because AMD’s Naples-era memory population guideline is no longer reachable on amd.com and the 2DPC figure rests on a secondary source. On SP5 there is no published figure at all. AMD’s population guide tabulates the first DIMM per channel only, and no AMD document read for this site states the data rate for a fully populated twenty-four-DIMM socket. That is a gap in the vendor documentation rather than a gap here, and it is a good reason to read the configured speed off a running machine rather than trusting any table.
The board can also be the binding constraint on speed rather than the processor. An HPE ProLiant DL385 Gen10 with EPYC 7002 processors tops out at DDR4-2933 in HPE’s tables. The DL385 Gen10 Plus takes the same processors at 3200. Same chip, same vendor, different board.
Module type is a whole-machine decision, and DDR5 took one option away
Four module types matter for capacity. What separates them mechanically is the subject of RDIMM, UDIMM and LRDIMM; what matters here is what each one does to the ceiling.
Unbuffered modules cap lowest and cap hardest. The address and command lines of every module hang directly off the controller, so the load rises with every module and every rank, and the practical limit is small. A PowerEdge R610 supports “up to 192 GB of RDIMM memory (twelve 16 GB RDIMMs)” and, in the same document, “up to 24 GB of UDIMM memory (twelve 2 GB UDIMMs)”. Eight to one, in one machine, with one table separating them. An ML350p Gen8 does it with slot count as well as capacity: 768 GB load-reduced across 24 slots, 384 GB registered across 24, and 128 GB unbuffered across 16. The unbuffered row does not even get all the slots.
Registered modules buffer address and command only, which puts one load per module on those lines instead of one per DRAM. That is what makes eight or sixteen slots per socket possible at all, and it is what nearly all second-hand server memory is. Browse registered ECC DDR4 or registered ECC DDR3 and you are looking at the bulk of the market.
Load-reduced modules buffer the data lines too, which is what lets a quad-rank module present a single load, and it is why they reach the capacities they do. The premium is real and so is the return: the DL580 Gen10 worked below is a case where load-reduced memory doubles the capacity and keeps the clock at the same time.
3DS parts stack dies behind one chip select, which is how DDR5 servers reach their top capacities without load-reduced modules at all. That is the change worth stating plainly. AMD’s EPYC 9004 guide lists no LRDIMM and no UDIMM on SP5, only RDIMM and 3DS RDIMM, with UDIMM, LRDIMM, NVDIMM-N and NVDIMM-P all named as unsupported, and the DDR5 server tables from the other vendors read the same way. If your model of how a server reaches four terabytes is load-reduced memory, that model retired with DDR4.
MRDIMM is the fifth type and is a speed feature rather than a capacity one: it multiplexes two ranks onto the bus to raise the effective data rate. It is new enough that this site’s DDR5 listings hold very few, which is why no facet page for the type is linked here yet. A facet with nothing behind it is a link that breaks, and the rule this site follows is to link the generation page until the inventory earns a page of its own.
Population rules, interleaving, and the configuration that costs two thirds of the bandwidth
This is the failure that shows up as no smaller number anywhere. The capacity is right, the modules are right, the machine boots, and it runs at a third of the speed it should.
Lenovo measured STREAM Triad across every population of a two-socket EPYC 7002/7003 machine, all at the same clock, and reported each as a fraction of the fully populated figure:
| DIMMs per socket | Pattern | Interleave sets | Bandwidth |
|---|---|---|---|
| 1 | one channel | 1 | 14% |
| 2 | one channel pair | 1 | 28% |
| 4 | two channel pairs | 1 | 54% |
| 6 | six channels (7003 only) | 1 | 71% |
| 8 | all eight channels, 1DPC | 1 | 100% |
| 12 | six channels at 2DPC (7003 only) | 1 | 71% |
| 12 | four channels at 2DPC, four at 1DPC | 2 | 35% |
| 16 | all eight channels, 2DPC | 1 | 100% |
Read the two twelve-module rows together. Going from eight modules to twelve can cost two thirds of the machine’s memory bandwidth, because twelve cannot form one interleave set: the controller builds two four-channel sets instead of one eight-channel one. Eight and twelve are not “enough memory” and “more memory”. They are a good configuration and a broken one, and the broken one is the one you land on by adding four modules to a machine that already worked.
Only 1, 2, 4, 6, 8, 12 and 16 modules per socket are sensible on that platform, and only 8 and 16 reach full bandwidth with identical modules. Genoa widened the rule rather than dropping it: EPYC 9004 forms interleave sets of 2, 4, 6, 8, 10 or 12 channels, and to form one “all channels are required to have the same DIMM type, the same total memory capacity and ranks.”
Mixing two capacities is cheaper than it looks, provided the balance holds. Lenovo measured about 3 per cent bandwidth loss for a near-balanced configuration, meaning two capacities, every channel carrying the same total capacity, and an even total rank count. Eight 64 GB plus eight 32 GB modules per socket reaches 768 GB for almost nothing against sixteen identical ones, and that is the configuration a used buyer most often arrives at by accident rather than by design.
The asymmetry to keep in mind is that balance is a per-channel property, not a per-socket one. Two capacities distributed so that every channel carries the same total is nearly free. The same two capacities distributed so that four channels carry 96 GB and four carry 64 GB is the 35 per cent row.
Reading a vendor population table
They all look alike, and the important part is never the grid. Supermicro’s table for a six-slot X11UP board:
1 DIMM DIMMA1
2 DIMMs DIMMA1 / DIMMD1
3 DIMMs DIMMC1 / DIMMB1 / DIMMA1
4 DIMMs DIMMB1 / DIMMA1 / DIMMD1 / DIMME1
5 DIMMs DIMMC1 / DIMMB1 / DIMMA1 / DIMMD1 / DIMME1
(unbalanced: not recommended)
6 DIMMs DIMMC1 / DIMMB1 / DIMMA1 / DIMMD1 / DIMME1 / DIMMF1
Four things to take from it. A row is a set of slots, not the first n slots, and four modules go in A1, B1, D1 and E1, skipping C1, so that each of the two memory controllers gets two channels rather than one getting three. A row can be legal and still be marked not recommended, which is the vendor saying that firmware will boot it and the bandwidth will be poor; that is where odd module counts live. Slot colour encodes order rather than capability, and white or blue is first, and it is the slot further from the processor. And the footnotes carry the rules the grid cannot express: which processor generations a row applies to, which module types it excludes, whether a capacity is one-per-channel only.
Two more things such a table will not tell you. First, whether a reliability
mode is quietly eating capacity, which is the next section. Second, what the
firmware actually detected, which is dmidecode -t 17 and the
commands in the compatibility guide.
Mixing rank counts
Where mixing is allowed at all, one rule is universal and the rest is a table.
The universal rule: the heavier module goes in the slot furthest from the processor, the white or blue one that is populated first. Intel’s guidance is to put the higher electrical load in the first slot, with quad-rank modules farthest from the processor. Cisco puts it in one line, “Higher rank DIMMs shall be populated on Slot 1.” Reversed, the channel trains slower or not at all, and the failure is a POST code rather than an error message.
The table is per-vendor and unforgiving. Cisco’s M7 matrix permits a 64 GB 2Rx4 in slot 1 with a 32 GB 1Rx4 in slot 2, and a 256 GB 8Rx4 with a 128 GB 4Rx4, and close to nothing else. Its 128 GB 2Rx4 on 32 Gb dies “cannot be mixed with any other memory DIMMs”. Its 48 GB and 96 GB modules on 24 Gb dies cannot be mixed with any other capacity, and 48 GB is one module per channel outright. The footnote that catches people: “When mixing two different DIMM densities, all 8 channels per CPU must be populated.”
And the rule is not the same rule at every vendor, which is the reason to read yours rather than the one you remember. AMD’s EPYC 9004 guide says “Do not mix x4 and x8 DIMMs within a memory channel.” Supermicro’s X11 guide, on Intel, says “x4 and x8 DIMMs can be mixed in the same channel.” Both are correct about their own platforms, and a forum answer quoting one at somebody running the other is the commonest way this goes wrong.
Capacity on the wrong socket is not the same capacity
On a machine with more than one processor, memory is attached per socket, and the total is a sum of parts that are not equivalent to each other. Cores on socket 0 reach socket 0’s memory directly and socket 1’s across the interconnect, at higher latency and lower bandwidth. The operating system is told about the split, and reports it:
$ numactl -H
available: 2 nodes (0-1)
node 0 cpus: 0-15 32-47
node 0 size: 193017 MB
node 1 cpus: 16-31 48-63
node 1 size: 193512 MB
node distances:
node 0 1
0: 10 21
1: 21 10
The distance matrix is a ratio rather than a measurement: 10 is the normalised local cost and 21 is the remote one, so remote memory on this machine costs roughly twice what local memory costs. Those figures come from the firmware’s ACPI tables, and they describe the topology rather than a benchmark of it.
The capacity consequence is a rule about module counts. Buy in multiples of the channel count times the socket count, because anything else leaves one node smaller than the other, and a process scheduled on the larger node’s cores that outgrows its local memory starts allocating remotely without telling anyone. A 2P machine with all its memory on one socket is the extreme case: half the cores now reach every byte they touch over the interconnect, and the machine performs worse than the same memory in a single-socket box.
Both vendors let you subdivide further. AMD exposes nodes-per-socket settings on EPYC, so one socket can present as one, two or four NUMA nodes depending on how the memory controllers are grouped, and Intel offers sub-NUMA clustering for the same purpose. Neither changes the capacity. Both change how the capacity is reported and how badly an unbalanced population behaves, because a finer partition makes a lopsided one more visible rather than less.
Four-socket machines add one more wrinkle worth knowing before buying one
cheaply. The interconnect is not fully meshed on every platform, so some
socket pairs are two hops apart rather than one, and the distance matrix above
grows entries at 31 or higher. On a four-socket box the difference between
filling all four sockets’ channels and filling two of them is not 2:1 in
bandwidth, it is worse than that for anything scheduled on the starved
sockets, which is a good reason to read the numactl -H output of a machine
before deciding what its memory is worth.
Mirroring, sparing and lockstep spend capacity on reliability
Three firmware settings can take memory away after the machine has counted it, and only one of them is usually described in a way that makes the cost obvious. All three live in the same BIOS menu, normally called memory RAS or memory operating mode, and the default on nearly every machine is independent mode, which costs nothing.
Mirroring halves it. In full memory mirroring every write goes to two channels and the reported capacity is half the installed capacity. A 768 GB machine reports 384 GB. There is no subtlety in the arithmetic and a great deal of value in the outcome, which is that an uncorrectable error on one channel is served from the other rather than halting the machine.
Address range mirroring is the version worth knowing about, because it
changes the trade from all to some. Instead of mirroring the whole of memory,
the firmware mirrors a requested range and leaves the rest unmirrored, so the
cost is the size of the range rather than half the machine. Linux drives it
through the boot parameter kernelcore=mirror, which places kernel allocations
inside the mirrored region and leaves user pages outside it. The kernel is the
part whose corruption takes the machine down, so a few gigabytes of mirror buys
most of the benefit at a fraction of the capacity. That is a reasonable default
for a machine whose job is to stay up, and it is invisible in every capacity
table.
Sparing reserves ranks, and the cost depends on the modules. Rank sparing holds one or two ranks per channel in reserve; when a rank accumulates correctable errors past a threshold, its contents are copied into the spare and it is retired. The reserved rank is not addressable, so the arithmetic is:
usable = installed × (R - S) ÷ R
R = ranks per channel
S = spare ranks per channel
That formula is mine, and it reproduces Supermicro’s own published figures exactly. Its tables show a single 64 GB quad-rank module reporting 48 GB under one-rank sparing and 32 GB under two, and a 16 GB dual-rank module reporting 8 GB:
64 GB, R = 4, S = 1 64 × 3 ÷ 4 = 48 GB matches
64 GB, R = 4, S = 2 64 × 2 ÷ 4 = 32 GB matches
16 GB, R = 2, S = 1 16 × 1 ÷ 2 = 8 GB matches
Three checks out of three, which is what makes the formula worth carrying rather than looking the table up. The cost of sparing is a quarter on a quad-rank channel and a half on a dual-rank one, which is the opposite of the intuition that more ranks means more overhead. And a single-rank module cannot support sparing at all, because retiring its only rank retires the module. Intel’s own board rules put rank sparing behind “at least 2 SR or DR DIMM installed, or at least one QR DIMM installed, on each populated channel”.
Lockstep costs bandwidth, not capacity, and this is the one people get backwards. In lockstep mode two channels are operated as a single wider interface, so a cache line is fetched across both at once and the error correction code spans the pair. Capacity is the sum of the two channels and nothing is reserved. What you lose is half your independent channels, so the achievable bandwidth falls roughly in half, and the population rules get stricter because the paired channels must match. The reason anybody accepted that trade is that a wider code word corrects more: lockstep is how x8 modules were given single-device correction on platforms that could otherwise only do it with x4.
Intel retired lockstep on the Xeon Scalable generation in favour of ADDDC, adaptive double device data correction, which survives a second device failure after the first by remapping at a finer granularity. Lenovo documents it as working “with x4-based memory DIMMs” and requiring “two DIMM ranks per channel”, plus a Gold or Platinum processor. A channel holding one single-rank x8 module can enable none of these modes, which is worth knowing before buying the cheapest modules for a machine whose whole point is that it does not fall over.
Check what is actually enabled before you go looking for missing gigabytes:
$ sudo dmidecode -t 16 | grep -E "Error Correction Type|Maximum Capacity"
Error Correction Type: Multi-bit ECC
Maximum Capacity: 3 TB
The RAS mode itself is not in SMBIOS. It is in the firmware setup screen, and on a machine bought second-hand it is whatever the previous owner left it at. A server that reports exactly half the memory you installed is almost always mirrored rather than broken, and that is a five-minute fix rather than a return.
What a refusal looks like, and how to read one
Memory training happens before there is any display output, so a machine that will not accept a module tells you about it in one of three ways, and only one of them is obvious.
It does not POST at all. The firmware failed to train a channel and stopped. On a server this normally surfaces as a fault light beside a specific DIMM slot, a POST code on a two-digit display, or a beep pattern on older hardware. The slot is the useful part: it names the channel, and moving the module to a slot on a different controller separates a bad module from a bad slot in one reboot. This is also the failure that memory test software cannot help with, because the machine never reached the point where software runs.
It POSTs with the module mapped out. This is the sneaky one. Everything works, the machine is stable, and the capacity is short by exactly one module’s worth. The discriminator is that SMBIOS still describes the module, because the firmware built those structures from the presence detect, while the kernel does not count it:
$ sudo dmidecode -t 17 | grep -c "Size: 32 GB"
12
$ grep MemTotal /proc/meminfo
MemTotal: 362251776 kB
Twelve modules of 32 GB described and eleven modules’ worth counted. A gap of exactly one module is a mapped-out module, not a memory hole, and the firmware normally logged why.
It POSTs at a lower speed than the modules are graded for. Covered above as
the rank derate, but there is a second cause worth separating: a mixed
population where one module’s presence detect offers a lower top grade drags
the whole channel, or the whole machine, down to it. Configured Memory Speed
against Speed in dmidecode -t 17, read per module rather than in aggregate,
shows which module is setting the pace.
The logs are where the reason lives, and on a server they are out-of-band rather than in the operating system:
$ sudo ipmitool sel list | grep -i memory
Correctable error thresholds also retire memory at runtime rather than at boot, which is why a machine can lose a rank’s worth of capacity months after it was last touched. On a used machine, an event log that has already been cleared is itself information about the seller.
The purchase advice that falls out of this is about timing rather than technique. Fit everything on arrival, inside the returns window, and read the count before you read anything else. A module that maps out on the bench in week one is a return; the same module discovered in month four is a loss. The same window argument runs through buying used RAM on eBay, and memory is the one component where it is easy to act on, because there is nothing to wear out and nothing to burn in.
Installed is not usable, and the gap is not a rounding error
Even with every ceiling above satisfied, the number the operating system reports is smaller than the number on the modules, and the difference is structural.
The low four gigabytes of physical address space are not all memory. Memory mapped I/O lives there: PCIe configuration space, device apertures, the local APIC, firmware reserved regions, and on machines with integrated graphics a stolen framebuffer. Every one of those regions is a hole where DRAM cannot be addressed. The firmware’s answer is to remap the displaced DRAM above the four gigabyte line, which is what the BIOS setting variously called memory remap or memory hole remapping does, and it is why the setting exists at all.
The residue is visible on any running machine:
$ grep MemTotal /proc/meminfo
MemTotal: 395806208 kB
$ sudo dmidecode -t 17 | grep -c "Size: 32 GB"
12
Twelve 32 GB modules is 384 GiB, and the arithmetic on the gap is:
384 GiB × 1,048,576 KiB = 402,653,184 KiB installed
395,806,208 KiB reported by the kernel
difference = 6,846,976 KiB = 6.53 GiB = 1.70 per cent
That gap is firmware-reserved regions, the ACPI tables, whatever the kernel set
aside for a crash kernel, and the memory map holes the firmware could not
reclaim. Its size is a property of your firmware and your kernel command line,
not a constant, so the 1.7 per cent above is an illustration and your own
machine’s dmesg e820 map is the authority. What is constant is the direction:
reported memory is always less than installed memory, never more, and a gap of
one or two per cent on a large server is normal rather than a fault.
Where it stops being normal is when the gap is a round number. Missing exactly
half is mirroring. Missing exactly one module’s worth is a module the firmware
failed to train and mapped out, which dmidecode -t 17 will show as installed
while /proc/meminfo does not count. Missing everything above 4 GiB is the next
section.
The 32-bit traps that still bite second-hand buyers
These are the ceilings nobody expects to meet in the 2020s, and they are all still reachable with hardware and software that is on sale today.
A 32-bit operating system cannot use memory above 4 GiB, and usually cannot use all of the first 4 GiB either. The address space below the line is shared with the memory mapped I/O described above, so a 32-bit install on a machine with 8 GB fitted typically reports somewhere between 3.2 and 3.5 GB, with the exact figure set by how much aperture the graphics and storage devices claimed. Physical Address Extension widens the physical address to 36 bits and was in every x86 processor from the Pentium Pro onward, so the hardware has not been the obstacle since 1995.
The obstacle was a licence. 32-bit Windows client editions cap at 4 GB regardless of PAE, while 32-bit Windows Server editions of the same era used PAE to reach 64 GB. That was a product decision rather than an architectural one, and it is the reason the folklore that “32-bit can only see 4 GB” is half right and has outlived every system it described.
The licence ceiling did not disappear with the move to 64-bit. Microsoft publishes per-edition physical memory limits, and the ones that bite a home lab are the client editions: Windows 11 Home stops at 128 GB and Windows 11 Pro at 2 TB. A second-hand two-socket server with 192 GB fitted, running the edition that came with the desktop it replaced, will report 128 GB and no error. The server editions sit in the tens of terabytes, above anything in this article, so the check is worth running only on the client side.
On Linux the equivalent limits are compile-time and generous. A 64-bit kernel
using four-level paging addresses 64 TiB of physical memory, the same 46 bits
the example machine’s processor reports, and five-level paging raises that to
4 PiB. The one that catches people is not a limit at all but a leftover: a mem=
parameter on the kernel command line from some past debugging session truncates
physical memory silently and survives reboots.
$ cat /proc/cmdline
BOOT_IMAGE=/vmlinuz-6.8.0 root=UUID=... ro quiet mem=64G
That trailing mem=64G is the whole bug, on a machine with 384 GB fitted, and
nothing else on the system will mention it.
Two more from the same era, both still live on used hardware. Older firmware without a memory remap or above 4G decoding option cannot relocate the DRAM displaced by the MMIO hole, so that memory is simply lost, and on machines of that vintage the setting is sometimes present and disabled by default. And a 32-bit hypervisor or a 32-bit storage driver inside an otherwise 64-bit stack imposes its own ceiling that the host firmware knows nothing about.
Persistent memory is why some headline capacities are unreachable with DRAM
A modern technical guide may quote a memory maximum that no quantity of DRAM can reach, because the figure includes persistent memory modules in the same slots. The distinction is not signposted, and the gap between the two kinds of number is large enough to be an expensive mistake.
Dell’s own comparison table is the clean example. Against the PowerEdge R650 it prints “Up to 16 Intel Persistent Memory 200 series (BPS) slots, 12 TB max”. The adjacent R660 column has no such line, and its 8 TB is a DRAM-only figure. Two machines, one table, and the larger number is not the same kind of number.
Three facts decide whether a persistent memory figure is available to you. The
modules are generation-locked: the 100 series pairs with second-generation Xeon
Scalable and the 200 series with third, and a module in the wrong machine does
not train. They occupy DIMM slots, so every persistent module is a DRAM slot you
no longer have, and the population rules require conventional DRAM alongside
them in a defined ratio rather than allowing a machine built entirely of
persistent memory. And Intel discontinued the product line, which means the
supply is entirely second-hand and the tooling is a niche: ipmctl reads and
provisions them, and a module that a previous owner locked with a passphrase
arrives unusable until it is either unlocked or securely erased.
Where this article’s arithmetic is concerned, the rule is short. If a machine’s headline capacity is more than slots multiplied by the largest DRAM module in its own supported list, the difference is persistent memory and it is not a DRAM ceiling. Check the machine’s own supported-DIMM table and multiply; the gap will be obvious.
The modern version of the same idea is CXL memory expansion, where capacity hangs off a PCIe-attached device rather than a DIMM slot, and it is supported on fourth- and fifth-generation Xeon Scalable and on EPYC 9004. It appears to the operating system as an additional NUMA node with higher latency than local DRAM. It is genuinely another route past the slot count, and it is not a route a second-hand buyer will be taking, because the devices are neither cheap nor plentiful. It is here so that a capacity figure quoted with CXL in the footnotes is recognisable as one.
Working the real ceiling out of three documents and the running machine
Every section above is one term. Here is the order to evaluate them in, which is chosen so that the cheapest document to read eliminates the most wrong answers.
Document one: the processor’s own specification. Intel’s ARK entry or the family datasheet, AMD’s product page or datasheet. It answers the per-socket maximum, the channel count, the maximum data rate, and the DRAM die technologies the controller supports. Multiply the per-socket figure by the sockets actually fitted, not by the sockets present. If the chip is an Intel Xeon Scalable part, check for an M or L suffix before believing anything above a terabyte per socket. This site’s processor index carries the same fields with the document each came from.
Document two: the system or board document. QuickSpecs for HPE, the technical guide for Dell, the motherboard manual for Supermicro, the product guide for Lenovo. It answers the slot count, the per-module-type ceilings, the chassis conditions, and the population order. Read the footnotes rather than the headline, because the conditions are always in the footnotes: how many processors the slot count assumes, which drive cage the figure assumes, whether the largest module is registered or load-reduced. This site’s compatibility pages record the vendor’s own maximum together with the conditions attached to it, per vendor for Dell, HPE, Lenovo and Supermicro.
Document three: the memory population rules. A separate document on every platform worth the name, and the one nobody opens. It answers the rank budget, the mixing rules, whether 3DS parts are supported and at what depth, the speed by population, and which capacities are one-module-per-channel only. It is where the rules that make a legal-looking configuration illegal actually live.
Then the machine itself, which settles what the previous four paragraphs cannot:
$ sudo dmidecode -t 16 | grep -E "Number Of Devices|Maximum Capacity"
Maximum Capacity: 3 TB
Number Of Devices: 24
$ sudo dmidecode -t 17 | grep -E "Locator:|Size:|Rank:|Configured Memory Speed:" |
grep -v "No Module Installed" | head -8
Locator: DIMM_A1
Size: 32 GB
Rank: 2
Configured Memory Speed: 2666 MT/s
Locator: DIMM_B1
Size: 32 GB
Rank: 2
Configured Memory Speed: 2666 MT/s
$ lscpu | grep -E "Socket|NUMA node\(s\)"
Socket(s): 2
NUMA node(s): 2
Four readings from that, and they are the four the documents cannot give you.
Number Of Devices is the true slot count including the empty socket’s share.
Rank is the package rank from the SPD, so a 3DS module reports 2 rather than
the logical count on its label, which is exactly what the chip-select budget
cares about. Configured Memory Speed against the module’s own Speed field is
the downshift, measured rather than predicted. And NUMA node(s) matching
Socket(s) confirms that both processors are present and that the slots wired
to the second one are live.
The answer is the minimum of the four, and on a second-hand machine it is usually the processor or the module type, not the slot count that people actually count.
Four machines, worked
HPE ProLiant DL380 Gen11. Headline 8 TB, from 32 × 256 GB DDR5 RDIMMs, which are octal-rank 3DS parts and not something you find loose. HPE’s own footnote adds “32 DIMMs only with 8SFF or 16SFF, 16 DIMMs maximum with 24SFF”, so ordering the dense drive cage halves the ceiling to 4 TB before a module is chosen. It also keeps the machine at one module per channel across its sixteen surviving slots, which is full 1DPC speed. The 8SFF machine reaches 8 TB, and HPE’s own speed table puts that population at 4400 MT/s. The dense chassis costs capacity; the dense population costs clock.
HPE ProLiant DL580 Gen10, four sockets, 48 slots. Three maxima out of one chassis: 1.5 TB as 24 × 64 GB RDIMM at 2933 (1DPC), 3 TB as 48 × 64 GB RDIMM at 2666 (2DPC), and 6 TB as 48 × 128 GB LRDIMM at 2933 (2DPC). Load-reduced memory doubles the capacity and keeps the clock, which is the case where paying the LRDIMM premium is unambiguously right. Above a terabyte per socket you also need M- or L-suffix processors, so check the chips before shopping for modules, and check the mezzanine kit before believing in 48 slots at all.
A second-hand two-socket EPYC 7002 box, 32 slots, 4 TB on the spec sheet. Start by working out where that 4 TB came from, because it is not AMD’s number:
board figure 32 slots × 128 GB LRDIMM = 4,096 GB
AMD's figure 2 sockets × 4 TB = 8,192 GB with 256 GB modules
The board’s stated maximum is half the socket’s, and the module it assumed is recoverable by division. Now three routes to roughly a terabyte. Sixteen 64 GB modules, eight per socket at one per channel, run at 2933 and 100 per cent of peak bandwidth. Thirty-two 32 GB modules reach the same terabyte at 2666, also 100 per cent, with every slot spent and nothing left for later. Twenty-four 32 GB modules, which is the configuration you land on by adding sticks to an eight-module machine, is 768 GB at 35 per cent, because twelve modules across eight channels cannot form a single interleave set and Rome has no six-channel mode to fall back on. The cheapest-looking purchase is the one that halves the machine.
The example machine, taken as far as it goes. Two Gold 6130 processors, 24 slots, twelve 32 GB dual-rank registered modules fitted, 384 GB reported.
where it is now 12 × 32 GB = 384 GB, 1DPC, full interleave
fill the empty slots 24 × 32 GB = 768 GB, 2DPC, full interleave
same slots, 64 GB 24 × 64 GB = 1,536 GB, at the processor ceiling exactly
same slots, 128 GB 24 × 128 GB = 3,072 GB, refused: 2 × 768 GB = 1,536 GB
Three of those four are reachable and the fourth is not, and the boundary is the processor rather than anything you can see in the chassis. The interesting line is the third: 24 × 64 GB lands exactly on the 1,536 GB ceiling, which is what a well-designed platform looks like from the inside, and it means the largest useful module for this machine is 64 GB even though the slots would take something bigger. Buying 128 GB modules for it is buying capacity the processors will refuse.
The cheapest route to a capacity is usually more modules, not bigger ones
Everything above is a constraint. This is the consequence, and it is the part that decides what you actually spend.
On the second-hand market the price per gigabyte rises with module capacity, and it rises steeply at the top of each generation. The reason is supply rather than silicon: the largest module of a generation was expensive when new, was bought in small numbers, is usually a stacked part, and is being retired later than the mainstream capacities that filled a decade of two-socket servers. The mainstream capacity of a generation is the one that arrives on the used market by the pallet.
You can check the direction on this site rather than taking it on trust. Open DDR4 16GB, DDR4 32GB and DDR4 64GB, sort each by price per gigabyte, and compare the cheapest row of each. The order is normally smallest first, and the gap between the mainstream capacity and the top capacity is usually a multiple rather than a margin. The same comparison on DDR3, between 8GB, 16GB and 32GB, is starker still, because DDR3 server memory has been leaving data centres for longer. Prices and inventory move every week, which is why the instruction here is to run the comparison rather than to quote a figure that will be stale.
So the arithmetic for a target capacity runs the other way from the way people usually do it. Take a sixteen-slot two-socket DDR4 server and a target of 512 GB:
route A 16 × 32 GB 2DPC, 4 ranks per channel, derated clock, zero slots left
route B 8 × 64 GB 1DPC, 2 ranks per channel, full clock, eight slots free
route C 4 × 128 GB four channels of eight populated per socket
Route C is wrong for a reason that has nothing to do with money: four modules cannot fill eight channels, so the machine runs on half its interleave and the bandwidth collapses in the way the Lenovo table describes. Between A and B, the cost difference is whatever the per-gigabyte gap between 32 GB and 64 GB modules is on the day, and the performance difference is one speed bin against eight free slots. Route A is normally cheaper and route B is normally faster and upgradable, and neither is the obvious answer, which is the point.
Two corrections to the “more modules is cheaper” rule, because it has limits.
The first is the rank budget. More modules of a smaller capacity means more ranks on the channel, and at some point the budget binds rather than the slot count. Sixteen dual-rank modules on an eight-channel socket is four ranks per channel, comfortably inside an eight-rank budget. Sixteen quad-rank modules is eight ranks per channel, which is the budget exactly, and on many boards it is also a speed bin lower and a module type commitment for the whole machine.
The second is the empty socket, and it is the biggest single lever on a used machine. Half the slots on a two-socket board are wired to the second processor, so a single-processor machine cannot use route A at all. On a platform old enough to be worth buying second-hand, the second processor is frequently cheaper than the four modules it unlocks the slots for, and it doubles the channel count and the bandwidth at the same time. Price the processor before you price the bigger modules. The processor index lists what each family takes so that you can confirm the pair you are matching is a pair.
Two more things worth doing before spending, both specific to this market. Compare against whole generations rather than one capacity, because a DDR3 server that is cheap to fill can be a better machine than a DDR4 server that is not: the DDR3 and DDR4 pages sort the same way, and the all-generations price table puts them side by side. And read auction listings as offers rather than prices. A bid that has not finished rising sorts to the top of any price-per-gigabyte ranking and is not a figure you can transact at, which is a property of how eBay works rather than a fault in the seller.
What cannot be known, and what this site records instead
Every section above has a boundary, and collecting them is more useful than scattering caveats.
The vendor’s maximum is a test result, not a measurement of the silicon. It tells you the largest configuration somebody validated. It does not tell you the largest that works, and the gap between those is a risk nobody quantifies.
Which die density a given firmware revision knows is not published. Vendors publish qualified module lists and release notes; they do not publish the reference code’s table of supported DRAM technologies, so whether a firmware update will lift a capacity ceiling is not answerable in advance.
The 2DPC data rate is missing from AMD’s SP5 documentation entirely. The population guide tabulates the first DIMM per channel and stops. This site records that as an open question against the platform rather than inferring a figure from the DDR4 generations.
Per-SKU capacity ceilings are frequently absent from the machine’s own documents. Lenovo’s SR650 V2 specification says only that “the operating speed and total memory capacity depend on the processor model and UEFI settings”, without naming a SKU tier, so the 8 TB figure is recorded as the top-of-range stated maximum rather than as a promise for every processor.
A module’s rank count is often not in its part number, which is why the label on the module settles it and a listing photograph of the heatspreader settles nothing. The vendor-by-vendor decoding is in ranks and 3DS modules.
Vendors contradict themselves inside single documents, and the R660’s two tables are only the clearest case. Where two vendor statements disagree, this site records both and says which was taken, rather than silently picking one.
Where a figure cannot be sourced, this site records the field as not stated rather than inferring it, and every compatibility record carries the documents it was built from. The reasoning is on the methodology page, and the short version is that a wrong capacity sends somebody to buy memory that will not work, which costs more than an empty field.
What to do with this on a listing page
- Count the processors before you count the slots. A two-socket board with
one chip fitted has half its slots dead, and no amount of memory changes
that.
lscpureporting one socket, ordmidecode -t 17showing twelve empty locators in a row, is the whole answer. On an older platform the second processor is often cheaper than the modules it unlocks. - Find the per-socket ceiling on the processor, not on the chassis. The processor index carries it per family with the source. If it is an Intel Xeon Scalable part above a terabyte per socket, look for the M or L suffix in the listing title, and assume its absence means absence.
- Decide the module type once, for the whole machine. RDIMM and LRDIMM cannot be mixed, 3DS and non-3DS cannot be mixed at two modules per channel, and DDR5 servers take no load-reduced memory at all. Start from registered ECC unless a vendor table has told you otherwise, and treat load-reduced as a decision about the machine rather than about the module.
- Work the rank budget, not the slot count. Two quad-rank modules fill an
eight-rank channel. Two dual-rank modules use half of it. The largest module
on any vendor’s list is nearly always a stacked part with rules of its own,
and the label tells you which, in the
nRxWstring rather than the capacity. - Prefer a population the interleave can use. One module per channel, or two per channel on every channel, and nothing in between. Twelve modules on an eight-channel socket is the configuration that costs two thirds of the bandwidth while looking like an upgrade.
- Buy the mainstream capacity of the generation, then check the ceiling again. Compare 16GB, 32GB and 64GB on price per gigabyte and expect the smallest to win, then confirm that the module count you need fits the slots, the ranks and the interleave rules before ordering.
- Read the listing for the organisation string, not the capacity. A 64 GB
module can be
2Rx4monolithic,4Rx4load-reduced or a 2-high stack, and those three are not interchangeable in any machine in this article. If the listing does not show it, ask for a photograph of the label. - Check the machine’s own report before and after.
dmidecode -t 16for the slot count and the firmware’s opinion,dmidecode -t 17for what is fitted and at what speed,/proc/meminfofor what the kernel kept. A machine reporting exactly half of what you installed is mirrored, not faulty. - Match the memory to the machine on this site rather than to the number in the title. The compatibility pages state each machine’s vendor maximum with the conditions attached, and server memory by generation narrows the catalogue to the modules those machines take.
- Treat a bid as an offer and a buy-it-now as a price. An unfinished auction sorts to the top of every price-per-gigabyte ranking on this site and on every other, and it is not a number you can pay today.
The capacity on the marketing page was true once, for one configuration, in a laboratory. The capacity you can reach is the smallest of eleven numbers, and after this article every one of them is something you can look up rather than something you discover when the modules arrive.
Related guides
- Ranks and 3DS: why two 64GB modules are not the same
- Server RAM vs desktop RAM: RDIMM, UDIMM and LRDIMM
- How to find out exactly which RAM your computer takes
- ECC RAM vs non-ECC: which do you actually need?
- Buying used RAM on eBay: what's safe, and how to prove it
- How much RAM do you actually need in 2026? Home and enterprise