RAMindex

Ranks and 3DS: why two 64GB modules are not the same

By Harry Saarinen · Updated

A rank is one full-width set of DRAM chips that answers to a single chip select, and 3DS means each of those chips is several dies stacked in one package and wired together by through-silicon vias. The quantity a memory controller budgets is ranks. It does not budget gigabytes, and it does not budget modules.

That is why four different 64GB DDR4 registered modules exist, why they carry four different speed grades, why two of them cannot go in the same channel as each other, and why the words “64GB DDR4 ECC REG 2666” separate none of them. It is also why a machine advertised at 1.5TB will refuse the modules that add up to 1.5TB, and why the module that works is often the more expensive one.

This article works the subject at four layers, because every piece of bad advice about ranks is advice that covers one layer and assumes the others. The die layer is what a rank is made of and what a stacked package does to it. The module layer is how the organisation code, the raw card and the buffers turn dies into something you can hold. The channel layer is chip selects, bus loading and the derating tables that fall out of both. The population layer is the vendor rules that decide which combinations a firmware will accept at all. Ranks are the only thing all four layers count in the same unit, which is exactly what makes the number worth understanding.

One housekeeping note, because it changes how to read the rest. The command output shown below is the shape of the output, carrying one consistent example module through the article, and it is not a capture from any particular machine. The example is a 32GB DDR4-2666 registered module, two ranks of x4 devices, in a two-DIMM-per-channel server. Where a figure is a vendor’s or a standard’s, the document is named. Where a figure is worked out here, the arithmetic is written out so you can check it, and the text says it is a derivation.

A rank is a chip select, and everything else follows

A module’s data bus is 64 bits wide, 72 with ECC on DDR4 and earlier, 80 on a DDR5 registered module. A rank is however many DRAM devices it takes to cover that width at once. Divide:

non-ECC, x8 devices      64 ÷ 8  =  8 devices per rank
ECC DDR4, x8 devices     72 ÷ 8  =  9 devices per rank
ECC DDR4, x4 devices     72 ÷ 4  = 18 devices per rank
ECC DDR5 RDIMM, x4       80 ÷ 4  = 20 devices per rank (10 per sub-channel)
ECC DDR5 RDIMM, x8       80 ÷ 8  = 10 devices per rank (5 per sub-channel)

Every device in a rank receives the same clock, the same address and the same command, and each contributes its own slice of the word. A read of one 64-bit word is eighteen x4 devices each handing over four bits, simultaneously, into eighteen different lanes of the same bus.

What separates one rank from the next is a single wire. The chip select, CS_n, is asserted for exactly one rank at a time, and a device that does not see its chip select asserted ignores the command on the bus entirely. That is the whole mechanism. Extra ranks widen nothing, because the bus is already fully covered by one of them. Only one rank per channel drives the data bus at any instant, so a second rank is capacity and overlap and never more bits per transfer.

This is the point at which most explanations stop, and it is the point at which the interesting part starts, because the chip select is not the only signal that is replicated per rank.

The signals that are per rank, and the ones that are not

On a DDR4 module the per-rank signals are CS_n, CKE and ODT. A dual-rank DDR4 module carries CS0_n and CS1_n, CKE0 and CKE1, ODT0 and ODT1, and the memory controller drives all six independently. The clock, the address bus, the command bus, the bank-group bits and the data bus are shared by every rank on the channel.

Three consequences follow from that split, and they are the reason rank count turns up in places that look unrelated to it.

Power management is per rank. CKE is the clock-enable, and dropping it puts one rank into a power-down state while the other keeps working. A controller managing a four-rank channel has four independent power states to schedule, and the exit latency from a power-down state is a real timing parameter it has to respect before that rank can accept a command again.

Termination is per rank. ODT switches on-die termination at the DRAM. On a multi-rank channel the controller terminates on the ranks it is not talking to, so that reflections from the idle stubs do not corrupt the transfer to the rank it is talking to. Every additional rank adds an entry to that termination matrix, and the matrix is not a firmware convenience: it is a set of resistor values the board vendor characterised for a specific number of ranks at a specific speed. This is one of the mechanisms by which rank count and speed grade are the same subject.

Addressing is not per rank, and on one rank it is deliberately scrambled. JEDEC’s module specifications allow a module’s odd ranks to be wired with several address and bank-group lines swapped, because the devices on the back of the PCB are mirrored relative to the ones on the front and mirroring the traces is shorter than routing around. The SPD declares which arrangement the module uses, and the controller has to un-mirror commands it sends to those ranks. So the controller does not merely count ranks. It has to know, per rank, how that rank’s address pins are wired, which is why a rank is a first-class object in the memory controller’s model of the world and a module is not.

DDR5 changed part of this. There is no CKE pin on a DDR5 device at all: power state entry moved onto the command/address bus, selected by the chip select. The rank is still the unit, but one of its three defining signals has been folded into another.

Reading the organisation code

The notation is nRxW: rank count, the letter R, then the width of one DRAM device. It is printed on the module label, it appears in every serious compatibility list, and it is absent from most listing titles.

Label Ranks Device width Devices on the module Where you see it
1Rx16 1 x16 4 Cheap desktop and SODIMM parts
1Rx8 1 x8 8, or 9 with ECC Desktop, ECC UDIMM, small RDIMMs
2Rx8 2 x8 16, or 18 with ECC Desktop, laptop, entry server
1Rx4 1 x4 18 Server only
2Rx4 2 x4 36 The commonest server RDIMM shape
4Rx4 4 x4 72, or 36 dual-die Quad-rank RDIMM and LRDIMM
2S2Rx4 4 logical, 2 package x4 36 stacked 3DS

Two spellings catch people out. An x16 device cannot build an ECC module, because 72 does not divide by 16, and Micron’s DDR4 RDIMM specification allows x4 and x8 only - so 1Rx16 next to “ECC REG” in a listing title is a typing error, not a rare part. And a stray D, as in 4DRx4, is the dual-die spelling some LRDIMM vendors use for a quad-rank x4 module; roughly thirty titles in this site’s listing corpus write it that way, which is enough to matter when you are searching and not enough to be worth a facet.

The stacked form counts twice over, and the two numbers are not the same kind of thing. The number before the S is package ranks: how many chip selects the module consumes. The number before the R is logical ranks inside each package rank: how many dies are stacked behind each of those chip selects. So 2S2Rx4 is two package ranks with two dies each, four logical ranks in total. The check is AMD’s own capacity ladder for EPYC 9004, where 2S2Rx4 is 128GB, 2S4Rx4 is 256GB and 2S8Rx4 is 512GB, which only works if you multiply the two numbers together.

That is the whole distinction in one sentence. A package rank costs a chip select; a logical rank does not.

Why the label is the only reliable copy of this

The organisation string is printed on the module because it is the one fact about a module that cannot be derived from anything else visible. Capacity is on the label too, and capacity plus generation plus form factor still leaves four constructions live at 64GB. The speed grade is on the label and is a property of the binning, not of the layout. Only 2Rx4 tells you, in five characters, how many chip selects the module will take and how wide its devices are, and both of those are load-bearing for whether it will work.

Which is why the single most useful thing to ask a seller for is a photograph of the label rather than of the heatspreader, and why a listing that shows a heatspreader with a vendor logo on it has shown you nothing. For server parts the label is a paper sticker on the PCB, usually on the side with fewer components, and it carries the part number, the organisation string, the JEDEC module designation and often the raw card revision. All four are useful. The heatspreader is a piece of aluminium.

Capacity is die density, device width and rank count in one equation

JEDEC’s DDR4 SPD annex does not contain a module capacity field. There is no byte anywhere in the 384-byte block that says “32GB”. The capacity is computed from four separate fields, and the formula in the annex is:

module capacity (bytes) = die density (bits) ÷ 8
                        × primary bus width ÷ device width
                        × logical ranks per DIMM

Worked on the example module carried through this article, a 32GB 2Rx4 DDR4 RDIMM built on 8Gb dies:

8 Gb ÷ 8            =  1 GB per die
64 bits ÷ 4 bits    = 16 data devices per rank
1 GB × 16           = 16 GB per rank
16 GB × 2 ranks     = 32 GB

The ECC devices are outside that sum, which is why a 72-bit module holds 64 bits’ worth of addressable memory and carries eighteen devices rather than sixteen. Two of the eighteen exist to store a checksum the memory controller computes, and they add nothing to the number in the listing title. The mechanism is in ECC vs non-ECC.

Intel’s Core Ultra 200S datasheet lists every DDR5 UDIMM construction its memory controller supports, and the whole table falls out of the same formula with three die densities:

Die Device organisation Devices Ranks Module
16 Gb 2048M x8 8 1 16 GB
16 Gb 1024M x16 4 1 8 GB
16 Gb 2048M x8 16 2 32 GB
24 Gb 3072M x8 8 1 24 GB
24 Gb 1536M x16 4 1 12 GB
24 Gb 3072M x8 16 2 48 GB
32 Gb 4096M x8 8 1 32 GB
32 Gb 2048M x16 4 1 16 GB
32 Gb 4096M x8 16 2 64 GB

Look at 32 GB. It appears twice, as 2Rx8 on 16 Gb dies and as 1Rx8 on 32 Gb dies, and so does 16 GB, once 1Rx8 and once 1Rx16. Same capacity, same generation, same slot, different module. On this exact processor with four DIMMs installed the single-rank part is rated 4800 MT/s and the dual-rank part 4400. Die density is not in the listing title and is rarely legible in the photograph.

The 24 Gb rows are also why DDR5 has capacities that are not powers of two. 12, 24, 48 and 96 GB are all 24 Gb dies, and a 48 GB module is 2Rx8 or 1Rx4 depending on who built it; AMD’s EPYC 9004 organisation list has both. Anyone who tells you a capacity implies a rank count is working from a memory of DDR3, when there were fewer die densities in production at once.

The same capacity, four different ways

Here is the 64GB DDR4 registered module, built four ways, with the arithmetic for each. All four are real constructions that shipped.

monolithic 2Rx4 on 16Gb dies
  2 GB per die × 16 data devices = 32 GB per rank × 2 ranks  = 64 GB
  36 packages, 36 dies, 2 chip selects

DDP (dual-die) 4Rx4 on 8Gb dies
  1 GB per die × 16 data devices = 16 GB per rank × 4 ranks  = 64 GB
  36 packages, 72 dies, 4 chip selects

3DS 2-high, 2S2Rx4, on 8Gb dies
  1 GB per die × 16 data devices = 16 GB per logical rank × 4 = 64 GB
  36 packages, 72 dies, 2 chip selects

LRDIMM 4Rx4 on 8Gb dies, data-buffered
  1 GB per die × 16 data devices = 16 GB per rank × 4 ranks  = 64 GB
  36 packages, 72 dies, 4 chip selects, plus 9 data buffers

Four modules, one capacity, one generation, one form factor, one notch position. They differ in die density, die count, package count, chip-select consumption, electrical loading, speed grade, price and which machines will take them. The listing title for all four is the same four words.

That is the whole argument of this article, and everything below is the detail needed to tell them apart and to know which one your machine wants.

What a second rank actually buys

A rank is not just storage. It is an independently addressable set of banks, with its own activation budget and its own refresh schedule, and the performance argument for two ranks is entirely about what a controller can overlap.

JESD79-4 gives each DDR4 device 16 internal banks arranged in four bank groups at x4 and x8, and eight banks in two groups at x16. JESD79-5 raised that to 32 banks in eight bank groups at x4 and x8. Because every device in a rank receives the same command, the rank’s bank count is the device’s bank count: a 2Rx4 DDR4 module has 16 banks per rank and 32 across the module, not 16 times eighteen.

A second rank therefore doubles the number of rows that can sit open at once. It also doubles two budgets that are counted per rank rather than per channel: the tRRD spacing between successive row activations, and the tFAW window that limits how many activations may start inside any rolling period. Both exist because activating a row is the most power-hungry thing a DRAM does and the charge pumps need time to recover. Neither limit knows the other rank exists, so a controller with two ranks available has two independent activation budgets and can start a row on one while the other is inside its window.

The rank switch is cheaper than the bank-group repeat

Micron’s 16Gb/32Gb 3DS DDR4 datasheet is the clearest published statement of what changing rank costs, because a stacked part forces the datasheet to tabulate read-to-read spacing three ways where an ordinary part tabulates it twice. The ranks it separates are the logical ranks inside a stack rather than two package ranks on a module, but the ordering is what matters. At DDR4-2666, where tCK is 0.750 ns:

same rank, same bank group        tCCD_L(SLR) = MAX(4nCK, 5ns)      →  7 clocks
different rank                    tCCD(DLR)   = MAX(4nCK, 3.748ns)  →  5 clocks
same rank, different bank group   tCCD_S(SLR) = 4nCK                →  4 clocks

Switching ranks costs less than hitting the same bank group twice, and more than moving to a fresh bank group in the rank you are already in. That ordering is the entire interleaving argument: a second rank gives the controller a third way to hide the array’s recovery time, ranked between the two it already had.

Run the same formulas at DDR4-3200, where tCK is 0.625 ns, and the gap widens: 5 ÷ 0.625 is 8 clocks for the same bank group, 3.748 ÷ 0.625 rounds up to 6 for a different rank, and tCCD_S stays at 4. That extrapolation is arithmetic on Micron’s published formulas rather than a figure from a speed bin table, and it is here to show the direction: the faster the bus, the more clocks a same-bank- group repeat costs, and the more there is for a rank switch to save.

One honest qualification that most explanations of rank interleaving skip entirely. The tCCD(DLR) figure above is a device timing, published because both logical ranks live inside one package on one interface. Between two separate package ranks there is an additional system-level cost that JEDEC does not specify at all: the bus turnaround, sometimes written tRTRS, during which one rank’s output drivers must have released the bus before the other’s take it. That number belongs to the board and the controller, it is set during memory training, and no datasheet publishes it because no DRAM knows it exists. So the published rank-switch penalty is a floor, and the real one on your machine is a floor plus a board-specific gap. This is the reason honest benchmark writeups of single-rank against dual-rank memory report a range rather than a figure.

Refresh is per rank, and that is where the second rank quietly earns most

Every DRAM row has to be refreshed before its capacitor leaks below the sense threshold. JESD79-4 sets the average refresh interval tREFI at 7.8 microseconds at or below 85 °C, halving to 3.9 microseconds above it. The refresh itself occupies the rank for tRFC, and tRFC scales hard with die density:

DDR4 die density tRFC1 Refresh overhead at tREFI = 7.8 µs
4 Gb 260 ns 3.3 per cent
8 Gb 350 ns 4.5 per cent
16 Gb 550 ns 7.1 per cent

The tRFC1 column is JESD79-4’s; the overhead column is a derivation, and it is just tRFC ÷ tREFI. Take the 16Gb row: 550 ÷ 7800 is 0.0705, so a channel built from 16Gb DDR4 dies spends about seven per cent of its time refusing to answer even when it is otherwise idle, and that figure doubles above 85 °C.

Refresh is issued per rank. On a single-rank channel, a refresh stalls the channel. On a two-rank channel the controller can refresh one rank while serving reads from the other, and the refresh overhead stops being a hole in the timeline and becomes something to schedule around. This is the least discussed and most reliable benefit of a second rank, because unlike bank-level parallelism it does not depend on the access pattern at all. Refresh happens whether your workload is sequential or random.

It is also why die density is a performance property and not only a capacity one. The move from 8Gb to 16Gb dies raised the refresh overhead of a rank from 4.5 to 7.1 per cent by the arithmetic above, which is one of several reasons a 64GB 2Rx4 module on 16Gb dies is not simply a doubled 32GB module.

DDR5 attacked this directly. JESD79-5 adds same-bank refresh, which refreshes one bank across all bank groups instead of the whole rank, so the rank stays addressable for the rest of its banks during a refresh. It also adds refresh management, a mechanism by which the controller tracks activation counts and the device demands extra refreshes when a row is being hammered. Both are per-rank mechanisms that make a rank less of a monolith, and both reduce the marginal value of having a second one.

Why the gain is workload-shaped, and why it shrank on DDR5

Two ranks help when the controller has enough outstanding requests to find something useful to do with the second one. They help most on workloads with many independent streams, and least on a single dependent chain of pointer chases, where there is nothing to overlap because the next address is not known until the last one returns.

Three structural changes have been eroding the benefit since DDR3, and they are worth stating because the folklore has not caught up:

  • DDR4 gave every rank four bank groups where DDR3 had none, so a controller already had a cheap way to avoid the same-bank-group penalty inside one rank.
  • DDR5 doubled that again to eight bank groups and 32 banks, so the in-rank escape routes doubled again.
  • DDR5 split the module into two independent 32-bit sub-channels, each with its own command bus, which is a second form of parallelism that exists whether or not a second rank is present.

None of that makes a second rank worthless. It does mean that a rule of thumb learned on DDR3, where dual-rank memory was a substantial and dependable win, overstates the case on DDR5, and that the correct answer on a modern platform is usually decided by the derating table rather than by the interleaving gain. Which is the next layer.

Single, dual and quad rank, and why quad rank stopped existing

Four rank counts have shipped on ordinary DIMMs, and each one exists for a different reason. Reading them as a ladder from worse to better is the mistake that sells people the wrong module.

Single rank is the cheapest way to build a given capacity once the die density is large enough to reach it, because it needs half the devices of a dual-rank part on the same dies. It also loads the data bus least, which is why single-rank modules hold the top row of almost every population table at two DIMMs per channel. What it gives up is the overlap: one set of banks, one activation budget, one refresh schedule, and a stall on the channel every time that rank refreshes.

Dual rank is the default, and on servers it is the commonest construction by a wide margin. It is where a vendor puts the capacity sweet spot, where the derating is mild or absent at one module per channel, and where the used market is deepest. If a machine’s population table does not push you elsewhere, 2Rx4 on a server and 2Rx8 on a desktop is the shape to buy.

Quad rank is a historical artefact and worth understanding as one. Before 16 Gb dies, the only way to build a large registered module was to put four ranks on it, and a four-rank module spends an entire desktop channel’s chip-select budget by itself. That is the origin of nearly every confusing rule in a five-year-old motherboard manual: the slot that will not take the module, the ceiling that halves when you fill the second slot, the speed that drops two bins. The Dell R620 figures further down are a quad-rank platform being honest about what quad rank costs.

Eight ranks never happened as eight chip selects. It happened as eight logical ranks behind two, which is a 4-high 3DS module, and that substitution is the reason quad-rank modules disappeared rather than being extended. Stacking delivered what quad rank was for, without spending the budget quad rank spent.

So the progression is not a ladder. Single rank and quad rank are opposite answers to the same question - how do I stay inside the channel’s electrical budget - separated by twenty years of die density, and dual rank is where almost everybody should be standing.

Dual rank is not dual channel, and the difference is a factor of two

This confusion is common enough to deserve its own arithmetic, because the two phrases sound alike and describe opposite things.

two channels    two 64-bit buses, driven independently   128 bits per transfer
two ranks       one 64-bit bus, two sets of devices       64 bits per transfer

Two channels double the width of the path between the processor and memory. Two ranks do not widen it at all; they give the controller a second set of banks to hide latency behind, on the same 64 bits. Doubling channels is worth close to double the peak bandwidth. Doubling ranks is worth a schedule improvement whose size depends on what you are running.

The practical consequence is a buying rule that catches people out on desktops. If your board has four slots and two channels, two modules populating both channels beat one module of twice the capacity every time, regardless of rank count, because the second module is a second channel and the second rank is not. Rank count is the tie-breaker between two options that already fill the channels, never a substitute for filling them.

Two ranks on one module, or two modules of one rank?

Same total rank count on the channel, and the answer is still not a wash.

On the command and address bus, the one dual-rank module is fewer loads than two single-rank modules on an unbuffered channel, because each module is a set of devices and a connector’s worth of stub. On a registered channel that difference mostly disappears, since each module’s register presents one load either way, and two registers is two loads against one.

On the data bus the two arrangements are nearly identical, because the DQ loading is counted per rank and both arrangements have two.

On slots, the single module leaves one free, which is the only difference that matters at the moment you want to upgrade. On failure, the single module is one part to replace rather than two, and on power it is one set of register and buffer components drawing current rather than two.

The tie goes to the single dual-rank module in almost every case, and the exception is worth stating: when the board’s population table rates two single-rank modules higher than one dual-rank module plus anything, take the table’s word for it. The table was measured on that board. This reasoning was not.

Loading: why ranks per channel caps the speed grade

A memory channel is a transmission line with stubs hanging off it, and every device connected to it is a capacitive load. Two different buses are loaded in two different ways, and confusing them is why “more ranks means slower” is stated far more often than it is explained.

The command and address bus is loaded per device. Every DRAM on the channel must see the same address and command, so on an unbuffered module the address bus fans out to all eighteen devices of a rank, all thirty-six of a dual-rank module, and all seventy-two if there are two such modules in the channel.

The data bus is loaded per rank. A given data line reaches exactly one device per rank, because each device owns its own four or eight lanes. So the DQ lines see one load per rank on the channel, not one per device.

Count it for the example module, ignoring for one paragraph the register that a server module puts in the way. A 2Rx4 DDR4 RDIMM carries 36 devices, so two of them in a channel is 72 devices’ worth of address-bus load and four loads on each data line. Now the unbuffered 2Rx8 desktop module: 16 devices, so two of them is 32 on the address bus and four loads per data line. The desktop channel has less than half the command-bus loading and exactly the same data-bus loading. Same rank count, wildly different device count, and that split is precisely why the two classes of module buffer different things.

What each buffer actually removes

A registered module puts a registering clock driver between the connector and the DRAM. The RCD receives address, command, chip select, clock enable and termination control, latches them, and re-drives them to the devices. From the controller’s side the entire module is one load on the command bus regardless of whether it carries 18 or 72 devices. The cost is a clock of latency, because the command is latched and re-issued one cycle later, and the DQ lines are untouched: a registered module does nothing at all about data-bus loading, which is why rank count still sets its speed grade.

A load-reduced module adds data buffers on top of the register, one per byte lane, nine on a 72-bit DDR4 LRDIMM and ten on an 80-bit DDR5 one. Now the data lines are buffered too, so the controller sees one load per module on every bus it drives. That is what lets an LRDIMM carry four ranks at a speed an unbuffered four-rank module could not reach. The cost is more latency again, and a data buffer on the path of every single byte that enters or leaves the module.

A 3DS package removes load a third way, and it is the only one of the three that works inside the package rather than on the PCB. More on that below; the short version is that stacking puts the extra dies behind a master die that owns the only external interface, so eight dies present one device’s worth of load.

The three are not alternatives so much as three different places to put the same repeater, and a high-capacity DDR5 server module uses two of them at once.

The chip-select budget

One chip select, one package rank, so the ceiling on ranks per channel is a pin count rather than a preference. Intel’s Core Ultra 200S datasheet states it in the signal list: DDR0_CS[3:0] is “Chip Select: (1 per rank) … There is one Chip Select for each SDRAM rank”, four of them per sub-channel. The support matrix then states the budget outright: a maximum ranks-per-channel of 2 with one DIMM per channel and 4 with two.

DDR4 counted the same way, and spent the top two pins twice over. A DDR4 registered DIMM has CS0_n and CS1_n plus two dual-function pins, CS2_n/C0 and CS3_n/C1, where the alternative function is a chip ID for a stacked part. Micron’s RDIMM specification notes that the upper two “are not used on UDIMMs”. That pin sharing is not a curiosity: it is the connector telling you that a quad-rank module and a 3DS module are competing for the same four pins, and using them for two different purposes.

Four is the desktop number. Server controllers get more. Intel’s S2600WT board rules allow “a maximum of 8 ranks … on any one channel, counting all ranks in each DIMM on the channel”, and eight is what makes a Dell PowerEdge R620 able to take “up to two quad-rank RDIMMs … per channel” at all.

Dell’s own manual then shows what that costs. Quad-rank RDIMMs run at 1333 MT/s with one DIMM per channel and 1066 with two, and the third slot in a channel is closed to them entirely: only single- and dual-rank parts may fill it, and filling it caps the system at 1333. That eight-rank budget is also what caps a server’s real capacity, worked through in why your server won’t take its maximum RAM.

The rule that looks arbitrary on a spec sheet is this arithmetic and nothing else. A quad-rank module can spend a desktop channel’s entire budget by itself, so the board will not let you add a second one. A channel supports a limited number of ranks, not a limited number of modules.

The derating tables, and what they are derating

Every platform publishes a table of speed against population, and every one of them is indexed by rank count somewhere. Two platforms, twelve years apart, in increasing order of how much the rank count costs:

Platform Population Rated speed
Intel Core Ultra 200S 4 DIMMs, single-rank 4800 MT/s
Intel Core Ultra 200S 4 DIMMs, dual-rank 4400 MT/s
Dell PowerEdge R620 quad-rank RDIMM, 1 DPC 1333 MT/s
Dell PowerEdge R620 quad-rank RDIMM, 2 DPC 1066 MT/s
Dell PowerEdge R620 3 DIMMs per channel, single or dual rank only 1333 MT/s maximum

The first pair is the cleanest statement of the effect anywhere in a consumer datasheet: same slot count, same generation, same capacity available, and 400 MT/s of difference decided by nothing but whether the modules are one rank or two.

What the table is derating is signal integrity on the data bus, and the mechanism is worth one paragraph because it explains why the penalty is a cliff rather than a slope. Each rank’s DQ pins are a stub hanging off the data line. A stub reflects. Two stubs reflect into each other. The controller compensates by training the timing and the termination for the population it finds, and the eye diagram it ends up with either meets the setup and hold requirements of the speed bin it wanted or it does not. There is no partial credit. It drops to the next bin down, whole.

This is also why the derating is a property of the channel, not of the module. A 1Rx8 module that runs at 4800 alone runs at 4400 next to a dual-rank neighbour, because what changed is the number of stubs on the shared line and the module has no way to be unaffected by that.

And it is why a listing that quotes a speed grade is quoting the module’s bin, not the speed you will get. The bin says the module was tested good at that rate. The population table says what the channel will actually be clocked at. Those are different numbers and the second one is usually lower, which is the same distinction between an advertised rate and a running rate covered in RAM speed: MHz against MT/s.

3DS: more dies behind one chip select

Once the chip-select budget is spent, the only way left to add capacity is to put more dies behind one of the chip selects you already have. Micron’s 3DS DDR4 datasheet describes the arrangement: “The 3DS device provides a stack of DRAM die with one die configured as the master and the remaining die in the stack configured as slave device(s). Each die functions as a different logical rank.”

The master die carries the only external interface. The slaves are reached through through-silicon vias, holes etched clean through a thinned wafer and filled with metal, which is why the technology is named after them. Then comes the load-bearing sentence, and it is the one to memorise:

“In contrast to conventionally stacked DDR4 (TwinDie), 3DS stacks only have one CS_n pin regardless of the number of die; die (logical rank) selection is accomplished by the state of the Chip ID (Cx) pin(s), which behave as rank address(es).”

A 2-high stack uses CS_n plus one chip ID; a 4-high stack adds a second; three chip-ID bits reach eight dies. The chip IDs are extra address bits, not extra chip selects, so the controller’s budget never notices them. On a registered module it is the register that drives the devices’ chip-ID inputs, from chip selects the host encoded, which is why the stack height a module can present is a property of the register’s mapping rather than of the connector’s pin count.

The second benefit is electrical, and it is the one that makes high capacity possible at speed rather than merely possible. Because the master buffers the slaves, in Micron’s words, “the electrical signal loading of the external interface is that of a single DDR4 SDRAM”. Eight dies, one load. Compare that with the alternative of building the same capacity out of more packages, which would need more devices on the address bus and more stubs on the data lines, and the reason high-capacity DDR4 and DDR5 server memory converged on stacking is immediate.

Three costs come with it, and vendors are quiet about all three.

Thermal. Eight dies in one package share one path to the heatspreader, and the dies in the middle of a stack are insulated by the ones above and below them. DRAM retention halves above 85 °C by the tREFI rule quoted earlier, so a hot stack refreshes twice as often, which makes it hotter. This is a large part of why stacked modules carry conservative speed bins and why server vendors publish minimum airflow requirements against specific module types rather than against the chassis as a whole.

Yield. A TSV stack is assembled from dies that are individually good, thinned and bonded; a bonding failure scraps every die in the stack. That cost is in the price, and it is why a 3DS module has never been the cheapest route to a capacity that a monolithic part can also reach.

Timing. Every array timing in a 3DS datasheet exists twice, once for the same logical rank and once for a different one, which is where the SLR and DLR suffixes in the read-to-read table above came from. Crossing dies inside a stack is not free; it is merely cheaper than crossing package ranks.

DDP is the other kind of stacking, and it is not the same thing

A dual-die package, which Micron calls TwinDie and the trade calls DDP, also puts two dies in one package. It is not 3DS and the difference is exactly the one Micron’s sentence above draws: a DDP package, in Micron’s description, “uses CS1_n, CKE1, and ODT1 to control the second die”. The second die has its own chip select, its own clock enable and its own termination pin. It spends a chip select and it presents two loads. The package is shared; nothing else is.

So a 64GB DDP module is a four-rank module in every sense the memory controller cares about, and it occupies four of the channel’s chip selects, while a 64GB 3DS 2-high module holds the same bits with the same die count and occupies two.

Samsung’s 64GB DDR4 LRDIMM M386A8K40BM2 is DDP rather than TSV, and the part number says so: the M in its package position is Samsung’s dual-die code, where a stacked part carries a digit instead.

Construction Chip selects Stack Ranks the system sees Real part
Monolithic 2 none 2 M393A8G40AB2-CWE, 64GB DDR4-3200
DDP / TwinDie 4 2 dies, one package 4 M386A8K40BM2, 64GB LRDIMM
3DS 2-high 2 2H TSV 4 M393A8K40B21, 64GB DDR4-2400
3DS 4-high 2 4H TSV 8 MTA144ASQ16G72PSZ, 128GB DDR4 RDIMM

Every row is 36 x4 packages on a 72-bit module, because eighteen packages cover the width and there are one or two package ranks of them. The 128GB part is a genuine eight-logical-rank module that spends two of the controller’s four chip selects, and Micron’s leading 144 counts its dies rather than its packages: 36 packages, four dies each.

Rows one and three are the pair in this article’s title. Two 64GB Samsung x4 registered modules: one is two ranks and graded DDR4-3200, the other is four logical ranks and graded DDR4-2400, because the stacked part was designed when 8 Gb was the largest die in volume and the monolithic part waited for 16 Gb. On a board that derates by rank at two modules per channel they do not perform alike, and in a machine already carrying dual-rank modules they may not both fit.

The notation nobody agrees on

Multiply the two numbers in the stacked form and you get what population tables care about: 2S2Rx4 is four logical ranks on two chip selects, 2S4Rx4 eight, 2S8Rx4 sixteen. Past that, the labelling is a mess in vendors’ own documents, and the mess is worth knowing because it is the single most common source of a mis-ordered module.

  • The same 2-high 64GB construction is a raw card to Supermicro, which lists “3DS RDIMM Raw Cards: A/B (4Rx4)”, and a “(4Rx4) 3DS RDIMM” in Lenovo’s option list, while listings of Samsung’s M393A8K40B21 sell it as “2S2Rx4”. One module, two spellings, and only one of them says it is stacked.
  • AMD’s EPYC 9004 Series Memory Population Recommendations (publication 58269) tabulates 2S2R as “4 ranks” and 2S8Rx4 as “16 ranks”, then tabulates 2S4R as “4 ranks” - arithmetic that works on two rows out of three.

The resolution when they conflict: count logical ranks for capacity limits and speed bins, count package ranks for chip selects and electrical load, and take the module’s own label over any secondary source. A bare 8Rx4 printed on a stacked part is already the logical count.

A word on raw cards, since Supermicro’s phrasing above depends on them and most buyers have never heard the term. A raw card is a JEDEC-registered PCB design: the layout, the device count, the rank arrangement and the component positions, identified by a letter and a revision, as in RB2. Server vendors qualify raw cards rather than part numbers, because two vendors’ modules built on the same raw card are electrically the same object. When a support list names a raw card and a listing names a part number, the raw card is the more specific of the two and it is printed on the module label alongside everything else.

A 64GB 3DS RDIMM and a 64GB LRDIMM are two answers to one problem

Both exist because a channel runs out of electrical budget before it runs out of address space. They solve it in different places, and the difference decides which machines take which.

A load-reduced module buffers the data lines on the module, in nine DDR4 data buffers sitting between the connector and the DRAM. The ranks behind those buffers are ordinary ranks with ordinary chip selects. The controller sees one load per module on the data bus and one on the command bus, and it still spends four chip selects on a four-rank part.

A 3DS registered module buffers nothing on the data lines. It reduces the load inside each package instead, by hiding half, three-quarters or seven-eighths of the dies behind a master depending on the stack height, and it reduces chip-select consumption at the same time by addressing those dies with chip IDs.

64GB LRDIMM (DDP 4Rx4) 64GB 3DS RDIMM (2S2Rx4)
Buffers on the command bus register register
Buffers on the data bus nine data buffers none
Load per data line one per module one per package rank, so two
Chip selects consumed four two
Logical ranks for capacity four four
Added latency register plus data buffer register only
Platform support needs explicit LRDIMM support needs explicit 3DS support

Read the last row twice. Neither is a superset of the other, and a machine that takes one is not thereby able to take the other. Support for load-reduced memory and support for stacked memory are two separate lines in a platform’s memory controller documentation, and there are machines that take 64GB LRDIMMs and refuse 64GB 3DS RDIMMs, and machines the other way around.

This has a history, and knowing it explains why DDR5 looks the way it does. DDR3 load-reduced modules carried a single memory buffer that performed rank multiplication: the buffer presented, say, eight physical ranks to itself as two logical ranks to the host, doing the chip-select translation in the buffer. That is the same idea as a chip ID, implemented on the PCB instead of inside the package, one generation earlier. DDR4 split the buffer into a register plus nine data buffers and kept load reduction as a module-level feature. Then DDR5 moved the whole job into the package: AMD’s EPYC 9004 platform supports no LRDIMM at all, because 3DS registered modules cover the capacities load reduction used to be needed for, at lower latency and with fewer parts to fail.

The practical consequence for anybody shopping second-hand DDR4 is that load-reduced modules and registered modules are two separate markets with different price behaviour, and the machine decides which market you are in before you look at either. That decision is made in RDIMM, UDIMM and LRDIMM, and the rank arithmetic here is what it rests on.

Where the ranks sit on DDR5

A DDR5 module is not one 64-bit channel. It is two independent 32-bit sub-channels, 40 bits each once registered ECC is added, and each sub-channel has its own command and address bus and its own chip selects. Intel’s desktop controller exposes four of them, DDR0 through DDR3, each with CS[3:0] and CA[12:0] of its own: two 64-bit channels, four sub-channels, four chip selects apiece.

The module’s rank 0 is twenty x4 packages, ten in each sub-channel, and each sub-channel selects its half with its own chip select. So a two-rank x4 registered module carries forty packages. Intel’s ECC module table counts it the same way at x8: ten devices for one rank, twenty for two. A 128GB 2S2Rx4 3DS part carries those same forty packages, because stacking changes the die count inside each package, not the number of packages.

That is worth stating plainly because it is where DDR5 module photographs mislead. A 128GB DDR5 RDIMM and a 32GB DDR5 RDIMM can carry an identical number of identically sized black rectangles. What differs is inside them.

One terminology note, because listings and forum threads use the words interchangeably: JEDEC’s word for DDR5 is sub-channel, not pseudo-channel. Pseudo-channels are an HBM arrangement in which the two halves share one command bus. DDR5’s sub-channels share nothing but the module they are soldered to, and that independence is the architectural point.

For server DDR5, AMD’s EPYC 9004 guide is the clearest published list of what exists:

Type Organisations Capacities
RDIMM 1Rx8, 1Rx4, 2Rx8, 2Rx4 16, 24, 32, 40, 48, 64, 80, 96 GB
3DS RDIMM 2S2Rx4 128, 192 GB
3DS RDIMM 2S4Rx4 256, 384 GB
3DS RDIMM 2S8R 512 GB, marked “pending ecosystem enablement”

The first row of that table is the entire non-stacked market and the other three are the entire stacked one, and the boundary between them falls at 96GB. Below it, ordinary registered modules. Above it, nothing but 3DS.

One line under the table costs more money than the table itself. EPYC 9004 supports no LRDIMM and no UDIMM at all, so on that platform high capacity goes through stacked registered modules or it does not happen, and that is now the general pattern for DDR5 servers rather than an AMD quirk. The population rules the same document attaches to those parts are in the mixing section below, and two of them are absolute.

The registering clock driver on a DDR5 module is defined by JESD82-511, revised since as JESD82-514 and JESD82-515, and it is a more capable part than its DDR4 predecessor under JESD82-31. It drives both sub-channels, it carries its own control interface for training and for error reporting, and on a stacked module it is the component that turns host chip selects into device chip IDs. It is also the reason a DDR5 registered module is not a passive object: there is firmware on it, and a platform’s qualified-module list is in part a list of register revisions it knows how to talk to.

The one place ranks become a speed feature rather than a constraint is the multiplexed-rank module, which reads two ranks at once and multiplexes them onto the bus, doubling the effective data rate at the connector for a first generation rated at 8800 MT/s. It is the first module type in the history of DDR for which more ranks is the point rather than the price, and it needs a controller built for it. Whether that is the buying decision you are making is settled in RDIMM, UDIMM and LRDIMM.

Compared against DDR4, none of this changes the arithmetic. It changes where the 40 bits sit and how many command buses there are. Ranks still cost chip selects, stacking still hides dies behind a master, and the derating table still indexes on the number of ranks the channel ended up with. The catalogue for it is at DDR5 memory, where the registered modules are the ones a server will take and everything else in the list is a desktop or laptop part.

Reading rank and organisation off a module you already own

Every module carries a small EEPROM holding its Serial Presence Detect data, and that block is where the firmware learns everything in this article before it trains the channel. For DDR4 the layout is JEDEC Standard 21-C, Annex L; for DDR5 it moved out into its own document, JESD400-5. Both are the authority, and both are readable from a running machine if the bus is exposed to it.

The DDR4 fields that matter, and where they live

Four bytes between them carry everything this article has discussed. They are not adjacent, they are not named helpfully, and one of them is the only place a stacked module admits to being stacked.

Byte Field Bits What it carries
4 SDRAM Density and Banks 3:0 Die density, from 256 Mb to 32 Gb
4 SDRAM Density and Banks 7:4 Bank address bits and bank group bits
6 SDRAM Package Type 7 Monolithic device, or not
6 SDRAM Package Type 6:4 Die count inside the package
6 SDRAM Package Type 1:0 Signal loading: multi load stack or single load stack
12 Module Organization 5:3 Package ranks per module
12 Module Organization 2:0 Device width: x4, x8, x16 or x32
12 Module Organization 6 Rank mix: symmetrical or asymmetrical
13 Module Memory Bus Width 2:0 Primary bus width, 8 to 64 bits
13 Module Memory Bus Width 4:3 Bus width extension, which is the ECC byte

Byte 6 is the one worth staring at. The signal loading field is where 3DS and DDP part company in the data rather than in the marketing. A dual-die package declares a die count of two and a multi load stack, because its second die has its own chip select and its own loading. A 4-high 3DS package declares a die count of four and a single load stack, because the master hides the other three. Those are different values in the same two bits, and they are the machine-readable version of the distinction this article spent a section on.

Now read the example module’s fields, with the derived capacity underneath:

byte  4    die density              8 Gb
byte  6    package type             monolithic
byte 12    package ranks            2
byte 12    device width             x4
byte 13    primary bus width        64 bits
byte 13    bus width extension      8 bits          (so 72 total, ECC)

capacity = 8 Gb ÷ 8 × 64 ÷ 4 × 2 = 1 GB × 16 × 2 = 32 GB

And the 128GB 3DS part built from the same 8 Gb dies:

byte  4    die density              8 Gb
byte  6    package type             non-monolithic, 4 dies, single load stack
byte 12    package ranks            2
byte 12    device width             x4
byte 13    primary bus width        64 bits

logical ranks = package ranks × die count = 2 × 4 = 8
capacity      = 8 Gb ÷ 8 × 64 ÷ 4 × 8 = 1 GB × 16 × 8 = 128 GB

The capacity formula multiplies by logical ranks. Byte 12 reports package ranks. Multiply by byte 12 alone on a stacked module and you get 32GB for a 128GB part: a quarter of the real answer, from reading the field that is actually there rather than the one the formula asks for. That is not a hypothetical mistake. It is exactly the error behind every listing that describes a 128GB 3DS module as 2Rx4, and it is why the die count in byte 6 is not optional reading.

Byte 12’s rank mix bit deserves a line of its own. An asymmetrical module has ranks of different capacities, which is legal and rare, and it exists because a vendor can build an odd capacity out of two unequal ranks. If you see it set, stop treating the module as two of anything.

What the tools print

Two commands, in increasing order of how much they know and decreasing order of how often they work.

sudo modprobe ee1004
sudo decode-dimms | grep -E 'Fundamental|Module Type|Package Type|Size|Ranks|Device Width|Part Number'
Fundamental Memory type          DDR4 SDRAM
Module Type                      RDIMM
SDRAM Package Type               Monolithic
Size                             32768 MB
Ranks                            2
SDRAM Device Width               4 bits
Part Number                      MTA36ASF4G72PZ-2G6

That reads the SPD itself, so it has the die count and the signal loading and everything else in the table above. The catch is access: DDR4 SPD EEPROMs hang off an I2C bus that on a server is usually owned by the baseboard management controller, and an operating system that cannot reach the bus gets nothing. On DDR5 the arrangement changed again, to an SPD hub device that sits between the host bus and a local bus carrying the module’s own power management and temperature sensors, so the tooling and the kernel driver are different and older guides’ instructions do not apply.

The command that always works knows less:

sudo dmidecode -t 17 |
  grep -E '^[[:space:]]+(Locator|Size|Type|Speed|Configured|Part Number|Rank):'
	Locator: DIMM_A1
	Size: 32 GB
	Type: DDR4
	Speed: 2666 MT/s
	Configured Memory Speed: 2400 MT/s
	Part Number: MTA36ASF4G72PZ-2G6
	Rank: 2

Three things about that output, and all three are traps.

Rank: is the package rank count, because SMBIOS takes it from a firmware field the firmware filled from the SPD’s package-rank byte. A 3DS module normally reports 2 there, not the eight logical ranks printed on its own label. The number is not wrong; it is answering the chip-select question rather than the capacity question, and it is the question a memory controller asks.

There is no device width field at all. Data Width: 64 bits and Total Width: 72 bits tell you the module has ECC, and nothing in the structure distinguishes x4 from x8. If you need to know whether a machine can do single device data correction, dmidecode cannot tell you and the label can.

Configured Memory Speed below Speed is the derating table, live. The first is the module’s bin. The second is what the controller settled on after counting the ranks it found. A machine showing 2666 and 2400 on the same line pair has just told you, without a manual, that its population dropped it a bin.

Between the two commands there is a workable procedure. Run dmidecode first, because it always works and it answers the two questions that matter most: whether the ranks add up to what you expected, and whether the channel is running at the speed you paid for. Run decode-dimms when the SPD bus is reachable and you need width, die count or stacking. Read the label when neither is available, which on a module you have not bought yet is always. For the rest of the machine-side procedure there is how to find out exactly which RAM your computer takes.

Where rank hides in a part number

Rank is often not in the part number, which is why the label matters and why a listing that quotes only a part number has still left you work to do. What is encoded varies by vendor, and the variation is not random: the vendors who sell mostly to system builders encode construction, and the vendors who sell mostly to end users encode rank.

Samsung fixes DIMM type, device width and construction at known positions. Take M393A8K40B21:

Field Values
DIMM type (chars 2-4) 393 RDIMM, 386 LRDIMM, 391 ECC UDIMM, 378 UDIMM, 471 SODIMM
Density and organisation (four chars, 8K40) first digit is the depth in G x72; K marks 8 Gb dies and G 16 Gb; last digit is the device width
That width digit 0 = x4, 3 = x8, 4 = x16
Package B flip chip (monolithic), M DDP, a digit for a TSV stack, 2 for 2-high

So …8K40B21 is a 2-high stack of x4 dies and …8G40AB2 is a monolithic x4. Rank itself you derive, from module depth divided by die depth times devices per rank. Samsung’s DDR5 modules changed the type field rather than the grammar: where a DDR4 registered module is M393, a DDR5 one is M321.

Micron encodes device count and construction but not rank. The number straight after MTA or MTC is the DRAM die count, so MTA9… is nine dies and MTA36… thirty-six, and on a stacked part it counts dies rather than packages. The family letters separate monolithic from stacked. Three parts from one family make the scheme legible at a glance:

MTA36ASF8G72PZ      64GB  monolithic RDIMM   36 dies in 36 packages
MTA72ASS8G72PSZ     64GB  3DS RDIMM          72 dies in 36 packages, 2 high
MTA144ASQ16G72PSZ  128GB  3DS RDIMM         144 dies in 36 packages, 4 high

Three different die counts, one package count, two capacities. The PSZ suffix and the SS and SQ family letters are what mark the stack; the leading number is what tells you how tall it is, if you already know the module has 36 packages.

Kingston ValueRAM states rank outright, as a letter immediately before the width: S single, D dual, Q quad. KVR32S22S8/16 reads as DDR4-3200, SODIMM, CL22, S8 for 1Rx8, 16 GB. KVR1333D3Q8R9S/4G is the quad-rank spelling on DDR3. This is the friendliest of the schemes and it belongs to the vendor selling to people who have to work it out for themselves.

Crucial’s server parts do the same thing in a different position: the letter group before the width digit carries the rank, D for dual and S for single, so a …RFD4… part is dual-rank x4 and a …RFS4… part is single-rank x4. It is a consistent scheme across their registered line, and it is not published as a decoder, which is a fair description of the state of part-number decoding generally.

And then the honest part. SK hynix, Hynix-branded modules resold under a system vendor’s number, and every white-box module built on somebody else’s raw card are not reliably decodable from the outside, and this site will not print a decoder for schemes the vendor does not document. Where a part number cannot be resolved, the answer is the label, the SPD, or a question to the seller. What this site does hold, and what it will show you, is the per-part page: the parts it has seen in listings, with the specifications it could source for them, at RAM by part number.

One more encoding, because it is on every server module and almost nobody reads it. The JEDEC module designation - PC4-2666V-RB2-11 and its relatives - carries the generation in PC4, the data rate in 2666, a latency grade in the letter after it, and then the module type: R registered, L load-reduced, E unbuffered ECC, U unbuffered. That single letter is the one that decides whether the module will POST in your machine, and it is printed on every label next to the organisation string. RB2 adds the raw card, B, at revision 2.

The mixing rules ranks impose

Ranks are where population rules come from, and population rules are where a bargain turns into a machine that will not boot. Four families of rule, in descending order of how absolute they are.

Rules that are physical. You cannot mix registered and unbuffered modules in a machine, or ECC and non-ECC on most platforms that want ECC, and on DDR4 the notch does not save you from trying. Those are covered in RDIMM, UDIMM and LRDIMM and are not rank rules, but they are the first filter and they are worth clearing before any of what follows applies.

Rules that vendors state as absolutes and do not agree on. AMD’s EPYC 9004 guide says “Do not mix x4 and x8 DIMMs within a memory channel.” Supermicro’s X11 guide, describing Intel platforms, says “x4 and x8 DIMMs can be mixed in the same channel.” Both are correct about their own platforms. Read the document for the board in front of you, not the one you remember, and treat any general-purpose advice about mixing widths, including advice from a forum thread about a superficially similar machine, as not applicable.

Rules about order. Where mixing rank counts is allowed at all, the near-universal instruction is to populate the slot furthest from the processor first and to put the module with the most ranks in it. Dell and HPE both write their memory population guidelines this way. The reason is termination: the far module sits at the end of the transmission line and its terminators are what the channel is tuned around, so the heavier, more heavily loaded module goes where the board expects the load. Reverse it and the machine may post, run, and fail memory tests under load, which is the worst of the three available outcomes.

Rules about symmetry. These are the ones people break by accident, because nothing refuses to boot. AMD’s guide states that channel interleaving “requires all channels to use the same DIMM type, total memory capacity, and ranks”. A machine with one channel carrying a dual-rank module and the rest carrying single-rank modules of the same capacity has the right amount of memory and the wrong shape of it, and the firmware responds by interleaving across a smaller set of channels rather than by complaining. You lose bandwidth silently. The same clause is why adding a single module to a balanced machine is usually the worst upgrade available: it is the one change that can leave you with more gigabytes and less throughput.

Stacking has its own symmetry rule, and it is absolute on the platform that states it: 3DS and non-3DS modules cannot be mixed in a two-DIMM-per-channel EPYC 9004 system. That one is worth remembering when a machine is being upgraded in stages, because the sensible-looking plan of buying two 128GB stacked modules now and moving the existing 64GB monolithic ones into the second slot per channel is exactly the plan the rule forbids.

Rank sparing spends a rank, and the arithmetic is unkind

Memory sparing sets a rank aside as a replacement, watches the correctable error rate on the others, and fails a degrading rank over to the spare before it produces something uncorrectable. It is a rank-level feature, so it costs a rank, and the cost is a fraction of the channel rather than a fixed quantity.

channel with two 2R modules    4 ranks, 1 spared   = 25 per cent of capacity
channel with one 2R module     2 ranks, 1 spared   = 50 per cent of capacity
channel with two 1R modules    2 ranks, 1 spared   = 50 per cent of capacity

That derivation is the reason rank sparing has entry requirements. Intel’s S2600WT board rules put it behind “at least 2 SR or DR DIMM installed, or at least one QR DIMM installed, on each populated channel”, because a channel with one rank on it has nothing to spare from. A channel holding a single 1Rx8 module cannot turn the feature on at all.

The other rank-dependent reliability features stack the same way. Lenovo documents Intel’s adaptive double device data correction, which survives a second device failure after the first, as working “with x4-based memory DIMMs” and requiring “two DIMM ranks per channel”, plus a Gold or Platinum processor. Read that as three conditions in series: the right processor, the right device width, and enough ranks. Miss any one and the feature is not merely degraded, it is absent from the firmware menu.

Single device data correction is the one that is purely about width rather than rank, and Lenovo’s ThinkSystem documentation states it as a requirement rather than a preference: “Single Device Data Correction (SDDC, also known as Chipkill, requires x4-based DIMMs)”. With x4 devices the 72-bit word reads as eighteen 4-bit symbols and a symbol-correcting code covers every bit one dead device contributes. A dead x8 device straddles two symbols and is not correctable by the same code. This is the single strongest argument for x4 server memory and it has nothing to do with rank count, which is why it is worth separating from the rank rules it is usually bundled with. The correction mechanism is in ECC vs non-ECC.

One repair mechanism is deliberately outside all of this. Post-package repair, added to the DRAM itself in JESD79-4, lets the firmware retire a failing row and substitute a spare one inside the device, permanently. It operates per device, so rank count neither helps nor hinders it, and it is the reason a machine that logged correctable errors on one address six months ago may never log them again.

Which machines take which stacking

Stacking support is a property of the memory controller, which lives inside the processor. The board’s manual is where you read it, but the board is not what decides it, and a board that takes a 256GB stacked module with one processor installed will refuse the same module with an older one in the same socket.

There is a way to work out which module a platform’s headline capacity was written against, and it needs nothing but division. The headline is the slot count multiplied by the largest module the processor qualifies, so dividing the headline by the slots gives you the module. Run it on six generations:

Xeon E5-2600 v4         1.54 TB ÷ 12 slots = 128 GB per module
Xeon Scalable 1st gen    768 GB ÷ 12 slots =  64 GB per module
Xeon Scalable, M SKUs    1.5 TB ÷ 12 slots = 128 GB per module
EPYC 7001                  2 TB ÷ 16 slots = 128 GB per module
EPYC 7002 and 7003         4 TB ÷ 16 slots = 256 GB per module
EPYC 9004                  6 TB ÷ 24 slots = 256 GB per module

(all per socket)

The arithmetic is mine; the headline figures and the slot counts are the vendors’. What it exposes is the thing the headline hides. The first-generation Xeon Scalable row is the interesting one: its standard processors were sold against a 64GB module, so the 128GB stacked parts the boards physically accept sit outside the ceiling the processor was specified to, and reaching 1.5TB meant buying an M-suffix part at a substantial premium for nothing but a larger number in a firmware table. That premium is invisible in a listing for the processor and decisive for the memory you can put behind it.

The EPYC rows show the other pattern. Each generation’s ceiling moved by doubling the stacked module rather than by adding slots, until Genoa added both at once, which is why second-hand 256GB stacked DDR4 modules are useful on one generation of AMD hardware and dead weight on the one before it.

Against that, four vendor statements about specific machines, which are the kind of thing worth checking before spending:

  • A Supermicro X11DPi-N takes a 256GB 3DS module only with second-generation Xeon Scalable installed; the same board with first-generation processors caps at 128GB stacked. Sixteen slots at 256GB is the 4TB the board is advertised with, and it is unreachable with the wrong processor in the socket.
  • A Supermicro X10SRL-F reaches 128GB load-reduced only with a 3DS part. Eight slots at 128GB is 1TB, and a non-stacked 128GB LRDIMM is not a thing that exists for it.
  • A Lenovo ThinkSystem SR630 is documented at 3TB, which is 24 slots at 128GB, written specifically against 128GB 3DS registered modules, while its load-reduced ceiling is 64GB per module. That is the opposite of the usual pattern, in which load-reduced parts reach higher than registered ones, and it is a direct consequence of DDR4-era stacking having overtaken data buffering.
  • A Dell PowerEdge R620 predates all of it and shows the pre-stacking world: quad-rank registered modules, two per channel maximum, and a speed penalty for using them.

The rule that survives all four: check the processor, then the board, then the module, in that order. Doing it in the other order produces a module that fits the slot, satisfies the board manual, and is not on the processor’s list. Where a machine’s ceiling depends on stacking or on rank, the compatibility pages on this site say so in words rather than leaving it to be inferred from a number, and the per-processor view is at RAM by processor.

What ranks cost in power, and in the minutes before the machine boots

Two costs of rank count that never appear in a comparison and both of which people meet in practice.

Standby current is per rank. A DRAM device draws current when it is not being accessed, for its refresh machinery, its delay-locked loop and its termination, and those currents are specified per device in every datasheet. A channel with four ranks of eighteen x4 devices has seventy-two devices drawing standby current where a channel with one rank has eighteen. This is a large part of why a fully populated server idles substantially higher than a half-populated one with the same total capacity in fewer, denser modules, and it is an argument for buying the larger module rather than two smaller ones that is entirely separate from the speed-grade argument. DDR5 moved voltage regulation onto the module itself, so on that generation the power behaviour of a module is a property of the module and not only of the board.

Training is per rank. Before a memory channel can run at speed, the firmware has to find, for every rank, the timing at which each data lane’s signal is centred: write levelling, read levelling, per-bit deskew, termination selection. The work scales with the number of ranks, not the number of modules, and on a two-socket server with twenty-four dual-rank modules that is forty-eight ranks to train across twelve channels. This is why a large server takes minutes to reach its own firmware screen and a desktop takes seconds, why the delay gets worse when you fill the second slot per channel, and why some server firmware caches training results and boots faster the second time. A machine that suddenly takes much longer to POST after a memory change has not developed a fault. It has been given more ranks to train, or a population it no longer has cached results for.

The older generations still use the same grammar

Most of this article’s examples are DDR4 and DDR5 because that is where the documents are. The arithmetic did not change.

DDR3 server memory is the generation where rank count was most visibly a constraint, because the die densities were small enough that high capacity meant many ranks. The largest DDR3 server modules were 4Rx4, quad-rank, spending four chip selects each; channels accepted two of them and derated when they did. DDR3 load-reduced modules answered that with a memory buffer that performed rank multiplication, presenting a module’s physical ranks to the host as a smaller number of logical ranks and doing the chip-select translation itself. That is the same trick as a 3DS chip ID, implemented on the circuit board a generation earlier, and the reason it moved into the package is that a buffer on the PCB cannot fix the loading of the devices behind it while a master die can.

For anyone buying second-hand server memory today this is not history. DDR3 registered and load-reduced modules are usually the cheapest ECC memory per gigabyte on this site, precisely because the machines that take them are old, and the rank rules are what decide whether a given pile of them will fill a given machine. The catalogue is at DDR3 memory, with registered and load-reduced split out, and the rank arithmetic in this article applies to all of it unchanged.

Further back, DDR2 and DDR modules use the same nRxW grammar with smaller numbers and the same chip-select budget, and they are still worth buying for a machine that needs them; there is simply less to say, because there was no stacking and the buffering options were fewer. Both remain catalogued, at DDR2 and DDR.

Why this site leaves the rank field empty so often

Rank is the field this site most often cannot fill, and the reason is structural rather than lazy. A seller listing a module writes a title from the things that sell it: capacity, generation, speed, and the words ECC and REG if they apply. The organisation string sells nothing, so it reaches the title on a minority of server listings and almost never on a desktop one.

The temptation is to infer it, and the inference is available: a 32GB DDR4 registered module is 2Rx4 far more often than it is anything else, so filling the field with the common case would be right most of the time. This site does not do it, for a reason the rest of the article has been building toward. A guessed rank count feeds the compatibility pages, where it becomes an assertion that a specific module fits a specific machine, and a wrong assertion there costs somebody a module and a return. An empty field costs them one question to the seller.

So the rule is the same one used everywhere else on the site: the field is filled when the title, the item specifics or a sourced part-number record states it, and left empty when it does not. That is why two apparently identical listings can show an organisation string on one row and a blank on the next. The blank is not a different module. It is a seller who wrote a shorter title.

Where a part number is present and resolvable, the part pages carry what could be sourced for it, which is frequently more than the listing did. Where it is not, the honest position is the one on the methodology page: record what is stated, derive only what follows arithmetically from what is stated, and leave the rest alone.

What a listing cannot tell you

Collecting the boundaries in one place is more useful than scattering caveats through the article.

Die density is not in the title and rarely in the photograph. A 32GB DDR5 UDIMM is 2Rx8 on 16 Gb dies or 1Rx8 on 32 Gb dies, and Intel’s own table rates those two differently in the same machine. Nothing short of the label or the SPD settles it.

Whether a stacked module is stacked is not in the capacity. A 64GB DDR4 registered module is monolithic, dual-die or 2-high stacked depending on when it was built, and the three consume different numbers of chip selects. The organisation string says so and the capacity does not.

A part number does not always decode. Samsung, Micron and Kingston publish enough structure to work with. Several other vendors do not, and modules built on a shared raw card and sold under a system vendor’s own number often decode to nothing at all.

Rank mixing rules are per platform and contradictory across platforms. AMD forbids mixing x4 and x8 in a channel; Supermicro permits it on Intel. There is no rule that is true everywhere, so a general answer to “can I mix these” is always wrong somewhere.

The speed you will get is not the speed on the label. The label is the module’s bin. The population table is the channel’s rate. dmidecode shows both, after you have bought it.

Whether a machine accepts stacking is a processor property, not a board property. Two machines with the same model number and different processors have different memory ceilings, and the listing for the machine quotes the better one.

Where this site cannot source a field, it records it as not stated rather than inferring it, for the reasons on the methodology page. A guessed rank count would put a module in the wrong compatibility list, which is worse than an empty field, because an empty field makes you go and look.

What to do on a listing page

  1. Read the organisation string before the capacity. 2Rx4, 1Rx8, 2S2Rx4: it settles chip-select consumption, device width and stacking in five characters, and it is the fact most listing titles leave out. If the listing does not show it, ask for a photograph of the paper label rather than of the heatspreader.
  2. Count ranks against the channel, not modules against slots. A channel supports a limited number of ranks. Work out what your machine’s budget is, subtract what is already in it, and only then look at capacities. The arithmetic for the whole machine is in why your server won’t take its maximum RAM.
  3. With one module per channel, prefer more ranks. Nothing derates, and the second rank buys refresh overlap that no access pattern can defeat. Compare what a given capacity costs in each organisation before deciding, starting from DDR4 server memory or DDR4 registered ECC.
  4. With two modules per channel, rank count picks the speed bin. A 1R and a 2R module of the same capacity are not interchangeable there, and the difference is the difference between two rows of the board’s population table. Check the table before deciding which to buy, not after.
  5. At the top capacities, buy the construction the processor lists. 3DS support is a memory-controller property. Divide the platform’s headline capacity by its slot count to see which module size the ceiling was written against, and if that module is stacked, a non-stacked module of the same capacity will not get you there. 64GB DDR4 is the capacity where this bites first.
  6. Do not assume two modules of the same capacity are the same part. Four different 64GB DDR4 registered constructions shipped. They differ in chip selects, loading, latency and speed grade, and their listing titles are identical.
  7. Match what is already in the machine, or replace all of it. Interleaving wants every channel to carry the same type, capacity and rank count. Adding one module to a balanced machine is the upgrade most likely to leave you with more memory and less bandwidth, and nothing will tell you it happened.
  8. Check the machine before the module, every time. The compatibility pages at memory by server and memory by processor start from the machine, which is the direction that works. Dell servers are the largest of those families today.
  9. Use the part page when a part number is all you have. RAM by part number holds what this site has been able to source for a part, and an empty field there means nobody could confirm it, not that it does not matter.
  10. Compare on price per gigabyte across the whole generation before narrowing to a construction. All RAM prices ranks everything the site has seen, and the cheapest route to a capacity is frequently a different rank arrangement than the one you started out looking for.

Two 64GB modules, four constructions, one listing title. The difference between them is two numbers printed on a sticker, and every machine that will refuse one of them has told you so in advance, in a table indexed by exactly those two numbers.

Related guides

Affiliate disclosure:We are a member of the eBay Partner Network and earn a commission from qualifying purchases made through links to eBay on this site. Prices and availability are captured periodically and may have changed - the live price is always the one shown on eBay.