Technical report · 8 Sep 2026

The Memory Wall

Sixty Years of Memory Scaling, 1966–2026: Why Arithmetic Became Almost Free and Moving a Bit Did Not

A technical report. Draft for publication, September 2026.


Reading guide. Part I states the problem in one page and checks the standard explanation. Part II is the sixty-year history, era by era. Parts III–VI go deep on the four "walls" that compound into the memory wall: the DRAM cell, the on-chip wire, the off-chip link, and the memory-system organization. Part VII covers the architectural response (caches everywhere), Part VIII the compute side of the divergence, and Part IX projections through 2040. Appendices contain derivations and data tables. All numbers are order-of-magnitude engineering figures unless a source is named; the references section identifies which figures come from which primary sources. Two companion figures are embedded below: Figure 1 (DRAM density, DRAM latency versus processor clock period, and their ratio, 1970–2026) at the end of the executive summary, and Figure 2 (energy per arithmetic operation versus energy per memory access) in Section 1.4. Tap either to open it full size.


Executive summary

Between 1966, when Robert Dennard invented the one-transistor DRAM cell, and 2026, when the first HBM4 stacks entered mass production, the density of memory improved by a factor of roughly thirty million: from 1,024 bits on Intel's 1103 die (1970) to 32 gigabits on a DDR5 die. Over the same period the cost per bit fell by a similar factor, about 10⁷.

But two other properties of memory barely moved:

The root cause is a mismatch between two kinds of objects. A transistor is a device whose speed and energy improve when it shrinks, because its parasitic capacitances shrink with it. A memory access is not a transistor: it is a long wire (a wordline, a bitline, a bus across a die, a trace across a package and a board) plus a charge-sensing operation plus a protocol. Long wires do not improve when the technology shrinks. Their capacitance per millimeter stays fixed at about 0.2 fF/µm, their resistance per millimeter gets worse as the square of the shrink factor, and their lengths are set by the size of the things they connect (dies, packages, boards), which did not shrink. When supply voltages stopped falling around 2005, the last lever on wire energy — the V² term in ½CV² — was lost as well.

Four distinct physical and organizational effects compound into what is loosely called "the memory wall":

  1. The cell wall. A DRAM cell must store enough charge (tens of femtocoulombs) to be sensed reliably and to survive radiation and leakage for 64 ms. That charge floor fixed the storage capacitance at 20–30 fF for thirty years and forced the capacitor to grow vertically into pillars with aspect ratios approaching 50:1. It also forces the access transistor to leak four orders of magnitude less than a logic transistor, which requires a high threshold voltage and a boosted wordline of ~3 V that cannot follow logic-voltage scaling.
  2. The on-chip wire wall. Reverse scaling of interconnect RC, copper's resistivity explosion below ~20 nm linewidths, the stall of low-k dielectrics at k≈2.5–3, and the end of voltage scaling together mean that the delay and energy of a fixed-length wire have been flat or worsening since roughly 2000.
  3. The off-chip wall. Driving a signal through a package, across a board, and into another package means charging picofarads instead of femtofarads and terminating transmission lines. Pins and package perimeter grew ~20× in forty years while transistor counts grew ~10⁵×. Every interface generation since 2013 (HBM, LPDDR5, GDDR7, HBM4) has been an exercise in shortening the wire rather than speeding up the transistor.
  4. The organization wall. To amortize the cost of the wires, memory is accessed in large rows (8 kilobits per chip per activation), refreshed periodically (a rising tax with capacity), and protected against interference that gets worse as cells get closer (Rowhammer). Each is a fixed or rising overhead that density scaling made worse, not better.

The architectural response has been the memory hierarchy. In 1968 the IBM System/360 Model 85 had one cache. In 1985 a personal computer had none. In 2008 three levels became standard. In 2026 a server has a micro-op cache, L1, L2, a die-spanning L3, optionally a stacked SRAM "V-Cache," on-package HBM or LPDDR, off-package DDR5, CXL-attached memory, NVMe flash, and networked storage: eight to ten tiers, each existing because the tier below it is too slow or too expensive to reach directly. The number of tiers is roughly the logarithm of the gap, and the gap keeps widening: on the AI side, peak accelerator FLOPS have grown about 3× every two years while memory bandwidth has grown about 1.6× (Gholami et al., 2024).

Projections through 2040 (Part IX) contain no physical mechanism that reverses the trend. Latency will not improve materially; DRAM row cycles will remain at 45–50 ns and 3D DRAM may make them slightly worse. Energy per bit will improve only where the wire gets shorter: hybrid-bonded 3D stacking (compute-to-memory distance from centimeters to tens of micrometers) and, later, optical interconnect (energy independent of distance) are the only two levers with an order of magnitude in them. Bandwidth will keep growing by widening interfaces and stacking dies, at rising cost: the 2025–2026 DRAM price shock, in which conventional DRAM contract prices roughly doubled in two quarters because HBM absorbed wafer capacity, is the first sign that the memory wall has acquired an economic dimension. The most reliable near-term relief will continue to come from algorithms that trade cheap arithmetic for scarce memory traffic: tiling, quantization, recomputation, and IO-aware kernels.

Three stacked log-scale panels, 1970 to 2026. (a) DRAM bits per die rise from 1 Kbit to 32 Gbit, about 3×10⁷×, with the doubling period slowing from 4× every three years to 2× every four to five years. (b) DRAM random-access time falls only about 10×, from 300 ns to 30 ns, and goes flat after 2005, while the processor clock period falls about 5,000×. (c) The ratio of DRAM access time to clock period rises from well under 1 in 1970, crosses 1 around 1980, and reaches 150–200 by 2026.
Figure 1. The divergence, 1970–2026. Density scaled; latency did not. A DRAM access cost about one clock period in 1980 and costs 150–200 today on the array timing alone — 300–500 in a loaded system. Approximate best-in-class values from datasheets, JEDEC timing bins and the Hennessy & Patterson DRAM survey.

Part I — The asymmetry, stated precisely

1.1 What scaling did for a transistor

Dennard's 1974 constant-field scaling rule (Dennard et al., IEEE JSSC, Oct. 1974) says: shrink every linear dimension of a MOSFET by a factor κ, reduce the supply and threshold voltages by κ, and increase the doping by κ. The electric fields inside the device stay the same, so the device keeps working, and every figure of merit improves:

Quantity Scaling with κ Meaning
Gate length, width, oxide thickness 1/κ smaller
Supply voltage V 1/κ lower
Transistors per unit area κ² denser
Gate capacitance C 1/κ smaller load
Drive current I 1/κ less current, but...
Switching delay ∝ CV/I 1/κ faster
Power per transistor ∝ IV 1/κ² cooler
Power density 1 constant — the miracle
Energy per switching event ∝ CV² 1/κ³ the deep source of cheap arithmetic

For thirty years (1975–2005) κ ≈ 1.4 arrived every two to three years. Cumulatively, from the 3-µm NMOS of 1980 to the 3-nm-class nodes of 2026, the energy to switch one logic gate fell by roughly four to five orders of magnitude. Arithmetic is made of gates and short local wires whose length shrinks with the gates. So arithmetic rode the whole curve.

Note what Dennard scaling does not say anything about: the wires that connect distant things, and the analog problem of storing and sensing charge.

1.2 What scaling did for a wire

Consider a wire of width W, thickness H, length L, made of a metal with resistivity ρ, surrounded by dielectric of permittivity ε. Shrink W, H and the spacing to neighbors by κ.

Quantity Scaling Why
Resistance per unit length, r = ρ/(WH) κ² cross-section shrinks quadratically
Capacitance per unit length, c ≈ 1 c depends on ratios (W/spacing, H/spacing), which are unchanged; c ≈ 0.15–0.25 fF/µm in every technology since the 1980s
RC delay per unit length, rc κ² worse
Local wire (length shrinks with κ): delay ∝ rc·L² 1 constant, while gates got 1/κ faster → wires fall behind by κ
Global wire (length fixed by die size): delay ∝ rc·L² κ² absolutely worse; relative to gate delay, κ³ worse
Energy to switch the wire, ½ c L V² ∝ L·V² falls only if V falls or L falls

This is the "reverse scaling" of interconnect identified by Mark Bohr at IEDM 1995 and analyzed exhaustively in Ho, Mai and Horowitz, "The Future of Wires" (Proc. IEEE, 2001). Its consequences:

1.3 What scaling did for a DRAM cell

A DRAM cell is a capacitor C_S holding a charge Q = C_S·V, read by dumping that charge onto a bitline of capacitance C_BL and sensing the resulting voltage step:

ΔV_BL = (V_core / 2) · C_S / (C_S + C_BL)

Sense amplifiers need ΔV_BL of roughly 100 mV to overcome their own offset and noise. Retention requires the charge to survive 64 ms of leakage. Radiation immunity requires the charge to exceed what an alpha particle or cosmic-ray neutron can deposit in the cell. Together these fix the charge floor at tens of femtocoulombs — a few tens of thousands of electrons — independent of the technology node. The industry rule for three decades was "30 fF per cell" and it has only recently relaxed toward 10–15 fF as bitlines got shorter.

So while the cell's footprint scaled with lithography (6F² at F ≈ 11–12 nm in 2026), its capacitance could not. The capacitor was pushed into the third dimension: trench capacitors dug into the substrate in the 1980s, stacked capacitors above the transistor in the 1990s, and today cylinders ~30 nm wide and more than a micrometer tall with zirconia-based high-k dielectrics a few nanometers thick.

The access transistor faces the mirror-image problem. To hold ~10 fC for 64 ms with acceptable loss, its off-state leakage must be below ~50 fA. A logic transistor at an advanced node leaks ~1–10 nA per micrometer of width, i.e., ~20–200 pA for a DRAM-sized device. The DRAM transistor must therefore leak 10³–10⁴× less than a logic transistor, which means a high threshold voltage (~0.8–1 V), a long effective channel (recessed or buried), and a wordline that swings to ~2.5–3 V (V_PP) to turn it on hard. The DRAM array's wordline voltage has not scaled since the 1990s. Every activation charges a kilometer-scale total length of wordline and bitline capacitance to voltages that logic left behind twenty years ago.

1.4 The budget of one operation versus one access

Mark Horowitz's ISSCC 2014 keynote, "Computing's Energy Problem," tabulated energies in a 45-nm, 0.9-V process (the numbers are per operation or per 64-bit access; see Figure 2):

Operation (45 nm) Energy
8-bit integer add 0.03 pJ
32-bit integer add 0.1 pJ
32-bit floating-point add 0.9 pJ
32-bit integer multiply 3.1 pJ
32-bit floating-point multiply 3.7 pJ
64-bit read from an 8 KB SRAM 10 pJ
64-bit read from a 32 KB SRAM 20 pJ
64-bit read from a 1 MB SRAM 100 pJ
64-bit read from DRAM (system level) 1,300–2,600 pJ
Horizontal log-scale bar chart of energy per operation at 45 nm, 0.9 V. Arithmetic bars run from a 0.03 pJ 8-bit integer add to a 3.7 pJ 32-bit floating-point multiply. On-chip SRAM reads run 10 pJ for 8 KB, 20 pJ for 32 KB and 100 pJ for 1 MB. An off-chip 64-bit DRAM read is 1,300–2,600 pJ, far off the right of every other bar.
Figure 2. The energy ladder (Horowitz, ISSCC 2014). Blue: arithmetic. Orange: on-chip SRAM, already wire-dominated — a 1 MB read costs 27× a 32-bit floating-point multiply. Red: off-chip DRAM, including the interface. A DRAM read costs ~350–700× a 32-bit floating-point multiply and ~50,000× an 8-bit add.

The DRAM access costs 350–700× the floating-point multiply and about 50,000× the 8-bit add. Between 45 nm (2008) and 5 nm-class nodes (2020s), the arithmetic energies fell another ~5–10×; the DRAM energy fell by a comparable factor only where the memory was moved onto the package (HBM). The ratio has not closed.

The time budget looks the same. A 2026 server core executes ~4–6 instructions per cycle at ~0.2 ns per cycle. A DRAM access that misses every cache takes 80–110 ns in a loaded system: 400–500 cycles, or 2,000+ instruction slots. Little's law says that keeping the core busy across such a miss requires ~2,000 instructions in flight; the largest reorder buffers hold 500–700. Out-of-order execution alone cannot bridge the gap; only caches, prefetching and massive multithreading (GPUs) can.

1.5 Checking the standard explanation

The explanation you were given (the "fact-checker" contrasting a 2D-versus-1D intuition pump with interconnect physics) is largely right. Three corrections and two additions:

  1. The 2D/1D intuition pump is not wrong, just incomplete. It correctly captures that transistor count grows as κ² while the length of the longest wires does not shrink, because die and package sizes did not shrink (die area stayed at the reticle limit of ~800 mm² for high-end parts; a DIMM is still about 13 cm long). The complete statement is: energy = length × (energy per length), and both factors stopped improving — length for geometric reasons, energy per length for the physical reasons the explanation lists.
  2. "Capacitance per unit length stays flat" is correct, with one lever that also stalled. c can be reduced only by lowering the dielectric constant. The industry went from SiO₂ (k = 3.9, through the 180-nm node) to fluorinated glass (3.6) to carbon-doped oxide (~2.9) to porous ultra-low-k (~2.4–2.6) and, at a few layers, air gaps. Porous dielectrics are mechanically weak and absorb moisture; k has been stuck near 2.5–3.0 since the 32/28-nm nodes (~2011). So even the "flat" c was supposed to fall and did not.
  3. The V_dd floor is a subthreshold-slope limit, not a thermal-noise limit. A MOSFET's off-current falls by at most one decade per (kT/q)·ln 10 ≈ 60 mV of gate voltage at room temperature, because carriers surmount the channel barrier with a Boltzmann distribution. Getting 10⁴–10⁵ of on/off ratio therefore costs 0.25–0.35 V of threshold, and the transistor needs another ~0.3–0.4 V of overdrive to be fast. That is the origin of the 0.6–0.9 V floor. Thermal noise (kT/C) would allow operation far lower — subthreshold circuits run at 0.2–0.3 V, and the theoretical minimum supply for a CMOS inverter with gain is only ~36 mV at 300 K (Swanson–Meindl) — but they are 10–100× slower. Also, ruthenium is used as a liner and, increasingly, as a barrier-free replacement conductor; tantalum nitride is the barrier that ruthenium replaces.
  4. Addition: the memory cell itself is a wall. Interconnect physics explains why moving a bit is expensive, but not why DRAM latency stalled at ~30 ns of access and ~48 ns of row cycle. That comes from charge sensing, the retention/leakage constraint on the access transistor, and the unscaled wordline voltage (Section 1.3, Part III).
  5. Addition: the wall is stacked. On-chip wires (Part IV) are one layer. Package and board signaling (Part V) is another with its own physics (transmission lines, termination, pad capacitance in picofarads). The organization of the memory system (Part VI: row overfetch, refresh, bank parallelism, Rowhammer) is a third. The hierarchy (Part VII) is the response to all of them at once.

Part II — Sixty years, era by era

2.0 Timeline

Year Event Why it matters for the wall
1953 Magnetic-core memory in MIT's Whirlwind (Forrester) Main memory ~1 µs; the standard for two decades
1962 Manchester Atlas: virtual memory and paging The idea of a hierarchy managed automatically
1965 Moore's "Cramming more components" article; Wilkes's "Slave memories" paper Density trend named; cache concept invented
1966–68 Dennard invents the 1T1C DRAM cell (IBM; patent filed 1967, granted 1968) The cell that would scale for sixty years
1968 IBM System/360 Model 85: first commercial cache (16–32 KB, 80 ns) against ~1 µs core The first machine built around a memory wall
1969 Intel 3101 (64-bit bipolar SRAM) and 1101 (256-bit MOS SRAM) Semiconductor memory becomes a product
1970 Intel 1103, 1 Kbit DRAM, 3-transistor cell, ~300 ns access Kills core by 1972–74
1973 Mostek MK4096 (4 Kbit): multiplexed row/column addresses (RAS/CAS) Halves pin count; locks in the two-step row/column access still used in DDR5
1974 Dennard scaling paper Thirty years of free transistor improvement begin
1976 Mostek MK4116 (16 Kbit) becomes the industry standard part Address multiplexing wins
1978–79 Alpha-particle soft errors discovered (May & Woods, Intel) Sets the critical-charge floor for cells
1979–81 64 Kbit generation; single +5 V supply; Japanese vendors take the lead DRAM becomes a commodity manufactured by the best fab, not the best designer
1984 Motorola 68020: first mainstream microprocessor with an on-chip (256-byte) cache Caches return to the chip
1985–86 Intel exits DRAM; 1 Mbit generation; trench (TI, IBM) vs. stacked (Hitachi, Fujitsu, NEC) capacitors Cell capacitor goes 3D because it cannot shrink
1989 Intel 80486: 8 KB on-chip L1 cache Every microprocessor now needs a cache
1990 Hennessy & Patterson, Computer Architecture: A Quantitative Approach The "processor–memory performance gap" chart
1992–93 Samsung becomes #1 in DRAM; first synchronous DRAM (Samsung KM48SL2000) Interfaces become clocked and pipelined
1995 Wulf & McKee, "Hitting the Memory Wall"; Pentium Pro with in-package L2; Alpha 21164 with on-die L2 The wall is named; two-level on-package hierarchies
1997 IBM ships copper interconnect (CMOS 7S); Tera MTA multithreaded supercomputer Wire resistance relief (one-time); latency tolerance via threads
2000 DDR SDRAM standard; Rambus RDRAM with Pentium 4 The Rambus–DDR war: cost and openness beat raw speed
2001 Ho, Mai & Horowitz, "The Future of Wires" The definitive interconnect-scaling analysis
2003 AMD Opteron/Athlon 64: memory controller integrated on the CPU die Removes one chip crossing (~30–40 ns) from every access
2004 Patterson, "Latency Lags Bandwidth"; Intel cancels Tejas; Dennard scaling ends Voltage stops falling → wire energy stops falling
2006 Berkeley View: "power wall + memory wall + ILP wall = brick wall" Multicore era begins
2007 DDR3; Intel 45 nm high-k metal gate; first iPhone (LPDDR) Mobile memory becomes its own scaling track
2008 Intel Nehalem: integrated memory controller + shared L3; buried-wordline DRAM (Qimonda) Three cache levels become standard
2009 Roofline model (Williams, Waterman, Patterson) Memory-boundedness gets a picture
2011–13 Micron HMC (2011); JEDEC HBM standard (Oct. 2013); Samsung 3D V-NAND; Haswell eDRAM L4 Memory goes vertical (stacking) and L4 appears
2012–13 Elpida bankruptcy; industry consolidates to Samsung, SK hynix, Micron Three suppliers for the world
2014 Horowitz's ISSCC keynote; Rowhammer paper (Kim et al.); DDR4 Energy numbers and cell-interference failures published
2015 AMD Fury: first product with HBM; Intel/Micron announce 3D XPoint Memory on the package; storage-class memory attempted
2017 Nvidia Volta tensor cores; O'Connor et al. "Fine-Grained DRAM" Arithmetic gets even cheaper; DRAM overfetch quantified
2019 CXL 1.0; UPMEM processing-in-memory DIMMs; Optane DIMMs ship Disaggregation and near-data compute, both as stopgaps
2020 DDR5 standard; Apple M1 with unified on-package LPDDR Memory on package goes mainstream
2021 Samsung HBM-PIM; EUV in DRAM (1α); Cerebras WSE-2 (40 GB on-wafer SRAM) Compute in memory; memory as the whole chip
2022 AMD 3D V-Cache (hybrid-bonded SRAM); HBM3 in Nvidia H100; Intel kills Optane; FlashAttention Cache scaling by stacking; storage-class memory fails; algorithms adapt
2024 HBM3E; GDDR7; Cerebras WSE-3 (44 GB); Gholami et al. "AI and Memory Wall" Compute 3×/2yr vs. bandwidth 1.6×/2yr documented
2025 JEDEC HBM4 spec (April); SK hynix 30-year DRAM roadmap (4F² VG, 3D DRAM); DRAM prices begin surging The road past 6F² is charted; the economic wall appears
2026 HBM4 mass production (Samsung Feb., Micron Q1, SK hynix); Nvidia Rubin; TSMC A16 with backside power; DRAM contract prices +90% QoQ in Q1 Where this report is written

2.1 Prehistory and the first wall (1950–1968)

Core memory was the first random-access memory that worked at scale. A core plane stored one bit per ferrite ring, read destructively by driving current through it, and cost about $1 per bit in the mid-1950s falling to a cent per bit by 1970. Its cycle time, ~1 µs, was set by the physics of magnetization reversal and the inductance of long drive lines — a wire problem before there were transistors to compare it with.

Processors improved faster. By the mid-1960s, transistorized CPUs had cycle times well under 100 ns while core stayed near 1 µs. Maurice Wilkes's two-page 1965 note "Slave Memories and Dynamic Storage Allocation" proposed a small fast memory that would automatically hold recently used words of the large slow one. IBM built it: the System/360 Model 85 (announced 1968) placed a 16 KB (expandable to 32 KB) "high-speed buffer" between an 80-ns processor and 1.04-µs core. J. S. Liptay's 1968 IBM Systems Journal paper on the Model 85 is where the word "cache" entered the vocabulary. The gap the cache bridged was 13×; the design goal was to make the machine behave as if all of memory ran at buffer speed, and simulations showed hit rates above 95% on real workloads.

The point of recording this is that the memory wall predates the microprocessor. It was there when memory was magnetic and processors were discrete transistors, because it is a property of distance, not of any particular technology.

2.2 The semiconductor memory revolution (1968–1980)

Semiconductor memory did not begin as DRAM. Intel's first products (1969) were static RAMs: the 3101 (64 bits, bipolar) and the 1101 (256 bits, MOS). SRAM was easy to use but consumed six transistors per bit.

Dennard's 1T1C cell — one transistor, one capacitor, invented at IBM in 1966 — was the minimum conceivable. Intel's 1103 (October 1970) used an intermediate three-transistor cell designed by William Regitz at Honeywell and productized by Joel Karp at Intel: 1,024 bits, PMOS, roughly 300 ns access, and notoriously finicky to use (it needed precharge clocks with tight timing, and it needed to be refreshed every two milliseconds). It was also about a cent per bit — the price of core, at a fraction of the space and power. By 1972 the 1103 was the best-selling semiconductor memory in the world and core was finished.

Two decisions made in 1973 still shape every memory access in 2026. Mostek's MK4096 (4 Kbit) was the first DRAM with a true 1T1C cell in volume and the first to multiplex its address pins: the row address is presented first with the Row Address Strobe (RAS), then the column address with the Column Address Strobe (CAS). Robert Proebsting's multiplexing cut the package from 22 to 16 pins and roughly halved the cost of the package and the board, at the cost of forcing a two-step access. Intel and TI initially refused to follow; by the 16 Kbit generation (Mostek's MK4116, 1976) multiplexing had won. The tRCD (RAS-to-CAS delay) that every memory controller still schedules around is Proebsting's 1973 decision, and the row-buffer structure it implies — open a whole row, then pick columns from it — is the reason DRAM access is a page-oriented, overfetching operation to this day.

The 64 Kbit generation (1979–81) brought a single +5 V supply and the arrival of Japanese manufacturers (Fujitsu, Hitachi, NEC, Toshiba, Mitsubishi) with yields and quality the American firms could not match. And in 1978–79 came the first sign that the cell had a physics floor: Intel's Tim May and Murray Woods traced a mysterious rise in soft errors in 16 Kbit parts to alpha particles emitted by trace uranium and thorium in the ceramic package material (the contamination was ultimately traced upstream to a mining region near the ceramic supplier, as the story is usually told). A single alpha particle deposits tens of femtocoulombs along its track in silicon. A cell holding less than a "critical charge" of that order would flip. From then on, DRAM designers had a number they could not scale below.

2.3 The microprocessor era and the gap (1980–1995)

For a few years the wall disappeared. An 8086 at 5 MHz had a 200 ns cycle and took four cycles per bus access; a 150-ns DRAM kept up with one wait state or none. Early PCs had no cache because they needed none. Cycle times fell fast, though: 80286 (1982) at ~100 ns, 80386 (1985) at 62 ns, 80486 (1989) at 40 ns, Pentium (1993) at 15 ns. DRAM access went from ~150 ns to ~70 ns in the same period. In cycles, a miss went from about 1 to about 10.

The caches came back. Motorola's 68020 (1984) carried a 256-byte instruction cache — tiny, but the first on-chip cache in a mainstream microprocessor. Intel's 386 relied on external SRAM caches (the 82385 controller, 1987); the 486 (1989) integrated 8 KB. RISC designs (MIPS R2000, 1986; SPARC) were built around split instruction/data caches from the start, because their whole premise — one instruction per cycle — was impossible without them.

DRAM in the same period moved from 256 Kbit to 16 Mbit, from NMOS to CMOS, and the capacitor went 3D. Two camps formed in the 1 Mbit and 4 Mbit generations: the trench capacitor etched into the substrate (Texas Instruments, IBM, later Toshiba and the Infineon/Qimonda lineage) and the stacked capacitor built above the transistor (proposed by Mitsumasa Koyanagi at Hitachi in 1978; adopted by Hitachi, Fujitsu, NEC, Mitsubishi, and Samsung). Both were responses to the same fact: at a 1-µm feature size a planar capacitor could no longer provide 30 fF. Stacked capacitors won by the 2000s because they scaled better with height. Meanwhile the manufacturing center of gravity moved from the US (Intel exited DRAM in 1985; Micron alone survived) to Japan (over half of world supply by 1986) and then to Korea (Samsung took the #1 position in 1992 and has held it, except briefly in 2025 when SK hynix's HBM business pushed it past Samsung in revenue).

In 1990 Hennessy and Patterson's textbook drew the chart that defined the field's anxiety: processor performance improving 35% per year until 1986 and 55% per year after; DRAM latency improving 7% per year. Two exponentials with different exponents diverge without limit. William Wulf and Sally McKee's four-page note "Hitting the Memory Wall: Implications of the Obvious" (Computer Architecture News, March 1995) did the arithmetic: with those rates, even a cache with a 99% hit rate would see its average access time dominated by misses within about a decade, and no amount of hit-rate improvement could help because the miss cost was growing without bound. They expected the wall to bite around 2005–2010 and asked, rhetorically, whether there was any way around it.

2.4 Naming the wall and building around it (1995–2005)

The industry's answer was: every trick at once.

Latency tolerance in the core. Out-of-order execution (Tomasulo's 1967 algorithm from the IBM 360/91, revived in the PowerPC 604, MIPS R10000 and Pentium Pro in 1995) lets independent instructions proceed while a load waits. Non-blocking caches (Kroft, 1981) allow multiple outstanding misses. Hardware prefetchers guess the next address. Simultaneous multithreading (Tullsen et al., 1995; Intel Hyper-Threading, 2002) lets another thread use the idle issue slots. The most radical answer was Burton Smith's Tera MTA (1997), which had no data cache at all and hid a 150-cycle memory latency by switching among 128 hardware threads every cycle. It was a commercial failure and a conceptual success: it is, in outline, how every GPU works.

More cache, closer. Intel's Pentium Pro (1995) put a 256 KB–1 MB L2 SRAM die in the same package on a private full-speed bus, an expensive dual-cavity ceramic package that Intel abandoned for the Pentium II cartridge (1997) and then re-absorbed onto the die (Celeron 300A, 1998; Pentium III Coppermine, 1999). DEC's Alpha 21164 (1995) had 96 KB of L2 on the die and a third level off it: the first three-level hierarchy in a microprocessor. Itanium 2 (2002) reached 3 MB of on-die L3, later 9 MB.

Faster interfaces, not faster cells. Synchronous DRAM (JEDEC, 1993; PC66/PC100 by 1998) put a clock on the bus and pipelined column accesses. Rambus's RDRAM ran a narrow bus at 400–800 MHz, and Intel bet the Pentium 4 platform on it (1999–2001); the Rambus–DDR war ended with double-data-rate SDRAM (JEDEC, 2000) winning on cost, openness and litigation fatigue. The key fact is that none of these interfaces made the DRAM array faster. Bandwidth per pin rose from ~100 Mb/s (PC100) to 400 Mb/s (DDR-400); tRCD stayed at 15–20 ns.

Copper and low-k. IBM shipped copper interconnect in 1997 (CMOS 7S); everyone followed by 2000. Copper's resistivity (1.7 µΩ·cm) is 40% below aluminum's, and low-k dielectrics started to appear at 130 nm. This was a one-time gift; Part IV explains why it ran out.

Integrating the memory controller. AMD's Opteron (2003) moved the memory controller onto the CPU die, with HyperTransport links between sockets. This removed a full chip crossing (the "northbridge") from every access — the single largest latency improvement of the decade, roughly 30–40 ns — and it is the reason modern latency is 80–100 ns rather than 130–150 ns. Intel followed with Nehalem in 2008.

And then voltage stopped falling. By 2004 Intel's 90-nm Prescott Pentium 4 dissipated over 100 W at 3.8 GHz; its successor Tejas was cancelled. Leakage current, which the subthreshold slope bounds from below, had made further threshold-voltage reduction impossible, and without threshold reduction the supply could not fall without losing speed. Dennard scaling ended. Transistors kept shrinking; their energy stopped falling as κ³ and fell as roughly κ instead. For wires — whose only remaining lever was V² — energy per millimeter stopped falling altogether.

2.5 Multicore and the bandwidth wall (2005–2015)

With frequency capped near 3–4 GHz, performance came from parallelism: dual-core in 2005–06, quad-core by 2007, and manycore GPUs in general-purpose use after CUDA (2007). Parallel cores multiply memory bandwidth demand. Rogers et al. ("Scaling the Bandwidth Wall," ISCA 2009) showed that with pin bandwidth growing far slower than core count, the number of cores that could actually be fed would saturate within a few generations unless caches grew or compression and other tricks were used.

The DRAM interface responded by getting wider and faster on the same array: DDR3 (2007) at 800–1,600 Mb/s per pin; DDR4 (2014) at 1,600–3,200; GDDR5 (2008) at 4–7 Gb/s for GPUs. A DDR3-1600 DIMM delivers 12.8 GB/s; a 2026 DDR5-6400 DIMM delivers 51 GB/s; a 2026 HBM4 stack delivers 2,000–3,300 GB/s. Bandwidth improved ~100× in fifteen years. Row-cycle time did not move.

Three developments in this decade set up everything that followed:

2.6 The AI era: the wall becomes the roofline (2015–2026)

Machine learning changed the shape of demand. A matrix multiply has high arithmetic intensity (many operations per byte fetched) when its operands are large and reused; a transformer decoder generating one token at a time has almost none — each weight is fetched from memory, used once, and discarded. Gholami et al. ("AI and Memory Wall," IEEE Micro, 2024) measured the twenty-year trend: peak server FLOPS up 60,000× (3.0× every two years), DRAM bandwidth up 100× (1.6× every two years), interconnect bandwidth up 30× (1.4× every two years). The ridge point of the roofline — the arithmetic intensity at which a machine becomes compute-bound rather than memory-bound — rose from ~1 FLOP/byte on vector supercomputers to ~300 FLOP/byte for FP16 on an H100 (2022: ~1 PFLOP/s against 3.35 TB/s) and over 1,000 FLOP/byte at FP4 on Blackwell (2024: ~9 PFLOP/s against 8 TB/s).

Every hardware development of the decade is an attempt to move memory closer or make arithmetic cheaper, and the second makes the wall worse:


Part III — The cell wall: DRAM in depth

3.1 Density: the curve that worked

Year Bits per die Lead parts / notes Cell technology
1970 1 K Intel 1103 (3T cell, PMOS, ~10 µm) planar, 3 transistors
1973 4 K Mostek MK4096 (1T1C, multiplexed address) planar 1T1C
1976 16 K Mostek MK4116 (three supplies) planar
1980 64 K Fujitsu, Hitachi, NEC, TI, Mostek (+5 V only) planar, ~2–3 µm
1983 256 K NEC, Hitachi, Toshiba planar, folded bitline
1986 1 M Toshiba, Hitachi, TI (1 µm) trench / stacked capacitors appear
1989 4 M Japanese leaders; Samsung enters top tier 3D capacitors universal; CMOS
1992 16 M Samsung #1; first SDRAM stacked capacitor dominant
1996 64 M Samsung, NEC, Hitachi hemispherical-grain poly to boost area
1999 256 M Samsung, NEC, Micron (0.18 µm) Ta₂O₅ / high-k dielectrics begin
2003 512 M DDR/DDR2 era (~90 nm) recess-channel array transistor (Samsung)
2006 1 G DDR2/DDR3 (~70 nm) ZAZ (ZrO₂/Al₂O₃/ZrO₂) dielectrics
2009 2 G DDR3 (~50 nm) buried wordline (Qimonda, then all)
2012 4 G DDR3/DDR4 (~30 nm, "3x nm") cylinder capacitors, aspect ratio >20
2015 8 G DDR4 (~20 nm, "2x/1x") saddle-fin access transistor
2018 16 G DDR4/DDR5 (1y/1z, ~16–17 nm) quad-patterning; aspect ratio >40
2021 16–24 G DDR5 (1α, ~14 nm; EUV enters) 6F² with F ≈ 14 nm
2024 32 G DDR5 (1β/1γ, ~12–13 nm) supporter lattices for pillars
2026 32 G (48–64 G sampling) DDR5, LPDDR5X/6, HBM4 (1c, ~11–12 nm) last 6F² generations; 4F² VCT prototypes

From 1 Kbit to 32 Gbit is a factor of 3.2 × 10⁷. Through the 1990s the industry delivered 4× every three years — Moore's original 1965 observation was, in fact, about memory. After 2010 the cadence slowed to 2× every four to five years, and the "1x/1y/1z/1α/1β/1γ/1c" node names conceal the fact that F has moved only from ~19 nm to ~11 nm in a decade. The remaining planar roadmap is one more node ("1d," ~10 nm, expected 2027–28) before the 6F² cell runs out (SemiAnalysis's VLSI 2025 summary; SK hynix roadmap, VLSI 2025).

3.2 Latency: the curve that did not

Access time is the interval from presenting a row address to receiving data (tRCD + CL in modern terms). Cycle time is the minimum interval between two accesses to the same bank (tRC = tRAS + tRP): the time to open a row, sense it, restore it and precharge the bitlines.

Year Generation Row access (ns) Column access (ns) Row cycle (ns)
1980 64 Kb 170–250 75 250
1983 256 Kb 150–220 50 220
1986 1 Mb 120–190 25 190
1989 4 Mb 100–165 20 165
1992 16 Mb 80–120 15 120
1996 64 Mb (EDO) 70–110 12 110
2000 256 Mb (PC133) 65–90 7 90
2004 1 Gb (DDR/DDR2) 55–70 5 70
2006 2 Gb (DDR2) 50–60 2.5 60
2010 4 Gb (DDR3-1600) tRCD ≈ 14; tRCD+CL ≈ 28 1 48.75
2018 16 Gb (DDR4-3200) tRCD ≈ 14; tRCD+CL ≈ 28 0.6 45.75
2021 16 Gb (DDR5-4800) tRCD ≈ 16; tRCD+CL ≈ 33 0.4 ~48
2026 32 Gb (DDR5-6400 / HBM4) tRCD ≈ 14–16; tRCD+CL ≈ 30 0.3 45–50

(1980–2006 rows adapted from the survey table in Hennessy & Patterson, Computer Architecture: A Quantitative Approach; 2010–2026 rows from JEDEC timing bins for the named speed grades. "Column access" here is the per-word transfer interval, which improved enormously through pipelining and burst transfer; the row numbers are the latency.)

Row access improved 6–8× in 46 years (Figure 1b). Row cycle improved 5×, all of it before 2010. Since then the DRAM array has been flat: every generation of DDR and HBM has raised the clock and widened the bus while the array underneath runs at the same speed, and CAS latency in nanoseconds has actually crept upward with DDR5.

3.3 Why the array is slow: a walk through one activation

  1. Precharge. Bitlines are equalized to V_core/2 (~0.5–0.55 V). The bitline is a wire hundreds of cells long with ~20–40 fF of capacitance.
  2. Wordline activation. The row decoder drives the wordline to V_PP (2.5–3 V). The wordline is a polysilicon/tungsten line with thousands of transistor gates hanging off it; modern designs use hierarchical (main/sub) wordlines and buried wordlines to cut its RC, but its RC delay is still several nanoseconds.
  3. Charge sharing. Each cell's ~10–20 fF dumps its charge onto its bitline. The bitline moves by ΔV = (V_core/2)·C_S/(C_S+C_BL) ≈ 100–200 mV. This is a passive RC process governed by the access transistor's on-resistance and the bitline capacitance.
  4. Sensing. A cross-coupled latch amplifies the differential to full rail. Sense-amplifier offset (tens of millivolts, from threshold-voltage mismatch) sets the minimum ΔV and hence the minimum C_S.
  5. Column access. The requested columns are gated from the sense amplifiers onto local I/O lines, then global I/O lines running across the die, then through the output drivers.
  6. Restore. The sense amplifiers write the full-rail value back into every cell in the row (the read was destructive). tRAS, ~30–35 ns, is dominated by this restore because the capacitor must be charged through the access transistor's resistance.
  7. Precharge again. tRP ≈ 13–15 ns.

Steps 2, 3, 4, and 6 are analog RC processes on long lines; step 5 is a global wire. Nothing in the list is a logic gate whose delay Dennard scaling would have reduced. This is why shrinking F from 1 µm to 11 nm — a factor of 90 — cut tRC by only 5×, and why DRAM vendors never marketed speed: the customer paid for bits.

3.4 The capacitor: forced into the third dimension

The stored charge must satisfy three constraints:

With V_core/2 ≈ 0.55 V, a 15 fF cell holds ~8 fC ≈ 50,000 electrons. The capacitance target fell from ~30 fF (1990s) to ~10–15 fF (2020s) only because shorter bitlines cut C_BL and better sense amplifiers cut the required ΔV.

Meanwhile the cell footprint went from ~1,500 µm² (1103) to ~0.001 µm² (6F² at F ≈ 12 nm): six orders of magnitude. Holding 15 fF in a footprint of ~(2F)² ≈ 600 nm² requires, with a k ≈ 35 dielectric ~6 nm thick, a plate area of about 0.3 µm² — 500× the footprint. Hence the cylinder: a pillar ~30 nm across and 1–1.5 µm tall, with both inner and outer surfaces used, held upright by lattice-like supporters so that it does not topple during processing, and lined with a ZrO₂/Al₂O₃/ZrO₂ dielectric and TiN electrodes. Aspect ratios approach 50:1. This is the physics that ends 6F²: the pillars cannot get taller and thinner indefinitely, the dielectric cannot get thinner without tunneling leakage, and higher-k materials tend to have smaller bandgaps and hence more leakage. The exits are 4F² vertical-channel-transistor cells (one more ~30% density step) and then 3D DRAM, in which cells are laid horizontally and stacked like 3D NAND — trading the aspect-ratio problem for a layer-count problem.

3.5 The access transistor and the wordline that would not scale

The transistor must deliver ~50 fA of off-current and enough on-current to restore the cell within tRAS. The industry solved short-channel leakage at 90 nm with the recess-channel array transistor (Samsung, 2003), then buried wordlines (Qimonda, 2008), then saddle-fin structures, which lengthen the channel below the surface so that a 20-nm-pitch device can behave like a 100-nm-channel device. The price is a wordline that must swing to V_PP ≈ 2.5–3 V (generated by on-die charge pumps at ~30–50% efficiency) and to a negative off-level (~–0.2 V) to suppress leakage. A logic wire swings 0.75 V; energy scales as V², so the wordline's per-unit-capacitance energy is ~16× that of a logic wire — before counting the pump's inefficiency.

3.6 Refresh: a tax that grows with density

JEDEC requires every row to be refreshed within 64 ms (32 ms above 85 °C, because leakage roughly doubles per 10 °C). A DDR4 chip issues an all-bank refresh every 7.8 µs; each takes tRFC ≈ 350 ns at 8 Gb and 550 ns at 16 Gb, during which the chip cannot serve requests — 4.5% to 7% of all time, and rising with capacity. Liu et al. (2012) showed that at 64 Gb per chip, unmodified refresh would consume roughly half of throughput and half of energy. DDR5 answers with per-bank and same-bank refresh, more banks (32 per chip), and on-die ECC to tolerate the weakest cells; HBM adds per-pseudo-channel scheduling. Each is an organizational patch over the same physics: charge leaks, and there are more cells every generation to top up.

3.7 Rowhammer and interference: density creates new failure modes

When rows sit ~20 nm apart, the field from an activated wordline and the charge injected by its transistors disturb neighboring cells; repeated activation drains them before their refresh. Kim et al. (2014) needed ~139,000 activations in 64 ms to flip a bit in DDR3; by 2020 (Kim et al., "Revisiting RowHammer," ISCA 2020) LPDDR4 parts flipped at ~4,800. Mitigations (target-row refresh, refresh-management commands in DDR5, per-row activation counters, and in 2024–26 in-DRAM trackers) all cost area, bandwidth or latency. The same story — variable retention time, sense-amp mismatch at small sizes, and the need for on-die ECC in DDR5 — shows density scaling actively degrading the parameters that determine speed and energy.


Part IV — The on-chip wire wall

4.1 Reverse scaling in numbers

Take the 2001 "Future of Wires" analysis forward to 2026. A minimum-pitch wire at the 180-nm node (1999) was about 250 nm wide and 500 nm tall; at a 3-nm-class node (2023) the tightest metal pitch is ~20–24 nm, so the wire is ~10–12 nm wide. The cross-section fell by ~500×, so resistance per length rose ~500× for the same metal. Capacitance per length stayed at 0.15–0.25 fF/µm. A 1-mm wire that had ~1 ns of unrepeated RC delay in 1999 would have ~500 ns today; in practice it is broken into segments a few tens of micrometers long with a repeater at each break, and its effective delay is set by the repeaters (a few tens of picoseconds per segment). A 2004 study (Saxena et al., IEEE TCAD) projected that repeaters alone could occupy tens of percent of a chip's cells by the 32-nm node; the estimate proved roughly right for high-frequency designs.

4.2 Copper runs out

Copper's advantage over aluminum was 40% in bulk resistivity (1.7 vs 2.7 µΩ·cm). At small dimensions two size effects erase it. The electron mean free path in copper is ~39 nm at room temperature; when a wire's dimensions approach that, electrons scatter off the surfaces (Fuchs–Sondheimer) and at grain boundaries, which become more frequent as grains are forced smaller (Mayadas–Shatzkes). At 10–15 nm linewidths, effective resistivity is 4–6× bulk. On top of that, copper needs a diffusion barrier (TaN) and a liner (Ta, Co, Ru) totalling 3–5 nm, which do not scale and consume half the cross-section of a 12-nm line. Net: a leading-edge M0/M1 line has ~8–10× the resistivity of the bulk copper the industry adopted in 1997.

The responses have been incremental: cobalt at the lowest layers (Intel, 2018); ruthenium as a liner and, at the tightest pitches, as a barrier-free replacement conductor (its bulk resistivity is 7 µΩ·cm, worse than copper's, but its mean free path is ~7 nm so it stops degrading where copper has already lost); molybdenum for similar reasons; subtractive etching instead of damascene fill; and air gaps at selected layers. Backside power delivery (Intel's PowerVia in 18A, shipping in 2026 products; TSMC's Super Power Rail in A16, entering production in 2026) moves the power grid to the back of the wafer, which relieves routing congestion and IR drop but does nothing for signal-wire RC — Intel's ISSCC 2025 paper noted that putting backside power inside an SRAM cell would actually enlarge it.

4.3 Low-k stalls

The capacitance term has only one knob, k. The industry moved from SiO₂ (3.9) to FSG (3.6, ~2001) to SiCOH (2.9–3.0, ~2004) to porous SiCOH (2.4–2.6, ~2009). Below about 2.4, porous films crack during chemical-mechanical polishing, delaminate in packaging, and absorb moisture that raises k back up. Progress stopped near 2.5. Air gaps (k ≈ 1) are used at a few layers where geometry permits. Effective k in a 2026 stack is about the same as in 2011.

4.4 Voltage stalls — the master lever

Everything above concerns delay and the c term. The energy of a wire is ½cLV², and after 2005 V froze. Supply voltage fell from 5 V (1980s) to 3.3 V (1993), 2.5 V (1997), 1.5 V (2000), 1.2 V (2005), 1.0 V (2010), and then ~0.75–0.85 V (2015–2026). From 1980 to 2005 V fell 4× (energy 16×); from 2005 to 2026 it fell ~1.4× (energy 2×). Had Dennard scaling continued, V in 2026 would be ~0.15 V and wire energy per millimeter ~30× lower than it is. The floor is the subthreshold slope (Section 1.5). Gate-all-around nanosheets (2022–2026) and stacked CFETs (2030s) improve electrostatics and may permit ~0.5–0.6 V; negative-capacitance and tunnel FETs that would break the 60 mV/decade limit remain laboratory devices after two decades of effort.

4.5 The consequence in one number

Energy per bit per millimeter on chip, from about 2005 to 2026: ~0.05–0.1 pJ, flat. Energy per 8-bit multiply-accumulate over the same period: from ~0.5 pJ to ~0.03 pJ, down ~15×. Crossing a 20-mm die (1–2 pJ/bit) is now 30–60× an 8-bit MAC per bit moved; a 64-bit word crossing the die costs the same as several thousand 8-bit MACs. This is why every 2026 accelerator is a grid of small tiles with local SRAM, and why "data movement" — not arithmetic — is the line item designers optimize.


Part V — The off-chip wall

5.1 Leaving the die changes the physics

On-chip, a signal charges femtofarads through resistive wires. Off-chip, it must drive a bond pad (hundreds of femtofarads), a package trace and ball (a picofarad-scale load and a transmission line), a board trace of centimeters, and a receiver with its own pad — and, at gigabit rates, the line must be terminated to a characteristic impedance (~40–50 Ω) to absorb reflections, which dissipates static current regardless of activity. A DDR4 I/O at 1.2 V burns roughly 10–20 pJ per bit at the interface alone; add the DRAM core (activation, sensing, on-die routing) and system-level figures of 20–40 pJ/bit for DDR4 and 40–70 pJ/bit for DDR3 were typical (O'Connor et al., 2017, and vendor power calculators).

Propagation delay, incidentally, is not the problem. A signal travels ~15–18 cm/ns in board dielectric; a DIMM 8 cm from the socket costs ~0.5 ns each way, a few CPU cycles out of ~450. The off-chip wall is charging capacitance and RC/protocol serialization, not the speed of light.

5.2 Pins do not scale

A chip's transistor count grows as κ² per node; its perimeter and package area do not. Rent's rule says required I/O grows as a power (~0.5–0.7) of gate count, so a chip is always short of pins. Intel's Socket 7 (1995) had 321 pins; LGA775 (2004) 775; LGA2011 (2011) 2,011; LGA4677 (2023) 4,677; LGA7529 (2024) 7,529. About 25× in thirty years, against 10⁴× more transistors. Each DDR channel consumes ~150 signal pins for 64 (72 with ECC) data bits; a server CPU with 12 channels devotes roughly a third of its package to memory. This is the "pin wall."

5.3 Every interface since 2013 shortens the wire

Interface Year Where the DRAM sits Distance Bits wide Per-pin rate Approx. system energy/bit
DDR3 2007 DIMM on board 5–10 cm 64/channel 0.8–1.6 Gb/s 40–70 pJ
DDR4 2014 DIMM on board 5–10 cm 64/channel 1.6–3.2 Gb/s 20–40 pJ
GDDR5/6 2008/2018 soldered next to GPU 2–4 cm 32/device 7–16 Gb/s 8–20 pJ
LPDDR4/5X 2014/2021 soldered or on package 0.5–3 cm 16–32/channel 4–8.5 Gb/s 5–15 pJ
DDR5 (incl. MRDIMM) 2020 DIMM on board 5–10 cm 2×32/channel 4.8–8.8 Gb/s 10–25 pJ
HBM2/2E 2016/2019 on interposer beside die 2–8 mm 1,024/stack 2–3.6 Gb/s ~4 pJ
HBM3/3E 2022/2024 on interposer beside die 2–8 mm 1,024/stack 6.4–9.6 Gb/s 2.5–3.5 pJ
HBM4 2026 on interposer, logic base die 2–8 mm 2,048/stack 8–13 Gb/s vendors claim ~20–40% below HBM3E (~2 pJ)
Hybrid-bonded SRAM (V-Cache) 2022 stacked on the die <100 µm vertical thousands of TSVs on-chip rates ~0.1–0.3 pJ (est.)
On-chip, 1 mm — same die 1 mm — — 0.05–0.1 pJ

(Energy figures are approximate system-level values — DRAM core plus interface — from O'Connor et al. 2017, vendor materials, and published estimates; they vary with access pattern and configuration. The "wide and slow" strategy of HBM is precisely a wire-length strategy: HBM4's 2,048 bits at ~8–13 Gb/s deliver 2–3.3 TB/s per stack across a few millimeters of silicon interposer, at per-pin speeds a DDR5 DIMM would find unremarkable.)

The row is monotonic: the closer the memory, the cheaper the bit. The improvement from DDR3 to HBM3E — roughly 15–20× — came almost entirely from geometry (distance from centimeters to millimeters) and modest voltage reduction, not from transistor scaling. The next factor of ten requires the next step in geometry: DRAM hybrid-bonded directly on top of logic, where the vertical distance is tens of micrometers.

5.4 Thermal coupling: the power wall meets the memory wall

Putting memory next to a 1,000-W GPU has a cost the tables above do not show. DRAM retention halves roughly every 10 °C above the JEDEC 85 °C threshold, doubling refresh. HBM stacks are limited in height (12-high in 2026, 16-high being qualified) partly by thermal resistance through the stack. Energy per bit inside the stack becomes heat that raises refresh and lowers margin, a feedback loop that caps how much bandwidth can be packed into a package.


Part VI — The organization wall

6.1 Overfetch: the 8-kilobit row

Because of the 1973 row/column design, a DRAM activation opens an entire row — 8 kilobits (1 KB) per ×8 chip, 8 KB across an 8-chip rank — into the sense amplifiers. A processor that wanted one 64-byte cache line with no row-buffer locality moved 8 KB of charge to get 64 B: an overfetch of 128×. Activation and precharge energy is on the order of a nanojoule per device per row. O'Connor et al. (MICRO 2017) measured that in an HBM2 device most access energy is spent in activation/precharge and on-die data movement rather than in the actual interface, and proposed "fine-grained DRAM" with narrower rows. Wider rows were originally chosen to amortize the wordline decoder and sense-amplifier area; they are an organizational choice that made density cheaper and random access more expensive.

6.2 Bank parallelism: buying latency tolerance with area

Since the array cannot get faster, the only way to raise throughput per chip is to have more independent arrays. DDR (2000) had 4 banks; DDR3 8; DDR4 16 in bank groups; DDR5 32; HBM3 up to 64 per stack across 16 pseudo-channels; HBM4 32 channels. Each bank carries its own sense amplifiers (which are large — sense amplifiers and decoders are a substantial fraction of a DRAM die) and its own timing constraints (tRRD, tFAW) to limit peak current. Memory controllers reorder hundreds of requests to exploit bank parallelism and row-buffer hits; the controller's queueing is itself a source of tens of nanoseconds of latency in loaded systems.

6.3 Where the 90 nanoseconds go

A 2026 server load miss, from core to DRAM and back, roughly:

Stage Approx. time
L1 miss, L2 lookup 3–4 ns
L3 lookup across the die-spanning mesh/ring, coherence check 20–35 ns
Memory controller queue and scheduling 5–20 ns (load dependent)
Off-chip transfer, PHY, DIMM 5–10 ns
DRAM tRCD + CL (row already closed) 28–33 ns
Return path 5–10 ns
Total ~80–110 ns

Only about a third is the DRAM array. Another third is on-chip wire and coherence machinery; the rest is queueing and interface. Every stage is wire or protocol; none is arithmetic.

6.4 The GPU's answer: don't wait

A datacenter GPU sees ~500–700 ns to HBM (measured by microbenchmarks; the long path includes address translation, crossbar and the HBM PHY). It does not try to hide this with reordering; it keeps ~100,000–250,000 threads resident and switches among them, exactly as the Tera MTA did in 1997. The cost is that the machine only works when there is that much parallel work — which is why batch size, not clock speed, is the first knob an inference engineer turns.


Part VII — The hierarchy: from no caches to caches everywhere

7.1 Sixty years of caches

Year Machine Hierarchy (sizes) Largest-cache latency Main memory latency
1965 Typical mainframe/mini registers → core — ~1 µs
1968 IBM S/360 Model 85 16–32 KB buffer (first cache) 80 ns 1.04 µs
1978 DEC VAX-11/780 8 KB cache 200 ns ~1.2 µs (effective)
1980 IBM PC / 8086-class none — ~200–300 ns
1984 Motorola 68020 256 B on-chip I-cache 1 cycle ~150 ns
1987 Intel 386 + 82385 32–64 KB external ~2 cycles ~150 ns
1989 Intel 486 8 KB on-chip L1 1 cycle ~100 ns
1993 Intel Pentium 8+8 KB L1; 256 KB board L2 ~3–5 cycles ~100 ns
1995 Pentium Pro / Alpha 21164 8+8 KB L1; 256 KB–1 MB L2 in package / 96 KB L2 on die + up to 64 MB L3 ~7 cycles ~130 ns
1999 Pentium III (Coppermine) 16+16 KB; 256 KB L2 on die ~7 cycles ~130 ns
2002 Itanium 2 16+16 KB; 256 KB; 3 MB L3 on die ~12 cycles ~150 ns
2006 Core 2 Duo 32+32 KB; 4 MB shared L2 ~14 cycles ~100 ns
2008 Intel Nehalem 32+32 KB; 256 KB; 8 MB shared L3 + IMC ~40 cycles ~65–80 ns
2010 IBM POWER7 + 32 MB eDRAM L3 ~25 ns ~100 ns
2013 Intel Haswell (Crystal Well) + 128 MB eDRAM L4 ~40 ns ~80 ns
2015 IBM z13 960 MB eDRAM L4 per node ~100+ ns ~200 ns
2020 Apple M1 192+128 KB L1; 12 MB L2; 16 MB SLC; LPDDR on package ~40 ns ~100 ns
2022 AMD EPYC Milan-X 32+32 KB; 512 KB; 768 MB L3 (3D V-Cache) ~14–15 ns ~100 ns
2023 AMD EPYC Genoa-X 1,152 MB L3 per socket ~14–15 ns ~100 ns
2024 Cerebras WSE-3 44 GB on-wafer SRAM, no DRAM ~1 cycle (local) —
2026 Typical server µop cache; 48–64 KB L1; 1–2 MB L2; 32–1,152 MB L3; optional HBM/LPDDR "L4"; DDR5; CXL; NVMe 10–40 ns (L3) 80–110 ns (DDR5), 200–300 ns (CXL)

The largest cache per socket grew from 16 KB to 1,152 MB — 72,000× — while its latency improved from 80 ns to ~15 ns, about 5×. The same law that governs DRAM governs SRAM: capacity scales, wires do not.

7.2 SRAM stopped scaling too

Node (year) High-density 6T bitcell (µm²) Scaling vs. prior
180 nm (1999) ~5.6 —
130 nm (2001) ~2.45 0.44×
90 nm (2004) ~1.0 0.41×
65 nm (2006) ~0.57 0.57×
45 nm (2008) 0.346 0.61×
32 nm (2010) 0.171 0.49×
22 nm (2012) 0.092 0.54×
14 nm (2014) 0.050 0.54×
10 nm / N7 (2017–18) 0.031 / 0.027 ~0.6×
N5 (2020) 0.021 0.78×
N3E (2023) 0.021 1.0× — no shrink
N2 (2025) 0.0175 0.83×
Intel 18A (2025–26) 0.021 (HD), 0.023 (HP) 0.88× vs Intel 3

(Sources: Intel and TSMC IEDM/ISSCC disclosures as reported; TSMC N2 and Intel 18A both quote ~38 Mb/mm² macro density at ISSCC 2025.)

A 6T cell is six minimum-size transistors that must read without flipping and write without failing across billions of copies. Its stability depends on the matching of threshold voltages, and mismatch grows as 1/√(W·L) as devices shrink (Pelgrom's law): random dopant fluctuation and line-edge roughness make small transistors unpredictable. The cell therefore cannot use the smallest transistors the process offers, cannot lower its supply voltage (V_min for SRAM is typically higher than for logic), and cannot shrink its contacts and wires below the pitch limits that also bound logic. From 2020 onward the bitcell has been shrinking by ~10–20% per node instead of the historical 50%, which is why:

7.3 Why the number of levels grows

A hierarchy level is worthwhile if it is several times faster than the level below and large enough to catch most misses from the level above. With a typical ratio of 4–6× per level and a total gap G between register and main memory, the number of levels is about log(G)/log(5). In 1980 (G ≈ 1–2) that is zero levels; in 1995 (G ≈ 20) it is two; in 2008 (G ≈ 200) three; in 2026 (G ≈ 500 to DDR5, ~1,500 to CXL, ~500,000 to NVMe) four to five cache levels plus memory and storage tiers. Every widening of the gap adds roughly one tier per factor of five. The tiers are not a design fashion; they are the shape of the gap.

7.4 The machinery of tolerance

Alongside the caches, the mechanisms that hide latency have grown to grotesque proportions:


Part VIII — The compute side: why arithmetic became almost free

For completeness, the other half of the divergence, in three phases:

  1. Dennard (1975–2005). Energy per gate switch fell as κ³, ~10⁴× over the period; clock frequency rose ~1,000× (from ~1 MHz to ~3 GHz); instruction-level parallelism added another ~4×.
  2. Post-Dennard scaling (2005–2016). Voltage froze; energy per gate fell roughly with capacitance (~κ per node); frequency froze; performance came from multiplying cores.
  3. Specialization and precision (2016–2026). Tensor cores, systolic arrays and reduced precision. An 8-bit integer multiply-accumulate costs ~0.2–0.25 pJ at 45 nm (Horowitz) and ~0.02–0.05 pJ at 5-nm-class nodes; a 4-bit operation less still. The most efficient 2026 accelerators execute on the order of 10¹⁵ low-precision operations per second per kilowatt.

Rough estimates of the ratio, for a 64-bit main-memory read versus a 32-bit multiply, over the decades:

Era 64-bit DRAM read (system level) 32-bit multiply (circuit level) Ratio
mid-1980s (5 V, ×1 DRAMs on a board) ~0.1–1 µJ ~2–5 nJ ~100×
2014 (Horowitz, 45 nm, DDR3) 1.3–2.6 nJ 3.7 pJ ~500×
2026, DDR5 ~1 nJ ~0.3 pJ ~3,000×
2026, HBM3E/HBM4 ~0.15–0.2 nJ ~0.3 pJ ~500–700×
2026, per 8-bit weight vs. 8-bit MAC ~16–24 pJ (HBM) ~0.03 pJ ~500–800×

The last row is the number that governs large-language-model inference: each weight fetched from HBM must be reused across ~500–1,000 multiply-accumulates (i.e., tokens in a batch) before the machine is compute-bound. Below that, the accelerator's arithmetic units are idle and the workload is "memory-bound" — which is the memory wall, restated for 2026.


Part IX — Projections, 2026–2040

9.1 What is already committed (2026–2029)

These are on published roadmaps with silicon in qualification or production:

9.2 Likely (2029–2035)

9.3 Speculative (2035–2040)

9.4 A quantitative projection

Metric 2026 2030 (projected) 2035 (projected) Basis
DRAM cell / node 6F², F ≈ 11–12 nm 4F² VCT, F ≈ 9–10 nm; first 3D DRAM 3D DRAM, 32–128 layers vendor roadmaps
Max bits per die 32 Gb (48–64 sampling) 64–96 Gb 128–512 Gb
DRAM row cycle (tRC) 45–50 ns 45–50 ns 45–55 ns cell physics unchanged
Loaded latency, core → DDR 80–110 ns 80–110 ns 70–110 ns wires and protocol
HBM bandwidth per stack 2–3.3 TB/s 4–6 TB/s 8–16 TB/s ~2× per ~3 years
HBM energy per bit ~2 pJ 1–1.5 pJ 0.3–0.8 pJ hybrid bonding, direct stacking
DDR channel bandwidth 51–70 GB/s 100–140 GB/s (DDR6) 200+ GB/s interface cadence
On-chip wire, pJ/bit/mm 0.05–0.1 0.05–0.1 0.03–0.08 c flat; V floor ~0.5 V
SRAM HD bitcell 0.0175–0.021 µm² ~0.015 µm² ~0.010 µm² 10–15%/node; CFET
8-bit MAC energy ~0.03 pJ ~0.015 pJ ~0.008 pJ process + design
Accelerator ridge point (low precision) ~1,000 FLOP/byte 1,500–2,500 3,000+ 3.0×/2 yr vs 1.6×/2 yr
Hierarchy tiers (server) 8–10 9–11 9–12 log of the gap
DRAM cost per bit ~2× the 2024 low volatile; 3D DRAM resets the curve slow decline supply/demand + 3D

9.5 What will not change

9.6 Scenarios


Appendix A — Derivations and reference physics

A.1 Dennard scaling. Scale L, W, t_ox, V by 1/κ and doping by κ. Saturation current I ∝ (W/L)·C_ox·(V−V_t)² ∝ 1·κ·κ⁻² = κ⁻¹. Gate capacitance C ∝ WL/t_ox ∝ κ⁻¹. Delay ∝ CV/I ∝ κ⁻¹·κ⁻¹/κ⁻¹ = κ⁻¹. Power ∝ IV ∝ κ⁻². Density ∝ κ², so power density is constant. Energy per switch ∝ CV² ∝ κ⁻³.

A.2 Wire RC. r = ρ/(WH) per unit length; c ≈ ε·f(W/S, H/S) per unit length, invariant under uniform shrink; parallel-plate estimate for a wire between neighbors and ground gives c ≈ 0.15–0.25 fF/µm for k = 3–4 at aspect ratios near 2. Distributed RC delay of an unrepeated line: t ≈ 0.38·r·c·L². With optimal repeaters (segment length ∝ √(R_dC_d/(rc))), delay becomes linear in L at ~ 2L·√(0.38·rc·R_dC_d) plus repeater delays.

A.3 Wire energy. E = ½·c·L·V² per transition (full swing). With c = 0.2 pF/mm and V = 0.8 V: 64 fJ/mm. Realized values 0.05–0.1 pJ/bit/mm including repeaters; ~1–2 pJ across a 20-mm die.

A.4 Copper size effects. Fuchs–Sondheimer surface scattering and Mayadas–Shatzkes grain-boundary scattering: ρ_eff/ρ₀ ≈ 1 + (3/8)(1−p)(λ/d) + grain term, with λ ≈ 39 nm for Cu at 300 K, d the wire dimension, p the specularity. At d ≈ 10–12 nm, ρ_eff ≈ 4–6ρ₀, and the barrier/liner (3–5 nm total) halves the conductive cross-section. Ru (ρ₀ ≈ 7 µΩ·cm, λ ≈ 7 nm) and Mo (ρ₀ ≈ 5 µΩ·cm, λ ≈ 11 nm) cross over below ~12–17 nm.

A.5 Subthreshold slope. I_off ∝ exp(q(V_gs−V_t)/(n·kT)); SS = n·(kT/q)·ln 10 ≥ 59.6 mV/decade at 300 K, with n = 1 + C_dep/C_ox ≥ 1 (practical 63–75 mV/decade). Five decades of on/off → V_t ≥ 0.3 V; adding ~0.3–0.4 V of overdrive for speed → V_dd ≈ 0.6–0.75 V. Minimum V_dd for a static CMOS gate with gain > 1: ~2(kT/q)·ln 2 ≈ 36 mV (Swanson–Meindl), usable only at very low speed.

A.6 DRAM sensing. ΔV_BL = (V_core/2)·C_S/(C_S+C_BL). C_S = 15 fF, C_BL = 30 fF, V_core = 1.1 V → ΔV ≈ 180 mV. Stored charge Q = C_S·V_core/2 ≈ 8 fC ≈ 5×10⁴ electrons. Retention: allowable leakage ≈ (Q/3)/64 ms ≈ 40 fA. Critical charge for soft errors: tens of fC (alpha) — the reason Q could not scale.

A.7 Capacitor geometry. C = ε₀·k·A/t. For 15 fF with k = 35, t = 6 nm: A ≈ 0.29 µm². A cylinder of diameter 30 nm and height h using both surfaces has A ≈ 2π(30 nm)h → h ≈ 1.5 µm; aspect ratio ≈ 50.

A.8 Refresh overhead. Fraction of time busy = tRFC/tREFI = 350 ns/7.8 µs ≈ 4.5% (8 Gb DDR4), 550/7,800 ≈ 7% (16 Gb); doubles above 85 °C.

A.9 Little's law for latency tolerance. In-flight instructions needed = issue width × miss latency (cycles) = 5 × 450 ≈ 2,250, versus reorder buffers of 450–650. GPUs substitute threads: ~10⁵ resident threads × (a few outstanding loads each) ≫ latency × bandwidth.

A.10 Tier count. Levels ≈ log(G)/log(ratio per level); G = 500, ratio = 5 → 3.9 levels of cache between register and DRAM.

A.11 Roofline ridge point. Ridge = peak FLOP/s ÷ peak bytes/s. H100: ~989 TFLOP/s FP16 dense ÷ 3.35 TB/s ≈ 295 FLOP/byte. B200: ~9 PFLOP/s FP4 dense ÷ 8 TB/s ≈ 1,100 FLOP/byte. Cray-1 (1976): 160 MFLOP/s ÷ 640 MB/s ≈ 0.25 FLOP/byte.


Appendix B — Glossary of the timing parameters that did not scale


References and further reading

Primary sources for the historical and physical claims above, in roughly chronological order:

  1. M. V. Wilkes, "Slave Memories and Dynamic Storage Allocation," IEEE Trans. Electronic Computers, EC-14(2), 1965.
  2. J. S. Liptay, "Structural aspects of the System/360 Model 85, II: The cache," IBM Systems Journal, 7(1), 1968.
  3. R. H. Dennard, "Field-effect transistor memory," U.S. Patent 3,387,286 (filed 1967, granted 1968).
  4. R. H. Dennard et al., "Design of Ion-Implanted MOSFETs with Very Small Physical Dimensions," IEEE J. Solid-State Circuits, SC-9(5), 1974.
  5. T. C. May and M. H. Woods, "Alpha-Particle-Induced Soft Errors in Dynamic Memories," IEEE Trans. Electron Devices, ED-26(1), 1979.
  6. D. Kroft, "Lockup-Free Instruction Fetch/Prefetch Cache Organization," ISCA, 1981.
  7. J. L. Hennessy and D. A. Patterson, Computer Architecture: A Quantitative Approach, 1st ed. 1990; 6th ed. 2017 (DRAM timing survey table, Ch. 2).
  8. M. S. Lam, E. E. Rothberg, M. E. Wolf, "The Cache Performance and Optimizations of Blocked Algorithms," ASPLOS, 1991.
  9. W. A. Wulf and S. A. McKee, "Hitting the Memory Wall: Implications of the Obvious," ACM SIGARCH Computer Architecture News, 23(1), 1995.
  10. M. T. Bohr, "Interconnect Scaling — The Real Limiter to High Performance ULSI," IEDM, 1995.
  11. D. M. Tullsen, S. J. Eggers, H. M. Levy, "Simultaneous Multithreading: Maximizing On-Chip Parallelism," ISCA, 1995.
  12. R. Ho, K. W. Mai, M. A. Horowitz, "The Future of Wires," Proc. IEEE, 89(4), 2001.
  13. V. Agarwal, M. S. Hrishikesh, S. W. Keckler, D. Burger, "Clock Rate versus IPC: The End of the Road for Conventional Microarchitectures," ISCA, 2000.
  14. P. Saxena et al., "Repeater Scaling and Its Impact on CAD," IEEE Trans. CAD, 23(4), 2004.
  15. D. A. Patterson, "Latency Lags Bandwidth," Communications of the ACM, 47(10), 2004.
  16. K. Asanović et al., "The Landscape of Parallel Computing Research: A View from Berkeley," UC Berkeley Tech. Rep. UCB/EECS-2006-183, 2006.
  17. B. Jacob, S. Ng, D. Wang, Memory Systems: Cache, DRAM, Disk, Morgan Kaufmann, 2008.
  18. S. Williams, A. Waterman, D. Patterson, "Roofline: An Insightful Visual Performance Model for Multicore Architectures," CACM, 52(4), 2009.
  19. B. M. Rogers et al., "Scaling the Bandwidth Wall: Challenges in and Avenues for CMP Scaling," ISCA, 2009.
  20. W. J. Dally, "Power, Programmability, and Granularity: The Challenges of ExaScale Computing," keynote, IPDPS, 2011 (energy-per-operation and per-wire figures).
  21. J. Liu, B. Jaiyen, R. Veras, O. Mutlu, "RAIDR: Retention-Aware Intelligent DRAM Refresh," ISCA, 2012.
  22. Y. Kim et al., "Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance Errors," ISCA, 2014.
  23. M. Horowitz, "Computing's Energy Problem (and What We Can Do About It)," ISSCC, 2014.
  24. M. O'Connor et al., "Fine-Grained DRAM: Energy-Efficient DRAM for Extreme Bandwidth Systems," MICRO, 2017.
  25. J. S. Kim et al., "Revisiting RowHammer: An Experimental Analysis of Modern DRAM Devices and Mitigation Techniques," ISCA, 2020.
  26. T. Dao et al., "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness," NeurIPS, 2022.
  27. A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney, K. Keutzer, "AI and Memory Wall," IEEE Micro, 2024 (arXiv:2403.14123).
  28. JEDEC, JESD79-4/-5 (DDR4/DDR5), JESD235/238 (HBM/HBM3), JESD270-4 (HBM4, April 2025).

Contemporary sources consulted for the 2025–2026 state of the industry:

  1. SK hynix, "SK hynix Presents Future DRAM Technology Roadmap at IEEE VLSI 2025" (press release, June 2025) — 4F² vertical-gate platform and 3D DRAM as the two pillars beyond 10 nm.
  2. SemiAnalysis, "Intel 18A Details & Cost, Future of DRAM 4F2 vs 3D, Backside Power Adoption…" (VLSI 2025 coverage) — 6F² scaling ends at 1d; storage-node-contact margin as the limiting factor.
  3. IEEE Spectrum, "Intel and TSMC Detail SRAM at ISSCC 2025" — 38 Mb/mm² on both 18A and N2; backside power does not shrink the bitcell.
  4. Micron, "HBM4" product page (2026) — 2,048-bit interface, >11 Gb/s per pin, >2.8 TB/s per stack, pJ/bit improvement versus HBM3E 12-high.
  5. Siemens EDA blog, "HBM3E and HBM4: IC design guide" (April 2026) — HBM4 architecture, 2.0–3.3 TB/s per stack, 16-high/64 GB stacks.
  6. Utmel, "HBM4 and the Shift to Customized AI Memory" (June–Aug. 2026) — JEDEC baseline 8 Gb/s and ~2 TB/s; Samsung 3.3 TB/s and Micron 2.8 TB/s products; all three vendors in production by August 2026.
  7. TrendForce press releases (Oct. 2025, June 2026) — conventional DRAM contract price increases; HBM wafer share of DRAM output (~18%/22%/30% for 2025/26/27); crowding-out dynamics.
  8. Findchips / Supplyframe Commodity IQ (June 2026) and Counterpoint/Gartner as reported — DDR5 contract price increases of ~50–57% QoQ in 1H 2026; DDR4 spot per-Gb prices exceeding HBM3E; Gartner's ~130% 2026 DRAM price-rise projection.
  9. IEEE Spectrum, "Interconnects Are in Need of a Major Overhaul" (IEDM 2022 coverage) — ruthenium, top-via, air-gap and backside-power directions.
  10. The Register, "TSMC says first 1.6nm chips coming in 2026" (April 2024) — A16 with Super Power Rail backside power.

Numbers not attributed to a specific source are the author's engineering estimates from the physics in Appendix A and public datasheets; they should be read as order-of-magnitude values.