The claim in one board
The argument fits on a blackboard, which is fitting, because the method comes from one: Reiner Pope's lecture on the Dwarkesh podcast, in which two numbers per chip and two numbers per model are enough to reconstruct how frontier labs train and serve.1 The two numbers per chip are its peak arithmetic rate and its memory bandwidth. Their ratio, the machine balance \(\beta\), says how many floating-point operations the chip must perform on every byte it fetches to stay busy.
That is the whole case. The rest of this report makes it quantitative in three directions: backward, to show what the hardware of 1986 actually looked like and why nobody then had a reason to think about arithmetic intensity; forward, to show how the balance drifted by three orders of magnitude and where the bytes of a modern training step go; and sideways, to catalogue the tricks that dig backprop out of the wall, each as an explicit trade of one resource for another. Every number in the text is computed live from the same model that drives the figures, so you can change the machine or the network and the story updates with it.
Two numbers per machine
Pope's roofline argument for inference goes like this. A decode step has to multiply a batch of \(B\) tokens by all the active parameters, which takes \(B \cdot N_\text{active} \cdot 2 / \text{peak}\) seconds of arithmetic, and it has to fetch every parameter once, which takes \(N_\text{total} \cdot \text{bytes} / \text{bandwidth}\) seconds of memory traffic. The step cannot be faster than the larger of the two. Setting them equal gives the batch size at which the chip is neither starved nor idle:
$$ B_\text{balance} \;=\; \frac{\text{peak}}{\text{bandwidth}} \times \frac{N_\text{total}}{N_\text{active}} \;\approx\; 300 \times \text{sparsity}. $$"On most GPUs this ends up being somewhere around 300," he says, and adds that it has "remained reasonably stable" from A100 to H100 to Blackwell: FLOPs and bandwidth both grew, and their ratio did not move much. He then reads off the practical consequence: a dense model wants a few hundred sequences in flight per weight fetch, DeepSeek's 8× sparsity wants a few thousand, and "generally people will go a little bit larger than this," two or three times, because real kernels do not reach the roofline.
The same two numbers describe training, with two changes. First, each weight is touched more than once per step: read in the forward pass, read again in the backward pass to propagate the error, and read and written once more to accumulate its gradient. Second, the arithmetic per weight triples, because the backward pass costs twice the forward. Per weight visit the intensity is the same as inference, but the number of tokens that must share one visit is the same \(\beta\)-sized quantity. We will call that number \(T\), tokens per weight visit: the number of tokens processed by one device between two reads of the same weight, which is the micro-batch size in tokens on that device.
Peaks are dense (no structured sparsity). β is the balance at the training precision. The drain time, capacity over bandwidth, is Pope's "the train departs every 20 ms": the time to read all of memory once, and hence the floor on a step that touches every weight and every saved activation.
Three things stand out in that table. The balance at bf16 has sat between 150 and 300 FLOP per byte since 2020, and at fp8 it doubles. The drain time has sat near 20–40 ms across four HBM generations. And the on-chip memory that would let a chip avoid the trip to HBM is tens of megabytes on a device with tens of gigabytes of weights, a ratio of a thousand, which is why the argument below is about the HBM boundary and not the register file.
The world backprop was born into
The Nature paper of October 1986 that made backpropagation a household algorithm2 describes networks of 7 to 630 parameters, trained on 4 to 100 patterns, for hundreds to thousands of sweeps: the symmetry detector (17 weights, 64 patterns, 1,425 sweeps), the family-tree network (630 weights, 100 triples, 1,500 sweeps), XOR. It gives two ways of applying the gradient, and the reason it offers for one of them is worth quoting because it is the paper's only remark about memory:
One way of using ∂E/∂w is to change the weights after every input-output case. This has the advantage that no separate memory is required for the derivatives. An alternative scheme, which we used in the research reported here, is to accumulate ∂E/∂w over all the input-output cases before changing the weights.Rumelhart, Hinton and Williams, 1986
So the published rule was full-batch gradient descent with momentum, and the memory the authors worried about was one accumulator per weight. The number of patterns that share one pass over the weights, the quantity this report calls \(T\), was a free choice, made for reasons of optimization: the whole training set in the Nature paper, blocks of 25 in Plaut, Nowlan and Hinton's CMU report of the same year, one word in NETtalk. Nobody argued for a value of \(T\) on hardware grounds, and within a decade the field's advice ran the other way: "stochastic learning is usually much faster than batch learning," LeCun and colleagues wrote in 1998, and Wilson and Martinez in 2003 called batch training "almost always slower than on-line training, often orders of magnitude slower."3 The only cost anyone counted was arithmetic. "The backward pass has the same computational complexity as the forward pass, and so it is not unduly expensive," says the PDP chapter; Hinton's 1989 survey estimates learning time on a serial machine as roughly cubic in the number of weights.
They could afford to count only arithmetic because of the machines in the table below. Every computer a connectionist could get time on in the mid-1980s, from a VAX in the department to a Cray at a national center, could deliver a byte from main memory for every fortieth to every fourth floating-point operation. Per-example backprop at fp32 needs a byte for every two or three operations. On those machines any choice of \(T\) ran the arithmetic units flat out, and the arithmetic units were slow: a VAX-11/780 with a floating-point accelerator ran the LINPACK benchmark at a seventh of a megaflop.
Peak is the vendor or derived peak at the stated format; sustained rates (LINPACK 100×100) were typically a tenth of peak on these machines. β = peak ÷ main-memory bandwidth. Sources and derivations are in the research notes in the repository.
Interconnections per second
The DARPA Neural Network Study of 1988, the field's first hardware survey, rated machines in "interconnections per second," the number of multiply-and-add operations on weights a machine could perform in a second, and implementation papers soon split it into CPS for the forward pass and CUPS, connection updates per second, for training.4 The numbers are worth a look because they show what fast meant then.
A connection update is one weight read, two reads of it in fact, a handful of multiply-adds, and a read-modify-write of the weight or its accumulator. Rating machines by it amounts to assuming that the arithmetic comes free with the memory traffic, which in 1988 it did. The study's authors already saw the seam: the quoted figure, they warned, "may or may not include the time necessary to obtain the variables that are to be multiplied and added." On an H100 the multiply-add takes a fraction of a picosecond of tensor-core time and the variables take a hundred picoseconds to arrive from HBM. The metric that made sense in 1988 would, applied today, describe a machine running at a fraction of a percent of its peak.
What reverse mode remembers
The other 1980s fact that matters is mathematical rather than technological. Backpropagation is reverse-mode automatic differentiation, which Linnainmaa described in 1970 and Werbos applied to neural networks in 1974.5 Reverse mode has a famous virtue, the cheap-gradient principle: the gradient of a scalar function costs at most a small constant times the function itself, regardless of how many inputs it has. It has an equally famous cost that the 1986 paper did not need to mention: to run the tape backward you must keep the tape. Every intermediate value of the forward pass must be stored until the backward pass consumes it, so memory grows with the length of the computation. Griewank's 1992 paper on logarithmic-memory reverse differentiation, the origin of what deep learning calls gradient checkpointing, exists because this cost was already biting in scientific computing by then.6 For a three-layer network with a hundred units it is a kilobyte. For a 126-layer transformer processing an 8,192-token sequence it is , which is more than half the size of the weights.
The drift
Figure 1 plots forty machines from the Cray-1 to Blackwell. Panel (a) is the quantity an algorithm feels: how many examples must share one fetch of a weight for the arithmetic to stay busy, \(\beta\, b_w / 6\). Panel (b) is the machine balance itself, and (c) and (d) are its numerator and denominator. Peak arithmetic has doubled roughly every eighteen months for fifty years. Memory bandwidth has doubled roughly every three years. The ratio has therefore risen by about a thousand times, and the number of examples a weight fetch must serve has gone from about one to about six hundred.
Table view
The people who watched it happen described it in the present tense. McCalpin, defining machine balance in 1995: "10 years ago, floating-point operations were considered quite expensive, often costing 10 times as much as an uncached memory reference. Today the situation is dramatically reversed, with the fastest current processors able to perform 100 or more floating-point operations in the time required to service a single cache miss." Wulf and McKee, the same year, with DRAM improving at 7 percent a year against a processor's 80: "as \(t_c\) and \(t_m\) diverge, \(t_\text{avg}\) will grow and system performance will degrade. In fact, it will hit a wall."7 Patterson's rule of 2004 is that in the time bandwidth doubles, latency improves by no more than 1.2 to 1.4×, and that DRAM bandwidth doubles every 2.9 years to a processor's 1.7.8 Gholami and colleagues measured the accelerator era: peak FLOPS up 3.0× every two years, DRAM bandwidth up 1.6×, interconnect bandwidth up 1.4×; Epoch AI's fit over 47 accelerators has FLOP/s doubling every 2.3 years and memory bandwidth every 4.1.9 Transistors shrink; pins, wires and the energy to drive them do not shrink with them. Kung had already shown in 1986 what the consequence is for matrix arithmetic: if the ratio of compute to I/O rises by \(\alpha\), the local memory needed to stay balanced rises by \(\alpha^2\), which is why every accelerator generation ships more SRAM and why it is never enough.
Neural networks met the wall first in inference. Jouppi and colleagues' 2017 paper on the first TPU, a chip with a ridge point of 1,350 operations per byte, reports that the production MLPs and LSTMs ran at batches of 64 to 200 examples per weight fetch and were "happily bumping their heads against the ceiling"; only the CNNs, with their spatial reuse, were compute-bound. What the drift did to training is visible in the dashed lines of panel (a) of Figure 1. Nothing about the algorithm's per-example intensity changed between 1986 and 2026. The hardware crossed it, in the mid-1990s for the machines connectionists used and a decade earlier for the Crays, and everything below is about the consequences and the workarounds.
What a training step moves
To be concrete we need a ledger for one step of a transformer, written per parameter and per token so that it applies at every scale. The constants come from counting the kernel sequence of a fused, FlashAttention layer with mixed-precision Adam; the companion notes derive each one and give ranges.10
Figure 2 is the roofline explorer. Pick a machine from either era, a network from either era, and a micro-batch, and it draws the roofline, places the training step on it, and reports the batch at which the step reaches the balance point. The presets are the cases the argument turns on.
Table view
Three floors and a ceiling
Solving \(I(T) \ge \beta\) gives the smallest micro-batch that leaves the wall:
$$ T_{\min} \;=\; \frac{12\, b_w\, d\, \beta}{72\, d \,-\, 200\,\beta} \;\xrightarrow{\;d\to\infty\;}\; 2\beta. $$On an H100 that limit is tokens per weight visit; on a B200 ; on an A100 . For a 7B-class model (\(d = 4096\)) the actual figure on an H100 is tokens, and for Llama-3 405B it is . With bf16 gradients and no fp32 accumulator the floor halves. This is the first floor, and it is the training-side twin of Pope's "batch of about 300 times sparsity": the same \(\beta\), multiplied by the number of times a weight is visited per token.
The second floor is a pole. The denominator of \(T_{\min}\) vanishes at \(d_{\min} = 200\beta/72\), which is on an H100. Below that hidden size the activation traffic alone, which grows with \(T\), exceeds the machine's byte budget, and no batch size helps. Only fusion (a smaller 200) or a slower chip does. This is why GPT-2-small-sized models get 20 to 30 percent of an H100 however they are batched, and why the 1986 networks in Figure 2 get essentially nothing.
The third floor is paid once per optimizer step rather than once per micro-batch. Adam in mixed precision moves about 30 bytes per parameter per step, so the step must contain at least \(5\beta\) tokens per device, on an H100, before the update costs less than the compute, and ten times that before it costs less than 10 percent. Data-parallel training adds the gradient all-reduce, about 4 bytes per parameter over the fabric; for a 7B model on 64 H100s over a single NIC the two together are worth token-equivalents of compute per device per step. Gradient accumulation exists to amortize exactly this term: it raises the tokens per step without raising the tokens per weight visit.
The ceiling is statistical. Past the gradient-noise scale, a bigger batch buys fewer steps per token: McCandlish and colleagues fit \(S = S_{\min}(1 + B_\text{noise}/B)\), and Kaplan and colleagues put the critical batch for language models at a few million tokens at pre-training losses.11 The window in which backprop is neither memory-bound nor statistically wasteful is
$$ N \times \max(T_{\min},\, T_\text{fixed}) \;\lesssim\; B \;\lesssim\; B_\text{crit}, $$and it narrows as the device count \(N\) grows or the model shrinks. Figure 6 draws it.
Table view
Backprop's own contribution: remembering
The weight-side wall is not specifically backprop's fault. Inference decode streams the same weights for a third of the arithmetic and is worse; any algorithm that touches every parameter per example has the same 1/T hyperbola. What is specifically backprop's is the tape. The activation of layer \(\ell\) is produced at step \(\ell\) of the forward pass and consumed at step \(2L - \ell\) of the backward pass, and in between it has to live somewhere. On a time-multiplexed device that somewhere is HBM, and the cost is capacity: \(34\, d\) bytes per token per layer, or for one sequence of Llama-3 405B against of weights. Batch makes it worse, linearly, while it makes the weight-side wall better. That is the tension every training configuration lives inside.
Table view
Figure 5 animates the tape. With everything stored, resident activations climb through the forward pass and drain through the backward pass, and the high-water mark is all \(L\) layers. Checkpointing every \(\sqrt L\) layers keeps only the checkpoints, then recomputes each segment forward when the backward pass reaches it: the high-water mark drops to about \(2\sqrt L\) and the arithmetic rises by a third (Chen and colleagues, 2016).12 Reversible layers, the Feistel construction that Pope ends his lecture on, store nothing but the final output and invert each layer on the way back: the high-water mark is constant and the arithmetic again rises by a third.13 "Spending more compute to save memory," as he puts it, and then: "spending more memory to save compute is generally profitable given where hardwares are." The two sentences describe the same trade from opposite sides of the balance point, and which one is profitable depends on \(\beta\).
Table view
Wulf and McKee ended their 1995 note with a question: "Are there any new ideas for how to trade computation for storage?" Checkpointing and reversibility are the answers backprop gives, and in energy the tape is cheap and in seconds it is expensive. Round-tripping a token-layer's saved activations through HBM costs \(68\,d\) bytes at 50 to 60 picojoules each; recomputing its forward costs \(24\,d^2\) FLOPs at about half a picojoule each. The two cross at \(d \approx 200\): above it, storing is the cheaper choice in joules, which is why full checkpointing is a capacity tool and not an efficiency tool. It lowers no traffic. Selective recompute, the FlashAttention variant that recomputes only the attention scores, is the exception: the \(s \times s\) score matrix has an intensity of order one through HBM, so recomputing it in on-chip memory removes \(O(s^2)\) bytes for \(O(s d)\) extra FLOPs, a trade that the balance makes favorable on every accelerator since Volta.14
The tricks, as trades
Each technique that makes backprop run efficiently on a modern accelerator gives up one resource to get another. Read the table with \(\beta\) in mind: every row is a way of converting the resource the hardware has in surplus, arithmetic, or the resource the algorithm has in surplus, samples, into the resource that is scarce, bytes.
Two of these rows deserve a second look because they show the wall from the statistical side. Pope, asked why training has a hard stop between forward and backward passes when inference does not, answers that "smaller is always better, actually, from an ML convergence rate perspective, because you're getting the freshest information from the gradient descent," and then: "from a total training time perspective, smaller is worse from a systems perspective. The optimum is the trade-off between those two." The systems side of that trade-off is \(\beta\). The batch sizes labs actually use, 4 to 16 million tokens for the largest runs, are not what the optimizer wants; they are what the optimizer will tolerate given what the hardware needs.
Table view
Concrete breakdowns
The table below runs the model on the cases that anchor the argument: the two classic 1980s runs on their own hardware, the 1987 network on 2022 hardware, and three modern pre-training runs at the micro-batch their hardware actually saw per weight visit. The point of the modern rows is not that they are memory-bound at the batch they use; they are not, or barely. The point is how they got there.
Utilization ceiling is what the roofline allows before kernel inefficiency; measured model FLOP utilization of the modern runs is 38–46 percent. The 1980s rows use fp32 numbers and per-pattern updates; caches are too small to hold the weights so every pattern streams them.
NETtalk, 1987
Eighteen thousand weights, updated after every word, on a VAX-class minicomputer. Per letter the simulator read a quarter of a megabyte of weights and did a hundred thousand FLOPs, an intensity of half a FLOP per byte on a machine whose balance is a few hundredths. The paper's own throughput figure, two letters a second for a ten-thousand-weight network, is the VAX's floating-point rate and two percent of its bus. The model in Figure 2 puts the ceiling at the arithmetic, and the measurement agrees. That the network learned to read aloud in a few days of such steps is the whole reason the method spread.
Zip codes, 1989
LeCun's convolutional network for handwritten digits had 1,256 units, 64,660 connections and 9,760 free parameters, and was trained for 23 passes over 7,291 digits, 167,693 presentations, on a SUN-4/260.17 Weight sharing reused each kernel weight at every position of the image, which raised the tokens per weight visit from one to the number of positions, about seven averaged across its layers. On a Sun-4 it was still arithmetic-bound, and the same structural fact is what makes convolutional networks comfortable at batch one on a GPU today: the image is the batch. The days the run took on a workstation, and the eighteen days Waibel's time-delay network took on an Alliant mini-supercomputer the same year, were spent on FLOPs, not bytes.
GPT-3, 2020
175 billion parameters, 96 layers, \(d = 12{,}288\), 2,048-token sequences, a batch of 3.2 million tokens, on V100s. One sequence per weight visit gives \(T = 2{,}048\), which on a V100 (\(\beta = \) ) puts the step at an intensity of FLOP per byte, above the balance point. The sequence is the batch: a transformer applies every weight to every token, so a single document supplies the reuse a 1986 MLP could only get from a thousand examples. This is the transformer's quiet win in the hardware lottery, and Pope's "batch of 300" rule is easy to satisfy in pre-training for exactly this reason. What GPT-3 could not do was fit: at 16 bytes per parameter its training state is on 32 GB devices, so the run existed only through tensor and pipeline parallelism, the escape hatch. Its measured utilization, 21 percent of peak by PaLM's later accounting, is what the first generation of that machinery achieved.
Llama-3 405B, 2024
Meta's report is unusually explicit: 16,384 H100s, 8,192-token sequences, tensor parallelism 8, pipeline parallelism 16, data parallelism 128 with FSDP, a batch schedule ending at 16 million tokens, no activation checkpointing, and 38 to 43 percent model FLOP utilization.15 Run the ledger. One sequence's tape at the selective-recompute policy is ; the weights are ; the training state is . Nothing fits on an 80 GB device without sharding, and after sharding the eight-way tensor parallelism leaves \(T\) per weight visit at one sequence, 8,192 tokens, which on an H100 is the floor \(T_{\min}\). The weight-side wall is paid. What remains is the price of paying it: 72 GB of tape per device at a micro-batch of one sequence (Figure 4), and an FSDP all-gather of 6 GB of weights per micro-batch over a 50 GB/s network card, which at sixteen micro-batches per step is two seconds of fabric time against six seconds of arithmetic. Meta attributes the drop from 43 to 41 percent when going from 8,192 to 16,384 GPUs to "the lower batch size per DP group": the same tokens spread over more replicas raise the fabric's share. The roofline ceiling for this configuration is ; the difference between that and 41 percent is the sum of the trades in the table above, plus bubbles and kernels that do not reach peak.
DeepSeek-V3, 2024
The mixture-of-experts case makes the batch floor visible in a new place. Each token activates 8 of 256 routed experts, so the arithmetic per token is that of a 37-billion-parameter model while a micro-batch that reaches every expert streams all 671 billion. Pope's rule gives a batch of \(300 \times 8 \approx 2{,}400\) tokens per expert visit before the weight fetch is amortized, and a single 4,096-token sequence gives each expert only 128. DeepSeek's answer was expert parallelism across 64 devices, so that each expert sees the pooled tokens of an entire group, a global batch that ends at 15,360 sequences (63 million tokens), FP8 weights that halve the bytes, and a bidirectional pipeline schedule, DualPipe, whose job is to overlap the all-to-all traffic of the expert layer with arithmetic at a compute-to-communication ratio the paper puts near one to one.16 The report gives no utilization figure; dividing its 2.66 million pre-training GPU-hours into the 6ND arithmetic gives about 35 percent of the bf16 dense peak, or 17 percent of the FP8 peak at which the GEMMs actually ran. The 2.8 million GPU-hours are the FLOPs; the paper's engineering sections are almost entirely about the bytes.
Objections
The thesis invites specific objections, and a report that used them as straw would be worth less than one that answers them.
At production batch sizes the GEMMs are compute-bound, so the wall is already solved.
Half true, and the half that is true is the point. Large pre-training runs sit at 40 to 46 percent MFU because a dozen trades have been made to put them there. The wall is "solved" the way a mortgage is: by paying every month. The interesting evidence is what happens when a trade is unavailable. Fine-tuning with small batches, reinforcement learning with per-step updates, online learning, small models below \(d_{\min}\), and anything that must run with \(T\) near one all fall straight back onto the hyperbola of Figure 2, and the report's single-token row shows how far: slower than compute-bound, at of peak.
The wall is a property of the hardware, not of backprop. Inference decode is worse.
Correct on both counts, and the report's title is "versus," not "caused by." Decode at batch one has a third of backprop's arithmetic per weight byte and no batch-independent activation term to hide behind. But decode has no tape and no optimizer state, so its wall is one-dimensional and batching removes it wholesale, which is what Pope's lecture is about. Backprop's wall has a second dimension, capacity, that batching makes worse. Backprop is the algorithm for which the two walls pull in opposite directions.
The sequence is the batch, so weight reuse comes free for transformers.
Yes, for pre-training a dense model at \(d \ge 1{,}000\) with sequences of a few thousand tokens. It is a lucky property of the architecture, not of the algorithm, and it fails in each of the directions the field is moving: sparse experts divide the effective \(T\) by the sparsity ratio, small models fall below \(d_{\min}\), and RL-style training with short rollouts and frequent updates shrinks both the sequence and the batch. A CNN got the same free reuse from image positions in 1989; an MLP on vectors, the network of the 1986 paper, never gets it.
Storing the tape is a property of reverse-mode differentiation in general, and checkpointing makes it cheap.
Reverse mode is backprop; the objection restates the thesis. Checkpointing makes the tape \(O(\sqrt L)\) for a third more arithmetic, which is cheap precisely because arithmetic is the surplus resource, and that is the report's claim in one sentence: the algorithm's native cost structure is wrong for the hardware, and the fix spends the resource the hardware has to spare. What checkpointing cannot do is lower traffic, and what nothing exact can do is remove the per-step optimizer and all-reduce terms that set the third floor.
The 1980s were not memory-rich; DRAM was tiny.
The distinction is between bandwidth per FLOP and capacity. Capacity was a few megabytes and it bounded the networks that could be tried, which is a different wall and a real one. But bandwidth per FLOP was two to three orders of magnitude higher than today, and the 1986 networks were small enough that capacity did not bind them. The report's claim is about the first quantity. It also does not claim the machines were fast: a VAX did a hundred thousand FLOPs a second, and the reason backprop was compute-bound was that its arithmetic was slow, not that its memory was good. That is precisely what changed.
Energy is what matters, and the HBM energy share at production batch sizes is modest.
True: at \(T\) in the thousands, HBM traffic is 10 to 20 percent of a training step's energy on an A100-class device, and the on-chip operand delivery inside the "FLOP" dominates. The wall in this report is stated in seconds, where the balance is 150 to 300, not in joules, where it is 70 to 140. Backprop's activation round trip is cheap in joules and expensive in seconds. In seconds the tensor cores idle while the elementwise kernels run, and the static power of the idle device is the largest energy term that batching cannot touch.
Alternatives to backprop are worse, so "subject to the memory wall" is not actionable.
Forward gradients, local losses, synthetic gradients and zeroth-order methods have not met backprop's accuracy at scale, and the report does not claim otherwise. What is actionable is the ledger: the three floors and the ceiling say which trade to make next for a given \(\beta\), \(d\), \(N\) and fabric, and they say what a hardware design with a different \(\beta\) would buy. The SRAM-heavy designs, Cerebras and Groq, have balances near 5 FLOP per byte against on-chip memory, back where the Cray was; whether the capacity they give up is worth it is the same trade, made in silicon.
What would change the picture
Three things would. A hardware trend that returned bandwidth per FLOP toward the 1980s ratio: HBM4 and the co-packaged memory of the Rubin generation move it a little, and the wafer-scale designs move it a lot, at a capacity cost. An algorithm that does not remember: reversible layers are the exact candidate, and they have been tried at ImageNet and language-model scale with small losses and awkward constraints; the honest state of the art is that nothing exact removes the tape for free, and nothing that removes it approximately has matched backprop. And a change in what is being trained: the drift toward reinforcement learning with per-rollout updates and toward sparse models pushes the effective \(T\) down, which is the direction of the wall, so the pressure is more likely to increase than to relax.
Pope's two-number method is, in the end, a way of noticing that hardware has a preferred algorithm, and that the preference is a single dimensionless ratio. Backpropagation was designed when that ratio was near one. It is now near three hundred, and the training stack of 2026 is the accumulated interest on the difference.
Appendix: parameters, sources, reproduction
Model parameters
| Symbol | Default | Source and uncertainty |
|---|---|---|
| \(b_w\), weight bytes per parameter per micro-batch | 12 | 2 bf16 reads + fp32 gradient read-modify-write; lean stacks 6 |
| \(a\), inter-kernel activation traffic per token per layer per unit \(d\) | 200 | kernel count of a fused FlashAttention layer, forward ≈ 52 d, backward ≈ 130 d; 120–300 depending on stack |
| saved activations per token per layer per unit \(d\) | 34 | Korthikanti et al. 2022, bf16, selective recompute; + 5 h s / d without |
| \(b_\text{opt}\), optimizer bytes per parameter per step | 30 | Adam mixed precision: read master, m, v, grad; write master, m, v, bf16 copy |
| \(b_\text{ar}\), all-reduce bytes per parameter per step | 4 | ring all-reduce of bf16 gradients, both directions |
| on-chip cache usable for re-blocking | ½ L2 | Hong–Kung bound: GEMM traffic ≥ MNK/√Z elements |
| HBM energy, all-in | 45–60 pJ/B | Dally 2023 (5 pJ/bit at the interface) + on-die transport; ±30% |
| FLOP energy, all-in at peak | 0.33–0.9 pJ | (TDP − static − HBM) ÷ peak per device; ±20% |
| \(B_\text{crit}\) | 4 M tokens | Kaplan et al. 2020 at pre-training losses; production runs use 2–16 M |
Method and limits
The model is a roofline: it bounds a step's time from below by the larger of its arithmetic time and its memory time, adds the elementwise phase as a separate memory-bound term, and counts bytes at one boundary, the device's main memory. It ignores kernel launch overhead, wave quantization, attention's quadratic term below a few hundred thousand tokens, and every inefficiency inside a GEMM, which is why a run's measured utilization sits below the model's ceiling. Historical machines are modeled with the same ledger using fp32 or fp64 numbers and the cache sizes of the era; their balance numbers depend on which peak one quotes, and the tables say which. Everything is reproducible: the JavaScript model that drives this page and the Python model in the repository share their constants, and the research notes under research/ record each number's provenance and the verification log against primary sources.
Notes and sources
- Reiner Pope with Dwarkesh Patel, "The math behind how LLMs are trained and served," Dwarkesh Podcast, 29 April 2026. dwarkesh.com/p/reiner-pope. Quotations in this report are from the published transcript.
- D. E. Rumelhart, G. E. Hinton and R. J. Williams, "Learning representations by back-propagating errors," Nature 323, 533–536 (1986); and chapter 8 of Parallel Distributed Processing, vol. 1 (MIT Press, 1986).
- Y. LeCun, L. Bottou, G. Orr and K.-R. Müller, "Efficient BackProp," in Neural Networks: Tricks of the Trade (Springer, 1998), §4.1; D. R. Wilson and T. R. Martinez, "The general inefficiency of batch training for gradient descent learning," Neural Networks 16 (2003). D. C. Plaut, S. J. Nowlan and G. E. Hinton, "Experiments on learning by back propagation," CMU-CS-86-126 (1986), for the blocks of 25. G. E. Hinton, "Connectionist learning procedures," Artificial Intelligence 40 (1989), §6.10, for the cubic estimate.
- DARPA Neural Network Study, MIT Lincoln Laboratory, ESD-TR-88-311 (1988), Part V ch. 3 for the definition and Table 3-2 for the rates; X. Zhang, M. McKenna, J. P. Mesirov and D. L. Waltz, "An efficient implementation of the back-propagation algorithm on the Connection Machine CM-2," NeurIPS 2 (1989); H. McCartor, "Back propagation implementation on the Adaptive Solutions CNAPS neurocomputer chip," NeurIPS 3 (1990), whose Table 1 collects the CUPS figures. T. J. Sejnowski and C. R. Rosenberg, "Parallel networks that learn to pronounce English text," Complex Systems 1 (1987), for the NETtalk rate. J. Dongarra's LINPACK table (netlib.org/benchmark/performance.pdf) for the VAX rating.
- S. Linnainmaa, master's thesis, University of Helsinki, 1970; P. Werbos, PhD thesis, Harvard, 1974. A. Griewank and A. Walther, Evaluating Derivatives (SIAM, 2nd ed. 2008) for the cheap-gradient principle and its memory cost.
- A. Griewank, "Achieving logarithmic growth of temporal and spatial complexity in reverse automatic differentiation," Optimization Methods and Software 1 (1992); A. Griewank and A. Walther, "Algorithm 799: revolve," ACM TOMS 26 (2000).
- W. A. Wulf and S. A. McKee, "Hitting the memory wall: implications of the obvious," ACM SIGARCH Computer Architecture News 23(1), 1995.
- D. A. Patterson, "Latency lags bandwidth," Communications of the ACM 47(10), 2004.
- A. Gholami, Z. Yao, S. Kim, C. Hooper, M. W. Mahoney and K. Keutzer, "AI and memory wall," IEEE Micro 44(3), 2024; arXiv:2403.14123.
- The companion note
backprop_memory_wall.mdin the repository derives every constant, with the energy price list in Dally's unit of the int8 add, and gives the A100 and H100 tables the figures here generalize. - S. McCandlish, J. Kaplan, D. Amodei and the OpenAI Dota team, "An empirical model of large-batch training," arXiv:1812.06162 (2018); J. Kaplan et al., "Scaling laws for neural language models," arXiv:2001.08361 (2020).
- T. Chen, B. Xu, C. Zhang and C. Guestrin, "Training deep nets with sublinear memory cost," arXiv:1604.06174 (2016). V. Korthikanti et al., "Reducing activation recomputation in large transformer models," arXiv:2205.05198 (2022) for the 34 d and selective recompute.
- A. N. Gomez, M. Ren, R. Urtasun and R. B. Grosse, "The reversible residual network: backpropagation without storing activations," NeurIPS 2017.
- T. Dao, D. Y. Fu, S. Ermon, A. Rudra and C. Ré, "FlashAttention: fast and memory-efficient exact attention with IO-awareness," NeurIPS 2022.
- Llama Team, AI @ Meta, "The Llama 3 herd of models," arXiv:2407.21783 (2024), sections 3.3–3.4.
- DeepSeek-AI, "DeepSeek-V3 technical report," arXiv:2412.19437 (2024), sections 3 and 4.
- Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard and L. D. Jackel, "Backpropagation applied to handwritten zip code recognition," Neural Computation 1(4), 1989: "All simulations were performed using the backpropagation simulator SN (Bottou and LeCun 1988) running on a SUN-4/260"; the 18-day Alliant figure for Waibel et al.'s TDNN is quoted there. J. D. McCalpin, "Memory bandwidth and machine balance in current high performance computers," IEEE TCCA Newsletter, December 1995, for the quotation in section 4; H. T. Kung, "Memory requirements for balanced computer architectures," ISCA 1986, for the α² result; N. P. Jouppi et al., "In-datacenter performance analysis of a tensor processing unit," ISCA 2017, for the TPU ridge point and workload table.
Additional sources, with the verification logs, are in the research/ directory of the repository: hardware specifications from vendor datasheets, the CUPS tables of the 1988 DARPA study and its successors, the classic papers' own descriptions of their hardware, and the training reports of GPT-3, PaLM, Llama 3 and DeepSeek-V3.