Notes toward a blog post · 31 August 2026

Backprop's days are numbered

Why backprop encodes a 1970 cost model, what it costs to keep it, and what would replace it — with the numbers.

Notes toward a blog post. Everything below is either a citation or a claim I'm prepared to defend. 31 August 2026.


The claim

I posted that in five years we won't be using backprop. Yann LeCun replied: "This is false."

So the stakes are set, and here is the argument.

Backprop encodes a cost model in which storing an intermediate value and reading it back later is free. That cost model was approximately true when reverse-mode automatic differentiation was invented in 1970, and roughly defensible when backprop was popularised in 1986. It is now wrong by four to five orders of magnitude. Every major systems technique in modern training — checkpointing, recomputation, pipelining, offloading, loss scaling, fused attention kernels — is a patch over that one wrong assumption, and the patches now cost more engineering than the algorithm they protect.

Note what this argument is not. It is not the biological-implausibility argument — cortex has no weight symmetry, backprop-through-time is not a thing brains can do, and so on. That argument is true and interesting, but it has a one-word rebuttal that everyone already knows ("planes don't flap"). My argument is about economics, not biology, and economics is checkable.


The gap, measured four ways

MeasureThenNowRatio
Energy: arithmetic vs. main-memory fetch (Dally, CACM 2022)"somewhat comparable"32-bit add 20 fJ vs. fetching two 32-bit words from DRAM 1.3 nJ64,000×
Energy: FP32 multiply vs. 32-bit DRAM read (Horowitz, ISSCC 2014, 45 nm)—3.7 pJ vs. 640 pJ173×
Scaling rate (Gholami et al., AI and Memory Wall, IEEE Micro 2024)—peak FLOPS 3.0×/2 yrs; DRAM bandwidth 1.6×/2 yrs; interconnect 1.4×/2 yrsover 20 yrs: FLOPS ×60,000, DRAM BW ×100, interconnect ×30
Roofline ridge point (FLOP per byte before you become compute-bound)~1 on a 1970s–80s machine~295 FLOP/byte on an H100 (989 TFLOP/s BF16 ÷ 3.35 TB/s)~300×

The one-sentence version: backprop was designed for a machine with a ridge point of 1 and is run on a machine with a ridge point of 300, and its defining behaviour — write down every intermediate, read them all back in reverse — is precisely the worst thing you can do on such a machine.

Bill Dally makes the same argument one layer down, about the RAM model itself, and writes my thesis for me in a single line (CACM 65(9), September 2022):

"This approximation was appropriate during the early days of computing when the costs of arithmetic and communication were somewhat comparable."

He is talking about RAM and PRAM as models of computation. I am extending the same observation up one level, from models of computation to learning algorithms.

The distance numbers

These are worth having because they make the cost geometric rather than just "memory is slow." From the same piece:

OperationEnergyTime
32-bit add20 fJ150 ps
Move two 32-bit words 1 mm on-chip1.9 pJ400 ps
Move 64 bits 40 mm (corner to corner, 400 mm² die)77 pJ16 ns
Go off-chip320 pJ6 ns/m
Read 64 b from an 8 KB SRAM sub-array0.64 pJ300 ps
Read 64 b from a 256 MB memory built of those sub-arrays58 pJ12.3 ns

The last two lines are the ones to remember. The bit cell is identical. 0.64 pJ becomes 58 pJ — a 90× penalty — purely for being far away. 57.4 of those 58 pJ are communication.

The software-side companion, in GPU clock cycles: register ~1, L1/shared ~30, L2 ~200, HBM ~500. Backprop's activation stash lives in the 500 tier by construction, because the whole point is that every forward activation must survive until the backward pass comes back for it.

A related claim I'll defend: analog and neuromorphic approaches buy you maybe a factor of three. Organising the data movement can buy several orders of magnitude. That is why this is a software problem, not a silicon problem.


Why backprop won, which is not the reason usually given

The usual story is that backprop falls out of the chain rule, so it is what calculus hands you. True, but shallow — and it doesn't explain the forty-year monopoly, because there were always other ways to get an update.

The real reason is that backprop is an interface.

We get backprop by enforcing separation of concerns. There is one group of people that provides a function, another that provides an optimizer, and they communicate through an API: evaluate the function, evaluate the gradient.

f(x) and ∇f(x) is the contract that let the modelling people and the optimization people stop talking to each other for four decades. Architectures could change without touching optimizers. Optimizers could change without touching architectures. That is an enormous amount of value.

It was worth its cost when arithmetic and memory cost about the same. At 64,000:1 it is not. And this reframing is what makes the "leaky abstraction" section below fall out for free: every workaround in the catalogue is a place where someone had to reach through the ∇f interface and touch the hardware directly.


Two problems, usually conflated

They have different fixes, and running them together is what makes most explanations of "backprop is memory-hungry" read as mush.

(a) Liveness. Every forward activation must stay alive from the moment it is produced until the backward pass reaches it — a lifetime of O(depth) — on a device whose fast memory holds O(1) layers' worth. Four layers, store four things. A hundred layers, store a hundred things. They don't fit in registers or cache, so they go to HBM, and every one of them is fetched back at 500-cycle, off-chip-energy prices.

(b) Sequential dependency. Layer k's update cannot be computed until layers k+1…n have been computed and then un-computed. The literature calls this backward locking or update locking; Jaderberg's synthetic-gradients paper named it in 2017. This is a latency and parallelism problem, not a memory problem.

(a) is the stronger one and the one that costs money. (b) is the one that has produced the more baroque engineering.


The catalogue: what it costs to keep backprop

Ordered by how cleanly each workaround is caused by backprop rather than by something else. This is the section that should convince a working ML engineer, because it is a description of their week.

Tier A — pure backprop tax

A1. Gradient / activation checkpointing. Store O(√n) activations, recompute the rest. The version Tim Salimans and I shipped in 2017 fit 10× larger models for ~20% extra compute time. The honest larger number is NVIDIA's: full activation recomputation costs 30–40% of execution time (Korthikanti et al., MLSys 2023). Their fix — selective recomputation plus sequence parallelism — gets overhead down to 2–4% and cuts activation memory ~5×, at the price of a hand-derived, architecture-specific formula for exactly which tensors to keep: sbh(34 + 5as/h) bytes per layer, where you must know sequence length, batch, hidden size and head count.

That is the leak, made visible. To make the gradient affordable you must abandon the idea that a gradient is a property of a function, and start reasoning about attention head counts. It is also, as of now, load-bearing infrastructure: gradient checkpointing is used in frontier training runs, and it is a hack.

A2. FlashAttention's backward recomputation. This is the strongest single piece of evidence.

FlashAttention never materialises the N×N attention matrix. In the backward pass it recomputes attention blocks in SRAM from a stored per-row log-sum-exp scalar. Result: 10–20× memory saving (linear in N rather than quadratic) and — the part that matters — the backward pass is 2–4× faster in wall-clock despite doing strictly more FLOPs (Dao et al., NeurIPS 2022).

Sit with that. Under backprop's assumptions, recomputing is strictly worse than remembering; that assumption is the entire justification for the backward pass storing anything at all. On real hardware, forgetting and recomputing is faster. The 1986 cost model is not merely inaccurate here — it is inverted. Everything else in this post is an argument. This one is a measurement, running in production, in every training stack in the world.

It also answers the obvious question of why nobody did this in 2018. Because that is how the array-programming abstraction was written: materialise the matrix, then differentiate it. The abstraction said store. The hardware said don't. It took until 2022 for anyone to ignore the abstraction.

A3. Loss scaling. In FP16 the gradients of a real network underflow to zero. The universal fix is to multiply the loss by 2^k before the backward pass and divide it out afterwards, with a controller that halves k on overflow. The number the chain rule computes is not a number the machine can represent, so we secretly multiply it by a power of two and hope. Cheapest example to explain; unarguably specific to backprop.

A4. Reversible architectures (RevNets, Reformer). Activations are reconstructed from the layer's outputs, giving O(1) activation memory in depth. Notice what happened: the architecture was constrained to fit the training algorithm's memory pattern. Nobody wanted a reversible residual block for modelling reasons. That is the definition of a legacy constraint — the tail wagging the dog.

A5. Activation offload to CPU / NVMe (ZeRO-Infinity and relatives). When activations don't fit in HBM, ship them across PCIe to host DRAM, or to an SSD, and ship them back for the backward pass. We move bytes off the accelerator, over a bus, onto a disk, so we can read them back in reverse order. Given 320 pJ just to leave the chip, this is the memory wall's reductio ad absurdum.

A6. Gradient accumulation. The batch you want doesn't fit because of activations, so you run it in slices and sum. Minor, but one more line of user-facing complexity that exists only to serve the backward pass.

Tier B — real, but only partly caused by backprop

B1. Pipeline parallelism. You split a model across devices for a reason that has nothing to do with backprop: the parameters don't fit. The backprop-caused part is the bubble, and it is caused precisely by the forward→backward dependency. GPipe's idle fraction is (p−1)/m for p stages and m micro-batches — with 8 stages and 8 micro-batches that is 47% of the machine doing nothing. Then look at what has been built to claw it back: micro-batching, 1F1B, interleaved 1F1B, zero-bubble schedules, and the detail that says it best — PipeDream's weight stashing, where each stage keeps several past versions of its own weights, because the backward pass arrives late and would otherwise be differentiating a model that no longer exists.

"We keep multiple copies of the model, from different points in time, so that the gradient arrives at the weights it was computed for" needs no hardware knowledge to recognise as a workaround, and it is a pure consequence of the backward dependency.

B2. Synthetic gradients / decoupled neural interfaces (Jaderberg et al., ICML 2017). The workaround that failed, which is why it is good evidence. It named update locking and attacked it by predicting the gradient rather than waiting for it. It does not scale. The point is that the field recognises the sequential dependency as a defect serious enough to warrant approximating the gradient away, and has been trying since 2017.

Tier C — examples that don't hold up

Batching. Mini-batching predates the 1986 paper, is used identically in methods with no backward pass at all, and exists for reasons unrelated to backprop's memory: gradient noise, and turning matrix-vector products into matrix-matrix products for arithmetic intensity. It is not evidence. (The exception is micro-batching inside a pipeline schedule — but that is B1.)

Optimizer-state sharding, 8-bit Adam, Adafactor. This is Adam's tax — two extra FP32 copies per parameter, 12Φ of the 16Φ bytes of model state — not backprop's. Related, but a different argument.


The GIL

Python's GILBackprop
Introducedearly CPython, ~1992reverse-mode AD 1970; popularised Oct 1986
The world at the timeone core; multicore was exoticarithmetic and memory cost about the same
Why it wona simple, correct interface: refcounting stays sound, C extensions stay easya simple, correct interface: f(x) and ∇f(x), so modellers and optimizer authors never have to meet
The workaround eramultiprocessing (PEP 371, Python 2.6, 2008), C extensions releasing the lock, subinterpreters (PEP 554/734)checkpointing (2016–17), pipelining (2018–19), ZeRO-Offload (2020), FlashAttention (2022), selective recompute (2022)
The tellyou pay a process boundary and a serialization step to use a core you already ownyou pay a recomputation and a memory-hierarchy round trip to use FLOPs you already own
RemovalPEP 703 accepted 24 Oct 2023; experimental in 3.13 (Oct 2024); officially supported in 3.14 (Oct 2025)—
Cost of removal5–10% single-thread slowdownunknown; that is the research programme
Elapsed~33 years40 years since Nature; 56 since Linnainmaa

The GIL took 33 years from introduction to supported removal. Backprop is at 40 years and counting. On GIL-time, backprop is already seven years overdue — and unlike the GIL, nobody has even written the PEP.

The second lesson is the honest one. The GIL was not removed when someone proved it was bad. Everyone knew it was bad by 2000. It was removed when one person shipped a working implementation whose cost was small enough to argue about. The argument was never the bottleneck. The artifact was.


"Even gods make mistakes — look at the platypus"

This was offered to me as a counter-argument. It is the best argument for my position that anyone has made against it.

The platypus's oddities are not mistakes. Egg-laying is the ancestral amniote condition; monotremes are simply the lineage that never lost it. The platypus is not an error — it is a frozen accident that survived because nothing in its environment ever selected against it. Isolated continent, low competition, an unchanging niche.

Which is exactly my argument about backprop. It is not here because it is right. It is here because for forty years the environment never punished it hard enough. Energy prices and the memory wall are the environment changing.

The platypus kept laying eggs because Australia never made it pay. Backprop keeps storing every activation because, until recently, nobody was paying the electricity bill.

The better vestigial examples

The recurrent laryngeal nerve. In every tetrapod it runs from the brain down the neck, loops around the aortic arch, and comes back up to the larynx a few centimetres from where it started — because in the fish ancestor the route was direct and the neck had not been invented yet. In the giraffe the detour is ~4.6 metres to travel a few centimetres. In sauropods with 14-metre necks, Wedel (2012) estimates single neurons at least 28 metres long; his paper is titled, perfectly, "A monument of inefficiency."

The analogy is exact: the nerve's route was fixed when the distance was zero. Backprop's memory pattern was fixed when the distance — in energy — was zero. Then the neck grew.

The vertebrate retina is wired backwards. Photoreceptors sit behind the nerve fibres and blood vessels, so light passes through the wiring, and the bundle has to punch a hole through the retina to leave — which is the blind spot. Cephalopod eyes evolved independently, are wired the sensible way, and have no blind spot. Same function, better design, different starting point. This is the strongest one for my argument, because it shows the good design is reachable — just not from where the vertebrates started.


What replaces it

The honest answer is that I don't know, which is the whole reason for the research programme. But the shape is constrained, and there is a precedent that does most of the persuading.

The shape. A GPU is a hierarchy: register ~1 cycle, L1 ~30, L2 ~200, HBM ~500. You want a computation that arranges itself to match — a lot of arithmetic per HBM fetch, a bit less per L2 access, less still per L1. Something that fits the vessel we were given. Backprop was designed from the point of view of abstract linear algebra, by people who by professional convention were not supposed to think about hardware at all.

The precedent, which is the best argument in the whole programme. In numerical PDE solving, you start with gradient descent, because it is the general-purpose method that works on anything. Then you exploit the structure of the operator and you get multigrid, which solves at multiple scales simultaneously and is enormously more efficient. Nobody solves Poisson with steepest descent any more.

So: what is the multigrid of learning?

That question is better than "backprop's days are numbered," because it is constructive rather than merely contrarian, and because it answers the reader's real objection — fine, but what else is there? — with a historical fact rather than a promise. It is also much harder to call false: the claim becomes the specialised method exists and has not been found yet, and the entire history of numerical analysis is on that side.

The direction I actually believe in is architecture-specific updates. Adam is a black box — it works no matter what the parameters are. Muon is less of a black box: it assumes the parameters are matrices, though it doesn't care how they're combined. Compositional Muon goes further and assumes the matrices are combined in the QKV pattern. That axis — updates that know what they're updating — is barely explored. I don't believe in optimizer research; I believe in that.

And there is already evidence in the small. Sparse parity is a task whose textbook solution is a small neural network trained with backprop. Point an agent at it with an energy budget and it solves it with Gaussian elimination — no backprop anywhere. Application-specific solutions that skip the gradient entirely already exist; we just haven't been looking for them, because for forty years finding them by hand was more expensive than paying the tax.

That last clause is the thing that changed. Software got cheap in 2025. The search is now affordable.


The evidence in the small

From cybertronai/sutro-problems, where energy is scored under a simplified version of Dally's model. Last week alone:

Problem / target26–27 Aug 202630 Aug 2026Improvement
sparse-parity, 20% secret recovery17,331,683 (Gray scan)1,317,480 (static-IS walk + staged layout)13.2×
sparse-parity, 100% secret recovery43,325,468 (full Gray walk)12,461,6103.5×
matmul 4×41,316 (naive)800 (outer product)1.6×

A 13× energy improvement in four days, on a task whose textbook solution is backprop, by methods that contain no backprop. Small, but dated, public, and verifiable by anyone.


What would actually settle it

Three things, roughly in order of how much they'd move the argument:

  1. Measure the tax directly. Take a NanoGPT-scale run, instrument one training step, and report the split: joules (or bytes moved) attributable to activation stores plus backward-pass activation reads, versus arithmetic. Nobody has published this cleanly. On any long-context configuration it will be well above half, and then the claim stops borrowing Dally's numbers and has one of its own.
  2. Post the benchmark, not the essay. "Beat backprop on WikiText at fixed loss, measured in data-movement distance." Per the GIL: the argument is not what ends these things, the artifact is.
  3. Ask the sceptics what would change their mind, in advance. A prediction with no agreed falsifier is entertainment. Five years is a long time to be vague.

Sources

ClaimSource
Arithmetic vs. communication, 20 fJ vs 1.3 nJ = 64,000×W. Dally, "On the Model of Computation: Point," CACM 65(9), Sept 2022
FLOPS 3.0×/2yr vs DRAM BW 1.6×/2yr; ×60,000 vs ×100 over 20 yrsGholami, Yao, Kim, Hooper, Mahoney, Keutzer, "AI and Memory Wall," IEEE Micro 2024 (arXiv:2403.14123)
FP32 multiply 3.7 pJ vs 32-bit DRAM read 640 pJ @45 nmM. Horowitz, "Computing's Energy Problem," ISSCC 2014
CPU–DRAM divergence; "memory wall" coinedWulf & McKee, "Hitting the Memory Wall," ACM SIGARCH CAN, Dec 1994
Backprop popularisedRumelhart, Hinton & Williams, Nature 323:533–536, 9 Oct 1986
Reverse-mode AD inventedLinnainmaa 1970 (MSc thesis); Werbos 1974 (PhD thesis)
10× larger models for ~20% extra computeSalimans & Bulatov, gradient-checkpointing, 2017
Full recompute 30–40%; selective 2–4%; 5× activation memory; sbh(34+5as/h)Korthikanti et al., "Reducing Activation Recomputation in Large Transformer Models," MLSys 2023 (arXiv:2205.05198)
10–20× memory saving; backward 2–4× faster despite more FLOPsDao, Fu, Ermon, Rudra, Ré, "FlashAttention," NeurIPS 2022 (arXiv:2205.14135)
Pipeline bubble (p−1)/m; weight stashing; interleaved 1F1BHuang et al., GPipe 2019; Narayanan et al., PipeDream, SOSP 2019; Narayanan et al., Megatron, SC 2021
Update / backward locking namedJaderberg et al., "Decoupled Neural Interfaces using Synthetic Gradients," ICML 2017
GIL removal timeline; 5–10% single-thread costPEP 703 (accepted 24 Oct 2023); PEP 779; CPython 3.13 / 3.14 release notes
Giraffe RLN ~4.6 m; sauropod neurons ≥28 mWedel, "A monument of inefficiency," Acta Palaeontologica Polonica 57(2), 2012
Distance-aware cost modelsGianinazzi et al., "The spatial computer," arXiv:2205.04934; and my notes
Energy records for sparse parity and matmulcybertronai/sutro-problems