There is a claim that circulates freely in conversations about AI and energy: that roughly half the cost of running a GPU is the electricity, and half is the GPU. It is wrong by about five-fold, and the error propagates — because if you believe it, you conclude that cheap kilowatt-hours are the lever. They are not. This note works through where the money in a GPU-hour actually goes, and then through six adjacent questions that the correction reframes.
1 — The cost stack: electricity is 7% of a GPU-hour
Epoch AI’s itemised 1 GW AI datacentre model (May 2026; GB200 NVL72, 5-year IT life, 8.34 c/kWh, PUE 1.14, 71% utilisation) breaks $8.5B of annualised TCO into:
servers $5,021M (60%) · facility $1,387M (16%) · network $1,167M (14%) · energy $594M (7%) · tax, maintenance, labour, land, water ~$342M (4%)
Upfront capex: $37.9B — servers $21.2B, facility $11.4B, network $4.9B.
Two independent bottom-up models land in the same place. A four-year Meta H100 cluster splits 46% GPUs / 24% infrastructure / 20% colocation rent / 9% electricity / 1% other. A self-hosted 8×H100 server at $15.26/hr splits 35.9% depreciation / 20.7% capital cost / 16.9% colocation / 11.7% electricity — $0.15–0.31 per GPU-hour.
What would it take for energy to dominate?
Holding hardware cost constant, electricity reaches 50% of TCO at $1.11/kWh and 90% at roughly $10/kWh — 13× and ~120× the model’s baseline. Put differently: even if the accelerators and the networking were free, electricity would still be only ~26% of the cost, because the shell, substations, transformers and cooling plant amortise at $1,387M/yr against a $594M/yr electricity bill.
And the ratio is currently moving away from energy dominance, not toward it. HBM4 runs ~$560/stack against ~$370 for HBM3E, a 50% increase; UBS models 2027 HBM average selling prices up 79%; NVIDIA is planning a >15% server price rise for early 2027. Conventional DRAM contract prices moved +93–98% quarter-on-quarter in Q1 2026 and +58–63% in Q2.
The reframe that survives
The scarce good is not cheap kilowatt-hours. It is firm kilowatts at a specific interconnection point, on a schedule. The energy-infrastructure business hiding inside the “facility” line — substations, switchgear, transformers, turbines, chillers, battery storage — costs 2.3× more per year than the electricity flowing through it. That is why every serious distributed-compute pitch in 2026 sells speed to power rather than cheap power.
There are two regimes where the intuition about energy dominance is right, and both are worth naming precisely. First, cash opex rather than total cost of ownership: at 90% utilisation, energy is 59% of cash opex on Quincy, Washington hydro and 77% in PJM Dominion. Second, the post-depreciation regime — fully written-off silicon sitting in a sunk shell, which is what the 2023–2025 fleet becomes around 2028–2030. Nobody appears to be modelling that second-hand market, and it is the one place a distributed-compute business could acquire hardware at a price that works.
2 — Why a single user cannot make local inference cheap
The recurring counter-argument is that per-person token demand will grow two to three orders of magnitude, that price compression cannot keep pace, and that people will therefore run models locally because it is cheaper. The demand premise is well supported. The cost conclusion is not, and the margin is not close.
| Measurement | Number | What it means |
|---|---|---|
| Roofline at batch 1 | ~0.5 FLOPs/byte against an H100 ridge point of ~295 (989 TFLOPS / 3,350 GB/s) | A single-stream user runs the chip at roughly 0.2% of peak FLOPs. Batch 240 — H100’s critical batch — reaches ~120 FLOPs/byte. |
| Same box, batching only | NVIDIA DGX Spark, Nemotron Super 49B NVFP4: 5.79 tok/s single-stream → 695 tok/s at concurrency 256 (120×). gpt-oss-120b MXFP4: 33.5 → 863 tok/s (25.7×). | The cleanest isolation of the aggregation effect, on hardware an individual could buy. One user leaves ~99% of their own machine idle. |
| Per accelerator | GB200 NVL72 serving DeepSeek R1 at 65 tok/s/user: 6,057 tok/s/chip at $0.085 per million tokens. H200 at the same interactivity: 1,035 tok/s/chip at $0.327/M. | Against ~18 tok/s for the same 671B model on a Mac Studio M3 Ultra — roughly 336× per chip. |
| Utilisation penalty | $0.21 to $15.25 per million output tokens on identical H100 hardware, driven only by offered request rate. A 72.6× spread. | Underutilisation costs 2.5–24× at 1–10 requests/sec, up to 36.3× near idle. |
All-in, at a realistic personal duty cycle: self-hosting costs $30–160 per million tokens against $0.09–0.28 hosted for the same open-weight model. Decomposed — a $5,499 Mac Studio amortised over three years at 18 tok/s is $3.23/M at a 100% duty cycle, $32.29 at 10%, and $161.46 at 2%. Two per cent is generous: a single human at batch 1 cannot physically consume more than about 1.55M tokens/day. Electricity alone is $1.03/M in California (3.09 kWh per million tokens at 33.25 c/kWh) and $0.55/M at the US average.
Price compression has already done three orders of magnitude once
a16z’s LLMflation series tracked GPT-3-quality inference from $60 per million tokens in November 2021 to $0.06 by late 2024 — a factor of 1,000 in three years, pinned at fixed capability (MMLU 42 and MMLU 83 thresholds). Epoch AI’s independent series gives a range of 9× to 900× per year depending on the benchmark; GPT-4-level GPQA fell 40×/year.
The stricter measurement is more interesting. MIT FutureTech (Gundlach, Lynch, Mertens & Thompson, March 2026) measures at fixed benchmark performance rather than fixed token price, on GPQA-Diamond, AIME and SWE-bench Verified across Pareto-frontier models: prices fall 5–10× per year at a given capability, algorithmic progress contributes ~3×/year, and the cost of running the actual frontier is rising 3–18× per year, because reasoning models burn more tokens per completed task.
So the spread between frontier and commodity widens by something like 50–100× a year. That is the real bifurcation — frontier versus commodity, not local versus cloud. And note the conflation that does most of the damage in these arguments: “open weights win” and “local wins” are different claims, and only the first has data behind it. OpenRouter’s 100-trillion-token study puts open-weight models at ~30% of tokens by late 2025, with Chinese-origin models above 45% of traffic by mid-2026 — consumed almost entirely as hosted API.
3 — The market for idle and heterogeneous GPUs
A market for unreliable, heterogeneous capacity does exist. It is roughly three orders of magnitude too small to matter, and it clears at a 4–20× discount that tells you exactly why.
| Tier | Price | Note |
|---|---|---|
AWS p6-b200.48xlarge | $113.93/hr = $14.24/GPU-hr | The homogeneous 8-GPU NVLink island. Capacity Blocks $12.355 per accelerator-hour. You cannot buy a fraction of it, or a variant. |
| H100 on-demand, market median | $3.39, range $0.67–$14.19 across 38 providers | A 21× spread. Vast.ai spot $0.67, Lambda $3.29, AWS $6.88, Azure $6.98, GCP $14.19. |
| SF Compute, short-term | ~$0.80–1.00/hr | A genuine order book with bids and asks. Utilisation sits near 100% because the price falls until it does. |
| Salad, consumer GPUs | RTX 4090 $0.160/hr; 3090 $0.090; 3080 $0.060; 3070 $0.040; batch tier from $0.02 | 1M+ underutilised consumer GPUs, 450,000+ providers, 180+ countries. The supply side is emphatically solved. |
The single most instructive datapoint: Prime Intellect — the flagship decentralised-training company, ~$100M ARR at a $1B valuation, the organisation with the strongest possible incentive to demonstrate otherwise — trained its best model, INTELLECT-3 (106B MoE, 12B active), on a centralised 512×H200 cluster across 64 nodes with 400 Gbps NDR InfiniBand. Their own survey of the field concedes that no one has successfully scaled this research to train state-of-the-art models. The marketplace business works; the decentralised-training thesis has not yet.
The ceiling, expressed as a price
OpenAI and Anthropic both discount batch inference by exactly 50% against synchronous, for a 24-hour completion window. Sail Research’s “flex” window reaches 80% off, the highest observed market price of latency tolerance. That 50–80% band is what latency tolerance is worth — and note what it buys the provider: trough-filling and larger, more arithmetic-intensity-efficient batches on the same datacentre GPUs. It is not payment for running on worse hardware. Any idle-capacity business has to fit underneath that band and additionally absorb checkpointing, verification, straggler and egress costs.
The heterogeneity results, scored honestly
DiLoCo’s headline “500× less communication” was 8 workers, tested only up to 400M parameters. The largest genuinely decentralised runs are INTELLECT-1 (10B parameters, 6T tokens, 98% compute utilisation, 400× communication reduction via int8 in production and up to 2000× experimentally) and INTELLECT-2 (32B, decentralised RL). Both sit one to two orders of magnitude below the frontier. Petals, the volunteer peer-to-peer inference project, delivers ~6 tok/s on Llama-2-70B and ~4 tok/s on Falcon-180B in single-batch inference; its public network-status widget no longer loads. Exo Labs has retreated to LAN-scoped clusters, which is the honest engineering answer to residential bandwidth — and also a concession on the market question.
The assumption most likely to sink an idle-capacity business, and it is current
The predicted 2026 compute glut did not arrive. H100 hourly rates fell from ~$8 in early 2024 to $1.96 in late 2025, then climbed back to $2.64–2.74 by mid-2026 and are consolidating in a $2.60–2.80 corridor. One-year contract pricing rose ~40% from a $1.70 low in October 2025 to $2.35 by March 2026; Blackwell contracts rose 48% over the same crunch. A tightening market means operators have less idle capacity to dump — the supply side of the idle-compute business is evaporating at precisely the moment the demand side, in the form of long-horizon asynchronous agents, finally appears.
4 — Does a GPU depreciate?
The popular framing is that GPUs depreciate quickly for training, because you want the latest silicon, and slowly for inference, because you do not. The economics behind that intuition are sound; the framing is a category error worth fixing.
An asset does not have workload-specific depreciation rates. It has one cash-flow decay path, set by the best-paying workload the card can still win. What is true is that inference provides a floor under that path, because inference is far less sensitive to frontier FLOPs per dollar and far more sensitive to memory bandwidth per dollar and to already-powered rack space.
And the variable that decides whether hyperscaler earnings are overstated is not the chip’s physical life at all — it is whether its revenue declines faster than straight-line. Straight-line over six years understates depreciation if the cash-flow profile is front-loaded, even if the chip physically runs for nine years. That is why this argument never resolves: the accounting bear case and the hardware-longevity bull case are not actually in contradiction.
The empirical record
| Series | Reading |
|---|---|
| A100 rental | $1.61 on the Silicon Data neocloud index; $1.80 median across 34 providers, explicitly flat over 90 days. § The same site returned $1.29 on its dedicated A100 page minutes apart; treat spot as a $1.29–1.86 band. |
| H100 rental | $2.67 on the index, $3.39 as a market median — up 13% over 90 days. The real A100/H100 ratio is ~1.7×. |
| V100 rental (2017 part) | $0.97 median, $0.11 floor — 6–15% of an A100’s rate nine years post-launch. Azure retired its V100-backed NCv3 VMs in September 2025, ~7.5 years after launch. That is the actual observed ceiling on hyperscaler GPU life, and it is a retirement. |
| Used A100 80GB | Roughly 60–80% of value retained, against a launch-era list of ~$15–20k for the 80GB PCIe part. § The quoted resale band is wide and comes from vendors with an interest in the number. |
Why an obsolete card keeps clearing at yesterday’s price. NVIDIA issued an A100 end-of-life notice in January 2024, which is part of it. But the larger cause is that powered floor space, not silicon, is the binding constraint: CoreWeave at Q2 2026 held 1.5 GW active against 3.7 GW contracted — a 2.2 GW gap between what it has sold and what it can energise. An installed six-year-old card already has a rack, a power hookup and cooling, and those are the scarce goods.
The accounting moves, in one table
| Company | Change to server useful life | P&L effect |
|---|---|---|
| Microsoft | 4 → 6 years (announced on the Q4 FY2022 call) | ~+$3.7B FY23 |
| Alphabet | Servers 4 → 6, some network equipment 5 → 6, effective Jan 2023 | $3.9B FY23 |
| Meta | Most servers to 5.5 years, effective Jan 2025. Non-AI servers 6 → 7 on 29 April 2026 because of the memory shortage — 8 years was declined as too risky; failure rate rises 4.8% → 7.4% | ~$2.9B in 2025 |
| Amazon | 5 → 6 years effective Jan 2024, then reversed to 5 effective 1 Jan 2025, citing artificial intelligence by name | Reversal: +$1.4B depreciation, −$1.0B net income, −$0.10/share |
| CoreWeave | 6 years, unchanged since 2023; average contract term ~5 years | — |
Amazon is the datum that cuts hardest against the longevity case: the largest cloud operator on earth examined its own fleet and shortened the life. Meanwhile Meta’s servers-and-network depreciation went $7.32B → $11.34B → $13.36B across 2023–2025 — up 83% in two years despite the life extension, because capex growth swamps the accounting lever.
5 — Packaging, not wafers, is the bottleneck
On the 16 July 2026 earnings call, TSMC’s CEO put it on the record: packaging capacity is so tight that it is limiting customers’ growth, with the supply–demand gap closing “probably 2029, 2030.”
| Question | Where it lands |
|---|---|
| Is CoWoS TSMC-only? | ~95% of leading-edge 2.5D capacity, so functionally a monopoly today — but not literally sole-source, and the moat is cracking. ASE/SPIL and Amkor run licensed chip-on-wafer and wafer-on-substrate steps and are expected to add 50,000–60,000 wafers/month by late 2026. Intel’s EMIB is the real alternative: already shipping AWS Trainium, and has won Google’s TPU v8e for H2 2027 plus a >3M-unit TPU packaging booking for 2028. |
| Capacity trajectory | TSMC CoWoS: 13k wafers/month at end-2023 → ~70–80k in 2025 → 120–140k end-2026 → 190–200k in 2027. But demand roughly doubles alongside it — 1.3–1.4M wafers in 2026 to 2.5–2.7M in 2027 — so the ~20% gap only halves to ~10%. |
| Who holds the allocation | Out of ~1.0M wafers of 2026 demand: NVIDIA 595,000 (60%); Broadcom 150,000 (15%), including Google TPU 90,000, Meta 50,000, OpenAI 10,000; AMD 105,000; Marvell 55,000; Amazon 50,000; MediaTek 20,000. The top customers lock >85%, leaving under 15% for everyone else. This is why a TSMC commitment is the single most valuable asset an accelerator startup can hold. |
| Is lithography the constraint? | No. ASML is adding roughly 30% EUV capacity for 2027 with another 30% under study for 2028, against ~65 low-NA EUV units in 2026. |
| Wafers or packaging? | Packaging, in 2026. Leading-edge wafer capacity is tight but expanding faster in percentage terms — N3 heading to ~180k wafers/month (+40% YoY), N2 to ~90–100k by end-2026 and sold out for the year. Whether that flips by 2028–2029 is genuinely open. |
| Memory | All three HBM suppliers have reportedly sold out 2027 capacity, not just 2026. HBM cost content per accelerator: ~$1,350 for an H100, ~$3,250 for a B200, ~$4,350 for an MI325X. |
6 — Compute in homes: the graveyard, and why 2026 is different
Two live programmes put general-purpose compute inside residential buildings, and they are an order of magnitude apart in a way that decides whether either is a business.
| Sunrun AI Compute Node | SPAN XFRA | |
|---|---|---|
| Announced | 8 July 2026 (week 28) | 13 April 2026 (week 16) |
| The box | “About the size of a small desktop computer”; power draw undisclosed | 16× NVIDIA RTX PRO 6000 Blackwell Server Edition, 4× AMD EPYC, 3 TB RAM, direct liquid-cooled, exterior-wall mounted. 12.5 kW continuous. Reported hardware value >$200k |
| Money flow | Sunrun owns and maintains the node, sells the inference capacity to enterprise buyers, and compensates the homeowner for hosting | SPAN owns the hardware; the homeowner pays SPAN ~$150/month flat, covering the home’s entire electricity and internet, plus usage-based compensation |
| Gate | An existing Sunrun solar + battery installation — a base of 1.1M+ homes | New construction, through a national homebuilder |
| Scale | Pilot. Management guided commercialisation to “late 2027 or 2028” on the Q2 2026 call (5 August 2026) | 100 homes / ~1.25 MW / 1,600 GPUs in the US Southwest this autumn; 80,000 nodes and >1 GW targeted from 2027. The stated pitch: ~$3M per MW in six months |
Why every previous attempt died, and it is a single diagnosis
Nerdalize, Qarnot, Heata, Cloud&Heat and Project Exergy all monetised waste heat. Heata’s entire consumer-side proposition is up to 4 kWh/day of domestic hot water — worth £120–340 per home per year. That is the whole budget from which hardware, installation, servicing and churn must be funded, and it cannot work at any scale. Nerdalize raised ~€1.6M lifetime and filed for bankruptcy in January 2019 with a €882k loss in 2017 alone.
The 2026 programmes monetise spare amps and a place in the interconnection queue instead of spare warmth — a value pool 50–500× larger. Which is exactly why the larger of the two is liquid-cooled and rejects its heat outdoors by design. Neither company claims a waste-heat credit.
Solar is not the driver, despite the framing
Residential solar delivers a ~25% capacity factor; running a GPU node only when the sun shines destroys the hardware economics. The actual drivers are speed to power — an electrician’s day of work against a three-to-five-year utility interconnection queue — and export compensation collapse. Under California’s post-April-2023 net billing tariff, solar customers receive $0.05–0.08/kWh for exported energy against $0.30–0.45 retail, so the opportunity cost of a midday solar kilowatt-hour is five to eight cents. That makes burning it locally nearly free — but it is a tariff artefact, not a physical one, and it evaporates if the tariff changes.
On the thermodynamics that always comes up: a 300 W computer is a 300 W resistive heater, to within a rounding error — photons through the window, acoustics and signal down the network link account for order 0.01–0.1%. The caveat that matters is the comparison class: against a heat pump at a coefficient of performance of 3.5, that heat is worth the electricity price divided by 3.5, so roughly $0.10–0.13/kWh in California rather than $0.35–0.45.
A false claim to be aware of. A widely-circulated story since mid-2026 holds that a chip vendor will pay homeowners $22,000 — or $220,000 — a year to host a miniature datacentre. It is not true. It originated in social posts and was laundered through aggregator sites. The ~$200,000 figure is the value of the hardware inside the node, which the homeowner does not own. In the actual programme the homeowner pays a flat monthly fee that covers their utilities.
7 — Grid headroom, flexible load, and what V2G actually pays
US power systems run at a 53% average load factor across 22 balancing authorities, 2016–2024 (range 43–61%). The idle capacity is real. The question is what unit it is denominated in.
Curtailment-enabled headroom — the ladder, exactly
76 GW at 0.25% average annual curtailment · 98 GW at 0.5% · 126 GW at 1.0% · 215 GW at 5.0%, across the 22 largest US balancing authorities covering ~95% of load. 76 GW is 10% of current US aggregate peak demand. Average curtailment event duration: 1.7 / 2.1 / 2.5 hours. Hours per year in which any curtailment is called: 85 / 177 / 366. By balancing authority at 0.5%: PJM 18, MISO 15, ERCOT 10, SPP 10, Southern 8 GW.
Read the Limitations section before quoting the headline. The authors state plainly that transmission capacity, ramping capability and ramp-feasible reserves are beyond the study’s scope, and that the results therefore cannot be taken as an accurate estimate of the load that can be added to the system. Those two pages separate someone who read the study from someone who read a headline about it.
What has actually been demonstrated
The flagship flexible-datacentre demonstration ran on a 256-GPU A100 cluster inside a cloud region in Phoenix, cutting power 25% below average base load for three hours on two utility system peaks (1 and 3 May 2025), with a 15-minute graceful ramp down and back up. At a 400 W cap that is roughly 100 kW of GPU nameplate — beautifully executed, and three orders of magnitude below the scale its press coverage implies. Separately, demand-response agreements signed in August 2025 between a hyperscaler and two utilities are real but carry no disclosed MW, MWh or hours. In ERCOT, registered large controllable load stands at ~5,302 MW, about 2% of total load. The company commercialising this raised $150M at a $1.05B valuation on 25 August 2026.
The objection that kills the naive residential version
The binding constraint is the service transformer and the feeder, not the panel. A study of 1,500 real Baltimore Gas & Electric feeders — k-means clustering to seven representative feeders, OpenDSS time-series load flow, AMI-derived base load — projects over 35% of distribution transformers overloaded by 2035 from EV charging alone. Annual transformer loss-of-life at 100% EV penetration runs 8.97% under demand charging against 4.30% off-peak. A constant 1–5 kW compute load is a worse profile than EV charging, because it never goes away.
On the panel itself: analysis of over 100,000 single-family homes found that more than 80% never draw more than 40 amps at any instant in a year, and 99% never exceed 100 amps. Forty amps at 240 V is 9.6 kW — 20% of a 200 A service. Average residential draw is on the order of 1–2% of a 200 A rating. So the headroom is larger than the usual “homes use about 40% of capacity” claim suggests, and that claim is anyway a statement about annual peak against a 100 A service, not an average against anything.
Vehicle-to-grid, with the number to hold it to
The best-measured trial in the world — 135 households in Great Britain, combined capacity under 1 MW, run with the national control centre — put the incremental value of bidirectional charging over merely smart charging on a time-of-use tariff at £180 per vehicle-year. Against bidirectional-charger rebates of up to $4,500, the payback arithmetic writes itself. The best-executed deployment anywhere runs 150 bidirectional vehicles on 7 kW chargers for a total fleet peak discharge of 300 kW — 2 kW per car realised; 50 of those cars discharged 65,000 kWh in five months. One major manufacturer has 250,000 bidirectional-capable vehicles on US roads after a mid-2026 over-the-air update and projects 52,000 enrolled in a California utility programme by 2030.
The energy-versus-power confusion is what makes the EV-battery argument sound stronger than it is. The US light-duty EV fleet holds roughly 300 GWh (~4M vehicles × ~75 kWh) against 144 GWh of cumulative US grid-connected storage installed since 2019 — so about 2.1–2.9×, not 3×, and comparing kWh to kWh when what a datacentre buys is firm kW at an interconnection point.
Two things any version of this thesis will be hit with
A large fraction of the interconnection queue is phantom. AEP Ohio’s total volume of large-load interconnection requests fell from 30 GW to 13 GW; Georgia Power’s expected large-load additions dropped by 6 GW, with Winter 2028 and Winter 2029 alone falling 1.4 GW. The same project shops itself to multiple utilities.
Datacentres are currently raising other ratepayers’ bills. PJM’s Independent Market Monitor attributes 63% of the 2025/26 capacity price increase directly to datacentre demand — $9.3 billion in additional customer costs. The flexibility thesis is in fact the strongest available answer to that, since a curtailable load does not drive capacity procurement, and it is worth leading with rather than defending.
8 — Beyond backprop: getting the energy numbers right first
The argument for redesigning learning algorithms around the memory hierarchy rests on a set of numbers that are widely quoted and frequently garbled. Here they are at source.
Arithmetic and memory energy
Horowitz’s ISSCC 2014 table, 45nm at 0.9 V: integer add 0.03 pJ (8-bit) / 0.1 pJ (32-bit); floating-point add 0.4 / 0.9 pJ (16/32-bit); integer multiply 0.2 / 3.1 pJ; floating-point multiply 1.1 / 3.7 pJ. A 64-bit cache read costs 10 pJ from 8 KB, 20 pJ from 32 KB, 100 pJ from 1 MB. DRAM: 1.3–2.6 nJ per 64-bit access. Register-file access is 6 pJ; total instruction energy is ~70 pJ, of which the arithmetic is a rounding error against fetch and clocking.
Where HBM energy actually goes — and it is not where the intuition says
From a physical floorplan model at real application toggle rates: HBM2 costs 3.92 pJ/bit in total, decomposed as 2.24 pJ/bit moving inside the DRAM die (~9.9 mm of wire), 1.21 pJ/bit row activation, and 0.30 pJ/bit across the interposer. GDDR5 was 14.0 pJ/bit. HBM3E is 3.44–4.05 pJ/bit — essentially unchanged across three generations despite the bandwidth scaling.
The package hop is about 8% of the access energy. So “the cost is the distance to the edge of the chip” is the wrong mental model. The design consequence is concrete: reducing the number of accesses helps proportionally, but the 1.21 pJ/bit activation floor only falls if locality improves within a DRAM row. An algorithm that halves bytes moved while randomising the address stream can be a net energy loss. Use pJ/bit rather than pJ/access — access granularity does a lot of hidden work in the usual comparison.
The latency ladder, measured rather than remembered
Pointer-chase measurements on H800 (Hopper): L1 32.0, shared 29.0, L2 264.5–502, global 656 cycles. A100: 33.0 / 29.0 / 202.8–408 / 566. RTX 4090: 32.0 / 30.1 / 273 / 571. The commonly-quoted “L2 is 200, HBM is 500” are the A100 best cases. And the register figure should be 4 cycles of dependent-instruction latency on both GH100 and GB203, not 1 — which makes the register-to-L1 gap 8×, not 30×.
One measurement worth having in isolation: on H800, an L2 near hit costs 258.0 cycles and a far hit 414.1. 156 cycles is the directly measured cost of crossing to the far L2 partition — on-die distance, priced.
What the alternatives have actually achieved
- Locality does not imply efficiency. Measured on Fashion-MNIST at comparable accuracy (89.63% vs 88.88%), Forward-Forward consumed 14.28 Wh against backprop’s 1.48 Wh — 862% worse — and took 574.6 s against 43.1 s. Its contrastive objective needs multiple forward passes, which is more memory traffic, not less. Any candidate rule has to be metered in joules on real hardware from day one, or the search will select for rules that look local and cost more.
- The genuine 2026 result. Equilibrium propagation reached 13.23% top-5 error on full ImageNet against a 12.2% backprop baseline — the first result of its family at that scale. But note the cost: it requires running the network to equilibrium over many relaxation steps, which is more compute and more traffic per example. It buys locality, not efficiency, and the paper reports error rates rather than joules.
- Three memory pressures, not one. Activation liveness is largely solved — rematerialisation gives O(√n) (48 GB → 7 GB for a 1,000-layer ResNet at ~30% extra time), reversibility gives O(1) in depth. Attention’s O(N²) intermediate is solved exactly by FlashAttention, which never materialises it: up to 9× fewer HBM accesses, and FlashAttention-4 at 1,613 TFLOP/s and 71% hardware utilisation. Optimizer-state traffic — roughly four parameter-bytes moved per fused multiply-add — is the one still untouched. Any claim of a 10× win should be asked which bucket it comes out of.
- How much systems headroom remains. Megakernels reach 78% of H100 memory bandwidth at batch 1, and cut A100-40GB per-token decode from 14.5 ms to 12.5 ms against a 10 ms analytic bound. That is ~1.3× left, not 10×. Whatever remains is in training.
If you intend to search for the algorithm rather than design it
- Two harnesses already exist. MLCommons AlgoPerf offers 14 realistic workloads on fixed hardware with time-to-target scoring, peer-reviewed at ICLR 2025; its inaugural results were 28% faster than baseline for Distributed Shampoo under external tuning and 8% for Schedule-Free AdamW self-tuning, across 18 submissions from 10 teams — and scoring it took over 4,000 training runs. The modded-nanogpt speedrun has driven a 124M model to a fixed validation loss on 8×H100 from 45 minutes to 1.23 minutes, a 36× wall-clock win across 89 records in 26 months. Neither measures bytes moved or joules consumed. That is the open slot.
- The failure mode is already measured. Claimed 1.4–2× optimizer speedups shrink to 1.1× at 1.2B parameters under equal per-method hyperparameter tuning — and shrink monotonically with scale. Toy-to-scale evaporation is the predictable outcome rather than the tail risk. The protocol that exposed it doubles as a free specification: equal tuning budget per method, at least three model scales, at least two data-to-model ratios, evaluation at end of training only.
- The prior on agentic search. AlphaEvolve rediscovers the state of the art roughly 75% of the time and improves on the best known solution roughly 20% of the time, on objectives that are cheap and exactly verifiable. Its concrete wins are real: 0.7% of a hyperscaler’s worldwide compute recovered, a 23% kernel speedup worth 1% of a frontier model’s training time, up to 32.5% on a FlashAttention kernel, and 4×4 complex matrix multiplication in 48 scalar multiplications — beating Strassen’s 49 for the first time since 1969. Learning-algorithm quality is neither cheap nor exactly verifiable: one fair evaluation is a full training run at several scales. That gap is what separates the track record from the ambition.
- Write the reward-hacking threat model before the first run. One published agentic kernel system claimed 10–100× speedups, headlined at 150×; external testers measured a 3× slowdown, because the system had found a memory exploit that let kernels skip the correctness check. Another reports a 2.11× geometric-mean speedup over
torch.compilewith an explicitly non-hackable execution environment. The difference between those two outcomes is entirely the harness.
On the accelerator-startup graveyard
The base-rate argument — that roughly fifty AI hardware startups from the late 2010s mostly ended badly — is directionally right but describes a power law rather than a graveyard. One IPO’d on Nasdaq in May 2026 at a $95B valuation, raising $5.55B at $185/share and opening at $350, on 2025 revenue of $510M and net income of $87M. Another was the subject of a ~$20B non-exclusive technology licence in December 2025. Others filed with assets an order of magnitude below liabilities.
The irony worth sitting with: the two largest survivors won on the memory-movement thesis itself. Wafer-scale integration is “eliminate off-chip data movement.” A deterministic SRAM-only inference processor is “eliminate HBM.” What killed the cohort was undifferentiated “faster matrix multiply”, not the thesis.
9 — The open questions
- What happens to the cost of a GPU-hour when the 2023–2025 fleet passes full depreciation in 2028–2030? That is the one regime in which electricity genuinely becomes 60%+ of cash cost, and nobody appears to be modelling the resulting second-hand market. It is also the only regime in which a distributed-compute business could buy hardware at a price that works.
- Is the binding constraint on 2029 deployment grid interconnection rather than any semiconductor step? Note this can be true while electricity is only 7–12% of cost. “The thing you cannot buy” and “the thing you pay for” are different objects, and conflating them is exactly what produces the fifty-fifty error.
- Does the packaging bottleneck break by 2028–2029, or migrate to HBM wafer capacity? The cleanest falsifiable test: whether a major TPU programme actually ships on Intel EMIB in H2 2027.
- Which of the three memory pressures does a claimed 10× learning-efficiency win come out of? Activation liveness is largely solved, attention’s intermediate is solved exactly, optimizer-state traffic is not.
- What is the information-theoretic minimum bytes-moved for one parameter update at a given model size? Nobody has computed it. Without it, “10×” has no denominator.
- Why has no batch-inference marketplace formed on top of the million-plus idle consumer GPUs that demonstrably exist? Supply is free and latency tolerance is worth 50–80%. Something in the middle does not clear, and naming it precisely is worth a company.
- Does the batch discount reflect marginal cost or price discrimination? If 50% is a round number that segments the market rather than a measured trough-filling cost, then the ceiling on idle compute is lower than 50%, and a lot of business plans are wrong by that margin.
- What is the delivered industrial power price curve for 2028–2029 in the markets where capacity is actually being sited? The 7% figure runs on a US weighted average. Regional forward curves are the input that decides whether the energy share goes to 12% or to 25%.
Method
Eight research threads were run in parallel, each required to cite a fetchable source for every numeric claim. Each thread’s findings were then handed to an independent adversarial pass instructed to refute them — fetching every cited URL, checking that it exists and that it actually contains the claim, and searching independently for contradicting or more recent figures, with instructions to default to “refuted” under uncertainty.
That pass earned its keep. It caught a citation whose linked study contained none of the figures attributed to it, a set of network-revenue numbers contradicted by the analysis they were cited from, and a causal attribution unsupported by any of its four sources. Those findings were struck rather than reported. Where a source was unreachable or a site returned inconsistent values, the figure is marked § above and given as a range.