Idle capacity and the cost of a GPU-hour

A research note, August 2026 (ISO week 35) · eight parallel research threads, each put through an adversarial verification pass against its own cited sources
energy is 7%, not 50%batch-1 is 0.2% of peakthe idle-GPU market clears at 4–20× off CoWoS is the bottleneck76 GW of curtailment headroom£180/vehicle-year for V2G

There is a claim that circulates freely in conversations about AI and energy: that roughly half the cost of running a GPU is the electricity, and half is the GPU. It is wrong by about five-fold, and the error propagates — because if you believe it, you conclude that cheap kilowatt-hours are the lever. They are not. This note works through where the money in a GPU-hour actually goes, and then through six adjacent questions that the correction reframes.

Every figure below was read at a primary or near-primary source. Figures marked § came back with a source that could not be re-fetched or that contradicted the finding, and are flagged in place rather than quietly dropped. Three fabricated or miscited numbers were caught by the verification pass and are not reported at all.

1 — The cost stack: electricity is 7% of a GPU-hour

Epoch AI’s itemised 1 GW AI datacentre model (May 2026; GB200 NVL72, 5-year IT life, 8.34 c/kWh, PUE 1.14, 71% utilisation) breaks $8.5B of annualised TCO into:

servers $5,021M (60%) · facility $1,387M (16%) · network $1,167M (14%) · energy $594M (7%) · tax, maintenance, labour, land, water ~$342M (4%)

Upfront capex: $37.9B — servers $21.2B, facility $11.4B, network $4.9B.

Two independent bottom-up models land in the same place. A four-year Meta H100 cluster splits 46% GPUs / 24% infrastructure / 20% colocation rent / 9% electricity / 1% other. A self-hosted 8×H100 server at $15.26/hr splits 35.9% depreciation / 20.7% capital cost / 16.9% colocation / 11.7% electricity — $0.15–0.31 per GPU-hour.

What would it take for energy to dominate?

Holding hardware cost constant, electricity reaches 50% of TCO at $1.11/kWh and 90% at roughly $10/kWh — 13× and ~120× the model’s baseline. Put differently: even if the accelerators and the networking were free, electricity would still be only ~26% of the cost, because the shell, substations, transformers and cooling plant amortise at $1,387M/yr against a $594M/yr electricity bill.

And the ratio is currently moving away from energy dominance, not toward it. HBM4 runs ~$560/stack against ~$370 for HBM3E, a 50% increase; UBS models 2027 HBM average selling prices up 79%; NVIDIA is planning a >15% server price rise for early 2027. Conventional DRAM contract prices moved +93–98% quarter-on-quarter in Q1 2026 and +58–63% in Q2.

The reframe that survives

The scarce good is not cheap kilowatt-hours. It is firm kilowatts at a specific interconnection point, on a schedule. The energy-infrastructure business hiding inside the “facility” line — substations, switchgear, transformers, turbines, chillers, battery storage — costs 2.3× more per year than the electricity flowing through it. That is why every serious distributed-compute pitch in 2026 sells speed to power rather than cheap power.

There are two regimes where the intuition about energy dominance is right, and both are worth naming precisely. First, cash opex rather than total cost of ownership: at 90% utilisation, energy is 59% of cash opex on Quincy, Washington hydro and 77% in PJM Dominion. Second, the post-depreciation regime — fully written-off silicon sitting in a sunk shell, which is what the 2023–2025 fleet becomes around 2028–2030. Nobody appears to be modelling that second-hand market, and it is the one place a distributed-compute business could acquire hardware at a price that works.

2 — Why a single user cannot make local inference cheap

The recurring counter-argument is that per-person token demand will grow two to three orders of magnitude, that price compression cannot keep pace, and that people will therefore run models locally because it is cheaper. The demand premise is well supported. The cost conclusion is not, and the margin is not close.

MeasurementNumberWhat it means
Roofline at batch 1~0.5 FLOPs/byte against an H100 ridge point of ~295 (989 TFLOPS / 3,350 GB/s)A single-stream user runs the chip at roughly 0.2% of peak FLOPs. Batch 240 — H100’s critical batch — reaches ~120 FLOPs/byte.
Same box, batching onlyNVIDIA DGX Spark, Nemotron Super 49B NVFP4: 5.79 tok/s single-stream → 695 tok/s at concurrency 256 (120×). gpt-oss-120b MXFP4: 33.5 → 863 tok/s (25.7×).The cleanest isolation of the aggregation effect, on hardware an individual could buy. One user leaves ~99% of their own machine idle.
Per acceleratorGB200 NVL72 serving DeepSeek R1 at 65 tok/s/user: 6,057 tok/s/chip at $0.085 per million tokens. H200 at the same interactivity: 1,035 tok/s/chip at $0.327/M.Against ~18 tok/s for the same 671B model on a Mac Studio M3 Ultra — roughly 336× per chip.
Utilisation penalty$0.21 to $15.25 per million output tokens on identical H100 hardware, driven only by offered request rate. A 72.6× spread.Underutilisation costs 2.5–24× at 1–10 requests/sec, up to 36.3× near idle.

All-in, at a realistic personal duty cycle: self-hosting costs $30–160 per million tokens against $0.09–0.28 hosted for the same open-weight model. Decomposed — a $5,499 Mac Studio amortised over three years at 18 tok/s is $3.23/M at a 100% duty cycle, $32.29 at 10%, and $161.46 at 2%. Two per cent is generous: a single human at batch 1 cannot physically consume more than about 1.55M tokens/day. Electricity alone is $1.03/M in California (3.09 kWh per million tokens at 33.25 c/kWh) and $0.55/M at the US average.

Price compression has already done three orders of magnitude once

a16z’s LLMflation series tracked GPT-3-quality inference from $60 per million tokens in November 2021 to $0.06 by late 2024 — a factor of 1,000 in three years, pinned at fixed capability (MMLU 42 and MMLU 83 thresholds). Epoch AI’s independent series gives a range of 9× to 900× per year depending on the benchmark; GPT-4-level GPQA fell 40×/year.

The stricter measurement is more interesting. MIT FutureTech (Gundlach, Lynch, Mertens & Thompson, March 2026) measures at fixed benchmark performance rather than fixed token price, on GPQA-Diamond, AIME and SWE-bench Verified across Pareto-frontier models: prices fall 5–10× per year at a given capability, algorithmic progress contributes ~3×/year, and the cost of running the actual frontier is rising 3–18× per year, because reasoning models burn more tokens per completed task.

So the spread between frontier and commodity widens by something like 50–100× a year. That is the real bifurcation — frontier versus commodity, not local versus cloud. And note the conflation that does most of the damage in these arguments: “open weights win” and “local wins” are different claims, and only the first has data behind it. OpenRouter’s 100-trillion-token study puts open-weight models at ~30% of tokens by late 2025, with Chinese-origin models above 45% of traffic by mid-2026 — consumed almost entirely as hosted API.

One refinement worth carrying: on high-end GPUs, batch-1 decode is kernel-launch-bound before it is memory-bandwidth-bound. Measured batch-1 decode reaches only ~27% of the memory-bandwidth floor on an H100, against ~81% on an L4 — which is why CUDA Graphs deliver a 1.259× speedup on H100 and only 1.028× on an L4. The bandwidth framing is right for Apple silicon and consumer cards and wrong in the details for datacentre parts.

3 — The market for idle and heterogeneous GPUs

A market for unreliable, heterogeneous capacity does exist. It is roughly three orders of magnitude too small to matter, and it clears at a 4–20× discount that tells you exactly why.

TierPriceNote
AWS p6-b200.48xlarge$113.93/hr = $14.24/GPU-hrThe homogeneous 8-GPU NVLink island. Capacity Blocks $12.355 per accelerator-hour. You cannot buy a fraction of it, or a variant.
H100 on-demand, market median$3.39, range $0.67–$14.19 across 38 providersA 21× spread. Vast.ai spot $0.67, Lambda $3.29, AWS $6.88, Azure $6.98, GCP $14.19.
SF Compute, short-term~$0.80–1.00/hrA genuine order book with bids and asks. Utilisation sits near 100% because the price falls until it does.
Salad, consumer GPUsRTX 4090 $0.160/hr; 3090 $0.090; 3080 $0.060; 3070 $0.040; batch tier from $0.021M+ underutilised consumer GPUs, 450,000+ providers, 180+ countries. The supply side is emphatically solved.

The single most instructive datapoint: Prime Intellect — the flagship decentralised-training company, ~$100M ARR at a $1B valuation, the organisation with the strongest possible incentive to demonstrate otherwise — trained its best model, INTELLECT-3 (106B MoE, 12B active), on a centralised 512×H200 cluster across 64 nodes with 400 Gbps NDR InfiniBand. Their own survey of the field concedes that no one has successfully scaled this research to train state-of-the-art models. The marketplace business works; the decentralised-training thesis has not yet.

The ceiling, expressed as a price

OpenAI and Anthropic both discount batch inference by exactly 50% against synchronous, for a 24-hour completion window. Sail Research’s “flex” window reaches 80% off, the highest observed market price of latency tolerance. That 50–80% band is what latency tolerance is worth — and note what it buys the provider: trough-filling and larger, more arithmetic-intensity-efficient batches on the same datacentre GPUs. It is not payment for running on worse hardware. Any idle-capacity business has to fit underneath that band and additionally absorb checkpointing, verification, straggler and egress costs.

The heterogeneity results, scored honestly

DiLoCo’s headline “500× less communication” was 8 workers, tested only up to 400M parameters. The largest genuinely decentralised runs are INTELLECT-1 (10B parameters, 6T tokens, 98% compute utilisation, 400× communication reduction via int8 in production and up to 2000× experimentally) and INTELLECT-2 (32B, decentralised RL). Both sit one to two orders of magnitude below the frontier. Petals, the volunteer peer-to-peer inference project, delivers ~6 tok/s on Llama-2-70B and ~4 tok/s on Falcon-180B in single-batch inference; its public network-status widget no longer loads. Exo Labs has retreated to LAN-scoped clusters, which is the honest engineering answer to residential bandwidth — and also a concession on the market question.

The assumption most likely to sink an idle-capacity business, and it is current

The predicted 2026 compute glut did not arrive. H100 hourly rates fell from ~$8 in early 2024 to $1.96 in late 2025, then climbed back to $2.64–2.74 by mid-2026 and are consolidating in a $2.60–2.80 corridor. One-year contract pricing rose ~40% from a $1.70 low in October 2025 to $2.35 by March 2026; Blackwell contracts rose 48% over the same crunch. A tightening market means operators have less idle capacity to dump — the supply side of the idle-compute business is evaporating at precisely the moment the demand side, in the form of long-horizon asynchronous agents, finally appears.

§ Quarterly revenue and GPU-count figures for several decentralised networks did not survive verification — one primary report was unreachable and one cited analysis actively contradicted the numbers it was cited for. Directionally these networks are very small; the specific figures are omitted rather than repeated.

4 — Does a GPU depreciate?

The popular framing is that GPUs depreciate quickly for training, because you want the latest silicon, and slowly for inference, because you do not. The economics behind that intuition are sound; the framing is a category error worth fixing.

An asset does not have workload-specific depreciation rates. It has one cash-flow decay path, set by the best-paying workload the card can still win. What is true is that inference provides a floor under that path, because inference is far less sensitive to frontier FLOPs per dollar and far more sensitive to memory bandwidth per dollar and to already-powered rack space.

And the variable that decides whether hyperscaler earnings are overstated is not the chip’s physical life at all — it is whether its revenue declines faster than straight-line. Straight-line over six years understates depreciation if the cash-flow profile is front-loaded, even if the chip physically runs for nine years. That is why this argument never resolves: the accounting bear case and the hardware-longevity bull case are not actually in contradiction.

The empirical record

SeriesReading
A100 rental$1.61 on the Silicon Data neocloud index; $1.80 median across 34 providers, explicitly flat over 90 days. § The same site returned $1.29 on its dedicated A100 page minutes apart; treat spot as a $1.29–1.86 band.
H100 rental$2.67 on the index, $3.39 as a market median — up 13% over 90 days. The real A100/H100 ratio is ~1.7×.
V100 rental (2017 part)$0.97 median, $0.11 floor — 6–15% of an A100’s rate nine years post-launch. Azure retired its V100-backed NCv3 VMs in September 2025, ~7.5 years after launch. That is the actual observed ceiling on hyperscaler GPU life, and it is a retirement.
Used A100 80GBRoughly 60–80% of value retained, against a launch-era list of ~$15–20k for the 80GB PCIe part. § The quoted resale band is wide and comes from vendors with an interest in the number.

Why an obsolete card keeps clearing at yesterday’s price. NVIDIA issued an A100 end-of-life notice in January 2024, which is part of it. But the larger cause is that powered floor space, not silicon, is the binding constraint: CoreWeave at Q2 2026 held 1.5 GW active against 3.7 GW contracted — a 2.2 GW gap between what it has sold and what it can energise. An installed six-year-old card already has a rack, a power hookup and cooling, and those are the scarce goods.

The accounting moves, in one table

CompanyChange to server useful lifeP&L effect
Microsoft4 → 6 years (announced on the Q4 FY2022 call)~+$3.7B FY23
AlphabetServers 4 → 6, some network equipment 5 → 6, effective Jan 2023$3.9B FY23
MetaMost servers to 5.5 years, effective Jan 2025. Non-AI servers 6 → 7 on 29 April 2026 because of the memory shortage — 8 years was declined as too risky; failure rate rises 4.8% → 7.4%~$2.9B in 2025
Amazon5 → 6 years effective Jan 2024, then reversed to 5 effective 1 Jan 2025, citing artificial intelligence by nameReversal: +$1.4B depreciation, −$1.0B net income, −$0.10/share
CoreWeave6 years, unchanged since 2023; average contract term ~5 years—

Amazon is the datum that cuts hardest against the longevity case: the largest cloud operator on earth examined its own fleet and shortened the life. Meanwhile Meta’s servers-and-network depreciation went $7.32B → $11.34B → $13.36B across 2023–2025 — up 83% in two years despite the life extension, because capex growth swamps the accounting lever.

Two corrections to figures that circulate: NVIDIA’s gross margin is 75.0% GAAP and non-GAAP for Q2 FY2027 (quarter ended 26 July 2026, $96.2B revenue, $89.0B data centre), guided down to 74.0% and bottoming at 71–72% in Q4 FY27 on memory costs. The ~82–88% number that gets quoted is a chip-level estimate for a single accelerator SKU — roughly $3,320 of bill-of-materials against a $25–40k street price — not a company figure. And the bear case on hyperscaler depreciation, ~$176B of understatement across the five largest operators for 2026–2028, targets the hyperscalers and neoclouds, not the chip vendor.

5 — Packaging, not wafers, is the bottleneck

On the 16 July 2026 earnings call, TSMC’s CEO put it on the record: packaging capacity is so tight that it is limiting customers’ growth, with the supply–demand gap closing “probably 2029, 2030.”

QuestionWhere it lands
Is CoWoS TSMC-only?~95% of leading-edge 2.5D capacity, so functionally a monopoly today — but not literally sole-source, and the moat is cracking. ASE/SPIL and Amkor run licensed chip-on-wafer and wafer-on-substrate steps and are expected to add 50,000–60,000 wafers/month by late 2026. Intel’s EMIB is the real alternative: already shipping AWS Trainium, and has won Google’s TPU v8e for H2 2027 plus a >3M-unit TPU packaging booking for 2028.
Capacity trajectoryTSMC CoWoS: 13k wafers/month at end-2023 → ~70–80k in 2025 → 120–140k end-2026 → 190–200k in 2027. But demand roughly doubles alongside it — 1.3–1.4M wafers in 2026 to 2.5–2.7M in 2027 — so the ~20% gap only halves to ~10%.
Who holds the allocationOut of ~1.0M wafers of 2026 demand: NVIDIA 595,000 (60%); Broadcom 150,000 (15%), including Google TPU 90,000, Meta 50,000, OpenAI 10,000; AMD 105,000; Marvell 55,000; Amazon 50,000; MediaTek 20,000. The top customers lock >85%, leaving under 15% for everyone else. This is why a TSMC commitment is the single most valuable asset an accelerator startup can hold.
Is lithography the constraint?No. ASML is adding roughly 30% EUV capacity for 2027 with another 30% under study for 2028, against ~65 low-NA EUV units in 2026.
Wafers or packaging?Packaging, in 2026. Leading-edge wafer capacity is tight but expanding faster in percentage terms — N3 heading to ~180k wafers/month (+40% YoY), N2 to ~90–100k by end-2026 and sold out for the year. Whether that flips by 2028–2029 is genuinely open.
MemoryAll three HBM suppliers have reportedly sold out 2027 capacity, not just 2026. HBM cost content per accelerator: ~$1,350 for an H100, ~$3,250 for a B200, ~$4,350 for an MI325X.

§ A widely-repeated claim that a major accelerator programme cut its 2026 production target 25% “because NVIDIA took all the CoWoS capacity” does not survive checking in that form. What is sourceable is an analyst revision: Fubon Research, relayed via Jefferies in December 2025, judged that CoWoS could not support the market’s speculated 4M TPUs for 2026 and modelled 3.1–3.2M instead — a ~20–23% haircut to an expectation, not to a published target. The sourced bottleneck story concerns a different customer’s reservation, not NVIDIA crowding anyone out. The causal “because” is unsupported.

6 — Compute in homes: the graveyard, and why 2026 is different

Two live programmes put general-purpose compute inside residential buildings, and they are an order of magnitude apart in a way that decides whether either is a business.

Sunrun AI Compute NodeSPAN XFRA
Announced8 July 2026 (week 28)13 April 2026 (week 16)
The box“About the size of a small desktop computer”; power draw undisclosed16× NVIDIA RTX PRO 6000 Blackwell Server Edition, 4× AMD EPYC, 3 TB RAM, direct liquid-cooled, exterior-wall mounted. 12.5 kW continuous. Reported hardware value >$200k
Money flowSunrun owns and maintains the node, sells the inference capacity to enterprise buyers, and compensates the homeowner for hostingSPAN owns the hardware; the homeowner pays SPAN ~$150/month flat, covering the home’s entire electricity and internet, plus usage-based compensation
GateAn existing Sunrun solar + battery installation — a base of 1.1M+ homesNew construction, through a national homebuilder
ScalePilot. Management guided commercialisation to “late 2027 or 2028” on the Q2 2026 call (5 August 2026)100 homes / ~1.25 MW / 1,600 GPUs in the US Southwest this autumn; 80,000 nodes and >1 GW targeted from 2027. The stated pitch: ~$3M per MW in six months

Why every previous attempt died, and it is a single diagnosis

Nerdalize, Qarnot, Heata, Cloud&Heat and Project Exergy all monetised waste heat. Heata’s entire consumer-side proposition is up to 4 kWh/day of domestic hot water — worth £120–340 per home per year. That is the whole budget from which hardware, installation, servicing and churn must be funded, and it cannot work at any scale. Nerdalize raised ~€1.6M lifetime and filed for bankruptcy in January 2019 with a €882k loss in 2017 alone.

The 2026 programmes monetise spare amps and a place in the interconnection queue instead of spare warmth — a value pool 50–500× larger. Which is exactly why the larger of the two is liquid-cooled and rejects its heat outdoors by design. Neither company claims a waste-heat credit.

Solar is not the driver, despite the framing

Residential solar delivers a ~25% capacity factor; running a GPU node only when the sun shines destroys the hardware economics. The actual drivers are speed to power — an electrician’s day of work against a three-to-five-year utility interconnection queue — and export compensation collapse. Under California’s post-April-2023 net billing tariff, solar customers receive $0.05–0.08/kWh for exported energy against $0.30–0.45 retail, so the opportunity cost of a midday solar kilowatt-hour is five to eight cents. That makes burning it locally nearly free — but it is a tariff artefact, not a physical one, and it evaporates if the tariff changes.

On the thermodynamics that always comes up: a 300 W computer is a 300 W resistive heater, to within a rounding error — photons through the window, acoustics and signal down the network link account for order 0.01–0.1%. The caveat that matters is the comparison class: against a heat pump at a coefficient of performance of 3.5, that heat is worth the electricity price divided by 3.5, so roughly $0.10–0.13/kWh in California rather than $0.35–0.45.

A false claim to be aware of. A widely-circulated story since mid-2026 holds that a chip vendor will pay homeowners $22,000 — or $220,000 — a year to host a miniature datacentre. It is not true. It originated in social posts and was laundered through aggregator sites. The ~$200,000 figure is the value of the hardware inside the node, which the homeowner does not own. In the actual programme the homeowner pays a flat monthly fee that covers their utilities.

7 — Grid headroom, flexible load, and what V2G actually pays

US power systems run at a 53% average load factor across 22 balancing authorities, 2016–2024 (range 43–61%). The idle capacity is real. The question is what unit it is denominated in.

Curtailment-enabled headroom — the ladder, exactly

76 GW at 0.25% average annual curtailment · 98 GW at 0.5% · 126 GW at 1.0% · 215 GW at 5.0%, across the 22 largest US balancing authorities covering ~95% of load. 76 GW is 10% of current US aggregate peak demand. Average curtailment event duration: 1.7 / 2.1 / 2.5 hours. Hours per year in which any curtailment is called: 85 / 177 / 366. By balancing authority at 0.5%: PJM 18, MISO 15, ERCOT 10, SPP 10, Southern 8 GW.

Read the Limitations section before quoting the headline. The authors state plainly that transmission capacity, ramping capability and ramp-feasible reserves are beyond the study’s scope, and that the results therefore cannot be taken as an accurate estimate of the load that can be added to the system. Those two pages separate someone who read the study from someone who read a headline about it.

What has actually been demonstrated

The flagship flexible-datacentre demonstration ran on a 256-GPU A100 cluster inside a cloud region in Phoenix, cutting power 25% below average base load for three hours on two utility system peaks (1 and 3 May 2025), with a 15-minute graceful ramp down and back up. At a 400 W cap that is roughly 100 kW of GPU nameplate — beautifully executed, and three orders of magnitude below the scale its press coverage implies. Separately, demand-response agreements signed in August 2025 between a hyperscaler and two utilities are real but carry no disclosed MW, MWh or hours. In ERCOT, registered large controllable load stands at ~5,302 MW, about 2% of total load. The company commercialising this raised $150M at a $1.05B valuation on 25 August 2026.

The objection that kills the naive residential version

The binding constraint is the service transformer and the feeder, not the panel. A study of 1,500 real Baltimore Gas & Electric feeders — k-means clustering to seven representative feeders, OpenDSS time-series load flow, AMI-derived base load — projects over 35% of distribution transformers overloaded by 2035 from EV charging alone. Annual transformer loss-of-life at 100% EV penetration runs 8.97% under demand charging against 4.30% off-peak. A constant 1–5 kW compute load is a worse profile than EV charging, because it never goes away.

On the panel itself: analysis of over 100,000 single-family homes found that more than 80% never draw more than 40 amps at any instant in a year, and 99% never exceed 100 amps. Forty amps at 240 V is 9.6 kW — 20% of a 200 A service. Average residential draw is on the order of 1–2% of a 200 A rating. So the headroom is larger than the usual “homes use about 40% of capacity” claim suggests, and that claim is anyway a statement about annual peak against a 100 A service, not an average against anything.

Vehicle-to-grid, with the number to hold it to

The best-measured trial in the world — 135 households in Great Britain, combined capacity under 1 MW, run with the national control centre — put the incremental value of bidirectional charging over merely smart charging on a time-of-use tariff at £180 per vehicle-year. Against bidirectional-charger rebates of up to $4,500, the payback arithmetic writes itself. The best-executed deployment anywhere runs 150 bidirectional vehicles on 7 kW chargers for a total fleet peak discharge of 300 kW — 2 kW per car realised; 50 of those cars discharged 65,000 kWh in five months. One major manufacturer has 250,000 bidirectional-capable vehicles on US roads after a mid-2026 over-the-air update and projects 52,000 enrolled in a California utility programme by 2030.

The energy-versus-power confusion is what makes the EV-battery argument sound stronger than it is. The US light-duty EV fleet holds roughly 300 GWh (~4M vehicles × ~75 kWh) against 144 GWh of cumulative US grid-connected storage installed since 2019 — so about 2.1–2.9×, not 3×, and comparing kWh to kWh when what a datacentre buys is firm kW at an interconnection point.

Two things any version of this thesis will be hit with

A large fraction of the interconnection queue is phantom. AEP Ohio’s total volume of large-load interconnection requests fell from 30 GW to 13 GW; Georgia Power’s expected large-load additions dropped by 6 GW, with Winter 2028 and Winter 2029 alone falling 1.4 GW. The same project shops itself to multiple utilities.

Datacentres are currently raising other ratepayers’ bills. PJM’s Independent Market Monitor attributes 63% of the 2025/26 capacity price increase directly to datacentre demand — $9.3 billion in additional customer costs. The flexibility thesis is in fact the strongest available answer to that, since a curtailable load does not drive capacity procurement, and it is worth leading with rather than defending.

8 — Beyond backprop: getting the energy numbers right first

The argument for redesigning learning algorithms around the memory hierarchy rests on a set of numbers that are widely quoted and frequently garbled. Here they are at source.

Arithmetic and memory energy

Horowitz’s ISSCC 2014 table, 45nm at 0.9 V: integer add 0.03 pJ (8-bit) / 0.1 pJ (32-bit); floating-point add 0.4 / 0.9 pJ (16/32-bit); integer multiply 0.2 / 3.1 pJ; floating-point multiply 1.1 / 3.7 pJ. A 64-bit cache read costs 10 pJ from 8 KB, 20 pJ from 32 KB, 100 pJ from 1 MB. DRAM: 1.3–2.6 nJ per 64-bit access. Register-file access is 6 pJ; total instruction energy is ~70 pJ, of which the arithmetic is a rounding error against fetch and clocking.

Where HBM energy actually goes — and it is not where the intuition says

From a physical floorplan model at real application toggle rates: HBM2 costs 3.92 pJ/bit in total, decomposed as 2.24 pJ/bit moving inside the DRAM die (~9.9 mm of wire), 1.21 pJ/bit row activation, and 0.30 pJ/bit across the interposer. GDDR5 was 14.0 pJ/bit. HBM3E is 3.44–4.05 pJ/bit — essentially unchanged across three generations despite the bandwidth scaling.

The package hop is about 8% of the access energy. So “the cost is the distance to the edge of the chip” is the wrong mental model. The design consequence is concrete: reducing the number of accesses helps proportionally, but the 1.21 pJ/bit activation floor only falls if locality improves within a DRAM row. An algorithm that halves bytes moved while randomising the address stream can be a net energy loss. Use pJ/bit rather than pJ/access — access granularity does a lot of hidden work in the usual comparison.

The latency ladder, measured rather than remembered

Pointer-chase measurements on H800 (Hopper): L1 32.0, shared 29.0, L2 264.5–502, global 656 cycles. A100: 33.0 / 29.0 / 202.8–408 / 566. RTX 4090: 32.0 / 30.1 / 273 / 571. The commonly-quoted “L2 is 200, HBM is 500” are the A100 best cases. And the register figure should be 4 cycles of dependent-instruction latency on both GH100 and GB203, not 1 — which makes the register-to-L1 gap 8×, not 30×.

One measurement worth having in isolation: on H800, an L2 near hit costs 258.0 cycles and a far hit 414.1. 156 cycles is the directly measured cost of crossing to the far L2 partition — on-die distance, priced.

A methodological caution that cuts against designing to these numbers: on a GPU these are latencies, which warp-level parallelism exists to hide. The binding constraints are bandwidth and picojoules per bit. An objective written in cycles optimises the wrong quantity; write it in bytes and joules.

What the alternatives have actually achieved

If you intend to search for the algorithm rather than design it

On the accelerator-startup graveyard

The base-rate argument — that roughly fifty AI hardware startups from the late 2010s mostly ended badly — is directionally right but describes a power law rather than a graveyard. One IPO’d on Nasdaq in May 2026 at a $95B valuation, raising $5.55B at $185/share and opening at $350, on 2025 revenue of $510M and net income of $87M. Another was the subject of a ~$20B non-exclusive technology licence in December 2025. Others filed with assets an order of magnitude below liabilities.

The irony worth sitting with: the two largest survivors won on the memory-movement thesis itself. Wafer-scale integration is “eliminate off-chip data movement.” A deterministic SRAM-only inference processor is “eliminate HBM.” What killed the cohort was undifferentiated “faster matrix multiply”, not the thesis.

For scale: the Landauer limit is kT·ln2 = 2.87×10−21 J at 300 K. A 3.7 pJ 32-bit floating-point multiply at 45nm is ~106× that; a 4 pJ/bit HBM access ~109×. Thermodynamics is nowhere near binding. Wires are.

9 — The open questions

  1. What happens to the cost of a GPU-hour when the 2023–2025 fleet passes full depreciation in 2028–2030? That is the one regime in which electricity genuinely becomes 60%+ of cash cost, and nobody appears to be modelling the resulting second-hand market. It is also the only regime in which a distributed-compute business could buy hardware at a price that works.
  2. Is the binding constraint on 2029 deployment grid interconnection rather than any semiconductor step? Note this can be true while electricity is only 7–12% of cost. “The thing you cannot buy” and “the thing you pay for” are different objects, and conflating them is exactly what produces the fifty-fifty error.
  3. Does the packaging bottleneck break by 2028–2029, or migrate to HBM wafer capacity? The cleanest falsifiable test: whether a major TPU programme actually ships on Intel EMIB in H2 2027.
  4. Which of the three memory pressures does a claimed 10× learning-efficiency win come out of? Activation liveness is largely solved, attention’s intermediate is solved exactly, optimizer-state traffic is not.
  5. What is the information-theoretic minimum bytes-moved for one parameter update at a given model size? Nobody has computed it. Without it, “10×” has no denominator.
  6. Why has no batch-inference marketplace formed on top of the million-plus idle consumer GPUs that demonstrably exist? Supply is free and latency tolerance is worth 50–80%. Something in the middle does not clear, and naming it precisely is worth a company.
  7. Does the batch discount reflect marginal cost or price discrimination? If 50% is a round number that segments the market rather than a measured trough-filling cost, then the ceiling on idle compute is lower than 50%, and a lot of business plans are wrong by that margin.
  8. What is the delivered industrial power price curve for 2028–2029 in the markets where capacity is actually being sited? The 7% figure runs on a US weighted average. Regional forward curves are the input that decides whether the energy share goes to 12% or to 25%.

Method

Eight research threads were run in parallel, each required to cite a fetchable source for every numeric claim. Each thread’s findings were then handed to an independent adversarial pass instructed to refute them — fetching every cited URL, checking that it exists and that it actually contains the claim, and searching independently for contradicting or more recent figures, with instructions to default to “refuted” under uncertainty.

That pass earned its keep. It caught a citation whose linked study contained none of the figures attributed to it, a set of network-revenue numbers contradicted by the analysis they were cited from, and a causal attribution unsupported by any of its four sources. Those findings were struck rather than reported. Where a source was unreachable or a site returned inconsistent values, the figure is marked § above and given as a range.

Figures were current as of late August 2026 (ISO week 35). Rental prices, allocation shares and capacity forecasts in this area move monthly; treat anything here as a reading rather than a constant, and re-check before acting on it.