Fact-check and research dossier: Yaroslav Bulatov on The Interlude Show

Every checkable claim from the 8 September 2026 recording, tested against sources as of the same day, with the strongest case for his thesis, the strongest case against it, who else is working the same problem, and what to ask next time.

Guest
Yaroslav Bulatov (Speaker 1) — Google Brain, early OpenAI, Meta, Together AI; runs the Sutro Group at South Park Commons
Hosts
Vatsal (Speaker 2) and Jake (Speaker 3)
Source
Transcript file 08sep28_-_vatsal_bajaj__spc__the_interlude_show_.txt, 09:00–10:14
Checked
Tuesday, 8 September 2026, via live web search; 57 claims scored

The central thesis in one picture

Nearly everything Bulatov argued — that backprop is energy-inefficient, that theoretical matmul results miss the point, that Dally's 2D model is the right one — rests on one physical fact: on a modern chip, arithmetic is nearly free and moving bytes is not. Bill Dally's 2022 numbers (from his Communications of the ACM "Point" essay) put it on a log scale:

Energy ladder: a 32-bit add versus data movement Log-scale axis from 10 femtojoules to 10 nanojoules. A 32-bit add costs 20 fJ. Moving its two operands 10 microns costs about 19 fJ, 1 millimetre costs 1.9 pJ, corner to corner on a 400 square millimetre chip costs 77 pJ, going off chip costs 320 pJ, and fetching the two words from DRAM costs 1.3 nJ, 64,000 times the add. 10 fJ 100 fJ 1 pJ 10 pJ 100 pJ 1 nJ 10 nJ energy per operation, log scale (Dally, CACM Sept 2022; 32-bit operands) 32-bit add: 20 fJ = moving its operands ~10 µm move 1 mm on-chip 1.9 pJ · ~95× the add corner to corner (40 mm) 77 pJ · ~3,900× leave the chip 320 pJ · ~16,000× fetch from DRAM: 1.3 nJ 64,000× the add

Bulatov's rule of thumb on air — that an add costs about the same as moving its two inputs ten microns — matches Dally's figures almost exactly (1.9 pJ per millimetre is 19 fJ per 10 µm, against a 20 fJ add). His term for the ratio you want to maximise is "compute to commute"; the literature calls the same idea arithmetic intensity, I/O complexity, or communication-avoiding algorithms.

Scorecard and how to read it

23Verified
20Mostly right — small detail off
3Needs nuance
3Inaccurate
4Couldn't verify
4Opinion or forecast
VerifiedMostly rightNeeds nuanceInaccurateCouldn't verifyOpinion / forecast

The pattern: Bulatov is reliable on the physics and the systems arguments — the memory-wall numbers, the ASML and CoWoS facts, the AlphaEvolve result, the Groq deal — and loose on dates and specs remembered from a distance (the Horowitz keynote year, the H800's cut spec, what it takes to reach 1% on MNIST). The three clear errors don't damage the core argument; the "needs nuance" items are where a sceptical listener would push back, and they're listed with the counter-evidence below.

Method. Each claim is quoted in paraphrase with its transcript timestamp, then judged against sources I fetched today (linked inline and listed at the end). Where I lean on standard literature I didn't re-fetch (e.g., the FlashAttention paper, LeCun's 2016 cake slide), I say so in the text. Transcription errors have been corrected in the paraphrases: "Amnesty" is MNIST, "Coulomb's conjecture" is Collatz, "Keith Metz" is Cade Metz, "Mitchell Nymas" is Mitchell Nahmias, "Southwark Commons" is South Park Commons, "Spark 7" is a SPARC workstation, "denying scaling" is Dennard scaling.

Claims: hardware and the memory wall

09:11Mark Horowitz gave a talk in 2012 called "The Computing Energy Problem", now 14 years old, predicting the death of Dennard scaling and that CMOS would stay.

Mostly right

The talk exists but is the ISSCC 2014 plenary, "Computing's Energy Problem (and what we can do about it)" — February 2014, so twelve years old, not fourteen. The content is as described: power became the binding constraint before Dennard scaling broke down, and no new device technology will rescue us; EETimes' coverage summarises his call for specialised silicon plus better algorithms. His prediction that CMOS persists has held.

09:12ASML's statistic: 95% of the machines it built in the last 30 years are still in use.

Verified

ASML's own 2023 story states that roughly 95% of lithography systems sold in the past 30 years remain active, with more than 5,000 machines in fabs as of end-2022, sustained by a refurbishment programme. Its product pages repeat that almost every system ever shipped is still in a customer fab.

09:19Transistors shrink in two dimensions, wires in one; so arithmetic got cheap much faster than data transfer. That's the "memory wall".

Verified

The conclusion is textbook and quantified in Dally's CACM Point essay (Sept 2022): arithmetic and logic costs fell by orders of magnitude while communication costs fell far more slowly, so a DRAM fetch of two 32-bit words now costs 64,000× a 32-bit add. Horowitz's 2014 table makes the same point at 45 nm. The "2D vs 1D shrink" geometry is a simplification — the deeper reason is that wire RC delay and energy per unit length don't improve with scaling — but it's a fair intuition pump.

09:20Heuristic: adding two numbers costs the same as moving them 10 microns across the chip.

Verified

Dally's numbers: a 32-bit add is 20 fJ; moving its two operands 1 mm costs 1.9 pJ, i.e. about 19 fJ per 10 µm. The heuristic is accurate to within rounding. Bulatov has posted the same arithmetic on X (June 2026 thread), citing the Parallel Explicit Communication Model.

09:20A chip is about 20 × 20 mm; a micron is a millionth of a metre, so ~20,000 microns across.

Mostly right

Ballpark. Dally's worked example uses a 400 mm² chip (20 × 20 mm); flagship AI dies are larger (an H100 is ~814 mm², close to the ~26 × 33 mm reticle limit). The arithmetic on microns is right.

09:20Movement is so dominant that arithmetic is effectively free — about 1% of energy — yet fast-matmul researchers count arithmetic.

Needs nuance

True whenever data comes from far away (see the ladder above) and true of the theory literature's cost model. But it is not true of a dense GEMM running on tensor cores with good tiling: at high utilisation the arithmetic units are a large share of chip power, which is exactly why kernels are tuned for arithmetic intensity (the roofline model). "1%" describes a memory-bound kernel or a global-memory access, not a well-blocked matmul. The directional claim — that asymptotic multiplication counts are the wrong objective — is widely shared.

09:21GPU libraries don't use Strassen; they use tiling. Discovered fast algorithms aren't used in practice.

Verified

cuBLAS/CUTLASS-style GEMMs are tiled classical multiplication. Strassen-family algorithms exist in research — D'Alberto (2023) argues Strassen still beats the AlphaTensor variants on real hardware — but they're not in mainstream GPU libraries. Notably, AlphaEvolve's only production-relevant matmul win was a 23% kernel speedup from smarter tiling, which supports Bulatov's point about what matters in practice.

09:22Six years ago Bill Dally proposed a computation model where data is laid out on a 2D grid and you pay for distance.

Mostly right

The essay is "We Must Extend Our Model of Computation to Account for Cost and Location", CACM September 2022 — four years ago. Dally proposes extending PRAM into a Parallel Explicit Communication Model with Manhattan-distance costs. A parallel academic formalisation appeared the same year: ETH Zurich's "The spatial computer" (Gianinazzi, Ben-Nun, Hoefler et al., May 2022) puts processors on a 2D grid with energy equal to distance travelled and proves matching upper and lower energy bounds for matrix multiplication and sorting. The older roots are Thompson's VLSI complexity (1979) and Hong & Kung's red-blue pebble game (1981), which Bulatov didn't mention.

09:25Stacking 16 layers at ~100 W/cm² each would need 1.6 kW/cm² of heat removal — so 3D is thermally limited.

Mostly right

Sound back-of-envelope: an H100 runs roughly 700 W over ~8 cm² (~85 W/cm²). Logic-on-logic stacking is thermally limited; the 12–16-high stacks that ship today are HBM (memory), which dissipates far less per layer. His conclusion — that even in 3D you exhaust the "nearby" volume quickly — is the standard argument.

09:26Low peak memory is sufficient but not necessary for low energy; you need a finer-grained measure than peak memory.

Verified

This is a direct consequence of distance-based cost models: energy tracks how far accesses travel, not the footprint. The Sutro repos operationalise it as a "data movement distance" (DMD) metric (ByteDMD, sutro-problems), which is the right follow-up to ask about.

10:02A signal can't cross a chip in one clock cycle, which limits how much knowledge one chip can hold.

Verified

Standard since the early 2000s (Agarwal et al., ISCA 2000, showed only a shrinking fraction of a die is reachable per cycle). Dally's figures: corner to corner on a 40 mm chip takes 16 ns versus 150 ps for an add — roughly 100 cycles at 6 GHz-class clocks.

Claims: backprop, learning, and toy problems

09:28Backprop was invented in the '80s (when it became useful); a Finnish dissertation in the '60s; you can go back to Newton's chain rule.

Mostly right

Seppo Linnainmaa's master's thesis is from 1970 (published 1976), not the 1960s; control-theory precursors (Kelley 1960, Bryson 1961, Dreyfus 1962) are the '60s work. Werbos (1974/1982) applied it to nets; Rumelhart, Hinton and Williams (1986) popularised it; LeCun proposed a version in 1985–87. The chain rule is usually credited to Leibniz (1676). Harry Law's June 2026 history covers the lineage.

09:09Hardware moved on for 40 years; the learning algorithms haven't, so we burn energy running 1980s algorithms.

Needs nuance

The core — a global gradient computed by a backward pass through every layer — is unchanged. Almost everything around it did change, and much of it was hardware-driven: mixed precision (FP16→FP8→FP4), FlashAttention's recompute-to-avoid-HBM trick, activation checkpointing, MoE, ZeRO sharding, Muon-style optimisers. Critics would say the field has been co-designing with hardware all along; Bulatov's rejoinder is that these are patches on a dependency structure that's intrinsically communication-heavy.

09:10Averaging gradients over a batch instead of summing them is a design mistake people take for granted.

Opinion

Sum versus mean is a learning-rate rescaling; the substantive question is how the step size should scale with batch size, where the literature has linear-scaling (Goyal et al. 2017) and square-root rules and no consensus that either convention is a "mistake". Worth asking him what he'd do instead.

09:27He had a public argument on X with Yann LeCun that backprop wouldn't exist in five years.

Couldn't verify

X threads aren't well indexed and searches didn't surface the exchange. What is on record: his pinned 23 Dec 2025 post saying gradients are energy-inefficient because of long-range dependencies and announcing the SF reading group; LeCun has publicly and repeatedly defended backprop's staying power. Ask for the link.

09:30Every gradient step must fetch activations from all 100 (or 1,000) layers, so backprop has a low "compute to commute" ratio.

Verified

This is exactly why activation memory dominates training and why gradient checkpointing exists (see below). The backward pass touches every layer's stored activations once per step; deeper and wider models scale that traffic.

09:34Mixture-of-experts trains higher-quality models but is "terrible for inference" because of small batches and poor utilisation.

Needs nuance

MoE inference is harder (all experts resident in memory, routing imbalance, small-batch utilisation), which is why DeepSeek built an expert-parallel load balancer. But DeepSeek-V3/R1, Mixtral and Kimi models are served at scale and MoE is now the default frontier architecture precisely because it lowers inference cost per token. "Terrible" is dated; "harder to serve well" is right.

09:35Somebody — Kimi — was doing "attention over depth", a possible route to disentangling layers.

Verified

Moonshot AI's Kimi team published "Attention Residuals" (16 March 2026): residual connections replaced by softmax attention over the outputs of all preceding layers, with a blocked variant for scale. They report a 1.25× compute advantage and gains across tasks when integrated into the 48B/3B Kimi Linear model; code is on GitHub. It's a striking data point for his "disentanglement in depth" idea, though it still trains with backprop.

09:36LeCun popularised backprop on a SPARC workstation; an iPhone is a million times more powerful, so you don't need big compute to study the core algorithm.

Mostly right

LeCun's 1989 zip-code network trained for three days on a Sun-4/260, an early SPARC workstation; Karpathy's 2022 reproduction ran in about 90 seconds on a laptop CPU. A Sun-4 sat around 1 MFLOPS; a current iPhone GPU is in the TFLOPS range, so "a million times" is defensible on peak arithmetic (and the wall-clock ratio above is a few thousand, because the 1989 code wasn't compute-bound in the same way).

09:37Sparse parity started the first AI winter: a New York Times article said the Perceptron would guide people to the moon, then Minsky showed it couldn't solve parity.

Mostly right

The NYT piece is real — New Navy Device Learns By Doing, 8 July 1958 — and it promised a machine that would walk, talk, see, write, reproduce itself and be conscious of its existence, and that perceptrons might be sent to the planets as space explorers (not "to guide people to the moon"). Minsky and Papert's Perceptrons (1969) proved single-layer perceptrons cannot compute parity/XOR (Hardt & Recht's account). Historians treat the funding collapse as multi-causal (the 1973 Lighthill report, DARPA cuts), but the book is widely credited as the trigger for neural nets specifically.

09:41In the group, agents routinely turn a neural-network baseline into a specialised solution 1,000× more energy-efficient.

Couldn't verify

The infrastructure is public — sutro-problems auto-scores sparse-parity submissions on accuracy, speed and energy (DMD), and SutroYaro lists 34 experiments — but I found no published measurement backing the 1,000× figure. A caveat he'd probably accept: sparse parity has an exact algebraic solver (Gaussian elimination over GF(2)), so beating gradient descent on it says little about generality. The interesting test is MNIST-without-labels, the group's third problem.

09:42At 1–2% MNIST error, linear classifiers are ruled out; at 1% you have to use convolutional networks.

Half wrong

Linear classifiers are indeed out (12% error, 7.6% with pairwise tricks). But sub-1% never required convolutions: on LeCun's own leaderboard a degree-4 polynomial SVM got 1.1% and a "virtual" SVM 0.8%; Simard's 2003 one-hidden-layer MLP got 0.7%; Cireşan et al. (2010) reached 0.35% with a plain deep MLP and distortions (table in Hardt & Recht). This matters for the Sutro benchmark design: the 1% target will not force structure the way he implied.

09:43Early OpenAI had ~50 people, no LLMs (Radford working alone), was bullish on RL; his A3C implementation needed ~1M frames to learn Pong.

Mostly right

Consistent with public history: OpenAI was a few dozen people in 2016–17; language modelling became a formal effort around GPT-1 (June 2018); RL and robotics dominated early. A3C on Pong in about a million frames is in the range reported for actor-critic methods of that era. The Schulman anecdote is unverifiable.

09:45LeCun's cake: unsupervised learning is the cake, supervised the frosting, RL the cherry.

Verified

LeCun's NIPS 2016 keynote slide (widely reproduced; not re-fetched here).

09:46He built gradient checkpointing with Tim Salimans at OpenAI because PixelCNN didn't fit in memory; it became a TensorFlow package.

Verified

OpenAI released the package in January 2018; the repo credits Salimans and Bulatov, cites Chen et al. (2016) "sublinear memory cost", and reports fitting >10× larger feed-forward models for ~20% more compute. The repo now lives under Bulatov's cybertronai GitHub org — the same org that hosts the Sutro projects.

09:46People later found checkpointing also sped up training, because of the memory wall.

Mostly right

Plausible and there's a famous analogue: FlashAttention (Dao et al., 2022) deliberately recomputes attention in the backward pass to avoid writing the attention matrix to HBM, and ends up faster — the canonical "compute is cheaper than commute" result. Recomputation also speeds training indirectly by enabling larger batches. Direct evidence that vanilla checkpointing is faster in isolation is thinner; NVIDIA's selective-recomputation paper (2022) reports small overheads rather than speedups.

09:47Save a checkpoint every 10 layers of 100 and recompute — you store ~28 Jacobians instead of 100.

Verified

That's the O(√n) scheme from Chen et al. (2016): √100 = 10 checkpoints plus up to 10 recomputed activations live at once; "28" is an illustrative constant.

Claims: matrix multiplication and agentic discovery

09:18In the last month Google released a result dropping the matrix-multiplication exponent — "I think 2.3".

Verified

Posted 17 August 2026: "Improving the matrix multiplication exponent with modern optimization and AlphaEvolve" (Google DeepMind with Alman, Vassilevska Williams and Zhou — the holders of the previous record). New bound ω < 2.371177, down from 2.371339, with AlphaEvolve used to refine the optimiser (summary). "2.3" is a loose rounding of 2.37; the improvement is 1.6 × 10⁻⁴ and purely asymptotic — a galactic algorithm, which reinforces his point that ω-chasing doesn't move practice.

09:18Naive matmul is about n³ (2n³ flops); Strassen is about n^2.8.

Verified

Strassen (1969): log₂7 ≈ 2.807. Correct.

09:18There is still no generic best algorithm for matrix multiplication.

Verified

ω is unknown (conjectured 2, best bound 2.371177 as of last month).

09:16Agents made algorithmic discovery reachable this year; discoveries at the frontier tend to be simultaneous.

Mostly right

"This year" is generous: AlphaEvolve (May 2025) improved the best-known solution on 20% of 50+ open problems and found a 48-multiplication 4×4 complex matmul, and last month's ω bound extends that. Simultaneous discovery is Merton's "multiples" (1961) — Strassen-era matmul and backprop itself are examples.

09:17"Yad" implemented all of Hinton's papers using a couple of billion tokens; the repo, under Cybertron AI, has animations for each paper.

Mostly right

Confirmed: cybertronai/hinton-problems (53 Hinton experiments, 1–3 May 2026) and schmidhuber-problems (58 experiments, 6–8 May), pure NumPy, built by Yad Konrad with Claude Code agent teams; his write-up reports 49 and 41 wall-clock hours, and Mark Saroufim cited the build in his MLSys keynote (21 May 2026). Bulatov's LinkedIn post says under $200 in three days; the build notes note that cache reads dominate token counts, which is how "billions of tokens" and "under $200" can both be true. I couldn't confirm the animations detail.

Claims: the AI industry and hardware startups

09:13About eight years ago Cade Metz wrote a New York Times piece on the explosion of AI hardware startups.

Verified

"Big Bets on A.I. Open a New Frontier for Chip Start-Ups, Too", NYT, 14 January 2018 (syndicated copy): at least 45 startups, five with $100M+ raised, $1.5B of VC in 2017.

09:13Cerebras is still around; Groq was recently "reverse acquired" by Nvidia; SambaNova is still around; most of the rest didn't make it.

Verified

Groq: on 24 December 2025 Nvidia took a non-exclusive licence to Groq's inference technology for about $20B and hired founder Jonathan Ross, president Sunny Madra and key engineers, while GroqCloud continues under a new CEO — analysts call it a "reverse acqui-hire" (Constellation, Techstrong). SambaNova: Intel's $1.6B takeover talks stalled in January; it then raised a $350M Series E in February 2026 with Intel participating and shipped the SN50 (DCD, Reuters via Yahoo). "Most didn't make it" fits the class of 2018 (Wave Computing bankrupt 2020, Nervana shut by Intel 2020, Graphcore sold to SoftBank 2024, Mythic nearly folded 2022 — from background knowledge).

09:14Mitchell Nahmias's optical-interconnect startup didn't make it; when he arrived at OpenAI they'd found algorithmic tricks that removed the need for fast interconnect.

Partly verified

Nahmias co-founded Luminous Computing (2018; CTO; Princeton neuromorphic-photonics PhD), which raised a $105M Series A in 2022 with Bill Gates among investors, and had pivoted in 2019 from optical compute to optical interconnect between chips, racks and data centres (Next Platform, VentureBeat). The shutdown and his move to OpenAI appear only in forum chatter in my searches; no mainstream confirmation.

09:14DeepSeek got "crippled" H800s with H100 compute but reduced memory bandwidth, and engineered around it.

Wrong detail

The H800 export variant keeps the H100's 80 GB HBM3 and its memory bandwidth; what was cut is chip-to-chip NVLink interconnect bandwidth (and FP64). DeepSeek-V3's paper cites ~160 GB/s effective NVLink versus 50 GB/s InfiniBand and built DualPipe and custom PTX-level communication to overlap the constrained interconnect (TechRadar, DeepSeek's open-infra index). The moral — software routed around a hardware limit — is right; the spec is not.

09:12The cost of software dropped at least 10×, maybe 100×, this year; hardware didn't.

Forecast / opinion

Anecdotes point both ways. For: Yad's 111 paper reproductions in ~90 agent-hours; StrongDM's "software factory" rule of $1,000 of tokens per engineer-day. Against: METR's 2025 randomised trial found experienced open-source developers 19% slower with AI tools (background knowledge); and the token bills below suggest cost moved to inference rather than vanished. No study supports a clean 10–100× yet.

09:58Uber declared it used twelve months' worth of tokens in four months — and the app didn't get better.

Verified (first half)

In April 2026 Uber's CTO said the company had exhausted its full-year Claude Code budget in four months, with 84–95% of ~5,000 engineers active monthly and individual bills of $150–$2,000 (TechFlow, Stocktwits/Yahoo). Sam Altman has called "spent my 2026 budget in Q1" a meme, and Microsoft cut internal Claude Code licences in May. Whether the Uber app improved is his opinion; the outcome-metrics question he raises (PRs vs. things people care about) is the productive one.

10:00His essay on AI companies needing an "AI god" narrative to justify valuations (title garbled as "God Golden GPUs").

Couldn't locate

His Substack launched around February 2026 but its index didn't render for me and search didn't surface a post by that name. Ask for the link; the title as transcribed is probably wrong.

Claims: energy, water, and the grid

09:51Data-centre water use is small compared with agriculture — California's lawns and almonds.

Mostly right

Nationally, yes: LBNL puts direct US data-centre water consumption at 17.4 billion gallons in 2023, about 0.3% of public supply, and about 228 billion gallons once power-plant cooling is counted — roughly 2% of US consumptive water use (MOST, CRS/LBNL summary). Agriculture is ~80% of California's developed water use (background knowledge). The counterpoint is local: in The Dalles, Oregon, Google's water use grew 316% while the town grew 12% (state-level data), and LBNL projects direct use rising to 38–73 billion gallons by 2028.

09:52He'd heard data centres use about 5% of electricity.

Verified

US: 176 TWh, 4.4% of national electricity in 2023 (LBNL 2024). The LBNL 2025 update projects 11.8% by 2030 (range 9.5–15.3%), up from the earlier 6.7–12% by 2028. Globally it's closer to 1.5% (IEA 2024, background knowledge). So "5%" is right for the US today and will be badly out of date within four years.

09:52His own electricity bill hasn't gone up much, which suggests data centres' impact hasn't been huge.

Contradicted regionally

A sample of one in California misses where the effect is. In PJM (13 mid-Atlantic and Midwest states, 65M people) capacity prices went from $28.92/MW-day (2024/25) to $329.17 (2026/27); the independent market monitor attributes 63% of the 2025/26 increase — $9.3B — to data centres (IEEFA). PJM retail prices rose ~49% in five years versus ~33% nationally (Industrial Info), and Q1 2026 wholesale costs were up 76% year on year (E&E News). The fairest counter-reading is E3's May 2026 review, which attributes only ~50% of the PJM rise to load growth. California's high rates are driven by wildfire and transmission costs, not data centres.

09:52Interconnection studies already throttle how much data-centre demand can come online.

Verified

Google says utilities quote four to ten years to interconnect, one as long as twelve (Network World); ~2,300 GW of generation and storage sit in US queues. PJM's December 2025 auction fell 6.6 GW short of its reliability target for 2027/28 (Introl).

09:53CoWoS packaging (attaching HBM next to the die) is the main bottleneck; only TSMC can do it; TPUs had to cut their allocation.

Mostly right

CoWoS-S and CoWoS-L are fully booked with 52–78-week lead times; Nvidia has ~60% of 2026 capacity (Silicon Analysts, summary of Morgan Stanley data). Google's 2026 TPU output was reportedly cut from ~4M to ~3M units because of CoWoS access, with a backlog of 3M+ unpackaged dies (BigGo/Fubon, Isaiah Research). Two corrections: CoWoS is TSMC's brand, but ASE/SPIL, Amkor and Intel Foundry offer comparable 2.5D packaging; and SemiAnalysis (April 2026) argues N3 front-end wafer capacity is now the binding constraint as CoWoS eases.

09:53In a couple of years chips stop being the bottleneck and grid capacity becomes it — so intelligence per joule is the lever.

Mostly right

Mainstream view now: Morgan Stanley (August 2026) estimates US AI data centres need 68 GW of new power by 2028 with only 30 GW covered — a 38 GW gap (summary); Satya Nadella has said the constraint is power, not compute; Gartner expects 40% of AI data centres to be power-constrained by 2027. It's a forecast, but a well-supported one — and it's the best "why now" argument for his research programme.

09:54Recursive self-improvement isn't a 1/(1−x) curve; humans have been doing it for 5,000 years; it just clears the current bottleneck faster.

Opinion

A bottleneck-shifting model rather than a singularity model. Compatible with the compute- and power-constraint literature above; incompatible with the strong RSI view. Not empirically decidable today.

Claims: limits of intelligence and history

10:01Chess was "the Drosophila of AI"; we had chess superintelligence in the '90s and people were still unsatisfied.

Verified

McCarthy's phrase (1990); Deep Blue beat Kasparov in 1997. Background knowledge, not re-fetched.

10:02Formulate any single task — ARC-AGI, say — and it gets solved far better than humans.

Mostly right

ARC-AGI-1 is saturated (multiple systems above 85%). On ARC-AGI-2, average individual human testers score roughly 53–60%, and by mid-2026 frontier systems (GPT-5.x, Gemini 3 Deep Think) post 85–92% on the semi-private set (leaderboard snapshot, analysis of the human baseline). Caveats: the "human panel" (≥2 people) still scores 100%; the 2025 Kaggle prize under a $0.20/task cost cap topped out at 24%, and the 2026 prize is still open. So "much better than a human" holds only when cost is ignored — which is precisely Bulatov's intelligence-per-joule point.

10:04Collatz: even/halve, odd/triple-plus-one; it always seems to reach 1, and there's no way to know except simulating.

Mostly right

Correctly stated and still open. Verified computationally for every starting value below 2⁷¹ (Bařina, J. Supercomputing 2025). One partial shortcut exists: Tao (2019) proved almost all orbits reach almost-bounded values — but there is no decision procedure, so his point stands.

10:05Rule 110 is known to have no shortcut: to know the state after a quadrillion steps you must run a quadrillion steps.

Mostly right

Predicting t steps of Rule 110 is P-complete (Neary & Woods, ICALP 2006), which rules out a fast parallel shortcut unless P = NC. That is believed but unproven, so "known" is slightly strong; "no shortcut is expected" is exact.

10:03Some problems — the three-body problem to enough accuracy — need more energy than the Sun provides; light speed and energy-per-operation are fundamental limits.

Verified

Chaotic dynamics make error grow exponentially, so precision costs are unbounded; Landauer's bound, the Margolus–Levitin bound and Bremermann's limit are the standard physical ceilings (background knowledge).

Claims: attention economy and psychology

10:10There's a film, "Good Luck Have Fun Don't Die", about people sucked into short-form video.

Verified

Gore Verbinski's Good Luck, Have Fun, Don't Die (Sam Rockwell; US release 13 February 2026) is a time-travel comedy about stopping a rogue AI, framed by critics as a satire of AI and technology addiction (RogerEbert.com, Deep Focus). His plot summary is loose but the theme is right.

10:10Ten years ago the NeurIPS conference was held in a casino in Nevada, where people were pulling slot levers.

Mostly right

NIPS 2012 and 2013 were held at Harrah's and Harveys casinos in Lake Tahoe, Nevada — thirteen years ago (background knowledge).

10:12Two reward pathways: dopamine gives anticipation, opioids give satisfaction; short-form video feeds only the first, so you end up drained.

Mostly right

A fair lay summary of Berridge and Robinson's "wanting versus liking" work: dopamine mediates incentive salience (wanting), while opioid hedonic hotspots mediate liking. Whether short-form video specifically fails to deliver "liking" is a hypothesis, not an established finding (background knowledge).

10:11TikTok has a billion users.

Verified

TikTok has reported over a billion monthly users since 2021 (background knowledge).

10:08In an economy where everyone has enough to eat, engagement is the last currency; AI is another lever for capturing attention, and the worst case is an AI that optimises pixels directly for engagement.

Opinion

A restatement of the attention-economy thesis (Tim Wu's The Attention Merchants; Herbert Simon's 1971 line about information consuming attention). His "people stuck in Claude Code feeding it tokens" version of the same failure mode is echoed by the "tokenmaxxing" coinage in the Uber coverage above.

Claims: the guest's bio (host intro, 09:08)

09:08Google Brain, early OpenAI, Meta, Together AI; led the first production deployment of deep learning at Google; wrote gradient checkpointing at OpenAI; runs the weekly Sutro Group at South Park Commons.

Mostly right

Employers match his X bio (South Park Commons; early OpenAI, Google Brain, Meta) and an AI Council profile listing him as principal researcher at Together AI, which also carries the "first production deployment" line and adds that he hired Ian Goodfellow as his intern. The deployment is the Street View house-number transcription system (Goodfellow, Bulatov, Ibarz, Arnoud, Shet, arXiv 1312.6082, Dec 2013), which read close to 100 million street numbers at human accuracy (MIT Tech Review). "First" is his own framing — Google's Android speech recogniser shipped a deep net in 2012 — so "one of the first" is safer. The reading group was announced in his pinned post of 23 December 2025; the Sutro repos live at github.com/cybertronai. Born 1980 in the USSR: consistent with his public profile, not independently checkable.

Strongest supporting arguments for his thesis

If you wanted to make Bulatov's case for him with the best evidence available today, these are the load-bearing pieces:

  • The cost model is not controversial among architects. Dally's 20 fJ add versus 1.9 pJ/mm move versus 1.3 nJ DRAM fetch, and Horowitz's 2014 table, are cited in every energy-efficiency keynote. The ETH "spatial computer" paper proves that under this model matmul's energy is bounded below by data distance, independent of the multiplication count — theory catching up with his intuition.
  • Practice already follows "compute is cheaper than commute". FlashAttention recomputes to avoid HBM traffic and gets faster; AlphaEvolve's real-world matmul win was a tiling change; Groq's LPU (now licensed by Nvidia for ~$20B) and Cerebras both bet on keeping weights in on-chip SRAM to avoid the trip to HBM. The market has priced data movement as the problem.
  • Backprop's dependency structure is the memory problem. Activation memory dominates training; checkpointing (his own package) and Attention Residuals' "block" variant both exist to manage traffic created by every-layer dependencies. His claim that you can't fix backprop without touching the architecture is consistent with where Moonshot's paper points.
  • Software really got cheap — at the research-prototyping scale that matters for algorithm search. 111 classic papers reproduced by agents in ~90 hours for under a few hundred dollars is the concrete existence proof for "shorten a 20-year plan to five". Enterprises burning annual token budgets in a quarter shows the same thing from the demand side.
  • Energy is becoming the binding constraint. PJM's 6.6 GW shortfall, Morgan Stanley's 38 GW gap by 2028, multi-year interconnection queues and 11.8%-of-US-electricity projections mean "intelligence per joule" is where the marginal capability will have to come from once CoWoS and N3 loosen.
  • Hardware bets are slow and often mistimed. The 2018 chip-startup cohort mostly failed or was absorbed; Luminous pivoted and faded; DeepSeek showed software can neutralise a hardware constraint in a single training run. His "why now is software, not hardware" argument has a strong track record behind it.
  • A community is forming around backprop-free, local learning for energy reasons. Forward-Forward (2022), Cascaded-Forward, Mono-Forward, predictive coding (Salvatori et al. 2026) and a May 2026 Hyperspherical Forward-Forward result passing 25% top-1 on ImageNet show the direction is live, if far behind.

Strongest counterarguments

  • Forty years of failed replacements. Every local or forward-only rule still trails backprop badly at scale: the best greedy local-learning ImageNet result in 2026 is ~25% top-1 (Hyperspherical Forward-Forward), and a 2025 rigorous evaluation of forward-only algorithms finds the energy savings real but the accuracy gap persistent. LeCun's position — that backprop's generality has beaten every biologically-motivated alternative — is the default prior for a reason.
  • Hardware is attacking the same wall, faster than he credits. CoWoS capacity is tripling in two years; HBM4, wafer-scale, SRAM-first accelerators, optical interconnect (Lightmatter, Ayar, Celestial) and 3D stacking all target data movement directly. If the wall moves, the payoff from an algorithm co-designed to today's wall shrinks.
  • "Arithmetic is 1%" overstates. At tensor-core utilisation on a well-tiled GEMM the datapath is a large share of chip power. His thesis survives (movement still dominates in aggregate), but the slogan invites a correct rebuttal from any GPU engineer.
  • RL as "stone soup" is hard to defend after 2025. DeepSeek-R1's reasoning came from large-scale RL with verifiable rewards on top of a pretrained base; o1-style test-time compute is RL-trained. He'd say that's the cherry on a pre-existing cake — but the cherry is what moved the frontier in 2025–26.
  • The toy-problem programme has a validity risk. Sparse parity has a closed-form solver; MNIST is reachable below 1% without convolutions. An agent that finds a 1,000× cheaper special-purpose solution is demonstrating specialisation, not a new learning rule. The group needs a task family where "specialised" and "general" solutions can't be told apart by the evaluator — worth raising on the show.
  • The environmental reassurance is partly wrong. National water and electricity shares are small today, but the price signal he appealed to has already fired in PJM, and LBNL's own 2030 projection nearly triples the share. "Capitalism will tell us" cuts against him in Virginia and Ohio.
  • Software-cost claims lack controlled evidence. The one randomised study of experienced developers (METR, 2025) found a slowdown; enterprise token bills show spend, not output. His own Uber anecdote is evidence for the sceptical reading.

Related groups and projects

Who is working the same problem, grouped by which part of his argument they touch.

Location-aware cost models and I/O complexity

WhoWhatLink
Bill Dally (NVIDIA / Stanford)Parallel Explicit Communication Model; the energy numbers everyone quotesCACM Point, Sept 2022
ETH Zurich SPCL (Hoefler, Ben-Nun, Gianinazzi)"The spatial computer": 2D-grid energy model with matching bounds for matmul, sorting, selectionarXiv 2205.04934
Demmel et al. (UC Berkeley); Hong & Kung lineageCommunication-avoiding algorithms and I/O lower bounds — the pre-history of "compute to commute"Background literature
Sutro Group (South Park Commons)Data-movement-distance (DMD) metric, sparse-parity and MNIST energy benchmarks, agent-run experimentsgithub.com/cybertronai

Backprop alternatives and depth restructuring

WhoWhatLink
Moonshot AI (Kimi team)Attention Residuals: attention over depth as a drop-in for residual sumsGitHub
Hinton; FF successors (Cascaded-Forward, Mono-Forward, Hyperspherical FF)Forward-only, layer-local training motivated by energy and biologyHFF (May 2026), survey
Predictive-coding community (Salvatori, van Zwol et al.)Energy-based local learning; JAX frameworks; matching backprop on CIFAR/Tiny-ImageNetPMC 2025, May 2026
Forward-only evaluation workRigorous energy-versus-accuracy comparisons of BP-free algorithmsarXiv 2511.01061
Feedback alignment, DNI, forward gradients, PEPITAEarlier lines removing weight transport or the backward pass (Lillicrap; Jaderberg; Baydin; Dellaferrera & Kreiman)Background literature

Measuring ML energy

WhoWhatLink
ML.ENERGY Initiative (U. Michigan; Chung, Chowdhury)Zeus measurement library, Perseus training optimiser, ML.ENERGY inference leaderboardBenchmark paper, Zeus
MLCommons MLPerf PowerWall-outlet power measurement for training and inference submissionsSee 2026 protocol survey
Hugging Face AI Energy ScoreFixed-batch energy ratings across model tasksSame survey

Agentic algorithm discovery

WhoWhatLink
Google DeepMind (AlphaEvolve) with Alman, Vassilevska Williams, Zhouω < 2.371177; 4×4 complex matmul in 48 multiplications; 23% Gemini kernel speedupAug 2026 paper, blog
Yad Konrad (Sutro)Agent-team reproductions of 53 Hinton and 58 Schmidhuber experiments in pure NumPywrite-up
OpenEvolve, Sakana AI (AI Scientist, ShinkaEvolve), FunSearchOpen and commercial evolutionary code-search systemsBackground knowledge

Hardware that bets against data movement

WhoWhatLink
Groq (tech now licensed to Nvidia)LPU with hundreds of MB of SRAM as primary weight storageDeal coverage
Cerebras, SambaNovaWafer-scale and reconfigurable-dataflow accelerators, both still independentSambaNova 2026
Lightmatter, Ayar Labs, Celestial AIOptical interconnect — the bet Luminous made too earlyBackground knowledge

Grid, water and energy analysts

WhoWhatLink
LBNL (Shehabi, Smith et al.)US data-centre energy and water reports; 2030 projections2025 update
PJM's market monitor (Monitoring Analytics); IEEFA; E3Attribution of capacity-price increases to data centres — and the dissentIEEFA, E3
SemiAnalysis; Morgan Stanley; Epoch AIChip-supply and power-gap forecastsSemiAnalysis

Follow-ups for the next episode

  • Get the DMD definition on the record. How is "data movement distance" computed for a Python function, and does it predict joules measured with something like Zeus or NVML? That's the bridge between his cost model and reality.
  • Ask for the 1,000× receipts. Which problem, what baseline, what energy measurement, and would the win survive on a task without a known algebraic solver?
  • Push on the MNIST target. Since SVMs and plain MLPs beat 1% error, what error threshold or constraint (parameter count? DMD budget?) actually forces structure?
  • Attention Residuals as a test case. Is Moonshot's attention-over-depth the "disentanglement in depth" he wants, or just a better residual? What would a backprop-free version look like?
  • FlashAttention framing. He never mentioned it, but it's the cleanest existing proof that recomputation beats data movement. Worth asking why the same trade hasn't been pushed further up the stack.
  • Correct the H800 point on air. The cut was NVLink interconnect, not HBM bandwidth — same moral, better facts.
  • The PJM price signal. Does a 10× capacity-price rise change his view that "capitalism will tell us" when data centres bite?
  • Hardware's counter-move. With CoWoS capacity tripling and optical interconnect shipping, how much of his advantage depends on today's wall staying put?
  • Links to collect: the LeCun thread; the Substack essay on AI-god narratives; the Cybertron animation demo.

Sources

  1. Dupont, Eisenberger, Kozlovskii, Mehrabian, Ruiz, See, Zhou, Alman, Vassilevska Williams, Balog — "Improving the matrix multiplication exponent with modern optimization and AlphaEvolve", arXiv 2608.16884 (17 Aug 2026). arxiv.org/abs/2608.16884; summary at alphaxiv.org
  2. Google DeepMind — AlphaEvolve blog post. deepmind.google
  3. Dally, W. — "Point: We Must Extend Our Model of Computation to Account for Cost and Location", CACM 65(9), Sept 2022. cacm.acm.org
  4. Gianinazzi, Ben-Nun, Ashkboos, Baumann, Luczynski, Hoefler — "The spatial computer: A model for energy-efficient parallel computation", arXiv 2205.04934. arxiv.org
  5. Horowitz, M. — "Computing's Energy Problem (and what we can do about it)", ISSCC 2014 plenary. PDF mirror; EETimes coverage via design-reuse.com
  6. ASML — "Revitalization through refurbishment" (Nov 2023). asml.com
  7. Kimi Team — "Attention Residuals", arXiv 2603.15031 (Mar 2026). arxiv.org; code github.com/MoonshotAI
  8. Groq–Nvidia deal coverage: Constellation Research; Techstrong.ai
  9. SambaNova: DCD (Jan 2026); Reuters via Yahoo (May 2026)
  10. Luminous Computing: The Next Platform (Mar 2022); VentureBeat
  11. H800 versus H100: TechRadar; DeepSeek open-infra index (DualPipe, EPLB). github.com/deepseek-ai
  12. Cade Metz, NYT, 14 Jan 2018 — syndicated at WRAL and SupplyChainBrain
  13. Uber token budget: TechFlow (May 2026); Stocktwits via Yahoo (June 2026)
  14. Yaroslav Bulatov: X profile; Substack; LinkedIn; AI Council profile; Digg summary of his June 2026 energy thread; "Reinventing AI From Scratch" interview (Mar 2026)
  15. Goodfellow, Bulatov, Ibarz, Arnoud, Shet — arXiv 1312.6082. hyper.ai; MIT Technology Review
  16. Gradient checkpointing repo. github.com/cybertronai/gradient-checkpointing
  17. Cybertron AI org (hinton-problems, schmidhuber-problems, SutroYaro, sutro-problems, ByteDMD). github.com/cybertronai; schmidhuber-problems; build notes; Yad Konrad's write-up
  18. LBNL — US Data Center Energy Usage Report, 2025 update. eta.lbl.gov; modelling page datacenters.lbl.gov
  19. Data-centre water: MOST Policy Initiative (Apr 2026); CRS/LBNL summary; state-level data
  20. PJM prices: IEEFA; E3 whitepaper (May 2026); E&E News (May 2026); Industrial Info (Jul 2026)
  21. Grid constraints: Network World; Introl; Morgan Stanley 38 GW summary (Aug 2026)
  22. CoWoS: Silicon Analysts; INDmoney summary; BigGo/Fubon on TPU output; Isaiah Research; SemiAnalysis (Apr 2026)
  23. MNIST leaderboard and history: Hardt & Recht, Patterns, Predictions, and Actions; Cireşan et al. 2010 arXiv 1003.0358
  24. Perceptron history: Open University; Hardt & Recht (above)
  25. Karpathy — reproduction of LeCun 1989. mirror
  26. ARC-AGI-2: ARC Prize 2025 results; leaderboard snapshot; human-baseline analysis
  27. Bařina, D. — "Improved verification limit for the convergence of the Collatz conjecture", J. Supercomputing 81 (2025). fit.vut.cz
  28. Neary & Woods — "P-completeness of Cellular Automaton Rule 110", ICALP 2006. Maynooth archive
  29. Backprop-free learning: arXiv 2511.01061; Hyperspherical Forward-Forward (May 2026); FF/CaFo/MF survey; predictive-coding framework (2025); closed-form predictive coding (May 2026)
  30. Energy measurement: ML.ENERGY Benchmark; Zeus; U-M story; 2026 joules-per-task survey
  31. D'Alberto — "Strassen's Matrix Multiplication Algorithm Is Still Faster" (2023). arxiv.org
  32. Good Luck, Have Fun, Don't Die: RogerEbert.com review; Deep Focus Review
  33. Law, H. — "Backpropagation is older than you think" (Jun 2026). learningfromexamples.com