Every checkable claim from the 8 September 2026 recording, tested against sources as of the same day, with the strongest case for his thesis, the strongest case against it, who else is working the same problem, and what to ask next time.
Nearly everything Bulatov argued — that backprop is energy-inefficient, that theoretical matmul results miss the point, that Dally's 2D model is the right one — rests on one physical fact: on a modern chip, arithmetic is nearly free and moving bytes is not. Bill Dally's 2022 numbers (from his Communications of the ACM "Point" essay) put it on a log scale:
Bulatov's rule of thumb on air — that an add costs about the same as moving its two inputs ten microns — matches Dally's figures almost exactly (1.9 pJ per millimetre is 19 fJ per 10 µm, against a 20 fJ add). His term for the ratio you want to maximise is "compute to commute"; the literature calls the same idea arithmetic intensity, I/O complexity, or communication-avoiding algorithms.
The pattern: Bulatov is reliable on the physics and the systems arguments — the memory-wall numbers, the ASML and CoWoS facts, the AlphaEvolve result, the Groq deal — and loose on dates and specs remembered from a distance (the Horowitz keynote year, the H800's cut spec, what it takes to reach 1% on MNIST). The three clear errors don't damage the core argument; the "needs nuance" items are where a sceptical listener would push back, and they're listed with the counter-evidence below.
09:11Mark Horowitz gave a talk in 2012 called "The Computing Energy Problem", now 14 years old, predicting the death of Dennard scaling and that CMOS would stay.
The talk exists but is the ISSCC 2014 plenary, "Computing's Energy Problem (and what we can do about it)" — February 2014, so twelve years old, not fourteen. The content is as described: power became the binding constraint before Dennard scaling broke down, and no new device technology will rescue us; EETimes' coverage summarises his call for specialised silicon plus better algorithms. His prediction that CMOS persists has held.
09:12ASML's statistic: 95% of the machines it built in the last 30 years are still in use.
ASML's own 2023 story states that roughly 95% of lithography systems sold in the past 30 years remain active, with more than 5,000 machines in fabs as of end-2022, sustained by a refurbishment programme. Its product pages repeat that almost every system ever shipped is still in a customer fab.
09:19Transistors shrink in two dimensions, wires in one; so arithmetic got cheap much faster than data transfer. That's the "memory wall".
The conclusion is textbook and quantified in Dally's CACM Point essay (Sept 2022): arithmetic and logic costs fell by orders of magnitude while communication costs fell far more slowly, so a DRAM fetch of two 32-bit words now costs 64,000× a 32-bit add. Horowitz's 2014 table makes the same point at 45 nm. The "2D vs 1D shrink" geometry is a simplification — the deeper reason is that wire RC delay and energy per unit length don't improve with scaling — but it's a fair intuition pump.
09:20Heuristic: adding two numbers costs the same as moving them 10 microns across the chip.
Dally's numbers: a 32-bit add is 20 fJ; moving its two operands 1 mm costs 1.9 pJ, i.e. about 19 fJ per 10 µm. The heuristic is accurate to within rounding. Bulatov has posted the same arithmetic on X (June 2026 thread), citing the Parallel Explicit Communication Model.
09:20A chip is about 20 × 20 mm; a micron is a millionth of a metre, so ~20,000 microns across.
Ballpark. Dally's worked example uses a 400 mm² chip (20 × 20 mm); flagship AI dies are larger (an H100 is ~814 mm², close to the ~26 × 33 mm reticle limit). The arithmetic on microns is right.
09:20Movement is so dominant that arithmetic is effectively free — about 1% of energy — yet fast-matmul researchers count arithmetic.
True whenever data comes from far away (see the ladder above) and true of the theory literature's cost model. But it is not true of a dense GEMM running on tensor cores with good tiling: at high utilisation the arithmetic units are a large share of chip power, which is exactly why kernels are tuned for arithmetic intensity (the roofline model). "1%" describes a memory-bound kernel or a global-memory access, not a well-blocked matmul. The directional claim — that asymptotic multiplication counts are the wrong objective — is widely shared.
09:21GPU libraries don't use Strassen; they use tiling. Discovered fast algorithms aren't used in practice.
cuBLAS/CUTLASS-style GEMMs are tiled classical multiplication. Strassen-family algorithms exist in research — D'Alberto (2023) argues Strassen still beats the AlphaTensor variants on real hardware — but they're not in mainstream GPU libraries. Notably, AlphaEvolve's only production-relevant matmul win was a 23% kernel speedup from smarter tiling, which supports Bulatov's point about what matters in practice.
09:22Six years ago Bill Dally proposed a computation model where data is laid out on a 2D grid and you pay for distance.
The essay is "We Must Extend Our Model of Computation to Account for Cost and Location", CACM September 2022 — four years ago. Dally proposes extending PRAM into a Parallel Explicit Communication Model with Manhattan-distance costs. A parallel academic formalisation appeared the same year: ETH Zurich's "The spatial computer" (Gianinazzi, Ben-Nun, Hoefler et al., May 2022) puts processors on a 2D grid with energy equal to distance travelled and proves matching upper and lower energy bounds for matrix multiplication and sorting. The older roots are Thompson's VLSI complexity (1979) and Hong & Kung's red-blue pebble game (1981), which Bulatov didn't mention.
09:25Stacking 16 layers at ~100 W/cm² each would need 1.6 kW/cm² of heat removal — so 3D is thermally limited.
Sound back-of-envelope: an H100 runs roughly 700 W over ~8 cm² (~85 W/cm²). Logic-on-logic stacking is thermally limited; the 12–16-high stacks that ship today are HBM (memory), which dissipates far less per layer. His conclusion — that even in 3D you exhaust the "nearby" volume quickly — is the standard argument.
09:26Low peak memory is sufficient but not necessary for low energy; you need a finer-grained measure than peak memory.
This is a direct consequence of distance-based cost models: energy tracks how far accesses travel, not the footprint. The Sutro repos operationalise it as a "data movement distance" (DMD) metric (ByteDMD, sutro-problems), which is the right follow-up to ask about.
10:02A signal can't cross a chip in one clock cycle, which limits how much knowledge one chip can hold.
Standard since the early 2000s (Agarwal et al., ISCA 2000, showed only a shrinking fraction of a die is reachable per cycle). Dally's figures: corner to corner on a 40 mm chip takes 16 ns versus 150 ps for an add — roughly 100 cycles at 6 GHz-class clocks.
09:28Backprop was invented in the '80s (when it became useful); a Finnish dissertation in the '60s; you can go back to Newton's chain rule.
Seppo Linnainmaa's master's thesis is from 1970 (published 1976), not the 1960s; control-theory precursors (Kelley 1960, Bryson 1961, Dreyfus 1962) are the '60s work. Werbos (1974/1982) applied it to nets; Rumelhart, Hinton and Williams (1986) popularised it; LeCun proposed a version in 1985–87. The chain rule is usually credited to Leibniz (1676). Harry Law's June 2026 history covers the lineage.
09:09Hardware moved on for 40 years; the learning algorithms haven't, so we burn energy running 1980s algorithms.
The core — a global gradient computed by a backward pass through every layer — is unchanged. Almost everything around it did change, and much of it was hardware-driven: mixed precision (FP16→FP8→FP4), FlashAttention's recompute-to-avoid-HBM trick, activation checkpointing, MoE, ZeRO sharding, Muon-style optimisers. Critics would say the field has been co-designing with hardware all along; Bulatov's rejoinder is that these are patches on a dependency structure that's intrinsically communication-heavy.
09:10Averaging gradients over a batch instead of summing them is a design mistake people take for granted.
Sum versus mean is a learning-rate rescaling; the substantive question is how the step size should scale with batch size, where the literature has linear-scaling (Goyal et al. 2017) and square-root rules and no consensus that either convention is a "mistake". Worth asking him what he'd do instead.
09:27He had a public argument on X with Yann LeCun that backprop wouldn't exist in five years.
X threads aren't well indexed and searches didn't surface the exchange. What is on record: his pinned 23 Dec 2025 post saying gradients are energy-inefficient because of long-range dependencies and announcing the SF reading group; LeCun has publicly and repeatedly defended backprop's staying power. Ask for the link.
09:30Every gradient step must fetch activations from all 100 (or 1,000) layers, so backprop has a low "compute to commute" ratio.
This is exactly why activation memory dominates training and why gradient checkpointing exists (see below). The backward pass touches every layer's stored activations once per step; deeper and wider models scale that traffic.
09:34Mixture-of-experts trains higher-quality models but is "terrible for inference" because of small batches and poor utilisation.
MoE inference is harder (all experts resident in memory, routing imbalance, small-batch utilisation), which is why DeepSeek built an expert-parallel load balancer. But DeepSeek-V3/R1, Mixtral and Kimi models are served at scale and MoE is now the default frontier architecture precisely because it lowers inference cost per token. "Terrible" is dated; "harder to serve well" is right.
09:35Somebody — Kimi — was doing "attention over depth", a possible route to disentangling layers.
Moonshot AI's Kimi team published "Attention Residuals" (16 March 2026): residual connections replaced by softmax attention over the outputs of all preceding layers, with a blocked variant for scale. They report a 1.25× compute advantage and gains across tasks when integrated into the 48B/3B Kimi Linear model; code is on GitHub. It's a striking data point for his "disentanglement in depth" idea, though it still trains with backprop.
09:36LeCun popularised backprop on a SPARC workstation; an iPhone is a million times more powerful, so you don't need big compute to study the core algorithm.
LeCun's 1989 zip-code network trained for three days on a Sun-4/260, an early SPARC workstation; Karpathy's 2022 reproduction ran in about 90 seconds on a laptop CPU. A Sun-4 sat around 1 MFLOPS; a current iPhone GPU is in the TFLOPS range, so "a million times" is defensible on peak arithmetic (and the wall-clock ratio above is a few thousand, because the 1989 code wasn't compute-bound in the same way).
09:37Sparse parity started the first AI winter: a New York Times article said the Perceptron would guide people to the moon, then Minsky showed it couldn't solve parity.
The NYT piece is real — New Navy Device Learns By Doing, 8 July 1958 — and it promised a machine that would walk, talk, see, write, reproduce itself and be conscious of its existence, and that perceptrons might be sent to the planets as space explorers (not "to guide people to the moon"). Minsky and Papert's Perceptrons (1969) proved single-layer perceptrons cannot compute parity/XOR (Hardt & Recht's account). Historians treat the funding collapse as multi-causal (the 1973 Lighthill report, DARPA cuts), but the book is widely credited as the trigger for neural nets specifically.
09:41In the group, agents routinely turn a neural-network baseline into a specialised solution 1,000× more energy-efficient.
The infrastructure is public — sutro-problems auto-scores sparse-parity submissions on accuracy, speed and energy (DMD), and SutroYaro lists 34 experiments — but I found no published measurement backing the 1,000× figure. A caveat he'd probably accept: sparse parity has an exact algebraic solver (Gaussian elimination over GF(2)), so beating gradient descent on it says little about generality. The interesting test is MNIST-without-labels, the group's third problem.
09:42At 1–2% MNIST error, linear classifiers are ruled out; at 1% you have to use convolutional networks.
Linear classifiers are indeed out (12% error, 7.6% with pairwise tricks). But sub-1% never required convolutions: on LeCun's own leaderboard a degree-4 polynomial SVM got 1.1% and a "virtual" SVM 0.8%; Simard's 2003 one-hidden-layer MLP got 0.7%; Cireşan et al. (2010) reached 0.35% with a plain deep MLP and distortions (table in Hardt & Recht). This matters for the Sutro benchmark design: the 1% target will not force structure the way he implied.
09:43Early OpenAI had ~50 people, no LLMs (Radford working alone), was bullish on RL; his A3C implementation needed ~1M frames to learn Pong.
Consistent with public history: OpenAI was a few dozen people in 2016–17; language modelling became a formal effort around GPT-1 (June 2018); RL and robotics dominated early. A3C on Pong in about a million frames is in the range reported for actor-critic methods of that era. The Schulman anecdote is unverifiable.
09:45LeCun's cake: unsupervised learning is the cake, supervised the frosting, RL the cherry.
LeCun's NIPS 2016 keynote slide (widely reproduced; not re-fetched here).
09:46He built gradient checkpointing with Tim Salimans at OpenAI because PixelCNN didn't fit in memory; it became a TensorFlow package.
OpenAI released the package in January 2018; the repo credits Salimans and Bulatov, cites Chen et al. (2016) "sublinear memory cost", and reports fitting >10× larger feed-forward models for ~20% more compute. The repo now lives under Bulatov's cybertronai GitHub org — the same org that hosts the Sutro projects.
09:46People later found checkpointing also sped up training, because of the memory wall.
Plausible and there's a famous analogue: FlashAttention (Dao et al., 2022) deliberately recomputes attention in the backward pass to avoid writing the attention matrix to HBM, and ends up faster — the canonical "compute is cheaper than commute" result. Recomputation also speeds training indirectly by enabling larger batches. Direct evidence that vanilla checkpointing is faster in isolation is thinner; NVIDIA's selective-recomputation paper (2022) reports small overheads rather than speedups.
09:47Save a checkpoint every 10 layers of 100 and recompute — you store ~28 Jacobians instead of 100.
That's the O(√n) scheme from Chen et al. (2016): √100 = 10 checkpoints plus up to 10 recomputed activations live at once; "28" is an illustrative constant.
09:18In the last month Google released a result dropping the matrix-multiplication exponent — "I think 2.3".
Posted 17 August 2026: "Improving the matrix multiplication exponent with modern optimization and AlphaEvolve" (Google DeepMind with Alman, Vassilevska Williams and Zhou — the holders of the previous record). New bound ω < 2.371177, down from 2.371339, with AlphaEvolve used to refine the optimiser (summary). "2.3" is a loose rounding of 2.37; the improvement is 1.6 × 10⁻⁴ and purely asymptotic — a galactic algorithm, which reinforces his point that ω-chasing doesn't move practice.
09:18Naive matmul is about n³ (2n³ flops); Strassen is about n^2.8.
Strassen (1969): log₂7 ≈ 2.807. Correct.
09:18There is still no generic best algorithm for matrix multiplication.
ω is unknown (conjectured 2, best bound 2.371177 as of last month).
09:16Agents made algorithmic discovery reachable this year; discoveries at the frontier tend to be simultaneous.
"This year" is generous: AlphaEvolve (May 2025) improved the best-known solution on 20% of 50+ open problems and found a 48-multiplication 4×4 complex matmul, and last month's ω bound extends that. Simultaneous discovery is Merton's "multiples" (1961) — Strassen-era matmul and backprop itself are examples.
09:17"Yad" implemented all of Hinton's papers using a couple of billion tokens; the repo, under Cybertron AI, has animations for each paper.
Confirmed: cybertronai/hinton-problems (53 Hinton experiments, 1–3 May 2026) and schmidhuber-problems (58 experiments, 6–8 May), pure NumPy, built by Yad Konrad with Claude Code agent teams; his write-up reports 49 and 41 wall-clock hours, and Mark Saroufim cited the build in his MLSys keynote (21 May 2026). Bulatov's LinkedIn post says under $200 in three days; the build notes note that cache reads dominate token counts, which is how "billions of tokens" and "under $200" can both be true. I couldn't confirm the animations detail.
09:13About eight years ago Cade Metz wrote a New York Times piece on the explosion of AI hardware startups.
"Big Bets on A.I. Open a New Frontier for Chip Start-Ups, Too", NYT, 14 January 2018 (syndicated copy): at least 45 startups, five with $100M+ raised, $1.5B of VC in 2017.
09:13Cerebras is still around; Groq was recently "reverse acquired" by Nvidia; SambaNova is still around; most of the rest didn't make it.
Groq: on 24 December 2025 Nvidia took a non-exclusive licence to Groq's inference technology for about $20B and hired founder Jonathan Ross, president Sunny Madra and key engineers, while GroqCloud continues under a new CEO — analysts call it a "reverse acqui-hire" (Constellation, Techstrong). SambaNova: Intel's $1.6B takeover talks stalled in January; it then raised a $350M Series E in February 2026 with Intel participating and shipped the SN50 (DCD, Reuters via Yahoo). "Most didn't make it" fits the class of 2018 (Wave Computing bankrupt 2020, Nervana shut by Intel 2020, Graphcore sold to SoftBank 2024, Mythic nearly folded 2022 — from background knowledge).
09:14Mitchell Nahmias's optical-interconnect startup didn't make it; when he arrived at OpenAI they'd found algorithmic tricks that removed the need for fast interconnect.
Nahmias co-founded Luminous Computing (2018; CTO; Princeton neuromorphic-photonics PhD), which raised a $105M Series A in 2022 with Bill Gates among investors, and had pivoted in 2019 from optical compute to optical interconnect between chips, racks and data centres (Next Platform, VentureBeat). The shutdown and his move to OpenAI appear only in forum chatter in my searches; no mainstream confirmation.
09:14DeepSeek got "crippled" H800s with H100 compute but reduced memory bandwidth, and engineered around it.
The H800 export variant keeps the H100's 80 GB HBM3 and its memory bandwidth; what was cut is chip-to-chip NVLink interconnect bandwidth (and FP64). DeepSeek-V3's paper cites ~160 GB/s effective NVLink versus 50 GB/s InfiniBand and built DualPipe and custom PTX-level communication to overlap the constrained interconnect (TechRadar, DeepSeek's open-infra index). The moral — software routed around a hardware limit — is right; the spec is not.
09:12The cost of software dropped at least 10×, maybe 100×, this year; hardware didn't.
Anecdotes point both ways. For: Yad's 111 paper reproductions in ~90 agent-hours; StrongDM's "software factory" rule of $1,000 of tokens per engineer-day. Against: METR's 2025 randomised trial found experienced open-source developers 19% slower with AI tools (background knowledge); and the token bills below suggest cost moved to inference rather than vanished. No study supports a clean 10–100× yet.
09:58Uber declared it used twelve months' worth of tokens in four months — and the app didn't get better.
In April 2026 Uber's CTO said the company had exhausted its full-year Claude Code budget in four months, with 84–95% of ~5,000 engineers active monthly and individual bills of $150–$2,000 (TechFlow, Stocktwits/Yahoo). Sam Altman has called "spent my 2026 budget in Q1" a meme, and Microsoft cut internal Claude Code licences in May. Whether the Uber app improved is his opinion; the outcome-metrics question he raises (PRs vs. things people care about) is the productive one.
10:00His essay on AI companies needing an "AI god" narrative to justify valuations (title garbled as "God Golden GPUs").
His Substack launched around February 2026 but its index didn't render for me and search didn't surface a post by that name. Ask for the link; the title as transcribed is probably wrong.
09:51Data-centre water use is small compared with agriculture — California's lawns and almonds.
Nationally, yes: LBNL puts direct US data-centre water consumption at 17.4 billion gallons in 2023, about 0.3% of public supply, and about 228 billion gallons once power-plant cooling is counted — roughly 2% of US consumptive water use (MOST, CRS/LBNL summary). Agriculture is ~80% of California's developed water use (background knowledge). The counterpoint is local: in The Dalles, Oregon, Google's water use grew 316% while the town grew 12% (state-level data), and LBNL projects direct use rising to 38–73 billion gallons by 2028.
09:52He'd heard data centres use about 5% of electricity.
US: 176 TWh, 4.4% of national electricity in 2023 (LBNL 2024). The LBNL 2025 update projects 11.8% by 2030 (range 9.5–15.3%), up from the earlier 6.7–12% by 2028. Globally it's closer to 1.5% (IEA 2024, background knowledge). So "5%" is right for the US today and will be badly out of date within four years.
09:52His own electricity bill hasn't gone up much, which suggests data centres' impact hasn't been huge.
A sample of one in California misses where the effect is. In PJM (13 mid-Atlantic and Midwest states, 65M people) capacity prices went from $28.92/MW-day (2024/25) to $329.17 (2026/27); the independent market monitor attributes 63% of the 2025/26 increase — $9.3B — to data centres (IEEFA). PJM retail prices rose ~49% in five years versus ~33% nationally (Industrial Info), and Q1 2026 wholesale costs were up 76% year on year (E&E News). The fairest counter-reading is E3's May 2026 review, which attributes only ~50% of the PJM rise to load growth. California's high rates are driven by wildfire and transmission costs, not data centres.
09:52Interconnection studies already throttle how much data-centre demand can come online.
Google says utilities quote four to ten years to interconnect, one as long as twelve (Network World); ~2,300 GW of generation and storage sit in US queues. PJM's December 2025 auction fell 6.6 GW short of its reliability target for 2027/28 (Introl).
09:53CoWoS packaging (attaching HBM next to the die) is the main bottleneck; only TSMC can do it; TPUs had to cut their allocation.
CoWoS-S and CoWoS-L are fully booked with 52–78-week lead times; Nvidia has ~60% of 2026 capacity (Silicon Analysts, summary of Morgan Stanley data). Google's 2026 TPU output was reportedly cut from ~4M to ~3M units because of CoWoS access, with a backlog of 3M+ unpackaged dies (BigGo/Fubon, Isaiah Research). Two corrections: CoWoS is TSMC's brand, but ASE/SPIL, Amkor and Intel Foundry offer comparable 2.5D packaging; and SemiAnalysis (April 2026) argues N3 front-end wafer capacity is now the binding constraint as CoWoS eases.
09:53In a couple of years chips stop being the bottleneck and grid capacity becomes it — so intelligence per joule is the lever.
Mainstream view now: Morgan Stanley (August 2026) estimates US AI data centres need 68 GW of new power by 2028 with only 30 GW covered — a 38 GW gap (summary); Satya Nadella has said the constraint is power, not compute; Gartner expects 40% of AI data centres to be power-constrained by 2027. It's a forecast, but a well-supported one — and it's the best "why now" argument for his research programme.
09:54Recursive self-improvement isn't a 1/(1−x) curve; humans have been doing it for 5,000 years; it just clears the current bottleneck faster.
A bottleneck-shifting model rather than a singularity model. Compatible with the compute- and power-constraint literature above; incompatible with the strong RSI view. Not empirically decidable today.
10:01Chess was "the Drosophila of AI"; we had chess superintelligence in the '90s and people were still unsatisfied.
McCarthy's phrase (1990); Deep Blue beat Kasparov in 1997. Background knowledge, not re-fetched.
10:02Formulate any single task — ARC-AGI, say — and it gets solved far better than humans.
ARC-AGI-1 is saturated (multiple systems above 85%). On ARC-AGI-2, average individual human testers score roughly 53–60%, and by mid-2026 frontier systems (GPT-5.x, Gemini 3 Deep Think) post 85–92% on the semi-private set (leaderboard snapshot, analysis of the human baseline). Caveats: the "human panel" (≥2 people) still scores 100%; the 2025 Kaggle prize under a $0.20/task cost cap topped out at 24%, and the 2026 prize is still open. So "much better than a human" holds only when cost is ignored — which is precisely Bulatov's intelligence-per-joule point.
10:04Collatz: even/halve, odd/triple-plus-one; it always seems to reach 1, and there's no way to know except simulating.
Correctly stated and still open. Verified computationally for every starting value below 2⁷¹ (Bařina, J. Supercomputing 2025). One partial shortcut exists: Tao (2019) proved almost all orbits reach almost-bounded values — but there is no decision procedure, so his point stands.
10:05Rule 110 is known to have no shortcut: to know the state after a quadrillion steps you must run a quadrillion steps.
Predicting t steps of Rule 110 is P-complete (Neary & Woods, ICALP 2006), which rules out a fast parallel shortcut unless P = NC. That is believed but unproven, so "known" is slightly strong; "no shortcut is expected" is exact.
10:03Some problems — the three-body problem to enough accuracy — need more energy than the Sun provides; light speed and energy-per-operation are fundamental limits.
Chaotic dynamics make error grow exponentially, so precision costs are unbounded; Landauer's bound, the Margolus–Levitin bound and Bremermann's limit are the standard physical ceilings (background knowledge).
10:10There's a film, "Good Luck Have Fun Don't Die", about people sucked into short-form video.
Gore Verbinski's Good Luck, Have Fun, Don't Die (Sam Rockwell; US release 13 February 2026) is a time-travel comedy about stopping a rogue AI, framed by critics as a satire of AI and technology addiction (RogerEbert.com, Deep Focus). His plot summary is loose but the theme is right.
10:10Ten years ago the NeurIPS conference was held in a casino in Nevada, where people were pulling slot levers.
NIPS 2012 and 2013 were held at Harrah's and Harveys casinos in Lake Tahoe, Nevada — thirteen years ago (background knowledge).
10:12Two reward pathways: dopamine gives anticipation, opioids give satisfaction; short-form video feeds only the first, so you end up drained.
A fair lay summary of Berridge and Robinson's "wanting versus liking" work: dopamine mediates incentive salience (wanting), while opioid hedonic hotspots mediate liking. Whether short-form video specifically fails to deliver "liking" is a hypothesis, not an established finding (background knowledge).
10:11TikTok has a billion users.
TikTok has reported over a billion monthly users since 2021 (background knowledge).
10:08In an economy where everyone has enough to eat, engagement is the last currency; AI is another lever for capturing attention, and the worst case is an AI that optimises pixels directly for engagement.
A restatement of the attention-economy thesis (Tim Wu's The Attention Merchants; Herbert Simon's 1971 line about information consuming attention). His "people stuck in Claude Code feeding it tokens" version of the same failure mode is echoed by the "tokenmaxxing" coinage in the Uber coverage above.
09:08Google Brain, early OpenAI, Meta, Together AI; led the first production deployment of deep learning at Google; wrote gradient checkpointing at OpenAI; runs the weekly Sutro Group at South Park Commons.
Employers match his X bio (South Park Commons; early OpenAI, Google Brain, Meta) and an AI Council profile listing him as principal researcher at Together AI, which also carries the "first production deployment" line and adds that he hired Ian Goodfellow as his intern. The deployment is the Street View house-number transcription system (Goodfellow, Bulatov, Ibarz, Arnoud, Shet, arXiv 1312.6082, Dec 2013), which read close to 100 million street numbers at human accuracy (MIT Tech Review). "First" is his own framing — Google's Android speech recogniser shipped a deep net in 2012 — so "one of the first" is safer. The reading group was announced in his pinned post of 23 December 2025; the Sutro repos live at github.com/cybertronai. Born 1980 in the USSR: consistent with his public profile, not independently checkable.
If you wanted to make Bulatov's case for him with the best evidence available today, these are the load-bearing pieces:
Who is working the same problem, grouped by which part of his argument they touch.
| Who | What | Link |
|---|---|---|
| Bill Dally (NVIDIA / Stanford) | Parallel Explicit Communication Model; the energy numbers everyone quotes | CACM Point, Sept 2022 |
| ETH Zurich SPCL (Hoefler, Ben-Nun, Gianinazzi) | "The spatial computer": 2D-grid energy model with matching bounds for matmul, sorting, selection | arXiv 2205.04934 |
| Demmel et al. (UC Berkeley); Hong & Kung lineage | Communication-avoiding algorithms and I/O lower bounds — the pre-history of "compute to commute" | Background literature |
| Sutro Group (South Park Commons) | Data-movement-distance (DMD) metric, sparse-parity and MNIST energy benchmarks, agent-run experiments | github.com/cybertronai |
| Who | What | Link |
|---|---|---|
| Moonshot AI (Kimi team) | Attention Residuals: attention over depth as a drop-in for residual sums | GitHub |
| Hinton; FF successors (Cascaded-Forward, Mono-Forward, Hyperspherical FF) | Forward-only, layer-local training motivated by energy and biology | HFF (May 2026), survey |
| Predictive-coding community (Salvatori, van Zwol et al.) | Energy-based local learning; JAX frameworks; matching backprop on CIFAR/Tiny-ImageNet | PMC 2025, May 2026 |
| Forward-only evaluation work | Rigorous energy-versus-accuracy comparisons of BP-free algorithms | arXiv 2511.01061 |
| Feedback alignment, DNI, forward gradients, PEPITA | Earlier lines removing weight transport or the backward pass (Lillicrap; Jaderberg; Baydin; Dellaferrera & Kreiman) | Background literature |
| Who | What | Link |
|---|---|---|
| ML.ENERGY Initiative (U. Michigan; Chung, Chowdhury) | Zeus measurement library, Perseus training optimiser, ML.ENERGY inference leaderboard | Benchmark paper, Zeus |
| MLCommons MLPerf Power | Wall-outlet power measurement for training and inference submissions | See 2026 protocol survey |
| Hugging Face AI Energy Score | Fixed-batch energy ratings across model tasks | Same survey |
| Who | What | Link |
|---|---|---|
| Google DeepMind (AlphaEvolve) with Alman, Vassilevska Williams, Zhou | ω < 2.371177; 4×4 complex matmul in 48 multiplications; 23% Gemini kernel speedup | Aug 2026 paper, blog |
| Yad Konrad (Sutro) | Agent-team reproductions of 53 Hinton and 58 Schmidhuber experiments in pure NumPy | write-up |
| OpenEvolve, Sakana AI (AI Scientist, ShinkaEvolve), FunSearch | Open and commercial evolutionary code-search systems | Background knowledge |
| Who | What | Link |
|---|---|---|
| Groq (tech now licensed to Nvidia) | LPU with hundreds of MB of SRAM as primary weight storage | Deal coverage |
| Cerebras, SambaNova | Wafer-scale and reconfigurable-dataflow accelerators, both still independent | SambaNova 2026 |
| Lightmatter, Ayar Labs, Celestial AI | Optical interconnect — the bet Luminous made too early | Background knowledge |
| Who | What | Link |
|---|---|---|
| LBNL (Shehabi, Smith et al.) | US data-centre energy and water reports; 2030 projections | 2025 update |
| PJM's market monitor (Monitoring Analytics); IEEFA; E3 | Attribution of capacity-price increases to data centres — and the dissent | IEEFA, E3 |
| SemiAnalysis; Morgan Stanley; Epoch AI | Chip-supply and power-gap forecasts | SemiAnalysis |