Interlude: claims, evidence, and what to test next
Yaroslav Bulatov’s Interlude Show discussion · Tuesday 2026-09-08 (week 37)
Research checked on the same day. Prepared with Codex. Claims are paraphrased; proposed experiments are this report’s analysis.
Yaroslav’s strongest argument survives scrutiny: the cost of learning depends on where information moves, and the value of AI depends on what people gain from it. Both deserve better measurements. Several supporting claims, however, need correction. Small memory does not guarantee low total energy. Reinforcement learning has produced substantial demonstrated results. More than 99% MNIST accuracy does not require a convolutional network. Rule 110 does not prove that every distant prediction requires literal step-by-step simulation. The neuroscience supports a distinction between wanting and liking, but not a brain with exactly two reward channels.
The most useful next step is an experiment that keeps the task and achieved quality fixed, compares strong existing methods with proposed alternatives, and measures the entire cost. A faster toy solver, a lower memory footprint, and a more satisfying life are different outcomes. The report connects them without treating one as proof of another.
The corrections to make first
| Claim or implication | Assessment | More defensible version |
|---|---|---|
| Small peak memory guarantees small energy use | Incorrect without further bounds | Small working sets can lower the cost of each access; runtime and the number of accesses still matter. |
| H800 has reduced memory bandwidth | Wrong bandwidth category | Its key restriction is GPU-to-GPU NVLink bandwidth; that is distinct from HBM bandwidth. |
| Each gradient addition fetches all layers | Misleading account of backprop | The reverse pass reuses intermediate derivatives and local saved values to obtain gradients throughout the network. |
| Checkpointing saves Jacobians | Wrong usual implementation | It normally retains selected activations and recomputes others; full Jacobian matrices are not normally stored. |
| Google recently reduced the matrix exponent | Supported, with a precise correction | A recent preprint reports an upper bound below 2.371177, a small advance on 2.371339. |
| MNIST at 99% needs convolutions | False | Non-convolutional multilayer perceptrons have exceeded that accuracy; training conditions still matter. |
| Reinforcement learning has not worked in practice | False as a general statement | It has demonstrated successes; sample efficiency and its incremental contribution remain important questions. |
| A thousandfold toy improvement means a thousandfold energy saving | Not established | Report whether the improvement is in operations, a model score, runtime, or measured joules. |
| Rule 110 requires a quadrillion steps to predict a quadrillion steps | Not the theorem | General prediction is P-complete; this does not establish that literal lower bound or exclude easy inputs. |
| Dopamine is anticipation, opioids are satisfaction, and these are the two pathways | Useful intuition, overstated biology | Wanting and liking are partly separable; several interacting mechanisms contribute to reward and wellbeing. |
How to read the verdicts. “Supported” refers to the scoped proposition, not every implication attached to it. “Incorrect” means a counterexample or source contradicts the statement. “Not established” means the evidence reviewed does not justify it; that is not proof of the opposite. Forecasts and personal experiences are identified separately. Hosts’ introductory claims and incidental conversation are outside this review.
1. Physical costs: what memory does and does not explain
Small working sets help; they cannot alone bound total energy
Claim: low peak memory is sufficient, though not necessary, for low energy use. Assessment: the sufficiency claim is false without execution bounds. A program can keep ten bytes in registers and loop for a very long time. Its peak memory remains unchanged while its energy consumption keeps growing. Small storage can make local access possible; it does not specify how many accesses or instructions occur, how long the hardware runs, or the cost of reading inputs and writing outputs.
The “not necessary” half is useful: a large dataset can have a small, frequently reused working set and infrequent expensive transfers. The quantity to measure is the distribution of actual accesses, not only the largest allocation.
An illustrative accounting model—not a universal hardware law—is:
Energy ≈ arithmetic + memory accesses + communication + baseline power × elapsed time
The coefficients depend on device, precision, voltage, memory tier, and layout. Peak memory influences some terms but cannot replace the sum. Report joules to reach a declared quality target, with a time limit and a stated measurement boundary.
The ten-micron comparison is real; “arithmetic is free” is not universal
Assessment: supported as a scoped heuristic, overstated as a whole-workload percentage. Dally’s Stanford AHA keynote explicitly compares the energy of an addition with moving its operands roughly 10 micrometres. His accompanying numbers are illustrative technology and operation estimates. This is a legitimate intuition for why locality matters, not a constant valid for every datatype or machine. Dally’s keynote, especially slides 8 and 13.
A costly memory access can feed many arithmetic operations when data is reused. Consequently, a ratio between one addition and one transfer does not establish that arithmetic accounts for 1% of an entire workload. Setting its cost to zero is an approximation that needs a declared regime and sensitivity checks. Dally’s published model retains an arithmetic cost. PECM paper.
Why the memory wall is not simply “2D transistors versus 1D wires”
Assessment: useful motivation, inaccurate physical explanation. Wires have width and thickness as well as length; their cross-sections do shrink. The more useful distinction is local versus global communication. Local connections can shrink with devices, while chip-scale distances need not; shrinking cross-sections can also increase resistance. Ho, Mai, and Horowitz analyze these different scaling behaviors directly. The Future of Wires.
This preserves the substantive claim—communication can become relatively expensive—without deriving it from a dimensionality argument that omits resistance, capacitance, and physical length.
Dally’s model and Horowitz’s argument: correct references, corrected dates
Assessment: the ideas were recalled substantially correctly; the dates need repair. Horowitz’s Computing’s Energy Problem was an ISSCC 2014 presentation, not 2012. It argued for the likely persistence of CMOS-like technology and better matching of applications and hardware; it did not establish that CMOS would last forever. Its familiar energy table describes 45 nm technology, not measurements of current GPUs. Horowitz paper.
Dally’s On the Model of Computation: Point appeared in 2022, rather than six years before this discussion. PECM associates computation and memory with locations and a distance function. Two-dimensional Manhattan distance is one useful instance; the framework can also represent other communication costs. A related, separately authored paper, The Spatial Computer, develops explicit two-dimensional models and energy/depth tradeoffs. Dally; Gianinazzi and colleagues.
H800 and DeepSeek: interconnect is not HBM
Assessment: an important category error inside an otherwise strong co-design example. DeepSeek’s own hardware paper describes the H800 SXM’s NVLink bandwidth as 400 GB/s versus 900 GB/s for H100. That concerns communication between GPUs, not a generic reduction of bandwidth to local HBM. “Same compute” also needs precision and operation qualifications; FP64 differs.
DeepSeek adapted parallelism, routing, and communication overlap to the hardware. Its engineers nevertheless identify remaining communication and memory-throughput limitations, including in MoE inference. This supports software/hardware co-design, not the conclusion that bandwidth has ceased to be a bottleneck. DeepSeek’s ISCA hardware analysis, sections 4 and 6.
3D stacking, heat, and long-lived equipment
Assessment: real constraints, several overstated extrapolations. Three-dimensional integration can shorten connections and improve locality. Sixteen-high HBM has been announced and developed; sixteen is not a universal ceiling on all forms of stacking. Stacked memory and stacked active logic also have different power densities. Imec on 3D integration; SK hynix’s 16-layer announcement.
The arithmetic 16 × 100 W/cm² = 1,600 W/cm² is correct if sixteen full-area layers each dissipate that power simultaneously. It is an assumed scenario, not a measured property of every stack. “Billions of times wider than high” is also an overstatement: even 20 mm divided by an extremely thin 1 μm layer is 20,000. Thermal design matters, but these comparisons do not establish that 3D locality gains must be negligible.
ASML’s 2025 annual report supports the claim that approximately 95% of systems sold over the preceding thirty years remained active. The statistic concerns durable, often upgraded equipment. It does not establish that all those machines will remain active for another thirty years or that wafer manufacturing excludes major architectural change. ASML annual report.
The Groq example also needs precise wording: its announcement describes a non-exclusive technology licence and key staff moving to NVIDIA, while Groq remained independent. That is not by itself evidence of startup failure. A claim that “most” hardware startups failed needs a defined cohort and definition of failure. Groq announcement.
2. Learning algorithms: the strongest counterexamples
The new matrix-multiplication exponent is a genuine result
Assessment: supported, with numerical and practical qualifications. The preprint submitted on Monday 2026-08-17 (week 34) reports an improvement from ω < 2.371339 to ω < 2.371177, using modern optimization and AlphaEvolve. The recollection of a recent Google-linked advance is therefore correct. It is a small refinement of an existing bound near 2.37, not a first jump from Strassen’s exponent to “2.3.” Recent preprint.
The familiar classical count of about 2n³ arithmetic operations is sound; Strassen’s exponent is log₂7 ≈ 2.80735. An asymptotic arithmetic upper bound does not establish a practical speed or energy win at a particular matrix size. Conversely, it is too strong to say fast arithmetic algorithms are never used practically: implementations of Strassen have demonstrated useful speedups. Tiling and Strassen can be combined. Strassen with BLIS.
Follow-up: Compare equal shapes, precision, and numerical tolerances using strong tiled library kernels and feasible fast variants. The result should be a map of where each method wins, not a universal ranking extracted from its exponent.
Backprop does not fetch all layers for each scalar gradient addition
Assessment: misleading if read literally; reasonable only as a statement about the complete training step. Reverse-mode differentiation propagates an error signal backward and reuses it to calculate derivatives for many parameters. Once a dense layer’s error vector δ is available, its weight-gradient contribution has the form δaᵀ, where a is that layer’s input. Each multiplication or addition in this outer product does not independently reread every other layer.
The entire backward pass must obtain relevant saved or reconstructed values throughout the network. That is a genuine storage and communication problem. But ordinary training generally calculates gradients for all trainable layers and updates them together; it does not optimize only the bottom layer in isolation. Depth alone does not determine the arithmetic-to-traffic ratio. Autograd and vector-Jacobian products.
Better wording: Backprop efficiently reuses derivative information, but its physical implementation must retain, move, or reconstruct intermediate values. Compare those costs with the costs of achieving the same learning quality by alternatives.
Checkpointing: activations, not full Jacobians
Assessment: the recomputation mechanism is substantially right; the stored object is mislabeled. Standard implementations do not normally materialize a full Jacobian for every layer. They retain activations and other values needed to apply local vector-Jacobian products. Checkpointing saves selected values and reconstructs missing portions during the backward pass.
The public TensorFlow package credits Tim Salimans and Yaroslav Bulatov jointly, supporting Yaroslav’s implementation contribution. The associated deep-network memory analysis is by Chen and colleagues; checkpoint scheduling has an older automatic-differentiation literature, including Revolve. For an idealized chain, square-root checkpointing reduces activation storage from O(n) to O(√n) with extra forward computation. This concerns activations, not all parameter and optimizer memory. Implementation; Chen and colleagues; Revolve.
Recomputation can lower runtime and energy if it avoids more expensive traffic, but a smaller allocation is not proof of either. FlashAttention supplies a concrete example of reducing HBM traffic through tiling and recomputation while preserving exact attention and backprop. Other work improves performance by reducing unnecessary recomputation. The optimum depends on the workload. FlashAttention; selective recomputation.
Removing backprop, removing update dependencies, and saving energy are different goals
Claim: eliminating backprop requires changing the dense layered architecture. Assessment: a research hypothesis stated too strongly. MeZO fine-tunes existing language-model architectures using a zeroth-order method and forward passes. That is a counterexample to architectural change being logically necessary. It is not proof of superior from-scratch pretraining or universally lower energy. MeZO.
The downstream dependency that Yaroslav describes is closely related to update locking. Synthetic-gradient methods predict downstream error information so modules can update without waiting for the full backward calculation. Reversible networks attack another part of the problem by reconstructing earlier activations. These are distinct research directions with distinct accuracy and cost tradeoffs. Synthetic gradients; RevNets.
The Kimi recollection is also supported: Attention Residuals learns an attention-weighted combination of earlier layer outputs. But it retains backprop, and attending to many earlier outputs is not automatically a reduction of dependencies. Attention Residuals.
The observation about small expert batches and inference utilization in MoE is valid. Calling MoE universally terrible for inference is not: DeepSpeed-MoE reports configurations with substantially better quality-matched inference cost and speed than dense alternatives. Standard MoE primarily sparsifies expert computation within layers, rather than skipping all network depth. DeepSpeed-MoE.
Sparse parity: a good diagnostic, an overstated origin story
Assessment: the mathematical example is sound; the single-cause AI-winter story is misleading. Parity of selected binary inputs is their sum modulo two. A linear threshold classifier cannot represent parity involving two or more relevant bits over all input combinations. Minsky and Seymour Papert analyzed parity limitations in Perceptrons. MIT Press edition.
That does not establish that today’s hidden-sparse-subset benchmark caused a decade-long shutdown of AI. The contemporary record contains broader arguments about scaling, research programs, and funding. The label “first AI winter” is itself contested by historians. Lighthill’s survey; McCarthy’s response; historical reassessment.
The thousandfold improvements: name the metric and the task
Assessment: public support exists for large toy-problem improvements; a general thousandfold energy saving is not established. A public Sutro project log reports 178–1,203× operation-count improvements in selected small cases and other runtime gains. It also records corrections to an earlier movement tracker. This supports taking the algorithmic discoveries seriously while checking which numbers remain comparable. SutroYaro discoveries.
The current public sparse-parity task specifies 32 input bits, 5 relevant bits, and 18 examples, scores exact secret recovery, and uses a simplified Dally cost model. That score is not a direct meter reading. Because 18 is less than 32, the linear system is underdetermined: ordinary elimination alone does not uniquely identify the secret without exploiting sparsity or additional search. Current task specification.
Follow-up: Separate noiseless full-rank recovery, underdetermined sparse recovery, corrupted-label recovery, and active queries. Irrelevant input bits are not the same as label noise. Giving a method permission to query new labels changes the information available and must be declared. Research on sparse noisy parity.
MNIST: an accuracy target is not an architecture requirement
Assessment: “99% needs convolutions” is false. A 2010 paper reports 99.65% test accuracy from plain multilayer perceptrons, trained with many deformed images. Convolution is therefore not necessary. Augmentation still introduces useful image structure, and the result does not establish that this method is the most energy efficient. Cireșan and colleagues.
The proposed interface—give a system labeled training data and the exact unlabeled test inputs, then ask it to return their labels—is a legitimate transductive learning task. It differs from training a predictor for unknown future inputs. Compare total fitting-plus-labeling cost and state whether test-input access is allowed. Hidden final labels and separate development data are essential if results are to generalize beyond repeated tuning on public MNIST. Transductive learning formulation.
Two smaller corrections: history and batch averaging
Backprop became influential in neural-network learning in the 1980s, including Rumelhart, Hinton, and Williams’s 1986 paper. Reverse-mode differentiation has earlier roots: Linnainmaa’s Finnish thesis dates to 1970, not the 1960s, and the earlier literature did consider storage costs. The story is richer than an algorithm invented without concern for computing resources. Landmark neural-network paper; historical account; Linnainmaa’s paper.
For a fixed batch of B examples, g_sum = B × g_mean. Under plain SGD, reducing the learning rate by B when using the sum produces the same update. Thus averaging is not intrinsically a design mistake. Adaptive optimizers, clipping, regularization, changing batch sizes, and finite precision require a fuller comparison; changing only the reduction convention can change those dynamics. Large-batch scaling research.
The unspecified old-workstation-to-iPhone “million times” comparison is not verified here: device, precision, workload, and throughput measure must be named. Small experiments can motivate new algorithms without relying on that ratio.
Reinforcement learning: poor sample efficiency is not failure to work
Claim: reinforcement learning has not been made useful in practice; successes may occur despite it, with prior learning doing 99.9% of the work. Assessment: the broad dismissal is contradicted; the attribution question is valid.
AlphaGo Zero learned superhuman Go through self-play reinforcement learning without human game examples. It still had engineered networks, search, game rules, a simulator, and substantial computation. Those ingredients constrain what its success proves; they do not remove the demonstrated role of learning from rewards. Original study.
DeepSeek-R1-Zero is a particularly useful test of the narrower argument. Its reinforcement-learning stage did not require preliminary supervised fine-tuning, but it started from a pretrained language model. The distinction matters: this is not language and reasoning learned from random initialization solely through sparse rewards. Subsequent work finds important reasoning behaviors already present in the base model, while also showing improvements from changing the RL method. Neither result supplies a measured “99.9%” allocation of credit. DeepSeek-R1; critical analysis of R1-Zero-like training.
Yaroslav’s Pong recollection is personal testimony, not an independently verified benchmark. A3C expands to Asynchronous Advantage Actor-Critic. Its published results support discussing sample efficiency, compute efficiency, and reward design separately. A3C paper.
Better wording: Sparse rewards can be an inefficient source of information. The research question is how much incremental capability RL buys over a strong starting policy, at the same total budget. Run ablations against the same base model, matched inference-time sampling, supervised alternatives, and an unchanged policy; include unsuccessful training runs in the cost.
3. Software progress and the evidence of value
Software becoming 10–100 times cheaper
Assessment: plausible for selected tasks or a personal workflow; unsupported as a general cost estimate. Generating a prototype, completing a correct change in an established codebase, and maintaining deployed software are different tasks. The denominator must include review, debugging, operation, and failures.
The experimental literature is heterogeneous. A controlled GitHub Copilot study reported 55.8% faster completion of one JavaScript HTTP-server task. METR’s early-2025 study of experienced maintainers found tasks took 19% longer with the available AI tools. Neither is a reliable estimate for every developer using current agents. METR’s later update explicitly says selection effects make the size of newer productivity gains hard to estimate. Copilot experiment; METR original experiment; METR update.
Follow-up: Pick a repeated research task and count independently verified results per human hour and per dollar, including abandoned attempts. Yaroslav’s “why now” thesis becomes testable if agents make exploring locality-aware algorithms materially cheaper under those measures.
Uber’s spending and the unchanged-app argument
Assessment: the spending story has a basis; an unchanged interface does not establish zero benefit. Uber’s CTO reported exhausting the annual AI budget by April. Contemporary report. A spending budget and a fixed quantity of tokens are different things.
There is also a substantive update. In its engineering post of Thursday 2026-08-27 (week 35), Uber reports 9.4 times as many weekly agent requests between February and August, with total AI spending relatively stable since April. It describes tracking quality, incident-resolution time, and cost per completed task alongside usage. These are corporate self-reports, not an independent causal demonstration of rider benefit. Uber’s engineering account.
Yaroslav’s question is stronger than the implied answer. Reliability, fraud prevention, cancellations, accessibility, or driver support can improve without changing the visible app. Conversely, more merged code can coexist with little customer value. Ask for predeclared customer outcomes, a comparison group where practical, and total operating cost. Token consumption alone settles neither case.
More output does not automatically mean greater welfare
Assessment: a sound measurement principle, not evidence that AI has no social benefit. Infant mortality, educational attainment, health, and life satisfaction are legitimate ends. They are distant and affected by many causes; demanding an immediate change in them from each software project would also be a poor evaluation design.
Use a causal chain with intermediate checks: a verified technical change, a better service, an outcome people value, then its distribution and durability. For a transport product, this might be fewer failed trips or better access to medical appointments. For a research agent, it might be a reproducible discovery that changes a subsequent experiment. Those are proposed evaluation targets, not measured benefits asserted here.
4. Datacenters, water, electricity, and bottlenecks
Agriculture is a useful comparison, not a local-impact exemption
Claim: datacenter water use is small relative to agriculture, including California almonds. Assessment: the broad scale intuition is reasonable; the exact comparison was not established. California’s water agency attributes about 80% of developed water use to agriculture, or about 40% when the denominator includes environmental uses. These are different denominators, and neither directly gives an almond-to-datacenter ratio. California DWR.
Berkeley Lab estimated 66 billion litres of direct US datacenter water consumption in 2023, plus nearly 800 billion litres indirectly associated with electricity generation. Both estimates concern all datacenters, not AI alone. Consumption is not the same as withdrawal. A sound comparison holds the geography, year, and accounting boundary constant. Berkeley Lab report, printed pages 55–58.
The key inference is geographic: a sector can be small nationally and still matter to a stressed local watershed or municipal supply. An appropriate follow-up is a facility-level water balance, including source, seasonal stress, cooling design, and the electricity supply’s water footprint.
“About 5% of electricity” needs a country and year
Assessment: roughly compatible with a historical US estimate, not a global or current universal number. Berkeley Lab’s 2024 report estimated US datacenters at 176 TWh, or 4.4% of US electricity, in 2023. Its current research overview describes a later projection range of 9.5–15.3% by 2030, with an 11.8% central estimate. These future values are scenarios, not observations. Historical report; Berkeley Lab overview.
The IEA’s 2026 update estimates 485 TWh globally in 2025, increasing to 950 TWh in 2030, approximately 3% of global electricity in that forecast year. It projects faster growth for AI-focused facilities. Again, total datacenter demand and AI demand are not interchangeable. IEA update.
A personal electricity bill and connection studies cannot establish harmlessness
Assessment: the inference is unsupported. A household bill combines its own usage, tariff design, taxes, fuel costs, and regulatory decisions. It does not isolate the marginal cost of a new datacenter elsewhere, nor costs that have not yet reached customers.
Large-load studies and interconnection processes are real safeguards. They do not prove that all reliability and cost-allocation problems have been solved. FERC’s Thursday 2026-06-18 (week 25) action required six regional grid operators to justify or reform large-load integration rules. Its explanation specifically discusses preventing other customers from bearing costs when promised demand fails to arrive. FERC action; cost-allocation explanation.
Better wording: Grid connection is constrained and regulated, but effects on reliability, local prices, infrastructure investment, and who pays must still be measured.
CoWoS, HBM packaging, and a moving bottleneck
Assessment: advanced packaging is a real constraint; “only TSMC can do it” confuses a branded technology with a broader capability. CoWoS is TSMC’s family of advanced packaging technologies. Integrating logic and high-bandwidth memory is not uniquely possible at TSMC: Samsung describes its own I-Cube/H-Cube approaches. This does not make vendors interchangeable for a particular accelerator, volume, yield, or qualification. TSMC CoWoS; Samsung H-Cube.
The specific claim that TPU production allocations were reduced was not confirmed here from a manufacturer’s primary disclosure. An analyst’s revised forecast should not be recast as an announced production cut.
The shift toward power constraints is plausible and already visible in the IEA’s discussion of grid delays. But grid supply is not fixed forever: new generation, transmission, storage, location changes, and flexible scheduling also change available capacity. At a fixed power budget, improved useful work per joule matters greatly; it is not the only way the system can grow. IEA analysis.
5. What computational limits actually establish
Rule 110: universality is not a literal no-shortcuts theorem
Assessment: incorrect as stated. Cook proved Rule 110 can perform universal computation. Neary and Woods established P-completeness of a finite prediction problem with the time horizon supplied in unary. This gives conditional evidence against a very efficient general parallel predictor, assuming P ≠ NC. It is not an unconditional theorem that a particular distant state takes exactly as many sequential steps to predict as to evolve. Cook’s proof; Neary and Woods.
A direct counterexample to the blanket phrasing is an all-zero initial state: Rule 110 leaves it unchanged, so its quadrillion-step state is known immediately. Particular input families can be easy even when general prediction is hard. Unbounded undecidable questions, finite worst-case complexity, sequential runtime, and energy lower bounds are four different claims.
Better wording: Some prediction problems resist general shortcuts. That still leaves room for faster methods on structured instances—the very opportunity algorithm discovery tries to exploit.
Collatz: unresolved does not mean every orbit must be naively simulated
Assessment: the unresolved-conjecture point is supported; the no-shortcut wording is too strong. The universal claim is that repeatedly halving an even positive integer, or applying 3n + 1 to an odd one, eventually reaches 1. The research reviewed does not provide a proof for all positive integers. Tao proved a significant but narrower almost-all result. Tao’s paper.
There are accelerated verification methods, and simple inputs such as powers of two can be certified without simulating every step. A published verification milestone below 2⁷¹ is evidence about a finite range, not a proof of the conjecture and not a claim about the exact live computational frontier. Bařina’s verification research.
Three-body chaos and “more energy than the Sun”
Assessment: the numerical comparison is not established. Sensitive dependence can make long-horizon trajectory prediction require extraordinarily accurate initial conditions. Research on selected massive black-hole triples illustrates precision barriers to reversing their trajectories; it does not derive a universal solar-energy cost for three-body simulation. Boekholt, Portegies Zwart, and Valtonen.
The claim needs masses, initial conditions, duration, output of interest, error tolerance, and a computational energy model. The Sun comparison also needs a time interval: energy and power have different units. More computation cannot recover initial information that was never measured. Yet some special trajectories are regular, and useful statistical predictions can survive when exact trajectories become inaccessible.
Follow-up: Compare the cost of predicting a precise trajectory with the cost of answering a decision-relevant question, such as which body escapes. Intelligence per joule depends on selecting the useful question as well as solving it efficiently.
Finite physics does not provide a calibrated ceiling on intelligence
Assessment: the physical constraints are real; the intelligence conclusion is an inference. Quantum mechanics, relativity, and thermodynamics constrain physical computation and communication. Lloyd’s analysis makes such bounds explicit for idealized devices. Ultimate physical limits to computation.
A signal failing to cross a chip in one clock cycle is a communication-latency constraint. It does not by itself bound all knowledge a system can store, access over many cycles, or distribute across machines. Nor does it establish that a particular modular architecture is mandatory, or that machines cannot broadly exceed human capability. Those conclusions need an additional theory connecting physical resources to performance.
The distinction between superhuman task skill and general capability is useful. But the suggestion that any chosen task will eventually be mastered is a forecast. ARC itself is a family of versioned tests: ARC-AGI-3 uses interactive environments and comparison with human action efficiency. Results from earlier static versions cannot be transferred to it. Chollet’s evaluation framework; ARC-AGI-3 specification.
6. Attention, pleasure, and satisfaction
Wanting and liking are separable; “two pathways” is too simple
Assessment: the distinction is supported, the exclusive biochemical mapping is not. Experiments in mice found that elevated dopamine could increase pursuit of sweet rewards without increasing hedonic taste reactions. Experiments in rats found that opioid stimulation in a small reward hotspot increased positive taste reactions; stimulation elsewhere could increase eating without the same pleasure effect. These are circuit-specific animal experiments, not direct measurements of human life satisfaction. Dopamine experiment; opioid-hotspot experiment.
Dopamine also participates in reward learning and prediction-error signaling. It is not simply a conscious anticipation meter. Opioids are not exclusively a satisfaction channel. Reward prediction-error research.
Better wording: The urge to continue and the pleasure or value obtained can diverge. That behavioral distinction is useful without claiming to read someone’s neurotransmitters.
Does rapid scrolling leave people less satisfied?
Assessment: supported in some experimental settings, not universally or through the claimed exclusive pathway. Tam and Inzlicht’s seven experiments included 1,223 participants. Key experiments found that switching among or within videos could increase boredom and reduce satisfaction. Results varied across extensions to different material and samples. They do not demonstrate permanent damage or establish that scrolling engages dopamine without opioids. Original study.
Another small experiment found that a TikTok condition impaired remembering a planned action in the tested task. That is a bounded prospective-memory result, not proof of general long-term memory loss. Chiossi and colleagues.
The broader engagement-versus-welfare concern also has causal evidence: a four-week Facebook-deactivation experiment increased a subjective-wellbeing index by 0.09 standard deviations, while reducing factual news knowledge. The tradeoff matters; these findings do not imply all online engagement is harmful. Allcott and colleagues.
AI that captures attention, including coding agents
Assessment: a plausible risk scenario, not an established inevitable outcome. The possibility of systems optimizing human attention more aggressively deserves testing. The imagined irresistible pixel pattern is speculation. Observing someone operate many agents does not diagnose addiction or show that the work is useless.
The revised MIT/OpenAI four-week study randomized 981 people across chatbot modalities and conversation topics. Most primary effects of those randomized assignments were not significant. Heavier voluntary use was associated with worse psychosocial outcomes, but duration was not randomized, and there was no no-AI control. The study does not establish that heavier use caused harm, and it did not test coding agents. Revised study.
Yaroslav’s proposed satisfaction tracking is a reasonable personal experiment, not a guaranteed defense against attention capture. Immediate pleasure, progress toward a chosen goal, the freedom to stop, and next-day endorsement can diverge. Measure all four before deciding whether more engagement is helping.
7. Connections that suggest better research
Agents may resolve an old objection to locality-aware programming
Dally’s published proposal had a direct counterargument: Uzi Vishkin argued that forcing programmers to reason about detailed locality imposes too much development burden. That is unusually close to Yaroslav’s “why now” thesis. If agents reduce the cost of searching layouts, schedules, and implementations, an old tradeoff may change. This is an inference connecting the arguments, not a finding either author already established about modern agents. Vishkin’s counterpoint; his department’s account.
Test: Compare verified kernel quality, engineering time, search cost, and measured execution energy for the same problems. A success would show that locality optimization became cheaper to obtain and worthwhile to deploy. It need not prove the entire software stack has become a hundred times cheaper.
Checkpointing connects the thesis to a mature optimization literature
The store-versus-recompute choice is already formalized in automatic differentiation and tensor rematerialization. Checkmate optimizes recomputation schedules using profiled costs. FlashAttention shows how changing execution can preserve the mathematical operation while changing its physical cost. These are strong baselines and research partners for the thesis, not obstacles to it. Checkmate; FlashAttention.
Test: Separate gains due to a better execution schedule from gains due to a different learning rule. If optimized backprop wins at equal quality, that is useful progress toward the energy goal. If a local learning method wins, report where and why, including the tasks where it loses.
Discovering a cheaper solver also has a cost
An agent may spend substantial energy searching for an implementation that is cheap to run. The relevant comparison depends on reuse. With search cost S, old per-run energy E_old, new per-run energy E_new, and R future uses, the proposed total-cost comparison is:
New total: S + R × E_new
Old total: R × E_old
Break-even: R > S / (E_old − E_new), if E_old > E_new
This is an accounting identity under the stated boundary, not a new experimental result. If the baseline also required a search, compare incremental search costs. If energy data are unavailable, report tokens, dollars, and human time separately instead of silently converting them into joules. Include failed candidates and evaluation costs. This makes one-off discoveries and widely deployed kernels comparable without pretending they have the same economics.
The same measurement error recurs at several scales
| Convenient proxy | What it leaves unresolved |
|---|---|
| Fewer arithmetic operations | Data reuse, communication, numerical quality |
| Smaller peak memory | Access count, runtime, input/output, total energy |
| Fewer joules per task | How many tasks get run and aggregate demand |
| More code or merged changes | Reliability and user benefit |
| More time interacting | Satisfaction, agency, and later endorsement |
This is the report’s main synthesis: keep the objective separate from the metric used to search for it. An improvement in a proxy is valuable evidence only after checking the next link. Efficiency can make more use affordable; it need not reduce total consumption. Engagement can reflect worthwhile work; it need not establish wellbeing.
8. A concrete follow-up agenda
These are proposed next actions; no experiment or outreach below has been performed.
1. Prepare a short correction sheet for the hosts
Prioritize the items that alter the mechanism: H800’s NVLink/HBM distinction; activations versus Jacobians; the per-add account of backprop; MNIST’s non-convolutional counterexample; RL’s demonstrated contribution; Rule 110’s actual theorem; and the qualified wanting/liking account. Preserve the recent matrix result with its exact bound. The dates of the Horowitz and Dally references can be corrected in accompanying notes.
Deliverable: a claim, corrected sentence, and one primary source for each item. The opening table provides the starting point. Personal forecasts should remain clearly labeled forecasts.
2. Run an energy-and-quality comparison that can falsify the thesis
Choose one small model and one accuracy target. Compare an optimized backprop baseline, selective checkpointing, and one well-specified alternative. Hold data access and evaluation fixed; report joules to target, wall time, peak memory, traffic by memory tier where measurable, numerical precision, and run-to-run variation. Include preprocessing, tuning, and evaluation in separately stated totals.
Success criterion: a repeatable energy improvement at comparable quality under the declared deployment conditions. Repeat on a second device or batch regime to reveal whether the result depends on one architecture. A lower proxy score alone does not pass.
3. Make the parity and MNIST tasks scientifically discriminating
For parity, publish n, k, sample count, label-noise process, allowed queries, hidden seeds, and whether success means prediction or exact secret recovery. Include strong algebraic/search baselines. Keep the underdetermined setting distinct from full-rank elimination.
For MNIST, permit competing architectures and state augmentation and test-input permissions. Separate inductive and transductive results. Final evaluation should use protected labels and a clearly separate development set. Require enough repeated runs to expose variance rather than selecting the best attempt.
Success criterion: a gain that survives a strong baseline and genuinely held-out cases under identical information access.
4. Test the Dally–Vishkin question directly
Give agents locality-sensitive problems with known-correct baselines. Record the human effort needed to specify, debug, verify, and maintain the resulting implementations. Sweep memory-access and arithmetic prices in the cost model, then compare predicted rankings with measured hardware rankings.
Success criterion: agents reduce the effort required to obtain robust locality gains. If rankings reverse when the arithmetic coefficient stops being zero, the benchmark has exposed an important limitation rather than a failed project.
5. Make the backprop prediction observable
Recover the original public statement and its starting date before assigning a deadline. Then define “gone”: absent from leading-model pretraining, no longer dominant in training expenditure, displaced for a specified class of tasks, or simply substantially changed in implementation. A model that uses depth attention but still differentiates backward does not automatically satisfy a replacement forecast.
Success criterion: two independent readers can decide whether the prediction passed using the same published rule. Track energy improvement independently so a useful technical result is not lost because a rhetorical prediction fails.
6. Evaluate AI work by progress and next-day endorsement
For a small personal pilot, compare matched work blocks with one active agent versus several. Predeclare the task and stopping time. Measure independently verified progress, total human time, interruptions, urge to continue, immediate satisfaction, and whether the result still feels worthwhile the next day. Treat a short pilot as exploratory, not a diagnosis or a study of brain chemistry.
Success criterion: more useful completed work without a persistent loss of satisfaction, relationships, or the ability to stop. This directly tests the optimistic and pessimistic futures discussed in the episode on a scale where Yaroslav can make an informed choice.
Scope and remaining uncertainty
This is a source-based review, not a reproduction of every experiment. It distinguishes published results, manufacturer or company self-reports, mathematical counterexamples, personal experiences, and forecasts. Recent preprints can change. A source supporting one component of an argument does not validate all its implications.
Yaroslav’s personal account of his work and the exact history of his five-year prediction are not independently established here. A specific hardware-startup survival rate, a manufacturer-confirmed TPU quota reduction, and a general thousandfold measured-energy gain were not established. No new GPU energy experiments or human-subject studies were performed for this report.
The research direction becomes stronger when its standards are explicit: less energy for a verified capability, less effort for a useful result, and technology that improves total life satisfaction.