← Meetings
Meeting notes · 7 September 2026 · South Park Commons

Sutro #30 — MNIST Unchained

The group has spent thirty sessions on toy problems — 4×4 matmul, 16×16 matmul, sparse parity — scored by how far bytes move rather than how many operations run. Session 30 picked the next target: MNIST, scored on the same physical grid, with the twist that the contestant designs the language, the memory layout and the scoring method.

Why memory movement, not arithmetic

The premise, after Bill Dally's Communication of the ACM argument for a new model of computation: transistors shrink in two dimensions, wires in one, so over four decades communication overtook arithmetic as the dominant cost. An algorithm is efficient when it does a lot of compute per micron of transport — when it works on data that is already nearby.

Deep learning's core learning algorithm dates to the 1980s and does not have that property. Backpropagation needs every activation fetched back for a single gradient update; with a hundred layers that is a hundred fetches per add. The group's premise is that a learning algorithm with a small memory footprint exists and has not been looked for, because nobody is optimising for joules.

Sense of scale. A die is 20–30 mm across; a micron is a thousandth of a millimetre. If moving a byte ten microns costs about what an add costs, then reaching across the chip for one byte is worth a few thousand adds. That ratio is the whole motivation.

Where the existing problems stand

ProblemStartBest nowFactor
4×4 matmul1,3166751.95×
16×16 matmul66,30063,8191.04×
Sparse parity, 20% band17,331,68386,753200×
Sparse parity, 100% band43,325,468392,666110×

The pattern worth noting: the bands that tolerate error fell furthest. Allowing a non-zero error rate opens far more room than exact solutions do — which is the whole reason accuracy targets, rather than exactness, are the right knob for a learning benchmark.

Whether matmul is saturating or the current generation of agents is saturating is genuinely open. Two 16×16 records landed the same afternoon as the meeting, so the honest answer is: not yet.

The MNIST design

Task

In: the training set, the training labels, and the unlabelled test set. Out: the test labels. One pass, eventually one kernel launch. There is deliberately no model — nothing in the specification requires a recognisable network, or weights, or a training phase. A submission is free to discover something that does not look like learning at all.

This is transduction: the whole test set arrives at once, which is a strictly easier problem than online prediction, and the design accepts that openly. The alternative — one test example at a time — is unrealistic on real hardware, where inputs are batched anyway.

Machine

Calibration

Moving one byte one step costs 1 femtojoule; the step pitch is one micron; a transfer takes 0.5 picoseconds. Matched against Dally's published figures these give the same energy and land within 3× on time — dimensions consistent with a 7 nm process, which is also the A100's node, so GPU measurements are a valid sanity check. Two that hold: an 8×8 matmul models to 0.1 J and measures 0.1 J on current silicon; an 8K matmul is ~1012 operations and measures close to 1 J.

Two consequences fall out of the model. A read is a round trip — the address travels out and the byte travels back — while a write goes one way, so writes are cheaper in time than reads, though equal in energy. And area means peak scratch occupancy: using a thousand cells at once means a thousand square microns of silicon.

Scoring is part of the challenge

A naive MNIST pass is on the order of 1015 byte-moves. Nothing scores that move-by-move. Rather than legislate a language, the design makes language choice and scoring implementation the contestant's problem, and adds time-to-score as a fourth scored category alongside time, energy and area.

Submissions are therefore ranked on a Pareto front rather than one number, with the classical VLSI time–area trade-off (quadruple the area, halve the time) left deliberately unfixed. The goal is a diverse population of strategies — a knowledge base later agents can build on — not a single winner.

Open questions

Provenance, and why releases matter

The sharpest exchange of the evening was not about hardware. Much of the recent record progress comes from agents, and agents build on each other's output, so errors compound. Someone proposed auditing the repository for hallucinations. The objection landed immediately:

Suppose the specification said to use Manhattan distance, and an agent "fixed" it to Chebyshev. How would an audit know whether the fixer was right?

It cannot, from the text alone. What makes it decidable is provenance — knowing which decisions came from a human and which from an agent — and that is an argument for shipping frozen, versioned releases. A released competition is not just human output; it is human output that was looked at several times and declared settled. Agents may build on it. They may not silently correct it. Continuous sessions, where everything stays mutable, lose that property.

The generalisation offered: the scarce skill over the next five years is not writing the algorithm. It is designing the mechanism — the objective, the environment, the feedback loop — then noticing when the search gets stuck and telling it to add structure. The analogy is evolution: you would not design the brain, you would design natural selection, and then watch for the moment modularity needs to appear.

Blue ocean

The strategic case for the whole group, stated plainly: every well-funded team is chasing capability. Almost nobody is chasing joules. If energy, grid capacity or cost becomes the binding constraint — and there are reasons to think one of them will — then a measurable reduction in movement per unit of useful work is worth a great deal, and it will have been worked on by a very small number of people.

Next