Sutro #30 — MNIST Unchained
The group has spent thirty sessions on toy problems — 4×4 matmul, 16×16 matmul, sparse parity — scored by how far bytes move rather than how many operations run. Session 30 picked the next target: MNIST, scored on the same physical grid, with the twist that the contestant designs the language, the memory layout and the scoring method.
Why memory movement, not arithmetic
The premise, after Bill Dally's Communication of the ACM argument for a new model of computation: transistors shrink in two dimensions, wires in one, so over four decades communication overtook arithmetic as the dominant cost. An algorithm is efficient when it does a lot of compute per micron of transport — when it works on data that is already nearby.
Deep learning's core learning algorithm dates to the 1980s and does not have that property. Backpropagation needs every activation fetched back for a single gradient update; with a hundred layers that is a hundred fetches per add. The group's premise is that a learning algorithm with a small memory footprint exists and has not been looked for, because nobody is optimising for joules.
Where the existing problems stand
| Problem | Start | Best now | Factor |
|---|---|---|---|
| 4×4 matmul | 1,316 | 675 | 1.95× |
| 16×16 matmul | 66,300 | 63,819 | 1.04× |
| Sparse parity, 20% band | 17,331,683 | 86,753 | 200× |
| Sparse parity, 100% band | 43,325,468 | 392,666 | 110× |
The pattern worth noting: the bands that tolerate error fell furthest. Allowing a non-zero error rate opens far more room than exact solutions do — which is the whole reason accuracy targets, rather than exactness, are the right knob for a learning benchmark.
Whether matmul is saturating or the current generation of agents is saturating is genuinely open. Two 16×16 records landed the same afternoon as the meeting, so the honest answer is: not yet.
The MNIST design
Task
In: the training set, the training labels, and the unlabelled test set. Out: the test labels. One pass, eventually one kernel launch. There is deliberately no model — nothing in the specification requires a recognisable network, or weights, or a training phase. A submission is free to discover something that does not look like learning at all.
This is transduction: the whole test set arrives at once, which is a strictly easier problem than online prediction, and the design accepts that openly. The alternative — one test example at a time — is unrealistic on real hardware, where inputs are batched anyway.
Machine
- A read-only input tape, a 2D grid used as scratch, a write-only output tape.
- A minimal instruction set — arithmetic plus read and write. Deliberately no loop instruction: loops belong to whatever higher-level language a contestant brings, not to the machine, because a physically plausible processor does not have one.
- Streaming rather than resident data. Fifteen trillion tokens are not held in memory in real systems, so MNIST should not be either.
Calibration
Moving one byte one step costs 1 femtojoule; the step pitch is one micron; a transfer takes 0.5 picoseconds. Matched against Dally's published figures these give the same energy and land within 3× on time — dimensions consistent with a 7 nm process, which is also the A100's node, so GPU measurements are a valid sanity check. Two that hold: an 8×8 matmul models to 0.1 J and measures 0.1 J on current silicon; an 8K matmul is ~1012 operations and measures close to 1 J.
Two consequences fall out of the model. A read is a round trip — the address travels out and the byte travels back — while a write goes one way, so writes are cheaper in time than reads, though equal in energy. And area means peak scratch occupancy: using a thousand cells at once means a thousand square microns of silicon.
Scoring is part of the challenge
A naive MNIST pass is on the order of 1015 byte-moves. Nothing scores that move-by-move. Rather than legislate a language, the design makes language choice and scoring implementation the contestant's problem, and adds time-to-score as a fourth scored category alongside time, energy and area.
Submissions are therefore ranked on a Pareto front rather than one number, with the classical VLSI time–area trade-off (quadruple the area, halve the time) left deliberately unfixed. The goal is a diverse population of strategies — a knowledge base later agents can build on — not a single winner.
Open questions
- Multicore. The simplest extension promotes every grid cell to a processor. But a real systolic array feeds its operand matrices in from different edges of the array — the original formulation was hexagonal, with three matrices rolling together — so single-point I/O is the wrong primitive. Either every edge cell becomes a tape, or the first version stays single-processor.
- Resolution ladder. Staged sizes, roughly 100× apart, so the smallest rung is solvable without exotic language tricks and the largest is not. Resolution and set sizes are being fixed against measured accuracy ceilings rather than chosen by eye.
- An analog track. MNIST has been run through a SPICE circuit — image in, labels out, energy measured directly. A digital and an analog competition could run side by side. Against that: a presentation to this group three weeks ago argued from first principles that analog does not work, and that argument deserves an answer before the track opens.
Provenance, and why releases matter
The sharpest exchange of the evening was not about hardware. Much of the recent record progress comes from agents, and agents build on each other's output, so errors compound. Someone proposed auditing the repository for hallucinations. The objection landed immediately:
Suppose the specification said to use Manhattan distance, and an agent "fixed" it to Chebyshev. How would an audit know whether the fixer was right?
It cannot, from the text alone. What makes it decidable is provenance — knowing which decisions came from a human and which from an agent — and that is an argument for shipping frozen, versioned releases. A released competition is not just human output; it is human output that was looked at several times and declared settled. Agents may build on it. They may not silently correct it. Continuous sessions, where everything stays mutable, lose that property.
The generalisation offered: the scarce skill over the next five years is not writing the algorithm. It is designing the mechanism — the objective, the environment, the feedback loop — then noticing when the search gets stuck and telling it to add structure. The analogy is evolution: you would not design the brain, you would design natural selection, and then watch for the moment modularity needs to appear.
Blue ocean
The strategic case for the whole group, stated plainly: every well-funded team is chasing capability. Almost nobody is chasing joules. If energy, grid capacity or cost becomes the binding constraint — and there are reasons to think one of them will — then a measurable reduction in movement per unit of useful work is worth a great deal, and it will have been worked on by a very small number of people.
Next
- Challenge specification circulating this week: instruction set, tape format, scratch and area accounting, accuracy targets, scoring and anti-gaming rules.
- Staged resolutions, with the smallest rung set by measured accuracy ceilings rather than convenience.
- Separate Pareto winners for time, energy, area and time-to-score.
- A decision on single-processor versus spatial multicore before the model is frozen.
- Session 31 — same format. The point of these evenings is not the talk; it is to get people competing.