The paper’s model from scratch: processors on a 2-D grid where a message costs its Manhattan distance in energy, with depth and wire-depth as the time proxies. Click through a message playground, the recursive-quadrant broadcast (and why an 8×8 square beats a 64×1 line), the Z-order parallel scan, and Cannon’s matrix multiplication with a live C = A·B check.
The classical backdrop and the counter-argument. The PRAM with its uniform shared memory (tree-sum, EREW/CREW/CRCW conflict lab, Hillis–Steele scan, pointer jumping, O(1) CRCW max), then Dally’s explicit-communication model — the simplified-dally-model worked example priced live, 15 vs 30 depending on placement — and a capstone that sums the same 16 numbers under PRAM, Dally, and the spatial computer.
For readers who want the math: energy, work, depth, and wire-depth of the spatial-computer algorithms (broadcast, reduce, scan, rank selection, sorting, cubic and Strassen matmul, plus FFT) re-derived under the single-core Dally model, set against the paper’s published parallel bounds — with derivation sketches so every row can be checked.
The source of truth for every bound quoted above. The tutorials cite its lemmas by number, so it reads well as a companion rather than a prerequisite.