Follow-ups · South Park Commons · Thu 10 Sep 2026

Where the hardware went, and how to find it again

Follow-ups from our conversation at South Park Commons, Thursday 10 September 2026.

You asked whether there is a text where you could learn algorithms as a beginner, but written in the seventies or eighties, back when the hardware was still in the room. There is, and the reason it disappeared is more interesting than the reading list.

That tradition did not die. It split into three literatures with different names, which is why it looks like it vanished if you go looking under "algorithms": VLSI complexity treats chip area and wire length as first-class resources, I/O complexity makes data movement the cost function, and communication-avoiding linear algebra proves lower bounds on how much data a computation must move. One free textbook carries all three as ordinary chapters, and it is the first item under "Why the algorithms ignore the hardware" below.

Everything here is free unless marked otherwise. I have put the load-bearing numbers into the sentences so you can decide what deserves an evening. I have also checked the things I told you, and corrected three of them at the end, because two were wrong in ways that would have mattered.


Start here

Five items. These are the ones that reframe the field rather than add to it.

On the Model of Computation: Point William J. Dally, CACM 65(9):30–32, 2022. Three pages, no computer science assumed, and the source of every constant I was quoting at lunch: a 32-bit add costs 20 fJ and 150 ps, moving its two operands one millimetre costs 1.9 pJ and 400 ps, and fetching them from DRAM costs 64,000 times the add. Read this first and the rest of the list becomes optional.

Computing's Energy Problem (and what we can do about it) Mark Horowitz, ISSCC 2014 plenary, pp. 10–14. Slide 33 of the companion deck, "Rough Energy Numbers (45nm)", is the most reproduced table in the field, and the CV² arguments around it are in your native vocabulary, so you can audit how the numbers were obtained rather than trusting them. Also contains the passage I mangled at lunch: his catch-22 on why a radically new device technology can never be funded, because "the larger the investment, the lower the risk the investors are willing to tolerate."

I/O Complexity: The Red-Blue Pebble Game Jia-Wei Hong and H. T. Kung, 1981. Eight pages that redefine the cost of an algorithm as the number of transfers between a small fast memory and a large slow one. This is the single idea that makes modern kernel design comprehensible, and it is forty-five years old.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré, 2022. The thing you had not heard of. Identical mathematics, identical output, several times faster, purely by respecting the memory hierarchy. It is a Hong and Kung argument applied to HBM and SRAM, not a numerics result, which is exactly why it belongs in this list rather than an ML one.

The role of optics in computing Rodney S. Tucker, Nature Photonics, 2010. Two pages by an optics person, for optics people, on why light keeps losing to electrons at everything except moving bits from A to B. The most useful inoculation available before spending a year on any "just use light" proposal, including your own.


Why the algorithms ignore the hardware

Models of Computation: Exploring the Power of Computing John E. Savage, 1998, free from the author since 2008. The literal answer to your question: a complete theory-of-computation textbook in which space-time tradeoffs, memory-hierarchy tradeoffs and VLSI models are chapters 10, 11 and 12 rather than omissions.

A Complexity Theory for VLSI C. D. Thompson, CMU PhD thesis, 1980. The founding document. Chip area and time are the resources, the result is a tradeoff curve (AT² ≥ cN² log² N for an N-point DFT), and the proofs are information-flow-across-a-cut arguments you will recognise on sight.

Introduction to VLSI Systems Carver Mead and Lynn Conway, 1980. The book that taught a generation of logic designers to think in microns and wire delays at the same time, and the reason the eighties produced a decade of algorithms papers where geometry was part of the answer.

Why Systolic Architectures? H. T. Kung, 1982. A hardware engineer arguing that the way to beat the memory bottleneck is to pump data through a regular array of cells. Direct ancestor of Google's TPU and of most accelerators you will meet.

The Input/Output Complexity of Sorting and Related Problems Alok Aggarwal and Jeffrey Scott Vitter, 1988. Turns Hong and Kung into a clean parameterised model with matching upper and lower bounds. The standard vocabulary for talking about any memory hierarchy.

Can Programming Be Liberated from the von Neumann Style? John Backus, Turing Award lecture, 1978. The canonical statement that the single wire between processor and memory is an intellectual constraint and not merely an engineering one. Explains why dataflow, streams and systolic arrays get reinvented every fifteen years.

Hitting the Memory Wall: Implications of the Obvious William A. Wulf and Sally A. McKee, 1995. Five pages of back-of-the-envelope arithmetic showing that two divergent exponentials make memory latency dominate everything eventually. The field has been living inside this argument for thirty years.

A New Golden Age for Computer Architecture John L. Hennessy and David A. Patterson, Turing lecture, 2019. Thirteen pages covering the gap you named directly: what x86, ARM and RISC-V are, why the difference matters, and why the end of Dennard scaling forced the field into domain-specific hardware. If you read one thing about computer architecture as a field, read this.

Computer Architecture: A Quantitative Approach Hennessy, Patterson and Kozyrakis, 7th edition, Dec 2025. Not free. The reference text. For you the order is Appendix B, then Chapter 2 on memory hierarchy, then Chapter 7 on domain-specific architectures. Skip the instruction-level-parallelism chapters on a first pass.

What Is Fast Matrix Multiplication? Nicholas J. Higham, 2022. Two pages answering the question you asked about why the asymptotically faster algorithm is not the one anyone runs. The reason is numerical rather than constant-factor: Strassen gives only a normwise backward error bound. (I said Strassen is n^2.7 at lunch. It is log₂7 = 2.8074.)

Matrix Multiplication in Quadratic Time and Energy? Gregory Valiant, ITCS 2024. A theorist asking what classical physics rather than the RAM abstraction permits, and getting n² polylog(n) time and energy. The paper most likely to convince you that your instincts are an asset here rather than a handicap.

A VLSI Circuit Model Accounting for Wire Delay Ce Jin, R. Ryan Williams, Nathaniel Young, ITCS 2024. Thompson's 1980 programme continued with propagation delay made explicit: Θ(n^(1/3)) delay for AND, OR and PARITY, Θ̃(n^(1/2)) for addition and multiplication. The physics you already know, as a live object inside complexity theory.


What data movement actually costs

Two tables. Everyone in the room is implicitly citing one or the other. One warning before you use them: Dally's 2022 article never states which process node its model assumes, so none of its constants can be scaled to another technology without guessing. From the implied SRAM macro density of 0.047 µm² per bit, it is 7 or 5 nm class, not 45 nm, which matters because people routinely pair his wire constants with Horowitz's 45 nm operation table and produce a 5x to 16x error.

Dally's cost model, CACM 2022, node unstated.

Operation Energy Time
32-bit integer add 20 fJ 150 ps
Move the two 32-bit operands 1 mm on chip 1.9 pJ 400 ps
Move 64 bits corner to corner, 400 mm² die 77 pJ 16 ns
Read 64 bits from an 8 KB SRAM sub-array 0.64 pJ 300 ps
Read 64 bits from 256 MB built of those sub-arrays 58 pJ, of which 57.4 pJ is communication 12.3 ns
Go off chip 320 pJ per 64 bits 6 ps/mm
Fetch two 32-bit words from main memory 1.3 nJ, or 64,000x the add

Normalised: 29.7 fJ per bit per mm on chip, 10 fJ per bit for a sub-array read, 5 pJ per bit off chip. Divide the add by the wire and you get the number I gave you, and that you wrote down: at roughly 10 µm, moving the operands costs as much as the arithmetic.

Three caveats on that number, because it is the one you will repeat and someone will check.

  1. No source has Dally stating it. It falls out of his numbers (20 fJ ÷ 1.9 pJ/mm = 10.5 µm). Call it a consequence of his model, not a rule of his.
  2. It is format-specific across a 70x range. Against Jouppi's 7 nm table: about 16 µm for an int32 add, 147 µm for an int8 multiply, 200 µm for an fp32 add, 689 µm for an fp32 multiply.
  3. In time rather than energy the crossover is 375 µm, 36x farther out. Energy and latency are not telling the same story, which is the single most useful thing in this section.

Per-operation energy in pJ, 45 nm versus 7 nm.

45 nm 7 nm
int8 add 0.03 0.007
int32 add 0.10 0.030
fp16 add 0.40 0.160
fp32 add 0.90 0.380
int8 multiply 0.20 0.070
int32 multiply 3.10 1.480
fp16 multiply 1.10 0.340
fp32 multiply 3.70 1.310
8 KB SRAM, 64-bit access 10 7.5
1 MB SRAM 100 14
DDR3/4 DRAM 1300 1300

45 nm from Horowitz, ISSCC 2014; 7 nm from Jouppi et al., TPUv4i, ISCA 2021, Table 2, the only public measured 7 nm per-operation table, from a team that shipped the silicon.

The headline is the last row. Logic got roughly 3x cheaper across five nodes, SRAM 2x to 7x, DRAM did not move at all, and wire energy per unit length improved less than 2x. Jouppi's own summary is the sentence to carry: "Logic improves much faster than wires and SRAM, so logic is relatively free." That divergence is the entire business case for everything in the photonics section below.

Two more, if the tables interest you:

Domain-Specific Hardware Accelerators Dally, Turakhia and Han, CACM 2020. Explains the thing a physicist finds least intuitive about computers: a general-purpose CPU spends orders of magnitude more energy on instruction fetch, decode and register file than on the add it is performing. Unlike the 2022 piece it states its 28 nm and 14 nm assumptions.

The Future of Wires Ron Ho, Kenneth Mai and Mark Horowitz, 2001. The canonical treatment of repeatered RC interconnect, and why an on-chip wire behaves nothing like the transmission line your instincts expect.


The ladder nobody has published

You asked, in effect, where light starts to win. I could not find a single public like-for-like comparison of optical and electrical links at matched boundaries. Everyone compares different boundaries. So here is the assembled version, with what is inside each number stated, because that is where all the disagreement lives.

Boundary Technology pJ/bit Latency Reach What is inside Confidence
On-chip wire repeatered, min-pitch 0.0297 per mm 400 ps/mm on-die wire and repeaters model
On-chip wire thick top metal, delay-optimal repeaters 5–15x higher 20–25 ps/mm on-die wire and repeaters measured
3D stack hybrid bonding < 0.05 — µm link only measured
Package UCIe advanced, CoWoS, 3 nm 0.29 3.5 ns one way ≤ 2 mm PHY plus adapter measured, ISSCC 2025
Package UCIe standard, organic, 3 nm 0.52 — ≤ 25 mm PHY plus adapter measured
Board NVLink 4 1.3 — in-rack link vendor
Board CEI-112G VSR SerDes 1.75 — 220 mm analog only consortium target
Board CEI-112G LR SerDes 4.9 tens of ns 1 m analog only, no FEC consortium target
Memory HBM3E, whole stack 3.4–4.05 — in-package array plus TSV plus PHY measured
Memory DDR5, system level ~80 — board controller and interface measured
Optical research transceiver, headline 0.120 — fibre modulator and receiver only measured
Optical the same device, laser and tuning restored ~0.240 — fibre plus laser, plus thermal tuning measured plus supplement
Optical co-packaged, Lightmatter Passage M1000 4.6 — fibre wall plug: 2.6 photonics and laser, 2.0 SerDes vendor, matched boundary
Optical co-packaged optics, general 5.6–6.9 ~10 ns per chiplet fibre including laser peer-reviewed
Optical linear-drive pluggable ~9.0 < 3 ns module including laser peer-reviewed
Optical DSP pluggable 17–21 8–10 ns plus 50–100 ns FEC module everything peer-reviewed

Where light wins

On energy, on chip: never, with anything that ships. The best shipping optical link measured at the wall plug is 4.6 pJ/bit. Divide by 29.7 fJ/bit/mm and the break-even is 155 mm of on-chip copper, about five reticle diagonals. No die is that big. The only optical numbers that break even inside a die are research transceivers with the laser and the thermal tuner removed from the budget: 120 fJ/bit gives 4.0 mm, and restoring the laser and tuning from the same paper's own supplement gives about 240 fJ/bit and 8.1 mm. That factor of two, from two omitted terms, is the whole controversy in miniature.

On energy, off the board: comfortably, and that is where the money is. Against a 4.9 pJ/bit long-reach SerDes or a 17–21 pJ/bit DSP pluggable, a 4.6 pJ/bit co-packaged link wins on energy and wins much harder on bandwidth density.

On latency: nowhere inside a rack, and this surprised me. Light in fibre is 4.9 ps/mm. Off-chip copper is 6 ps/mm. The velocity advantage is 1.2x, essentially nothing. All the optical latency is fixed overhead: roughly 10 ns per chiplet excluding time of flight, so about 20 ns for a link with a chiplet at each end, plus 50 to 100 ns per hop if you run KP4 forward error correction. Recovering 20 ns of overhead at a 1.1 ps/mm velocity edge takes 18 metres.

The inversion worth remembering. On chip at 400 ps/mm, a 42 mm reticle diagonal costs 16.8 ns of copper latency but only 1.25 pJ/bit. So on-chip optics can win latency and lose energy, while board and rack optics wins energy and loses latency. The two arguments point in opposite directions at different scales, which is why the debate never resolves.

How these numbers get compared unfairly

You will meet all of these. They are ordered by how much damage they do.

  1. The laser is omitted. Papers report modulator drive energy, essentially CV²/4, and call it the modulator energy. The laser that supplies the light is 10x to 100x that number and lives in a supplementary section if anywhere. Plasmonic modulators at 0.07 fJ/bit are the extreme case: they get there by burning 5 to 12 dB of optical insertion loss, which the laser pays for.
  2. Thermal tuning is omitted, and it is static. A microring needs 1 to 2 mW of heater power whether or not bits are flowing. At 200 Gb/s that is 10 fJ/bit and invisible. At 10 Gb/s it is 200 fJ/bit and dominant. At 10% link utilisation, multiply by ten. Every number on the ladder assumes 100% utilisation, and a copper wire has no static term at all, so the utilisation penalty is asymmetric and always favours copper.
  3. The baseline is a pluggable, not copper. Every co-packaged-optics headline multiple is measured against pluggable transceivers, which carry a long-reach SerDes driving 15 to 30 cm of board plus a DSP. Most of the claimed saving comes from shortening copper, not from optics beating copper. None of these are optics-versus-copper comparisons at all.
  4. Four boundaries circulate under one label. Modulator drive only; transmit and receive front ends; photonics plus laser; and wall plug including the electrical SerDes. Comparing the first to the last is a 20x to 40x error, and it happens constantly.
  5. "Off-chip" is not one number. It spans hybrid bonding at under 0.05 pJ/bit to DDR5 at about 80, and the ordering does not follow distance monotonically. My own cost model collapses all of this into a single 5 pJ/bit constant, which is roughly right for one long-reach SerDes end and roughly 20x wrong for an interposer hop. This is the flaw in my model you are best placed to fix.
  6. Energy per bit falls with line rate for any fixed static power. A ring at 400 Gb/s and the same ring at 10 Gb/s differ 40x in reported fJ/bit purely through the denominator. Always ask for the line rate.
  7. HBM's pJ/bit is not an interconnect number. Roughly 74% of it is inside the DRAM stack. Replacing the interposer link with anything, optics included, touches about 14% of it.
  8. Cooling and power supply are excluded almost everywhere. The one source that carries them turns a 67% device-level advantage into a 36% system-level one.

What is genuinely unknown


The version of your thesis that survives

You said there is no EDA tool that co-designs photonics and electronics. I went looking, because it is the most interesting claim either of us made, and the answer has two halves.

For room-temperature photonics, the claim is about a decade out of date. Cadence shipped a Virtuoso-based integrated electronic-photonic environment with Lumerical in December 2015. Synopsys launched OptoCompiler in September 2020 as "the industry's first unified electronic and photonic design platform", with Inphi reporting roughly 4x faster schematic-to-tapeout. The existence proof is eleven years old: 70 million transistors and 850 photonic components on one 45 nm SOI die running a dual-core RISC-V processor that communicates optically (Sun et al., Nature 528:534, 2015), with the photonic design-rule cleaning written in Cadence SKILL.

The steelman is better than the original and I would switch to it. From Shekhar et al., Nature Communications 15:751 (2024), seven of the field's most senior people: "mature PDKs, and abstraction languages are still in very early stages" and "Third-party IP support is mostly non-existent thus far." There is no photonic RTL, no logic synthesis, no synthesis-targetable standard cell library, no design-for-test standard, and no third-party IP market. That is defensible and specific. "There are no tools" is neither.

And then there is the half where you are simply right. Add the word cryogenic and the claim holds completely:

If you are hunting for a problem, that is a real hole, it is exactly the intersection of the three things you spent four years doing, and almost nobody else is standing in it.


Things proposed at lunch that already exist

The most enjoyable section to research. In one case the prior art is a NeurIPS oral, and in another the exact experiment was already run inside the competition being critiqued.

A neural network that is literally a netlist

This was Ping's idea at the table: something differentiable at the gate level, hardwired for one task, sold as a single product.

Convolutional Differentiable Logic Gate Networks Petersen, Kuehne, Borgelt, Welzel, Ermon, NeurIPS 2024, oral. 86.29% on CIFAR-10 from 61 million logic gates, 29x fewer gates than the previous best, and the trained artifact drops straight onto an FPGA. On MNIST: 98.46% at 4 ns per image. The 2022 original, Deep Differentiable Logic Gate Networks, has the trick itself, which is how you make AND, OR and XOR differentiable enough to train by gradient descent and then discretise without losing everything.

The gap, and it is the interesting part. Nobody has ever taped one out, and there is no measured joules-per-inference figure for a differentiable logic gate network anywhere in the literature. Petersen's papers report no power, no energy and no area, and say so explicitly. The adjacent LUT-based line, LogicNets, PolyLUT and NeuraLUT, reports latency and LUT counts down to 3 ns and publishes no watts either. The closest thing to silicon is Silicon Aware Neural Networks (Apr 2026), a SkyWater 130 nm post-layout simulation, not a fabrication: 97% on MNIST at 41.8 million inferences per second and 83.88 mW, which works out to 2.0 nJ per inference. That single derived number is, as far as I can tell, the only energy figure the entire field has, and it comes from a simulation.

So the idea is not merely alive, it is a NeurIPS oral with an unclaimed energy axis. That is precisely the gap the group I mentioned exists to fill.

Training against a surrogate instead of running the simulator

Your suggestion, and it had already been tried inside the competition you were critiquing.

SURROGATE_STUDY.md in the spicenn2 repo. The surrogate matched ngspice on calibration but reported gradients blowing up to 1e11, which turned out to be an artifact of the surrogate rather than physics. The resolution they landed on is the general answer to your question: the fast model trains, the slow model scores.

LASANA: Large-Scale Surrogate Modeling for Analog Neuromorphic Architecture Exploration Ho, Boyle, Liu, Gerstlauer, 2025. The published state of the art, with numbers worth arguing from: 82x speedup on MNIST crossbars, 1321x on spiking MNIST, under 7% energy error. It also carries a warning that settles the design question by itself: its MNIST accuracy came out higher than the SPICE baseline it was imitating, which disqualifies a surrogate as a scorer by construction.

Xyce, Sandia. The answer to "SPICE will not scale" that does not require giving up ground truth: GPLv3, written from scratch for MPI, exercised on million-transistor circuits across hundreds of processors.

circulax, gdsfactory, 2026. Differentiable circuit simulation in JAX with DC, AC, transient and harmonic balance, ingesting Verilog-A compact models and mixing electronic and photonic components in one netlist. The bridge between designing a circuit and training one.

On your verification question, which was the sharper one: the competition's current answer is its energy rule, which meters every driven pin including image, label, clock and bias, "because counting only the main supply would miss circuits powered through their inputs." That is exactly the gaming instinct you were pointing at, already anticipated.

Storing bits in a loop of fibre

The paper you read is Who Needs DRAM? We Have Fiber, Atmer, Yao, Voigt and Kaxiras, Uppsala, posted 9 July 2026. Worth nailing the citation before arguing about it, and worth knowing two things: it has no related-work section, and the authors describe it themselves as a first-order feasibility study.

Two pieces of prior art it does not cite:

The comparison that decides it is not pJ/bit at the receiver. Every stored bit transits roughly 56,000 amplifiers per second, each adding spontaneous emission noise, against DRAM's 15 to 30 refreshes per second. That ratio is the argument.

And the honest state of the art in genuine optical memory, which will read as familiar physics after Yale: coherent photonic RAM in long-lasting sound waves, Geilen, Becker and Stiller at the Max Planck Institute for the Science of Light, 2023, via stimulated Brillouin scattering. Storage time 3.5 ns. Phonon lifetime 10.5 ns. The gap to DRAM retention is the whole problem in one number.


Three things I got wrong at lunch

Horowitz's talk is ISSCC 2014, not 2012. No 2012 talk of that title exists. His later work is Life Post Moore's Law, and the central proposal, an "apps store for hardware" modelled on the 1980s ASIC revolution, starts around slide 37. He also says, in as many words, that chiplets are not the answer.

The chip-scale optical comb market is real and I understated it with you. You said maybe one or two companies had launched in the past two or three years. At least five ship or sample named products, and the oldest is much older than three years: Pilot Photonics in Dublin, whose founders have been at multi-wavelength laser communication since 2005, selling the Lyra OCS 1000 comb source; Enlightra in Lausanne, out of stealth in December 2025 with $15M, shipping ATLAS at 8 to 32 comb lines per output; Xscape Photonics, whose FalconX launched 11 March 2026 with $37M and a 128-wavelength roadmap; Ayar Labs SuperNova, driving 16 wavelengths into 256 carriers; and Octave Photonics. Worth knowing alongside that: the industry standardised on 8, 16 and 32 wavelengths in the O-band, not hundreds, which tells you what the customers actually believe.

On "millions of wavelengths", the arithmetic does not get there, and the reason is elegant. The whole low-loss window is about 59 THz. The best system ever built across it is NICT's June 2024 record: 1,505 channels over 37.6 THz, requiring six different doped-fibre amplifier chemistries plus Raman. Deployed C band alone is 80 to 96 channels on the 50 GHz grid. A million lines in 50 THz means 50 MHz spacing, below laser linewidth. The binding constraint is Shannon with a nonlinear noise floor rather than anything thermodynamic, and here is the part I enjoyed: the Kerr nonlinearity that limits the channel count is the same four-wave mixing that generates the microcomb. The two halves of the argument cut against each other.

Where you were right and I was not: Celestial's architecture, which you guessed unprompted as an optical interposer and backplane, and it is exactly that. Lumentum and Coherent on indium phosphide shipping discrete multi-wavelength sources as chiplets rather than a monolithic comb, also exactly right. And the chip was indeed called Jalapeño. Your recall beat mine.


One open question, and one favour

The question, which is the joint artifact if you want one. Nobody publishes delivered optical power per wavelength together with wall-plug efficiency at the operating point, which means no optical energy-per-bit number in the table above can be independently audited. You are one of a small number of people who could publish that pair, and it would be cited immediately.

The favour. I could not work out what put the Uppsala fibre-memory paper back into circulation in September, two months after it was posted. There is no version 3, it has zero citations, and the only Hacker News submission was in July. If you remember where you saw it, tell me, because that is one message and it beats any amount of further inference.

We meet Mondays at 18:00. Bring a number.