GRADIENT DISSENT / INTERACTIVE NOTES 01

Inside “Don’t Drop Dropout”

Same network.
Different paths.

During training, each sequence randomly skips whole transformer blocks. Deeper blocks disappear more often. As training ends, every block comes back.

Explore the paper’s preferred recipe below. These are sampled masks and schematic activations, not a trained model or a timing benchmark.

The path explorer

One model. Four sequences. Twelve blocks.

TRAINING
Execute attention + FFNIdentity bypassSelect any block to inspect it ↓
Per-sequence paths through twelve shared transformer blocksEach row is a sequence. Filled cells execute both branches; dashed cells pass the residual stream through unchanged. Select a cell for its probability and scaling.

Each row gets its own mask. All tokens within that sequence share it. Columns use the same model weights.

Active blocks in this batch—Realized Bernoulli sample
Expected active depth / sequence—An expectation, not a fixed count
Average masks omitted over training20%80% ÷ 2 across depth ÷ 2 across time

Zoom into one block

Sequence A · Block 12

BYPASS
Drop —Survive —Mask M —Branch multiplier —

z = h + M / ρ · Attention(h)
h′ = z + M / ρ · FFN(z)

ρ = 1 − p. The same M multiplies both branches. Layer normalization and the paper’s additional CompleteP depth scaling are omitted from this schematic; the displayed multiplier isolates dropout.

The depth × time recipe

Start shallower. End at full depth.

pℓ,t = pmax × ℓ / (L − 1) × (1 − t / (T − 1))

Block labels are 1–12; the formula uses ℓ = 0–11. So block 1 is always kept. The cursor shows training progress; inference settings do not change the training recipe.

01

A whole block, per sequence.

The three moving dots stand for token activations in one sequence. They share every keep/drop decision. Attention and FFN are retained or bypassed together; individual neurons are not being erased.

02

The residual stream survives.

A dropped block is an identity map: h′ = h. A retained block adds two scaled updates. Each sequence trains a sampled subnetwork of the same shared model, with newly drawn masks in each batch.

03

Missing neighbors become familiar.

This gives blocks practice working without some other blocks. It motivates depth robustness at inference, but does not guarantee that every smaller subnetwork works well. See what our experiments found →

Why scaling is not a guarantee of the same expected output

Inverse survival makes a single mask multiplier have mean 1. But here the FFN’s input already depends on the attention branch’s shared mask. The expected output of the whole block can differ from dense inference.

For a scalar example with h = 1, Attention(h) = h, FFN(z) = z, and survival ρ = ½: a dropped block outputs 1; a kept block uses multiplier 2 and outputs 9. Their equally likely average is 5. Dense inference uses multiplier 1 and outputs 4.

Read the exact counterexample in the review →
Paths are not performance measurements. Avoiding block arithmetic requires an implementation that actually skips it. Computing a branch and then multiplying it by zero saves no branch work. Animation speed is illustrative; block counts do not measure end-to-end FLOPs, latency, or energy.