Zoom into one block
Sequence A · Block 12
ρ = 1 − p. The same M multiplies both branches. Layer normalization and the paper’s additional CompleteP depth scaling are omitted from this schematic; the displayed multiplier isolates dropout.
Inside “Don’t Drop Dropout”
During training, each sequence randomly skips whole transformer blocks. Deeper blocks disappear more often. As training ends, every block comes back.
Explore the paper’s preferred recipe below. These are sampled masks and schematic activations, not a trained model or a timing benchmark.
The path explorer
Each row gets its own mask. All tokens within that sequence share it. Columns use the same model weights.
Zoom into one block
ρ = 1 − p. The same M multiplies both branches. Layer normalization and the paper’s additional CompleteP depth scaling are omitted from this schematic; the displayed multiplier isolates dropout.
The depth × time recipe
Block labels are 1–12; the formula uses ℓ = 0–11. So block 1 is always kept. The cursor shows training progress; inference settings do not change the training recipe.
The three moving dots stand for token activations in one sequence. They share every keep/drop decision. Attention and FFN are retained or bypassed together; individual neurons are not being erased.
A dropped block is an identity map: h′ = h. A retained block adds two scaled updates. Each sequence trains a sampled subnetwork of the same shared model, with newly drawn masks in each batch.
This gives blocks practice working without some other blocks. It motivates depth robustness at inference, but does not guarantee that every smaller subnetwork works well. See what our experiments found →
Inverse survival makes a single mask multiplier have mean 1. But here the FFN’s input already depends on the attention branch’s shared mask. The expected output of the whole block can differ from dense inference.
For a scalar example with h = 1, Attention(h) = h, FFN(z) = z, and survival ρ = ½: a dropped block outputs 1; a kept block uses multiplier 2 and outputs 9. Their equally likely average is 5. Dense inference uses multiplier 1 and outputs 4.
Read the exact counterexample in the review →