GPT: a real tradeoff
Keeping 8 of 12 blocks increases CE by 4.199 after dense training, versus 0.093 / 0.116 after constant / decreasing ILD. But full-model CE is worse with dropout. Flexibility costs full-model quality at this training budget.
Research synthesis · 9 September 2026
A model can learn to work with fewer layers. The important question is whether that gives us a better predictor, a useful family of predictors, or a cheaper computation.
Our experiments support useful subnetworks beyond language models. They also show that the metric and the pattern of missing layers can change the conclusion.
See the evidence ↓ Read the paper · PDF ↗Don’t Drop Dropout studies randomly omitting whole layers during training, with more dropout in deeper layers and less dropout as training progresses. At the end, the full network is used again; useful shorter paths can remain.
Think of one trained model serving several depth budgets. This could matter for early exits or for a smaller draft inside self-speculative decoding. Its value is the ability to choose among useful paths after training.
This has clear precedents. The paper’s contribution is a detailed modern language-model recipe and empirical study, building on stochastic depth, LayerDrop and growing subnetworks.
Schematic · not a measured trajectory
Color shows each layer’s probability of being active, not one sampled mask. Decreasing ILD: maximum dropout 0.8 → 0; average masking over training is 20%.
Cross-entropy (CE) scores how much probability a model assigns to correct answers; lower is better. Accuracy counts correct top predictions; higher is better. Neither can be replaced by “how much did pruning change the score?”
● Dense training● Constant ILD● Decreasing ILD
| Training recipe | Full | Pruned | Pruned − full |
|---|
Lines connect two measured endpoints; they do not represent training or intermediate depths. CE units: nats per target. GPT accuracy is token accuracy and is not comparable to image accuracy. No head refitting or calibration after pruning. Charts show means; the primary paired intervals appear below.
Keeping 8 of 12 blocks increases CE by 4.199 after dense training, versus 0.093 / 0.116 after constant / decreasing ILD. But full-model CE is worse with dropout. Flexibility costs full-model quality at this training budget.
Keeping 12 of 18 residual blocks gives 14.19% accuracy after dense training, versus 51.01% / 50.62% after ILD. The primary relative-CE intervals are wide and include zero; the large observed benefit needs more seeds.
Dense CE improves most when the last four blocks are removed, from 3.079 to 2.244. Constant ILD improves less, from 2.713 to 2.155, yet ends with lower mean CE. A smaller improvement does not imply a worse pruned predictor.
We fixed the primary mask and endpoint before the final comparisons: effect = (CEpruned − CEfull)ILD − (CEpruned − CEfull)dense. Negative means a lower signed CE change. When pruning improves CE, it can mean a larger improvement rather than “less damage.”
| Architecture | Constant ILD | Decreasing ILD |
|---|
Intervals cover variation across training seeds, not dataset or tuning uncertainty. Accuracy and random-mask comparisons are secondary. ViT’s mean pruned CE favors ILD, but its paired pruned-CE intervals include zero. Inspect individual seeds, all metrics and all masks →
For ViT, keeping the first eight blocks yields 46.68% dense accuracy. Keeping eight blocks scattered through the network yields only 17.52–24.15% across three fixed masks. Constant ILD retains 44.90–46.75% under those scattered masks.
Both retain the same depth. They ask different questions. Choose a random mask in the chart to see the change.
The first block is never dropped by the ILD training distribution. Removing it at test time leaves ViT near 1% accuracy and GPT near 1–3% token accuracy. Robustness has a training-distribution boundary.
Decreasing dropout also leaves useful subnetworks at the final dense step. We did not test a prolonged dense continuation, and no schedule wins everywhere.
Randomly shortened residual networks in vision.
Transformer subnetworks that can be pruned on demand.
Growing random subnetworks; depth dropout and early-exit self-speculation.
A larger causal-LM recipe study combining granularity, schedules and scaling.
The significant claim is that structured depth noise can be useful beyond classic overfitting prevention. Our cross-architecture audit supports conditional usefulness, while showing that pruning resilience, predictor quality and training efficiency need separate evidence. The local toy studies also include smaller dense models: a reusable subnetwork does not automatically beat training a small model directly.
Inverted dropout preserves the expected value of one fixed residual update. That does not establish the same property for a transformer block whose attention and MLP share a mask.
A linear scalar counterexample is enough. With two unit residual branches and 50% survival, the full block outputs 4, but the expected training output is 5.
This corrects an expectation argument under the printed implementation. It does not invalidate the paper’s measured results.
The connection to the Sutro meeting’s locality agenda is concrete: per-example skipping may remove arithmetic while still requiring almost every layer’s weights. With independent 50% dropout and a batch of 32, a layer is needed somewhere with probability 1 − 0.5³² ≈ 100%. That is a probability calculation, not a measured transfer or energy saving.
Compare per-example, grouped and whole-batch masks at matched predictive quality. Actually bypass work, then measure weight traffic, packing overhead, peak live memory, time and energy. Include a directly trained smaller dense model.
Predictive behavior under missing blocks. Training computed each branch before masking, so it demonstrates no saved training FLOPs, latency or energy. The paper’s nominal block-FLOP savings likewise need a separate systems measurement before becoming a wall-clock or joule claim.
How to use these results
The next decisive scientific checks are longer GPT training, vision recipes with smaller train–held-out gaps, validation-only calibration diagnostics, more seeds, and smaller dense controls at A100 scale.
Our GPT saw 0.843 token presentations per parameter, far below the paper’s typical 20. Both vision models use CIFAR-100; eight of nine learning-rate winners are at a grid boundary. Numerically reproduced CPU-dependent initialization variants qualify vision seed pairing. These limits apply to our experiments, not to the authors’ runs.