GRADIENT DISSENT

Research synthesis · 9 September 2026

What does depth
robustness buy us?

A model can learn to work with fewer layers. The important question is whether that gives us a better predictor, a useful family of predictors, or a cheaper computation.

Our experiments support useful subnetworks beyond language models. They also show that the metric and the pattern of missing layers can change the conclusion.

See the evidence ↓ Read the paper · PDF ↗
3 A100 architectures27 final + 27 tuning runs3 paired final seeds$27.81 metered snapshot
01

The valuable idea is flexibility.

Don’t Drop Dropout studies randomly omitting whole layers during training, with more dropout in deeper layers and less dropout as training progresses. At the end, the full network is used again; useful shorter paths can remain.

Think of one trained model serving several depth budgets. This could matter for early exits or for a smaller draft inside self-speculative decoding. Its value is the ability to choose among useful paths after training.

This has clear precedents. The paper’s contribution is a detailed modern language-model recipe and empirical study, building on stochastic depth, LayerDrop and growing subnetworks.

Schematic · not a measured trajectory

Color shows each layer’s probability of being active, not one sampled mask. Decreasing ILD: maximum dropout 0.8 → 0; average masking over training is 20%.

02

Watch the full model and the pruned model.

Cross-entropy (CE) scores how much probability a model assigns to correct answers; lower is better. Accuracy counts correct top predictions; higher is better. Neither can be replaced by “how much did pruning change the score?”

● Dense training● Constant ILD● Decreasing ILD

Mean score across three independent training seeds. Exact numbers and intervals are in the full report.
Training recipeFullPrunedPruned − full

Lines connect two measured endpoints; they do not represent training or intermediate depths. CE units: nats per target. GPT accuracy is token accuracy and is not comparable to image accuracy. No head refitting or calibration after pruning. Charts show means; the primary paired intervals appear below.

GPT: a real tradeoff

Keeping 8 of 12 blocks increases CE by 4.199 after dense training, versus 0.093 / 0.116 after constant / decreasing ILD. But full-model CE is worse with dropout. Flexibility costs full-model quality at this training budget.

ConvNeXt: promising transfer

Keeping 12 of 18 residual blocks gives 14.19% accuracy after dense training, versus 51.01% / 50.62% after ILD. The primary relative-CE intervals are wide and include zero; the large observed benefit needs more seeds.

ViT: the metric reversal

Dense CE improves most when the last four blocks are removed, from 3.079 to 2.244. Constant ILD improves less, from 2.713 to 2.155, yet ends with lower mean CE. A smaller improvement does not imply a worse pruned predictor.

The primary endpoint and its uncertainty

We fixed the primary mask and endpoint before the final comparisons: effect = (CEpruned − CEfull)ILD − (CEpruned − CEfull)dense. Negative means a lower signed CE change. When pruning improves CE, it can mean a larger improvement rather than “less damage.”

Paired mean effect [95% Student-t interval], n = 3, df = 2; six unadjusted primary intervals.
ArchitectureConstant ILDDecreasing ILD

Intervals cover variation across training seeds, not dataset or tuning uncertainty. Accuracy and random-mask comparisons are secondary. ViT’s mean pruned CE favors ILD, but its paired pruned-CE intervals include zero. Inspect individual seeds, all metrics and all masks →

03

Which layers disappear matters.

For ViT, keeping the first eight blocks yields 46.68% dense accuracy. Keeping eight blocks scattered through the network yields only 17.52–24.15% across three fixed masks. Constant ILD retains 44.90–46.75% under those scattered masks.

Both retain the same depth. They ask different questions. Choose a random mask in the chart to see the change.

A boundary, not arbitrary resilience

The first block is never dropped by the ILD training distribution. Removing it at test time leaves ViT near 1% accuracy and GPT near 1–3% token accuracy. Robustness has a training-distribution boundary.

Decreasing dropout also leaves useful subnetworks at the final dense step. We did not test a prolonged dense continuation, and no schedule wins everywhere.

04

What is new—and what already existed?

2020

LayerDrop

Transformer subnetworks that can be pruned on demand.

2024

RaPTr & LayerSkip

Growing random subnetworks; depth dropout and early-exit self-speculation.

2026

Don’t Drop Dropout

A larger causal-LM recipe study combining granularity, schedules and scaling.

The significant claim is that structured depth noise can be useful beyond classic overfitting prevention. Our cross-architecture audit supports conditional usefulness, while showing that pruning resilience, predictor quality and training efficiency need separate evidence. The local toy studies also include smaller dense models: a reusable subnetwork does not automatically beat training a small model directly.

05

A useful result can have an incorrect explanation.

Inverted dropout preserves the expected value of one fixed residual update. That does not establish the same property for a transformer block whose attention and MLP share a mask.

A linear scalar counterexample is enough. With two unit residual branches and 50% survival, the full block outputs 4, but the expected training output is 5.

This corrects an expectation argument under the printed implementation. It does not invalidate the paper’s measured results.

Mask off · probability ½1 → 1 → 1
Mask on · probability ½1 → 3 → 9
Expected training output½ × 1 + ½ × 9 = 5
Dense evaluation1 → 2 → 4
Read the full mathematical and reporting audit →
06

For energy research, the missing measurement is movement.

The connection to the Sutro meeting’s locality agenda is concrete: per-example skipping may remove arithmetic while still requiring almost every layer’s weights. With independent 50% dropout and a batch of 32, a layer is needed somewhere with probability 1 − 0.5³² ≈ 100%. That is a probability calculation, not a measured transfer or energy saving.

The experiment to do next

Compare per-example, grouped and whole-batch masks at matched predictive quality. Actually bypass work, then measure weight traffic, packing overhead, peak live memory, time and energy. Include a directly trained smaller dense model.

What our A100 study establishes

Predictive behavior under missing blocks. Training computed each branch before masking, so it demonstrates no saved training FLOPs, latency or energy. The paper’s nominal block-FLOP savings likewise need a separate systems measurement before becoming a wall-clock or joule claim.

How to use these results

Evaluate the family of predictors,
then price the computation.

The next decisive scientific checks are longer GPT training, vision recipes with smaller train–held-out gaps, validation-only calibration diagnostics, more seeds, and smaller dense controls at A100 scale.

Our GPT saw 0.843 token presentations per parameter, far below the paper’s typical 20. Both vision models use CIFAR-100; eight of nine learning-rate winners are at a grid boundary. Numerically reproduced CPU-dependent initialization variants qualify vision seed pairing. These limits apply to our experiments, not to the authors’ runs.