Skip to slides
GRADIENT DISSENT / SLIDES

Cerebras · arXiv:2609.05275 · posted week 36, 2026 · nine authors

Don't Drop
Dropout.

Randomly skipping whole transformer blocks during pre-training should come back into language-model recipes: configured right, it saves up to a quarter of training compute at similar loss and leaves a model that can run with fewer layers at inference.

→ / space / swipe to advance · o for contents · f fullscreen. Every chart is computed live from the paper's stated formulas; charts marked schematic show shape, not the paper's data.

Background

Stochastic depth, 2016: skip a block, let the residual carry the stream.

During training each residual block is skipped at random with some probability; the residual connection passes the stream through untouched. At inference, every block runs.

Two things make this interesting for language models:

  1. Structured sparsity. Unlike neuron-level dropout, skipping a block removes real compute instead of zeroing numbers inside a dense kernel.
  2. No block is load-bearing. A model trained with blocks randomly missing learns not to depend on any single one — exactly what early exit, layer skipping and self-speculative decoding need.

Dropout of all kinds left pre-training recipes because single-epoch training on huge data doesn't overfit and activation dropout hurt. The authors' thesis: layer dropout was thrown out with it, and the reported degradations came from bad configurations, not anything fundamental.

One residual block: kept or skipped
Press sample: with drop probability 40% the branch either runs (scaled by 1/0.6) or the identity carries the stream.

The setup

~2,400 runs on CS-3 systems, four knobs.

271M → 8.2Bdecoder-only transformers, Cerebras's Celerity family
20 TPPtokens per parameter for most runs — the compute-optimal ratio
~2,400training runs, most of them small hyperparameter sweeps

Knob 1 · the scaling factor

Scale a surviving block by 1 / keep-probability — because that is what CompleteP already says.

When a block survives, its output can be scaled up to compensate for the times it is absent, and codebases disagree on how. The paper's first finding: 1/p is the right choice. CompleteP scales each residual branch by 1/depth; a model with dropout has a smaller effective depth, so scaling surviving branches by 1/p keeps residual-stream updates the right size.

Payoff: the optimal learning rate, batch size and weight decay stay put across dropout rates — tune once, transfer. Verified with coordinate checks and sweeps on the smallest model.

Effective depth and survivor scale against keep probability
Effective depth p·L shrinks linearly; the survivor scale 1/p grows so that the residual update size stays where CompleteP put it. Computed, not measured.

Knob 2 · granularity

Drop whole blocks, and decide per sequence — not per batch.

Whole block vs sub-block. Dropping attention and feed-forward together beats dropping them independently in their tables — by a few thousandths of a nat.

Per sequence vs per batch. Per sequence is clearly better: 2–5 percentage points of loss at the same rate. Per batch means an entire step either has the layer or doesn't, which is very noisy. The two classic implementations — the original stochastic-depth code and fairseq's LayerDrop — are per batch, so this alone may explain some of the old bad results.

Which of eight sequences run which of twelve blocks in one training step
One training step, eight sequences × twelve blocks. Filled = block runs, hollow = identity bypass. The last column is the active depth each sequence saw. Sampled masks, not a trained model.

Knob 3 · distribution across depth

At equal FLOPs, non-uniform beats uniform; protect the early layers.

Three shapes: uniform, the same rate everywhere; increasing, from zero at the first layer linearly up to a maximum at the last; alternating, where only every other layer is ever dropped.

Compared at equal FLOPs savings, non-uniform wins. Alternating is best at the smallest size, increasing is best at the larger sizes, and they recommend increasing as models scale. Increasing is in fact the 2016 default: lean on the late layers.

Per-layer drop probability for uniform, increasing and alternating distributions at matched mean
Per-layer drop probability; each shape has the same mean, i.e. the same expected block-FLOPs saved. Hover a bar for its value.

Knob 4 · schedule over time — the headline

Decrease the rate to zero over training. Increasing it was backwards.

Constant, increasing from zero, or decreasing from the maximum to zero by the end. Decreasing wins at every size and increasing is clearly worst: at 20% savings, increasing costs 6–7% loss versus 1.5–3% for decreasing.

Increasing across depth + decreasing across time, at 5% savings, puts the 900M model 0.002 nats below the dense baseline — which they call the first time anyone has beaten dense with fewer training FLOPs.

Their reading: a decreasing schedule is stochastic model growing, a curriculum — the network starts at a fraction of its effective depth and grows into it. Note the reversal: LayerSkip (same first author) and Progressive Layer Dropping before it both increased dropout over time.

Drop rate of the last layer over training, and the depth-by-time grid of the recipe
Top: the last layer's drop rate through training for the chosen schedule. Bottom: the depth × time grid under the recipe (increasing in depth × chosen schedule); darker = dropped more often. The marker is the current progress.

Inference · “elastic depth”

A dense model breaks the moment one layer is missing. A dropout-trained model bends.

  • Static early exit (first layers, then the unembedding): dense loss blows up after a single dropped layer; dropout-trained models degrade gradually, higher rates giving flatter curves.
  • Interior skipping: alternating is the best distribution; increasing is best for early exit — so the two benefits trade off.
  • Frozen model + small early-exit adapters (Balcony-style): dropout-trained models start ahead and stay ahead.
  • Self-speculative decoding — a subset of the model's own layers drafts, the full model verifies: 1.2–1.4× at the small sizes, where dense gets essentially nothing, because in a dense model no cheap subset is a good drafter.
Measured losses at reduced depth for the 3.9B and 8.2B models
Measured (Table 5): validation loss at full depth versus with layers removed, dense against dropout-trained. Hover a bar for the number.

Scaling · the hero runs

Rates far beyond anything used before — the last layer skipped 80–99% of the time at the start.

ModelpmaxBlock FLOPs savedDense lossDropout lossVerdict
1.8B0.6015%1.8491.836slightly better than dense
3.9B0.8020%1.7321.745slightly worse, ~0.75%
8.2B0.9925%— none1.663no dense baseline at all

All three: increasing across depth, decreasing across time, 20 TPP. On the 500M model a tokens-per-parameter sweep from 2 to 64 finds the dropout–dense gap shrinks as data grows. Self-speculative decoding: 1.54× for the 3.9B dropout model vs 1.02× dense; 1.55× for 8.2B.

99%of sequences skip the last layer at step 0
24.75%average block FLOPs saved over the run (pmax/4 under the recipe)

The recipe

Five settings.

  1. Scale surviving blocks by 1 / keep-probability. It matches CompleteP, so learning rate, batch size and weight decay transfer across rates.
  2. Drop whole blocks — attention and feed-forward together.
  3. Decide per sequence, never per batch.
  4. Rate increasing linearly with depth: zero at the first layer, maximum at the last.
  5. Rate decreasing linearly to zero over training, with larger maximum rates for larger models — 0.6 at 1.8B, 0.8 at 3.9B, 0.99 at 8.2B.

Caveats, quickly · 1 of 2

No seeds, no error bars, anywhere.

ResultSize of the effectRuns per cellSurvives seed noise?
Per sequence beats per batch2–5% loss1yes — large
Decreasing beats increasing schedule~4 pts of loss at 20% savings1yes — large
Whole block beats sub-blocka few mnats1unproven
Non-uniform beats uniform in depth2–10 mnats1unproven
Beats the dense baseline2 mnats (900M); worse at 500M1unproven, and overstated

The whole-block result, the distribution result and the beat-the-baseline result all rest on gaps of 2–10 thousandths of a nat at one run each — seed noise at these sizes.

The “first time” claim is wrong historically: the original stochastic-depth abstract already claimed better accuracy in less time. And their own table shows the 500M dropout model a hair worse than dense.

Caveats, quickly · 2 of 2

What “25% faster” is counting.

  • The 8.2B run has no dense baseline, so the headline “25% at preserved accuracy” is asserted, not shown.
  • The loss-vs-FLOPs figure compares a finished dropout run against an unfinished dense run whose learning rate hasn't decayed. The fair comparison is a separate shorter dense run or a shallower dense model; they do neither.
  • FLOPs, never wall clock — and non-embedding FLOPs at that; at the smallest size most parameters are embedding matrices for a 128k-token vocabulary.
  • The decreasing schedule is early dropout — the 2023 “Dropout Reduces Underfitting” idea applied to layers — and it isn't cited. Since they call it stochastic model growing, the missing baseline is deterministic growing: train shallow, then stack.
  • The abstract and introduction disagree on model range and token count; the train-loss percentages in their first table don't match their own numbers.
Fraction of a layer's weights read per step: FLOPs versus bytes
For our purposes the accounting is FLOPs, not bytes moved: the preferred per-sequence variant still streams every layer's weights every step, because a layer is read whenever any sequence keeps it.

Verdict

What's durable, and what isn't.

Durable

  • A principled scaling rule (1/p, from CompleteP) that makes hyperparameters transfer across dropout rates.
  • Per sequence beats per batch.
  • Increasing-in-time schedules are bad; decreasing ones are good.
  • A multi-billion-parameter model can train with its last layer active 1% of the time and still work.

Not established

  • That dropout beats dense at equal compute.
  • That the elasticity is more than not collapsing: a 3.9B model with half its layers skipped (loss 2.13) scores about like a 500M dense model.
  • Anything about wall clock or energy: the whole accounting is FLOPs. The per-sequence variant they prefer streams every layer's weights every step.
1 / 14