Start here: what the paper and experiments mean → · Session transcript · How dropout works → · Slides: the talk → · MNIST A100 speed results → · Which digits can skip layers? → · 50% training dropout: 48-run experiment → · Run MNIST in your browser →
Skip to report
GRADIENT DISSENT
Critical review · arXiv:2609.05275v1

Gradient
dissent.

A promising recipe for training with fewer active layers. A weaker case for unchanged accuracy, exact mean preservation, or energy savings.

This review examines Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference, by Elhoushi and colleagues. It connects the work to stochastic depth, progressive subnetworks, and Sutro's search for learning algorithms that move fewer bytes.

24.75%modeled nonembedding FLOPs saved at the largest dropout rate
1.55×reported self-speculative speedup; a selected workload and draft search
3 testsexact mask enumeration, local execution timings, and multi-seed training

The original 27-page PDF, including appendices, is the source of the paper audit. Local toys test mechanisms and scope; they do not reproduce billion-parameter language-model training.

Source: paper, Table 5, p. 15. Savings calculated from its Eq. 9: 0.99 / 4 = 0.2475.

02 / 16

The useful result survives a narrower reading.

The contribution is a systematic recipe study for causal LMs, combining training efficiency with useful depth flexibility.

ClaimWhat the paper supportsWhat remains open
Training can use fewer FLOPsFewer active transformer blocks. Several favorable loss–compute comparisons.End-to-end elapsed time, optimizer cost, energy and hardware portability.
Accuracy is preserved1.8B validation loss improves: 1.849 → 1.836.3.9B loss worsens: 1.732 → 1.745. No dense 8.2B control in Table 5.
Inference becomes flexibleMuch more robust to depth removal; selected self-speculative decoding gains.Skipping is generally lossy. Verification preserves the chosen target model, not a different dense-trained baseline.
Scaling is understoodUseful ablations and encouraging model/data trends.No fitted universal dropout scaling law. Largest runs are about 20 tokens per parameter.

More than 2,400 runs is an unusually broad experimental effort. That breadth is valuable, but it is not a replacement for repeated seeds on headline comparisons, matched large-model controls, or a complete cost ledger.

Sources: §5, §§9–11, Tables 4–5, Table A.3.

03 / 16

A skip is a different path through the network.

Every sequence samples which transformer blocks it executes. A dropped block passes the residual stream through unchanged. Kept residual branches are amplified by the reciprocal survival probability.

hℓ+1 = hℓ + mℓ / (1 − pℓ) · fℓ(hℓ)

A real speedup requires actually skipping computation. Multiplying a fully computed output by zero gives the stochastic objective but keeps the expensive work.

Illustration · sampled masks

Blue: execute. Outlined: skip. Four illustrated sequences, eight blocks; these are sampled diagrams, not a model trace.

In a transformer, attention and FFN are separate residual branches. The paper recommends sharing a mask between them, while scaling each branch. That detail will matter when we examine the expectation argument.

Source: paper, Eqs. 1–7 and §6.1, pp. 4–6.

04 / 16

Start shallow. Finish dense.

The preferred schedule drops later layers more often, then gradually removes dropout during training.

pℓ,t = pmax · ℓ/(L − 1) · (1 − t/(T − 1))

Exact expected finite-grid average for L,T > 1 and equal block costs. Realized Bernoulli counts fluctuate. The average over depth halves the maximum; decreasing time halves it again. “80% dropout” therefore means 20% average active-block savings here.

Source: paper, §7, Eqs. 8–9, pp. 8–10. Interactive values are calculated, not measured.

05 / 16

New evidence, with a long lineage.

Expanded history: stochastic depth’s authors, later successes, and the novelty of this paper →

Stochastic Depth · Huang et al.

Randomly bypass residual blocks during training; survival varies with depth. Establishes the mechanism.

LayerDrop · Fan et al. / Progressive Layer Dropping · Zhang & He

Transformers become robust to depth removal. Scheduled dropping accelerates language-model training in an earlier regime.

RAPTR: Progressive Subnetworks · Panigrahi et al.

Random subnetworks progressively grow toward the full network. A close precursor to the “stochastic model growing” interpretation, absent from the reviewed paper's references.

LayerSkip · Elhoushi et al.

Layer dropout plus early-exit training supports self-speculative inference. The present work studies whether the base-model accuracy penalty can be reduced without an auxiliary early-exit loss.

CompleteP · Dey et al.

Transfers hyperparameters over width and depth. This paper combines that framework with dropout scaling and an extensive recipe sweep.

Don't Drop Dropout

The strongest novelty is the combined empirical study in modern decoder-only models: mask granularity, schedules, transfer, and downstream depth flexibility.

LayerDrop already compared whole-layer and sublayer dropping (§6); it also used alternating layers as an inference pruning rule, which differs from alternating-only training dropout. The older results are not interchangeable replications: architectures, data reuse, objectives and training budgets differ. RAPTR strengthens the conceptual lineage without erasing the value of this new study.

06 / 16

The same mask creates a cross-term.

Mathematical correction · exact enumeration

Take a scalar attention branch a·x and FFN branch b·z. Apply Eq. 6 with the same Bernoulli mask m and ρ = 1 − p to both branches. Even this linear example contradicts the unqualified expectation equality at the end of §5.

z = x + (m/ρ)ax   ;   y = z + (m/ρ)bz
E[y] = (1 + a + b + ab/ρ)x
yeval = (1 + a + b + ab)x
E[y] − yeval = abx · p/(1 − p)

Here x = 1, a = 0.5 and b = 0.4, matching the saved enumeration. The expectation averages both masks exactly. There is no sampling noise, optimizer choice, or nonlinearity to blame.

For one aggregate residual branch, E[mf(h)/ρ | h] = f(h) is valid. With independently sampled masks, this particular linear cross-term also cancels. Neither fact establishes equality for a general deep nonlinear network.

This corrects a justification, not the measured losses. The empirical scaling recipe can work well despite a false exact-equality explanation. Actual production code should be checked against the published Eq. 6.

Source: paper, Eq. 6 and final paragraph of §5, p. 5. Derivation and enumeration: local experiment source.

07 / 16

Matching a mean does not match the objective.

For a single scalar residual y = x + mθx/ρ, even the correctly preserved output mean is insufficient. Expected squared error includes a dropout-dependent penalty:

E[(y − t)²] = (x + θx − t)² + (p/ρ) θ²x²

With x = 1 and target t = 2, dense training prefers θ = 1. The dropout objective prefers θ = ρ. Its gradient and curvature change as dropout increases.

This is why an activation coordinate check is useful evidence for scale control, but cannot prove unbiased dense gradients, optimizer equivalence, or universal hyperparameter transfer.

Analytic curves, p = 0.5. Blue: dense objective; orange: dropout objective. Expected-loss convention has no factor of ½.

The paper itself acknowledges that transfer deteriorates at aggressive dropout and was mainly checked with constant schedules. The recommended decreasing schedule deserves its own transfer tests. Tests should include moments of updates and residual changes, not only activation magnitude after ten steps.

Sources: paper, §11, p. 16 and Appendix B, pp. 24–25. Objective identity is an independent derivation.

08 / 16

The headline combines different operating points.

Reported data · Table 5
SizeMax pFLOPs saved*Dense lossDropout lossCE changeSpec. speedup
1.8B0.6015%1.8491.836−0.70%1.34×
3.9B0.8020%1.7321.745+0.75%1.54×
8.2B0.9924.75%—1.663unknown1.55×

*Nonembedding active FLOPs. The paper rounds the final row to 25%. CE changes here are calculated from rounded printed values; all rows use ILD+DTS and 20 TPP.

Validation loss is not downstream accuracy

Assuming the roughly 1.8B model is the one labeled 1.9B in Table A.3, dropout is lower on 9 of 11 downstream metrics and tied on two, despite its better validation loss. This is a descriptive audit, not a statistical significance test.

Small gains need error bars

Table 3 reports 503M at +0.03% loss change for the favored 5%-saving schedule, while the surrounding text says it beats the baseline. The 906M improvement is tiny (reported −0.06%). Seeds and precision matter here.

At 3.9B, the CE difference of 0.013 corresponds to approximately 1.31% higher perplexity if CE is in nats. It may be a worthwhile efficiency trade-off, but “no loss” is too strong. At 8.2B, comparison to a dense model of the same size is simply unavailable in the reported table.

Sources: Table 3, p. 10, Table 5, p. 15, Table A.3, p. 25.

09 / 16

A controlled toy asks a smaller question.

Locally measured · synthetic regression

Can the schedule ranking survive outside the paper's setup? A small residual network learns a synthetic teacher under dense, constant, increasing and decreasing dropout. Learning-rate selection uses separate tuning seeds; evaluation uses held-out seeds.

Loading measured results…

Scope: this is a test of a proposed mechanism and its generality. It is not a reproduction of the paper's causal LMs, corpus, CompleteP recipe or Cerebras training. Training computes then masks. Its expected active-block budget is a conceptual model of ideal skipping, not measured FLOPs saved; it excludes heads, gathering, optimizer work and kernel overhead. Raw wall time is recorded separately.

Sources: complete local results JSON; reproduction scripts, raw runs and checks.

10 / 16

Zeroing a result does not save its computation.

Locally measured · CPU execution benchmark

Three implementations can represent dropout: run all rows and mask outputs; gather active rows, run them and scatter back; or share a mask for the batch and bypass entire blocks. The first two can be compared with exactly the same masks. Batch masks induce a different noise distribution.

Gathering reduces active matrix arithmetic, but costs indexing, copies and smaller matrix kernels. Whole-batch skipping can avoid an entire weight read; sequence masking usually cannot. Which wins depends on tensor shapes, implementation and hardware.

Timing samples characterize this machine and workload. They are not energy measurements or GPU speed forecasts. The displayed local benchmark measures forward plus backward, excluding optimizer steps, rather than the paper's full training system.

Sources: local timing samples and equivalence checks; paper, §6.2, p. 7.

11 / 16

Less arithmetic can leave weight traffic intact.

Cost model · not a hardware measurement

With independent sequence masks, a layer is needed by at least one sequence with probability 1 − pᴮ. For p = 0.5 and B = 32, that is 99.99999998%. Half the active work disappears, but almost every layer is still needed in every batch.

The speed model is 1/[f + (1−f)(1−p)], assuming perfect utilization and no new overhead. “Layer touched” is a probability, not measured DRAM bytes: caches, sharding and weight residency change actual transfers. Here p is a uniform rate, not the schedule's maximum rate.

For Sutro, record activation traffic, parameters and optimizer state, gather/scatter distance, and peak scratch occupancy separately. Layer dropout changes the executed graph; it does not automatically shrink stored weights or optimizer state.

Sources: paper, §6.2; Dally, 2022, cost and location of computation; Sutro meeting #30. Equations are independent cost-model calculations.

12 / 16

A fast draft is useful only if enough tokens survive.

Layer dropout makes a subset of the model a better candidate drafter. The full model verifies the draft. Correct rejection correction preserves the full target distribution; ordinary layer skipping or early exit does not.

Draft γ tokens→Verify with full model→Accept a prefix + one token

Illustrative stationary, independent acceptance model: expected output tokens = Σᵢ₌₀^γ αⁱ; speedup = that sum / (γ·draft cost + verify cost). Real acceptance is correlated and verification cost depends on hardware and cache behavior. This calculator is not a fit to Table 4.

The reported speedups follow searching draft layer subsets, with Table 4 evaluated on XSUM at draft length five. Search cost, workload transfer, inference hardware, batch size, timing uncertainty and serving conditions need a fuller report before adoption. A lossless decoder also preserves any quality gap already present in the dropout-trained target.

Sources: paper, §8.2.2 and Table 4; Leviathan et al., 2023; Draft & Verify, 2024.

13 / 16

Make the dropout mask a locality decision.

The most interesting extension is to optimize the bytes and distance avoided by a skipped block, alongside the accuracy it costs.

The Sutro group document frames the goal as energy-efficient learning through joint algorithm and hardware design. Meeting #30 proposes an MNIST challenge with streamed input, a 2D scratch grid, and Pareto objectives for energy, time, area and scoring time.

A useful hypothesis

Share masks across groups of nearby sequences or microbatches. Larger groups reduce opportunities for independent masking, but can avoid whole weight transfers and remote gathers. Search the group size jointly with depth and time schedules.

A complementary comparison

Checkpointing recomputes activations; reversible networks reconstruct them; dropout omits selected work and changes the training objective. Compare their movement–accuracy frontiers rather than treating all memory reduction as equivalent.

An honest boundary

Dropout retains backpropagation and full persistent model state. It may improve an incumbent learning algorithm without addressing the more radical question of whether a different learner can avoid its communication pattern.

Training plus deployment

Score total cost as Etrain + Esearch + Q·Einfer. A model that spends slightly more training energy can win over its lifetime if reduced-depth inference pays back across enough queries. For the one-shot MNIST task, Q must be the actual test workload, not hypothetical future deployment.

Related primary work: Chen et al., sublinear-memory training; Gomez et al., reversible networks; Raposo et al., Mixture-of-Depths. Grouped-mask locality is a proposed experiment, not a demonstrated result.

14 / 16

A small experiment program with clear failure conditions.

  1. Reconcile the implementation with Eq. 6.Compare shared subbranch scaling, independent masks, and one aggregate block mask. Check output moments and gradients on fixed activations before training. A recipe may remain useful after the equality claim is corrected.
  2. Optimize an actual cost–accuracy frontier.On the frozen MNIST task, compare dense full and smaller models, per-sequence, grouped and per-batch masks. Use identical data order, equal hyperparameter search, and no tuning on test labels; unlabeled transductive access follows the frozen rules. Plot energy proxy, time and area at matched accuracy.
  3. Measure movement on the same computation.Instrument executed reads/writes and placement. Count gather/scatter, parameters, optimizer state and retained activations. Publish the scorer version and execution trace summary; do not replace the official geometry or metric.
  4. Test the paper's scale claim directly.Add a dense 8.2B model with matched data and tuning, repeated runs at the decisive points, and longer token budgets. Distinguish equal steps, equal active FLOPs, equal wall time and equal energy.
  5. Validate the inference operating point.Select draft layers on held-out development prompts. Evaluate acceptance and latency across domains, batch sizes and output lengths. Include search and adapter cost in an amortized training-plus-inference ledger.

The MNIST proposal above is not an official challenge submission. The supplied meeting notes describe a specification still being finalized; use an approved version before reporting a challenge score.

15 / 16

Worth reproducing. Not yet a universal default.

What is significant

A substantial controlled sweep turns a familiar regularizer into a plausible compute and deployment strategy. Whole-block masks, depth-dependent dropout, decreasing schedules and careful scaling form a concrete recipe worth testing. Depth robustness is a credible secondary benefit.

What needs correction

The unqualified expectation identity fails for the preferred shared-mask equations. Some prose exceeds the tables: small-model “wins,” largest-model accuracy preservation, and generic inference guarantees need narrower wording.

What needs stronger evidence

Same-size largest-model controls, training wall time and joules, broader inference measurements, reproducibility details, and longer-budget accuracy. The schedule is a promising empirical choice, not a proven universal optimum.

The connection worth pursuing

Sutro can ask a sharper question than “did FLOPs fall?”: which masking pattern avoids the most physical movement at a fixed level of predictive quality? That experiment could favor a different granularity from this paper.

Assessment: a useful empirical contribution whose practical idea is stronger than parts of its explanation. The local counterexample limits the mathematical claim; the local measurements reveal implementation costs; neither substitutes for an LLM-scale reproduction.

16 / 16

Sources and provenance.

Paper claims, reported numeric values, independent derivations and measured toy results are kept distinct. Public primary sources are linked below. The linked Google document was read for context; its raw contents are not republished.

  1. Elhoushi, M. et al. Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference. arXiv:2609.05275v1, 4 Sep 2026; extended version of an ICML 2026 paper, per the arXiv record. Full 27-page PDF; §§4–11 and Appendices A–C.
  2. Huang, G. et al. Deep Networks with Stochastic Depth. ECCV, 2016.
  3. Fan, A., Grave, E. and Joulin, A. Reducing Transformer Depth on Demand with Structured Dropout. ICLR, 2020.
  4. Zhang, M. and He, Y. Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping. NeurIPS, 2020.
  5. Panigrahi, A. et al. Efficient Stagewise Pretraining via Progressive Subnetworks. 2024 preprint; ICLR, 2025.
  6. Elhoushi, M. et al. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding. ACL, 2024.
  7. Dey, N. S. et al. Don't Be Lazy: CompleteP Enables Compute-Efficient Deep Transformers. NeurIPS, 2025.
  8. Liu, H., Bauer, J. and Manning, C. D. Drop Dropout on Single Epoch Language Model Pretraining. Findings of ACL, 2025.
  9. Leviathan, Y., Kalman, M. and Matias, Y. Fast Inference from Transformers via Speculative Decoding. ICML, 2023.
  10. Zhang, J. et al. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding. ACL, 2024.
  11. Raposo, D. et al. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models. 2024.
  12. Chen, T. et al. Training Deep Nets with Sublinear Memory Cost. 2016.
  13. Gomez, A. N. et al. The Reversible Residual Network: Backpropagation Without Storing Activations. NeurIPS, 2017.
  14. Dally, W. J. On the Model of Computation: Point: We Must Extend Our Model of Computation to Account for Cost and Location. Communications of the ACM, 65(9), pp. 30–32, September 2022.
  15. Sutro #30 — MNIST Unchained. Meeting notes, 7 Sep 2026. Context and proposed benchmark, not a hardware specification.
  16. Sutro Group: top level. Read 9 Sep 2026 through the connected document service. Context: energy, data movement and joint algorithm/hardware design.

Reproducibility: scripts and raw outputs · machine-readable results · extended review. No paper figures are reproduced. Diagrams and plots are original, and analytical simulations are labeled.

1 / 16