GRADIENT DISSENT

Gradient dissent.

Extended critical review, experiments, and connections to energy-efficient learning.

Technical review

Critical technical review: Don't Drop Dropout

Paper: Mostafa Elhoushi et al., Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference, arXiv:2609.05275v1, 4 September 2026. PDF, paper page. The arXiv record identifies this as a slightly extended version of an ICML 2026 paper. Reviewed 9 September 2026. Page references below are printed PDF pages; all 27 pages, including the appendix, were read. Key equations and tables were also checked in rendered pages.

Assessment: A useful empirical recipe study with credible evidence for better training/inference tradeoffs in its tested model family. The most valuable result is that structured depth noise can be annealed away while leaving useful pruning robustness. The evidence does not establish a general prescription for frontier LLMs, an across-the-board accuracy improvement, or 25% end-to-end training acceleration. Several mathematical and reporting statements need correction. The empirical results can remain useful despite those problems.

This audit separates reported measurements, exact consequences of the printed equations, and hypotheses requiring new experiments. Local toy experiments elsewhere in this repository are mechanism checks, not reproductions of the authors' 2400+ LLM-scale runs.

What the recipe actually is

Each training sequence receives a random subset of transformer blocks. Attention and FFN in a block share the same Bernoulli mask. Within a retained block, the paper's Eq. 6 scales both residual branches by inverse survival probability, 1/rho, where rho = 1-p. A CompleteP depth multiplier is applied in addition. The preferred dropout rate increases linearly with layer index and decreases linearly with training progress:

p(l,t) = p_max * l/(L-1) * (1-t/(T-1)).

Thus expected active depth starts at (1-p_max/2)L and reaches L. Average omitted nonembedding block FLOPs equal p_max/4 under equal-cost blocks, constant work per step, and actual conditional execution. In particular, p_max=0.99 means 24.75% average block FLOPs omitted, not 99% of total training omitted. The last block initially survives only 1% of sequences, but average initial model depth is still 50.5% of full depth. [Paper §§4-7, Eqs. 6, 8-9, pp. 4-10; Table A.2, p. 24.]

There are two distinct goals. One is to preserve full-depth model quality using fewer active block computations. The other is to train a full model whose submodels remain useful. The best schedule for one need not be best for the other: the paper itself shows alternating dropout is better for some intermediate skipping patterns, while increasing dropout across depth is better for early exit.

What the measurements support

Result Reported evidence Defensible interpretation
Decreasing dropout beats constant/increasing dropout Table 3, p. 10: matched nominal average sparsity across 271M, 503M, 906M Strong within this model family and search grid; the differences between schedules are much larger than several claimed differences from dense training.
1.8B full-depth quality Table 5, p. 15: loss 1.849 dense vs 1.836 dropout, 15% block FLOPs saved A promising improvement of 0.013 cross-entropy, or 0.70% relative CE; statistical certainty is unavailable without seed variation.
3.9B full-depth quality Table 5: loss 1.732 dense vs 1.745 dropout, 20% block FLOPs saved An explicit quality/compute tradeoff: +0.013 CE, +0.75% relative CE, about +1.31% perplexity. It is not an equal-token improvement.
8.2B largest run Table 5: dropout only, loss 1.663, 24.75% nominal block FLOPs saved Demonstrates a large run can train successfully. There is no matched dense 8.2B result in the supplied PDF, so preserved baseline quality at the headline saving is unmeasured.
Depth robustness Table 5, Fig. 5, Figs. A.2-A.3 Large, persuasive reduction in degradation relative to pruning dense-trained models. This remains a quality tradeoff relative to the intact dropout-trained model.
Self-speculative acceleration Tables 4-5, pp. 14-15; XSUM explicitly specified for Table 4, draft length 5 Useful reported speedups after searching layer subsets. Not a generic guarantee across workloads, batch sizes, or hardware.
Transfer beyond 20 TPP Fig. 8, p. 15 The displayed sweep is a 503M model over 2-64 TPP. This is modest evidence for duration transfer, not a fitted trillion-token scaling law.

The low-dropout 906M Table 3 result is particularly small: displayed losses 1.953 dense and 1.951 dropout, with a printed relative delta of -0.06% based on unrounded values. At 503M, both displayed losses are 2.110 and the printed delta is +0.03%, so the §7.2 assertion that both 503M and 906M beat dense at 5% savings contradicts the table for 503M. Do not turn these results into a strong statistical claim without repeated seeds.

Downstream performance is mixed. The model called 1.8B in Table 5 is called 1.9B in Table A.3. Assuming these rows are intended to match, the dropout model loses on 9 of 11 listed tasks, ties 2, and wins none, despite its improved validation CE. This does not establish significant degradation on each task, because uncertainty estimates are absent, but it contradicts a casual inference that better CE necessarily improved capabilities. For 3.9B the task comparison is 7 wins and 4 losses, including OBQA 0.356 to 0.324. The 8.2B task row again lacks a dense control. [Table A.3, p. 25.]

Connections to earlier work and the actual contribution

Primary source What it already established What this paper adds or must distinguish
Huang et al., Stochastic Depth, ECCV 2016 Training randomly shortened residual networks and evaluating the deep model; training-time savings and improved generalization in vision. Modern causal LM recipe and scale exploration. Random block omission and depth-dependent survival are inherited ideas.
Fan et al., LayerDrop, ICLR 2020 Structured transformer dropout enables shallow subnetworks; §6 compares sublayer and whole-layer dropping. §3.2.2 uses alternating layers as an inference pruning rule. A larger causal-LM granularity study. Claims of being the first to compare whole layers and sublayers need narrowing. Its alternating training distribution differs from LayerDrop's alternating inference rule.
Zhang & He, Progressive Layer Dropping, NeurIPS 2020 Transformer training acceleration from increasingly dropping more layers during training; published abstract reports 25% FLOP and 24% wall-clock savings on BERT. Evidence that the opposite temporal direction works better for this causal-LM setting. Training savings without accuracy collapse were already known.
Panigrahi et al., Efficient Stagewise Pretraining via Progressive Subnetworks / RaPTr, 2024 Random subnetworks grow progressively to the full model; depth and width variants; explicit critique of increasing dropout; BERT and UL2-1.6B experiments; theory for learning complexity and stage transitions. This is the closest omitted precursor. The present paper studies a continuous depth/time schedule, per-sequence granularity, parameterization, and inference elasticity in causal LMs. The general growing-subnetwork principle is not new.
Elhoushi et al., LayerSkip, ACL 2024 Depth-increasing layer dropout plus shared early-exit loss; self-speculation by drafting early and verifying with remaining layers; experiments in Llama models. Removes the need for auxiliary early-exit loss during the main pretraining recipe and measures additional granularity/schedule choices. The inference motivation and broad mechanism are continuous with LayerSkip.
Zhang et al., Draft & Verify, ACL 2024 Layer-subset drafting and full-model verification, without further neural training; measured gains on Llama-2. Dropout pretraining can improve the space of useful draft subsets. It is not the invention of lossless self-speculation.
Dey et al., CompleteP, NeurIPS 2025 Width/depth parameterization for hyperparameter transfer and complete feature learning; residual branch scaling with inverse depth. Treats retained depth as another parameterization axis and empirically favors inverse survival scaling. That is a plausible extension, not a complete theorem about arbitrary dropout.
Bergsma et al., Power Lines, NeurIPS 2025 AdamW weight-averaging timescale tau=B/(eta*lambda*D) and empirical data/batch scaling. Supplies much of the optimizer-transfer machinery; Table A.2 attempts to combine these rules with dropout but has a data-exponent inconsistency discussed below.
Hillier et al., STLM Engineering Report: Dropout, 2024 Decreasing activation-dropout schedules on small LMs; its §2 applies masks to embedding and attention outputs. Schedule motivation transfers, but activation dropout is a different intervention from skipping blocks.
Liu et al., Drop Dropout on Single-Epoch Language Model Pretraining, ACL Findings 2025 Activation dropout on attention/MLP outputs worsens several single-epoch LM evaluations, including early dropout. The new result does not refute this finding: structured layer omission changes computation and inductive bias differently. Both findings can hold simultaneously.

The strongest fair contribution is an empirical reconciliation and practical combination: choose per-sequence whole-block masks, a depth-increasing distribution, a time-decreasing schedule, and suitable residual/optimizer scaling, then measure both full-model and submodel quality. It is valuable that a familiar regularizer has a use even when avoiding classic overfitting is not the goal.

Critical issue 1: the expectation argument fails for the preferred block implementation

Confidence: high; exact algebra from the printed equations. Scope: the justification, not a claim that the reported experiments are false.

At the end of §5 (p. 5), the paper says r_eval=1 ensures equality of evaluation activations and expected training activations. There is a valid local statement for Eq. 2 when the whole residual branch f(h) is held fixed:

E_M[h + (M/rho) f(h) | h] = h + f(h).

But the preferred transformer configuration is different: Eq. 6 scales attention and FFN separately and §6.1 sets their masks equal. Consider one scalar transformer-shaped block with fixed input x, linear attention a*x, linear FFN b*z, and shared M ~ Bernoulli(rho):

z_train = x + (M/rho)*a*x
y_train = z_train + (M/rho)*b*z_train
        = [1 + (M/rho)*(a+b) + (M/rho)^2*a*b]*x
E[y_train] = [1+a+b+a*b/rho]*x

y_eval = [1+a+b+a*b]*x
E[y_train]-y_eval = a*b*(1-rho)/rho*x

This fails even with linear sublayers and a deterministic block input. For a=b=x=1, rho=0.5, evaluation gives 4 and the training expectation gives 5. The reason is that the FFN input depends on the same mask multiplying its output; one cannot replace each mask independently by its expectation.

Independent attention/FFN masks recover equality in this linear example. Another locally unbiased construction is to evaluate the dense composite block and scale its aggregate residual once. Neither implies global equality in an arbitrary nonlinear deep network: generally E[f(H)] != f(E[H]). Moreover, matching logits in expectation would still not match expected cross-entropy.

The actual implementation may differ from the simplified equations; the supplied PDF does not resolve that ambiguity. A useful author clarification is precisely where inverse survival scaling is applied. This distinction can change the comparison between whole-layer and sublayer dropout, since the two then differ in both mask correlation and effective within-block interaction strength.

Critical issue 2: hyperparameter transfer is empirical and incomplete

Confidence: high for the limited evidence; medium for extrapolated failure mechanisms.

The coordinate checks in Fig. 2 and Appendix B measure activations for ten steps of a 40-layer model. The plotted y-axis is mean absolute activation, while Appendix B describes a Frobenius norm. Stable activation magnitude alone does not verify the stated maximal residual stream update condition, gradient noise, update direction, or long-run optimal learning rate. Fig. 3 provides useful evidence for a few learning-rate/batch/AdamW-timescale sweeps, but §11 explicitly acknowledges weaker transfer at high dropout and that transfer was mainly tested with constant schedules.

For a single inverted Bernoulli multiplier Z=M/rho, E[Z]=1, E[Z^2]=1/rho, and Var(Z)=(1-rho)/rho. Thus rho=0.01 preserves the first moment while creating multiplier variance 99 and a retained multiplier 100. Mean activation stability is compatible with rare large updates. Per-sequence sampling and large batches can soften this, but do not turn it into an exact effective-depth equivalence. This is a reason to inspect update norms, quantiles, gradient clipping, and optimizer-state dynamics; it is not proof that the large run was unstable.

A second issue is the printed weight-decay transfer rule. Table A.2 says:

B = B_base * m_D^0.4
eta_hidden = eta_base / m_d
D = D_base * m_D
tau = tau_base * (TPP/TPP_base)^(-0.5)

Substituting into the Power Lines definition lambda=B/(eta*tau*D) gives:

lambda_hidden = B_base*m_d / (eta_base*tau*D_base*m_D^0.6).

Table A.2 instead prints m_D^0.4 in the denominator. The resulting ratio is m_D^0.2; at m_D=16, the printed rule gives approximately 1.74 times the weight decay implied by the other rows. This appears to be a table error or an undocumented change to the scaling rule. It cannot establish what the code actually ran. [Paper Table A.2, p. 24; Power Lines §2.2 Eqs. 2 and 4.]

Critical issue 3: FLOP savings are narrower than training acceleration

Confidence: high.

Footnote 11 (p. 8) explicitly defines subsequent FLOPs to mean nonembedding FLOPs. Equations 8-9 assume skipping a proportion of identical-cost block executions saves that proportion of block work. They omit unchanged vocabulary projection, softmax, embeddings, optimizer work, routing/gather/scatter overhead, and communication. The first-page graphic nevertheless says up to 25% faster training; there is no accompanying end-to-end training timing table in the supplied PDF.

If q is the fraction of dense compute in unchanged components and s is the block-compute saving, then ideal total-compute saving is (1-q)s, before added implementation overhead. For illustration, with unchanged work 20% and block saving 25%, total compute falls 20% and ideal compute-bound speedup is 1/0.8=1.25x. This is illustrative arithmetic, not a measurement of the authors' system.

Per-sequence dropout is especially sensitive to implementation. Masking the output after computing every sequence saves no block arithmetic. Gathering active sequences can save arithmetic but reduce matrix dimensions and utilization. The authors discuss compute-bound conditions in §6.2; their statement that equal FLOPs should yield similar speedup is a conditional expectation, not a benchmark. Hardware-specific performance on CS-3 need not transfer to a GPU cluster.

There is also a small exact-budget issue for alternating dropout: all three small models have odd layer counts, 13, 17, and 23. The formula is p_mean=floor(L/2)*p_max/L, not exactly p_max/2. With p_max=0.2, nominal 10% sparsity is actually 9.23%, 9.41%, and 9.57%. Table 2 groups these with exact 10% settings. This can slightly favor ALD by giving it more compute; it probably cannot explain every schedule effect, but matters for very small differences.

Critical issue 4: matched-cost curves are not a tuned compute frontier

Confidence: high that the control is not demonstrated; medium that it changes the conclusion.

Figure 9 compares loss at equal cumulative block FLOPs along runs trained to the same token endpoint. A dropout run has seen more tokens and reached a later fraction of its scheduled training at a fixed FLOP count. That is a real operational advantage, but a dense checkpoint from a longer run may not have received an appropriate terminal learning-rate cooldown for the smaller comparison budget.

The clean experiment trains a dense control with the same total FLOP budget, a schedule ending at that budget, separately tuned hyperparameters, and repeated seeds. It should compare that control with (a) the dropout run and (b) alternative dense model sizes trained to completion. The supplied PDF does not show that complete reoptimization. Accordingly, Fig. 9 supports useful trajectory efficiency, not a proved global Pareto frontier.

Likewise, 20 tokens per nominal parameter is a conventional dense-model reference point, not proof that stochastic-depth models remain compute-optimal there. The optimum allocation between stored parameters, active depth, and tokens can change. [Paper §§3, 7, 9-10; Fig. 9; compare Hoffmann et al., Training Compute-Optimal Large Language Models.]

Critical issue 5: inference claims conflate different comparisons

Confidence: high.

Static skipping is approximate inference. In Table 5 the 3.9B dropout model's full CE is 1.745. Early exit at 75% depth produces CE 2.143, about 48.9% higher perplexity; alternating-layer skipping produces 2.129, about 46.8% higher perplexity. These are much better than pruning a dense-trained model, but not negligible differences from the intact dropout model. Even the 8.2B 75%-depth exit increases perplexity about 12.1% from 1.663 to 1.777.

Correct speculative verification can preserve the chosen target model's distribution (or greedy output under the applicable algorithm), as established by the cited speculative-decoding literature. It does not restore the quality of an independently trained dense baseline. If dropout changed target quality, lossless verification preserves that changed target. [Paper §8.2.2; Draft & Verify.]

Section 8.2.2 searches draft layer subsets using Bayesian optimization, genetic algorithms, hill climbing, and simulated annealing, then selects the highest speedup. The supplied PDF does not give enough detail about inference hardware, batch size, precision, timing methodology, search budget, or separation of search/evaluation examples to independently reproduce Tables 4-5. A held-out evaluation after identical search budgets is needed; selection bias is a possibility, not an established error.

Key takeaway 4 and the conclusion describe the 1.55x figure as zero-shot speedup, although it arises from the paper's own post-training/search-based self-speculative category. It can be weight-frozen without being an unconditional, zero-configuration property. The §10 claim that dropout is a prerequisite for useful self-speculation is too strong: Draft & Verify already measured gains without such training. Also, a speedup of 1.02x is a small improvement, not the regression described in §10.

Finding 6 says average dropout predicts both early exit and layer-skipping robustness. Average dropout is a useful scalar exposure measure, but the paper's ALD/ILD comparison shows it is not sufficient: the location of missing layers and the inference pattern matter. A practical predictor should include per-layer exposure and the inference mask, not just the mean.

Critical issue 6: the scaling extrapolation is not established

Confidence: high.

Section 9 discusses predicting trillion-token behavior, but provides no fitted loss surface, exponent estimates, prediction interval, or held-out large-budget scaling test. Figure 8 displays only the 503M family across 2-64 TPP, while the largest reported runs remain at 20 TPP. Many architecture and deployment axes remain untested: MoE, alternative normalization/positional schemes, longer contexts, heavily overtrained models, and inference robustness after SFT or continued training. The authors acknowledge several of these limitations.

The 8.2B model proves feasibility at p_max=0.99, not monotone improvement of tolerated dropout with scale. Model width, depth, total parameters, dataset size, and chosen maximum dropout all change together in Table 5. Separate factorial sweeps are required to isolate what governs the tolerance.

Reproducibility is limited by the supplied PDF: §3 calls the data a diverse natural-language/code corpus without enough mixture details or a public artifact reference; Appendix A gives architecture and symbolic transfer rules rather than a complete numeric recipe, seeds, optimizer settings, context lengths, and validation protocol. There is no consolidated manifest for the claimed 2400+ runs. This review does not assert that these artifacts exist nowhere; it states that they are not supplied with enough specificity in the reviewed PDF.

Additional reporting checks

These are secondary to the scientific issues above.

Highest-value next experiments

  1. Resolve Eq. 6 scaling placement. Compare shared-mask/per-sublayer scaling, independent masks, and shared-mask/aggregate-block scaling. Match expected block work and optimize each recipe equally. Log evaluation/Monte Carlo mean discrepancies and full/pruned quality.
  2. Retune dense fixed-budget controls. End both LR schedules at the same actual compute budget and repeat across seeds. Compare a smaller dense model as well as the same nominal architecture.
  3. Measure systems costs. Separate active block FLOPs, total arithmetic, tokens/s, accelerator-hours, optimizer/communication overhead, and end-to-end time to target loss. Benchmark actual skipped execution against dense masking.
  4. Probe noise and updates. At high p_max, record per-layer survival counts, gradient/update norm quantiles, clipping events, optimizer moments, and weight decay under skipped steps. Test whether per-sequence masks help through independent averaging.
  5. Test retention and deployment tradeoffs. Measure early exit and noncontiguous skipping after a dense tail, SFT, and long continued pretraining. Evaluate searched draft subsets on held-out domains, batch sizes, and generation lengths.
  6. Add missing scaling controls. Train a matched dense 8.2B model, vary depth and width independently at fixed parameter/compute budgets, and test high TPP with a predeclared predicted loss gap.

Questions the authors could answer concretely

The work is worth following up because its large pruning-robustness gains and consistent schedule ablation have practical value. The strongest current conclusion is that well-configured stochastic depth can be a useful training/inference tradeoff in this family of causal LMs. Its larger claims need stronger controls, reproducibility details, and corrected mathematical exposition.

Local experiments

Local experimental findings

Source: Don't Drop Dropout, arXiv:2609.05275v1, especially Eqs. 2, 4, 6–9 and Sections 5–7. All results below were run locally; they are not values copied from the paper. The reproducible script, raw data, environment, and limitations are in experiments/. No LLM-scale replication or energy measurement was attempted.

1. Mean preservation is conditional; the shared mask creates a stronger counterexample

For a single residual update, with a fixed incoming activation, inverted scaling correctly gives

EM[x+Mf(x)/ρ∣x]=x+f(x),ρ=1−p. \mathbb E_M[x + M f(x)/\rho \mid x] = x + f(x),\qquad \rho=1-p.

That statement does not imply equality between a deterministic evaluation network and the mean of a composed stochastic network. Nonlinear transformations do not generally commute with expectation. More strikingly, the paper recommends sharing the same mask between attention and FFN (Eq. 6, Section 6.1). Even affine attention and FFN then give a discrepancy within a single transformer-style block:

z=x+(M/ρ)ax,y=z+(M/ρ)bz. z=x+(M/\rho)ax,\quad y=z+(M/\rho)bz.

Since M2=MM^2=M ,

E[ytrain]=(1+a+b+ab/ρ)x,yeval=(1+a+b+ab)x. \mathbb E[y_{train}]=(1+a+b+ab/\rho)x,\qquad y_{eval}=(1+a+b+ab)x.

The bias is abpx/ρabp x/\rho . Our exact enumeration with x=1,a=.5,b=.4,p=.4x=1,a=.5,b=.4,p=.4 gives expected training output 2.233333 versus evaluation output 2.100000. The isolated first branch still has mean exactly 1.500000, matching its evaluation output. Thus this is not a failure of the usual isolated-branch dropout identity.

With independent masks, the affine example becomes unbiased. Replace the second branch by f2(z)=bz2f_2(z)=bz^2 , however, and independent masks produce mean 2.466667 versus deterministic 2.400000. The exact bias is ba2x2p/ρba^2x^2p/\rho .

Interpretation: the unqualified assertion at the end of Section 5 that evaluation features equal expected training features needs a conditional, isolated-residual-branch qualification. The recommended shared mask makes the distinction particularly consequential. This exact counterexample does not invalidate the empirical observation that inverted scaling helps hyperparameter transfer or model quality.

2. An unbiased branch prediction does not imply an unbiased dense-objective gradient

For the one-branch scalar model y=x+(M/ρ)axy=x+(M/\rho)ax and half-MSE target tt ,

E[12(y−t)2]=12((1+a)x−t)2+12pρ(ax)2. \mathbb E\left[\tfrac12(y-t)^2\right] =\tfrac12((1+a)x-t)^2+\tfrac12\frac p\rho (ax)^2.

The expected gradient therefore includes (p/ρ)ax2(p/\rho)ax^2 , even though the mean prediction is exactly unbiased. At x=1,a=.5,t=.7,p=.4x=1,a=.5,t=.7,p=.4 , the added penalty is .083333, and the expected gradient with respect to aa is 1.133333 versus the dense gradient .800000. For the shared two-sublayer example above, the corresponding gradients are 3.925926 and 1.960000.

This is an exact dropout-induced objective change, not merely zero-mean noise added to the dense gradient. In more general models, a Taylor expansion expresses analogous effects through the loss Hessian and feature-noise covariance. Connecting the paper to gradient-noise or curvature discussions should retain both changed objective/mean-gradient terms and covariance terms. The paper does not claim a general equivalence of objectives; this experiment identifies a common overinterpretation of its scaling argument.

3. Skipped work, elapsed time, and optimizer updates are separate questions

The local CPU benchmark retains exactly half of the sequence-block or batch-block work. It uses six residual GELU MLP blocks, float32, one thread, precomputed balanced masks, four warmups, and nine randomized-order repetitions of sixteen executions. These are forward-plus-backward medians; the optimizer is excluded.

Execution Small tensor [8,16,32] Speedup Medium tensor [32,32,64] Speedup
Dense .3663 ms 1.000× 2.5431 ms 1.000×
Per-sequence, compute then mask .4280 ms .856× 2.7204 ms .935×
Per-sequence, gather active sequences .3810 ms .962× 1.6455 ms 1.545×
Per-batch, true whole-block skip .2241 ms 1.635× 1.3219 ms 1.924×

The dimensions are [batch, sequence length, hidden width]. Each MLP expands hidden width by two. Forward-only and batch-gather controls are included in the raw results. Forward and input/parameter gradients match to floating-point tolerance between implementations using the same masks, including all-dropped and all-kept layers. The batch and sequence mask patterns are different models, so they are not asserted to have identical outputs.

Interpretation: compute-then-mask does all block matmuls and adds overhead. Actual sequence gathering becomes useful for the larger local shape but fails to deliver a speedup for the small shape. Batch skipping is cheaper here, while the paper's accuracy results favor sequence masks. This is a concrete quality/implementation tradeoff and supports the paper's own distinction between masking and actual sparse execution. It neither measures nor disputes Cerebras speedups. No inference from these measurements to joules is justified.

There is also an optimizer boundary detail. After one identical scalar AdamW update with lr=.1 and weight decay=.1, the parameter is .890000. On a dropped step, supplying an explicit zero gradient moves it to .814094, because momentum/weight decay still act. Supplying grad=None leaves it at .890000 under the tested PyTorch AdamW semantics. Consequently, identical forward outputs and mathematical zero gradients do not guarantee identical training trajectories for a true-skip implementation. A replication needs an explicit policy for dropped-layer optimizer state and weight decay. This is an implementation requirement, not evidence about the unpublished details of the paper's training system.

4. A controlled regression toy separates optimization, budget, and depth elasticity

The task uses fresh iid Gaussian 8D inputs with a fixed nonlinear analytic teacher. A trainable tanh embedding feeds six residual 24-dimensional tanh-linear blocks and a scalar readout. A five-block dense model provides a smaller-architecture control. This is an online, noiseless toy problem, not an overfit finite training dataset and not a transformer.

All three dropout treatments have 20% mean dropout. Constant ILD uses maximum .4; increasing/decreasing schedules use maximum .8. The schedules exactly follow the paper's discrete endpoints, and the code verifies that ILD+DTS mean dropout equals pmax/4p_{max}/4 . Each configuration and budget regime gets five learning rates × two tuning seeds. All choose the interior rate .03. Final results use six separate evaluation seeds and a fixed 2,048-example test set. Entries below are mean MSE [95% Student-t interval across seeds].

Configuration Matched 240 steps Matched expected active sequence-block budget
Dense, six blocks .01925 [.01500,.02350] .01925 [.01500,.02350]
Constant ILD .02382 [.01970,.02794] .01583 [.01283,.01883]
Decreasing ILD .02674 [.02010,.03339] .01563 [.01303,.01823]
Increasing ILD .02708 [.02244,.03172] .01862 [.01657,.02066]
Dense, five blocks .02174 [.01036,.03311] .01486 [.00779,.02193]

At matched steps, all dropout treatments are worse than dense6; the paired 95% difference intervals exclude zero. At matched expected active work, dropout gets 300 steps and dense5 gets 288, compared with dense6's 240. All then have mean losses below dense6, but every paired interval against dense6 includes zero. For decreasing ILD, the difference is −.003618 with interval [−.007292,+.0000556]. It would be misleading to call this a decisive superiority result. Dense5 has the lowest mean and substantial uncertainty; excluding that control would overstate the case for layer dropout.

The training code computes all rows before applying masks. Its active sequence-block budget is a conceptual model of ideal skipped work, not measured training FLOPs or wall-clock compute. It excludes the embedding, head, optimizer, and routing overheads. The separate execution benchmark above tests the runtime consequences of real skipping.

The inference result is stronger within this toy. At the matched active budget, retaining only the first four blocks gives MSE .19204 for dense6, .02172 for constant ILD, .02168 for decreasing ILD, and .02351 for increasing ILD. The dense5 control at depth four has .06414. These numbers support a depth-robustness benefit in this task. Constant and decreasing schedules give almost identical depth-four means, so this experiment provides little reason to prefer the decreasing schedule for elasticity.

Interpretation: schedule rankings depend on the training regime and budget. Our toy does not establish a universal decreasing-schedule optimum, and it does not falsify the paper's empirical schedule comparisons on its own language-model family. It illustrates why “same steps,” “same active work,” “same elapsed time,” and “best smaller dense architecture” answer different questions. It also independently demonstrates the plausible depth-elasticity mechanism.

Limits and next experiments

The six-seed intervals characterize initialization, data-stream, and mask variation on one fixed teacher and test set; they do not capture variation over datasets or architectures. The LR grid is finite, only LR is independently retuned, and intervals are not adjusted for multiple comparisons. Neither dropout schedules nor CompleteP can be declared generally optimal from these experiments. Balanced benchmark masks intentionally remove random active-count fluctuations and use precomputed masks; real routing and distributed load imbalance add costs.

A stronger replication would use a small decoder-only transformer with the paper's shared attention/FFN mask, equal searches over LR/batch/weight decay, both zero-versus-absent-gradient optimizer policies, shared token streams, a tuned smaller dense baseline, repeated seeds, and actual end-to-end throughput at fixed validation loss. For the gyroscope connection, directly measure dense-versus-mask expected gradients, gradient covariance, residual magnitudes, and top curvature directions through the dropout schedule. That would distinguish annealed regularization, gradient-noise averaging, and residual-stream stabilization instead of treating them as the same mechanism.

Energy and Sutro

Layer dropout through Sutro's energy and locality lens

Research note, 9 September 2026. This note connects the supplied paper, Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference (arXiv:2609.05275v1, 4 September 2026), to the energy/data-movement mission in the linked Google document and the public Sutro #30 meeting notes. Page references below are the printed pages of the paper PDF. The meeting describes a proposed MNIST challenge; it is not a frozen benchmark specification. The proposed experiments below are research proposals, not reported results or official submissions.

The useful connection

Layer dropout is a useful, relatively conservative baseline for Sutro: make the existing learning algorithm do less work, then ask whether the work removed was physically expensive. It is not a replacement for backpropagation, and it does not establish that reducing FLOPs reduces joules proportionally. Its strongest connection to the group's agenda is precisely the tension between statistically good sparsity and physically cheap sparsity.

The paper itself identifies this tension in §6.2, p. 7. Independently dropping a layer for individual sequences improves loss in its experiments, whereas dropping a layer for an entire batch can avoid loading that layer's weights. The authors argue that similar FLOP savings should yield similar speedups when training is compute bound, and suggest per-device or per-microbatch masks as an unexplored middle ground. Sutro can test that conditional claim under an objective in which location and transport explicitly matter.

The interesting research question is therefore: At equal accuracy, which granularity and schedule of omitted computation minimizes data movement, peak live state, and total energy? The paper optimizes some of the statistical dimensions; a Sutro experiment can add the physical dimensions.

Five deductions worth testing

1. Half the active FLOPs can mean almost no reduction in weight loads

Let a layer be independently dropped with probability p for each of B examples. If its weights must be loaded whenever at least one example uses the layer, then

P(layer used by batch) = 1 − p^B.

This is an exact probability calculation under independent masks, not an energy measurement. With p = 0.5 and B = 32, half the example-layer computations disappear in expectation, but the layer is used with probability 0.999999999767. In contrast, a single mask shared by the entire batch uses the layer with probability 0.5.

Drop probability Batch size Expected active example fraction Probability weights are needed
0.2 32 0.8 effectively 1
0.5 32 0.5 0.999999999767
0.9 32 0.1 0.965663
0.99 32 0.01 0.275020

The distinction is particularly relevant when weight traffic dominates. With unchanged weight bytes and fewer examples processed per layer, arithmetic intensity against those bytes decreases. A model that was compute bound before dropout can move toward a bandwidth bottleneck afterward. That is not a contradiction of the paper; it limits the regime in which its compute-bound argument applies.

There are important qualifiers. The formula concerns whether weights are needed, not how many physical transfers occur. Cached or permanently local weights may require no expensive reload. Conversely, an optimizer may still touch an unused layer for momentum or decoupled weight decay. A faithful implementation must state what happens when a layer has no gradient; silently skipping its optimizer step can change the algorithm. Training framework semantics, physical placement, and optimizer traffic belong in the experiment manifest.

2. Activation savings and model-state savings are different resources

An actually bypassed function does not need its internal forward intermediates for that example's backward pass. This can reduce saved activations. Multiplying a computed output by zero does not achieve that benefit: the costly function has already executed. The paper explicitly distinguishes these implementations in §4, p. 4.

The model's weights remain allocated, and adaptive optimizer state generally remains allocated too. Thus a dropout network with P parameters still has an O(P) parameter-state floor even if its active computation is much smaller. A batch-wide zero mask may avoid touching state on a particular step; it does not by itself release that state for the rest of training. Per-sequence gathering can also allocate packed copies, index maps, and full-batch residual buffers, so saved-activation bytes should be instrumented rather than inferred from the dropout rate.

This is a direct link to the meeting's area objective. If area is represented by peak occupied scratch cells, a schedule that becomes dense near the end can still reach approximately the dense peak activation demand. Lower average active depth does not imply proportionally smaller peak area. A useful schedule comparison should therefore display both cumulative movement and maximum live state over training.

The distinction also motivates an optimizer control. Adafactor reduces second-moment storage for matrices through row/column statistics. That attacks persistent optimizer state, unlike layer dropout's omission of selected forward/backward computations. Its relevance is complementary, not evidence that combining them will preserve the paper's accuracy findings.

3. A gather is a physical algorithm, not a free indexing operation

The paper's computationally efficient formulation selects only active sequences. On hardware, that may require gathering those sequences into a compact batch, evaluating the block, and scattering updates back. The same logical mask can have different transport costs under different layouts.

A minimal accounting identity is

E_sparse − E_dense = −E_omitted_work + E_mask + E_pack/scatter + ΔE_other.

ΔE_other includes changes in memory hierarchy behavior, synchronization, utilization, and the optimizer. This identity provides a crossover criterion: sparse execution helps only when the omitted work saves more energy than the added overhead. It does not assign hypothetical joules to a CPU timing measurement.

For a grid model, separately accumulate Σ(bytes transferred × prescribed distance). Count the forward gather/scatter and the inverse movement of gradients. State whether selected examples remain packed between successive layers or return to a canonical layout each time. Keeping them packed can avoid some copies but may cause later masks to require a new permutation; the mapping itself must be represented and costed.

Proposed extension: share masks within physically nearby groups of examples, and sweep group size. If there are G independent groups, the probability a shared layer is used somewhere is 1 − p^G. Larger groups create more opportunities to omit an entire block of execution, but reduce mask diversity. Local weight replicas could improve transport while increasing occupied area. Neither grouped masks nor replication is automatically beneficial; this is a natural Pareto problem.

4. Skipping graph layers need not shorten physical distance

Two inference paths may execute the same number of layers but have different layouts. A static prefix exit omits a suffix; scattered layer skipping may keep communicating across the full physical depth of a pipeline. Bypassing arithmetic at a spatially assigned layer does not teleport the activation to the next active layer.

By contrast, a compiler that statically removes and repacks unused layers may reduce both model state and communication distance. Compare at least two deployment assumptions: resident elastic model with bypasses, and specialized compact model produced for a fixed skip pattern. They serve different use cases. For a single execution engine that streams weights, omitted weight loads may be more important than the distances between logical layers.

The paper measures static early exit, intermediate skipping, and self-speculation (§8, pp. 10–14). Those results support depth robustness in its models. They do not identify which placement of an elastic model minimizes byte-distance. A visual demonstration should show logical edges and physical positions separately.

5. The optimum dropout schedule depends on the lifetime workload

The paper trains with layer dropout and then normally evaluates the full model, with optional reduced-depth execution or self-speculative decoding. Its decreasing dropout schedule improves base-model loss, while its inference results show a different tradeoff across training schedules (Table 4, p. 14). Therefore “best schedule” needs a workload objective.

Use a lifecycle expression such as

E_total(N) = E_fit + E_adaptation/search + N × E_prediction.

Here N counts predictions or tokens using a common unit. A choice that costs more to prepare can win after enough deployment use. For two alternatives, the crossover is (E_fit,A − E_fit,B)/(E_prediction,B − E_prediction,A) when A has higher preparation cost and lower per-prediction cost. If either numerator or denominator has the opposite sign, interpret the comparison directly rather than reporting a spurious positive crossover.

The MNIST meeting proposes one end-to-end transductive computation from labeled training data and unlabeled test data to test labels. In that workload there is one concrete total cost; training and test inference cannot be amortized over a hypothetical infinite deployment. The boundary for reusable code, pretrained weights, development-time search, and instance-specific adaptation must come from the challenge rules. Report those costs separately until the boundary is frozen.

Static classification also has no direct analogue of the paper's lossless autoregressive self-speculative verification. An early-exit MNIST classifier can be useful, but its error/energy curve must be measured. Do not import a 1.55× text-decoding speedup into that task.

Connections to existing work

Checkpointing: more arithmetic can improve the physical objective

Chen et al., Training Deep Nets with Sublinear Memory Cost (2016) trade recomputation for reduced activation storage, with an O(sqrt(L)) checkpointing construction for an L-layer chain and roughly one additional forward pass under the paper's assumptions. This gives an essential control: dropout reduces executed computations; checkpointing can increase computation while reducing expensive storage/transport. A FLOP-only score can rank these in the opposite order from an energy score. Recompute stochastic masks consistently during checkpoint replay; drawing a new mask changes the function whose gradient is being computed.

Reversible networks: preserve backpropagation, change what must be stored

Gomez et al., The Reversible Residual Network (2017) reconstruct layer activations during backpropagation using a reversible architecture; storage for the reversible blocks does not grow with their depth. This is relevant to the group's concern about activation traffic, but reversibility does not eliminate parameter state, recomputation, nonreversible boundary operations, or numerical issues. The appropriate question is whether a reversible architecture shifts the same accuracy/energy/area frontier further than stochastic depth does.

Local learning: a more direct challenge to long-distance credit assignment

Mostafa, Ramesh and Cauwenberghs (2017) train layers with local errors from auxiliary random classifiers and explicitly discuss reducing custom-hardware communication. Nøkland and Eidnes (2019) develop supervised local losses for image learning. These methods target the dependence on errors arriving from later layers, whereas dropout still backpropagates through the active subnetwork. Include a local-loss baseline for MNIST, accounting for auxiliary classifiers, label transport, and any accuracy cost. “Local error” is not synonymous with zero communication or proven lower joules.

Mixture-of-Depths: an informative contrast, not an interchangeable baseline

Raposo et al., Mixture-of-Depths (2024) use learned top-k token routing with a fixed capacity. This creates known tensor sizes but adds routing and token selection. Unlike whole-sequence dropout, it also changes which tokens can participate in attention. Its stochastic-routing control performs poorly (§4.1); that does not refute the new paper because granularity, attention context, schedule, and tuning differ.

MoD's sequence-wide top-k decision is noncausal, requiring special handling at autoregressive inference (§3.5). A transductive MNIST workload can legally inspect the whole unlabeled test set if the frozen rules permit it. This creates an interesting opportunity for globally budgeted image routing without that particular causality problem. Its sorting, buffering, and transport costs still count. The proposal is to compare random, confidence-based, and learned allocation under equal total physical budgets, not to assume learned routing wins.

The meeting's claims also need an audit

The motivation for charging communication is sound; the illustrative numbers and statements should not become unexamined physical laws.

Proposed MNIST experiment protocol

  1. Pin the problem before claiming a benchmark result. Record the released specification version, exact dataset bytes and splits, permitted precision, input/output placement, allowed processors, distance metric, tape traversal rules, and scoring boundary. The September 7 notes explicitly leave I/O/multicore choices and resolution rungs open. Until a release is available, name the implementation an exploratory surrogate and publish its assumptions beside results.

  2. Start with comparable predictors. Use a small residual MLP or CNN that can actually bypass whole residual blocks. Compare tuned dense training; per-example dropout; batch-shared dropout; and spatial-group masks. Match architecture, initialization, data order, and optimizer family, while tuning learning rate for each configuration as the paper motivates. Include a smaller dense network at comparable executed work: “same architecture, fewer FLOPs” does not establish the optimal architecture at that budget. Add checkpointing, a reversible architecture, and a local-loss baseline only as separately specified comparisons.

  3. Separate the statistical experiment from the systems experiment. First establish accuracy and loss under equal examples seen and under equal executed arithmetic budgets. Then use identical saved masks/inputs to compare dense-zero masking, actual bypass, gather/scatter, and grouped layouts. Check forward values and gradients against the mathematically equivalent masked implementation, including empty-active-set behavior. Treat any deliberate optimizer change as its own ablation.

  4. Record three kinds of cost without mixing them. Record operation counts and byte-distance under the declared model; wall time and peak allocation on the local implementation; and energy only if measured by a documented instrument or estimated by explicitly stated coefficients. A Mac CPU timer is not a GPU power meter. Include masks, index arrays, packing buffers, gradients, optimizer state, intermediate predictions, test-set adaptation, and output transport within the declared boundary.

  5. Test the strongest locality hypothesis. At fixed expected active depth, sweep mask-sharing group size and batch size. Evaluate both contiguous placement and a declared alternative layout; compare static prefix exits with scattered skips before and after specialization/repacking. Log unique weights touched, persistent bytes, saved activation bytes, total transferred bytes, byte-distance, and peak scratch occupancy. This makes a failure of FLOPs-to-energy proportionality observable and explains it.

  6. Use an accuracy frontier, not one cherry-picked rate. Hold test labels out of configuration selection. Use validation splits and multiple paired seeds; publish run-level measurements and dispersion. Show accuracy against energy, time, peak area proxy, and scorer time, and mark nondominated configurations. Never compare one method at an easier target without displaying the difference. Where a frozen specification defines an accuracy threshold, report cost to reach that threshold with the stated success criterion.

  7. Price the scorer. Implement exact event tracing at the smallest rung and compare it with aggregated analytical accounting on the same executions. Publish agreement checks and scorer runtime. If costs are approximate at larger scales, label the approximation and validate its error. The meeting's time-to-score objective makes the compiler's ability to summarize repeated structured operations a research result in its own right.

Falsifiable hypotheses and useful outcomes

The paper supplies a strong question for Sutro: whether training a model to tolerate missing computation also makes it tolerate locally organized missing computation. Answering that requires an algorithm, layout, compiler, and cost model together—the co-design objective in Sutro's research agenda.

Core references

  1. Elhoushi, M. et al. Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference. arXiv:2609.05275v1, 4 Sep 2026; extended version of an ICML 2026 paper, per the arXiv record. Full 27-page PDF; §§4–11 and Appendices A–C.
  2. Huang, G. et al. Deep Networks with Stochastic Depth. ECCV, 2016.
  3. Fan, A., Grave, E. and Joulin, A. Reducing Transformer Depth on Demand with Structured Dropout. ICLR, 2020.
  4. Zhang, M. and He, Y. Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping. NeurIPS, 2020.
  5. Panigrahi, A. et al. Efficient Stagewise Pretraining via Progressive Subnetworks. 2024 preprint; ICLR, 2025.
  6. Elhoushi, M. et al. LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding. ACL, 2024.
  7. Dey, N. S. et al. Don't Be Lazy: CompleteP Enables Compute-Efficient Deep Transformers. NeurIPS, 2025.
  8. Liu, H., Bauer, J. and Manning, C. D. Drop Dropout on Single Epoch Language Model Pretraining. Findings of ACL, 2025.
  9. Leviathan, Y., Kalman, M. and Matias, Y. Fast Inference from Transformers via Speculative Decoding. ICML, 2023.
  10. Zhang, J. et al. Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding. ACL, 2024.
  11. Raposo, D. et al. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models. 2024.
  12. Chen, T. et al. Training Deep Nets with Sublinear Memory Cost. 2016.
  13. Gomez, A. N. et al. The Reversible Residual Network: Backpropagation Without Storing Activations. NeurIPS, 2017.
  14. Dally, W. J. On the Model of Computation: Point: We Must Extend Our Model of Computation to Account for Cost and Location. Communications of the ACM, 65(9), pp. 30–32, September 2022.
  15. Sutro #30 — MNIST Unchained. Meeting notes, 7 Sep 2026. Context and proposed benchmark, not a hardware specification.
  16. Sutro Group: top level. Read 9 Sep 2026 through the connected document service. Context: energy, data movement and joint algorithm/hardware design.

Additional primary sources are linked beside the corresponding arguments above. The private document is cited for context; its raw text is not included.