Skip to results
GRADIENT DISSENTCode & raw runs ↗

Independent A100 experiment · verified final measurements

Does depth robustness survive a change of architecture?

Pruning lowers mean ViT CE for all recipes, most for dense training; ConvNeXt's primary CE intervals include zero.

We train GPT, a vision transformer, and a convolutional network with two layer-dropout recipes and a dense baseline. The test is whether removing a fixed fraction of residual blocks causes less damage—and whether the retained model is actually better.

3 architectures27 final runs3 paired seeds per comparison9 fixed interventions per model27 validation-only tuning runs

Scientific interpretation

Useful pruned models do transfer beyond GPT in these experiments, but the conclusion depends on the architecture, deletion pattern and metric. For both image architectures, the predefined two-thirds-depth models have higher mean accuracy under ILD than under dense training. The predefined relative-CE endpoint tells a more qualified story: GPT clearly favors ILD; ConvNeXt has large observed gains with wide intervals; ViT's early-exit comparison reverses because pruning improves the dense model's CE more. All intervals discussed here use three paired training seeds and are unadjusted for multiple comparisons; accuracy and additional masks are secondary outcomes.

ConvNeXt supplies the strongest practical image-classification contrast. Retaining 12 of 18 residual blocks reduces dense accuracy from 51.34% to 14.19%. Constant ILD retains 51.01% from a 54.62% full model, and decreasing ILD retains 50.62% from 54.12%. Nevertheless, the primary CE-damage effects are −4.795 [−10.705, 1.116] and −4.815 [−10.913, 1.283] nats: all three seed differences favor ILD, but both intervals include zero. The secondary paired pruned-accuracy gains are +36.82 [13.55, 60.10] and +36.43 [9.60, 63.26] percentage points. These are substantial observed benefits, with considerable uncertainty about their magnitude. ILD pruning slightly improves CE while reducing accuracy by about 3.5 points, illustrating why the two metrics should remain separate. Learned LayerScale stage means reach roughly 0.04–0.16 from an initialization of 10⁻⁶; residual scales remaining near their initial values cannot explain this result, although gamma alone does not measure branch importance.

ViT exposes a limitation of interpreting relative pruning damage as absolute predictor quality. Dense CE falls from 3.079 to 2.244 when retaining the first eight blocks; constant ILD goes from 2.713 to 2.155, and decreasing ILD from 2.561 to 2.209. Thus ILD receives less CE improvement from pruning, producing primary effects of +0.278 [0.226, 0.330] and +0.483 [0.389, 0.578]. Its pruned models still have lower mean CE and higher accuracy: 50.57% and 48.77%, versus dense 46.68%. Paired pruned-CE intervals include zero; secondary accuracy gains are +3.89 [2.04, 5.75] and +2.09 [1.52, 2.67] points. Dense ViT's large CE improvement accompanies only a 0.36-point aggregate accuracy decline. This is consistent with confidence or generalization problems, but calibration, logit magnitudes and individual prediction agreement were not measured. Low final minibatch losses alongside much larger held-out losses support concern about generalization, rather than an explanation that the vision models never learned.

The deletion pattern changes the ViT conclusion. With the three predefined random masks that retain the first block and eight blocks in total, dense accuracy averages only 17.52–24.15%; constant ILD gives 44.90–46.75%, and decreasing ILD 35.37–42.23%. These masks also favor both ILD recipes on relative CE damage, with all six secondary unadjusted intervals below zero. Prefix exit and intermediate deletion therefore probe different behavior at the same retained depth. Conversely, deleting the first block—a deletion never sampled by ILD training—leaves ViT at roughly 1% classification accuracy and GPT at only 1–3% token accuracy, and reduces ConvNeXt accuracy to 24.93%/17.85% under ILD, versus 37.56% for dense. The recipe does not create robustness to arbitrary missing blocks.

The 124M-parameter GPT arm reproduces the pruning effect with a clear full-model tradeoff. Prefix-eight CE damage falls from 4.199 nats for dense training to 0.093 with constant ILD and 0.116 with decreasing ILD; the paired effects are −4.106 [−5.390, −2.823] and −4.083 [−5.364, −2.802]. However, full-model CE worsens from 3.512 to 3.645/3.736, and token accuracy falls by 1.47/2.42 points. Robust subnetworks came with lower full-model quality at this budget. Each GPT run saw only 0.843 token presentations per parameter, so these measurements do not settle the paper's much longer training regime.

Decreasing-dropout training ends with useful subnetworks, but a universal advantage for that schedule does not appear. Constant ILD has better mean full and primary-pruned CE in GPT and ConvNeXt. Decreasing ILD has ViT's best full CE, while constant ILD has better prefix-eight CE and accuracy; at the more aggressive four-block exit, the ViT accuracy ranking changes again. These schedule comparisons are descriptive. Equal average masking does not isolate temporal order, because maximum probabilities and noise distributions also differ. The final step is dense, with no prolonged dense continuation. Stochastic depth in CNNs, LayerDrop and growing-subnetwork methods already provide precedents; the contribution here is a controlled, limited comparison of this recipe across the specified settings.

The next decisive checks would repeat vision training in a regime with a smaller train–held-out gap, add validation-only calibration diagnostics, and include directly trained smaller dense controls. Eight of nine learning-rate winners lie at a grid boundary, only one tuning seed was used, and both vision models share CIFAR-100. Separate initialization audits permit only explicitly reproduced numerical variants for the vision models; their tiny differences do not establish identical trajectories or negligible final-score effects. These limits belong to this experiment, not to the paper's reported runs. Training executed every branch before masking, so this study measures robustness and predictive quality, not saved training work, latency or energy.

Download the interpretation and scientific-input digest

Pruning damage and useful prediction are separate questions.

Primary effect = (CEtwo-thirds − CEfull)dropout − (CEtwo-thirds − CEfull)dense

Diamonds are paired means; small dots are individual training-seed differences. Bars are 95% Student-t intervals with 2 degrees of freedom. Six primary intervals are shown across this report, without multiplicity adjustment. Seed intervals omit dataset and learning-rate-selection uncertainty.

Keep the intact model and the pruned model in view.

RecipeFull CEPrimary CEFull accuracyPrimary accuracyFull CE vs dense
paired effect [95% CI]
Primary CE vs dense
paired effect [95% CI]

A smaller pruning penalty can coexist with worse predictions. Absolute full and pruned losses are therefore reported alongside robustness. For a CE difference, negative favors the dropout-trained model. When both full-to-pruned CE changes are negative, a positive paired effect means a smaller CE improvement, not worse absolute pruned CE. We do not interpret an interval crossing zero as equivalence. Symmetric t intervals are untruncated; accuracy bounds outside 0–100% are not observed accuracies.

All six primary effects and their three paired seed values
Architecture / dataRecipePaired effect [95% CI]Three seed differencesWhat the interval says
ViT · CIFAR-100Constant ILD+0.2777 [+0.2259, +0.3296]+0.2805, +0.2557, +0.2971The interval is entirely above zero: a higher CE change on pruning than with dense training.
ViT · CIFAR-100Decreasing ILD+0.4833 [+0.3887, +0.5780]+0.5228, +0.4469, +0.4803The interval is entirely above zero: a higher CE change on pruning than with dense training.
ConvNeXt · CIFAR-100Constant ILD-4.7947 [-10.7050, +1.1156]-6.0409, -6.2919, -2.0513The interval includes zero: the direction is unresolved by these three seeds; this is not an equivalence result.
ConvNeXt · CIFAR-100Decreasing ILD-4.8148 [-10.9129, +1.2833]-6.2250, -6.2391, -1.9802The interval includes zero: the direction is unresolved by these three seeds; this is not an equivalence result.
GPT · WikiText-103Constant ILD-4.1062 [-5.3895, -2.8228]-4.6689, -3.9963, -3.6533The interval is entirely below zero: a lower CE change on pruning than with dense training.
GPT · WikiText-103Decreasing ILD-4.0830 [-5.3643, -2.8016]-4.6454, -3.9715, -3.6320The interval is entirely below zero: a lower CE change on pruning than with dense training.

Nine interventions, fixed before final evaluation.

The same original readout is used for every retained network. No fine-tuning, adapters, output calibration, or inverse-survival rescaling is applied after pruning. These are nine specified masks, not an exhaustive subset search.

Visible training recipes

Interventions are categorical. Equal retained counts can remove different blocks; no interpolation between masks is implied. A negative change from full depth is allowed. Accuracy is token top-1 for GPT and image-class top-1 for vision; raw values should not be averaged across families.

Numeric table for every displayed point
InterventionRecipeKept blocksMean [95% CI]Three seed measurements

Inspect the retained blocks.

RecipeRaw CE [95% CI]CE change [95% CI]Accuracy % [95% CI]Accuracy change pp [95% CI]

Transfer across model families, under a declared budget.

ViT · CIFAR-100

A 12-block, width-768 ViT-style classifier, with 4×4 patches and mean patch pooling. There is no class token. The primary intervention keeps the first eight blocks and the original final readout. This changes architecture and modality relative to GPT, but shares CIFAR-100 with the CNN experiment.

NVIDIA A100-SXM4-40GB · BF16 autocast; FP32 master parameters/AdamW states; TF32 allowed

ConvNeXt · CIFAR-100

A ConvNeXt-style CNN with 18 residual blocks in stages of 3, 3, 9, 3. The primary mask retains stage prefixes of 2, 2, 6, 2; downsampling transitions remain. This is stagewise thinning, not early exit. Native 32×32 CIFAR images are resized to 64×64; interpolation adds no observed information.

NVIDIA A100-SXM4-40GB · BF16 autocast; FP32 master parameters/AdamW states; TF32 allowed

GPT · WikiText-103

A randomly initialized, tied-head GPT-2-style decoder with 12 blocks and width 768, trained on WikiText-103 with GPT-2 tokens. 0.843 token presentations per parameter is far below the reviewed paper's 20-token-per-parameter principal regime. Fixed document-contained test windows are not canonical full-corpus perplexity.

NVIDIA A100-SXM4-40GB · BF16 autocast; FP32 master parameters/AdamW states; TF32 allowed

Here, “transfer” means an effect replicating across settings; no pretrained weights transfer between families. The two image architectures use the same CIFAR-100 dataset and split. They test transfer across architecture, not replication across two independent vision datasets. All models are trained from scratch; GPT here is a short-run architectural baseline, not a reproduction of the paper's large-language-model pretraining.

FamilyParametersSteps × batchTraining exposure per runInputFixed test panel
ViT · CIFAR-10085,219,68410,240 × 2562,621,440 image presentations
58.25 equivalent passes with replacement
32 × 32 pixels10,000 images
ConvNeXt · CIFAR-10027,897,02810,000 × 5125,120,000 image presentations
113.78 equivalent passes with replacement
64 × 64 pixels10,000 images
GPT · WikiText-103124,439,8083,200 × 32104,857,600 token presentations
0.843 tokens / parameter
1,024-token context251,904 target tokens
246 actual windows

Token and image counts are presentations, not unique examples. Training samples with replacement. “Equivalent passes” means presentations divided by training-set size, with replacement; it does not mean shuffled epochs or guaranteed complete passes. The 10,000 official CIFAR test images are kept out of tuning; GPT uses the actual recorded document-contained test windows and target count shown above. A requested maximum is only a cap; valid spans can be fewer.

Two recipes with the same expected masking rate

Increasing layer dropout (ILD) drops later blocks more often and always retains the first. Constant ILD uses maximum probability 0.4. Decreasing ILD starts at maximum 0.8 and linearly reaches zero. Both schedules imply a 20% expected masking rate for example-block residual contributions, averaged over depth and training steps; realized Bernoulli masks vary. Decreasing ILD reaches exactly one zero-dropout terminal step, with no extended dense continuation.

These schedules differ in maximum probability and variance as well as temporal order. Their comparison does not isolate time ordering alone. Training uses compute-then-mask execution, so this expected 20% masks residual contributions after execution; it is not saved execution or a measured 20% compute reduction.

Matched comparisons, independently tuned recipes

Within each family, recipes share training exposure, initialization seeds, data streams, optimizer family and evaluation panels. Each selects a learning rate using full-depth validation CE only, with equal grids and a separate tuning seed. No test or pruned metric chooses the learning rate.

The result compares the complete recipe plus its chosen learning rate. Final three-seed intervals condition on the selection made from one tuning seed; they do not include search uncertainty.

Learning-rate grids, selections and boundary diagnostics
FamilyRecipeLearning-rate gridSelected LRGrid boundary?Tuning seedsTraining horizon
ViT · CIFAR-100Dense0.0001, 0.0003, 0.0010.0001yes1000matched
ViT · CIFAR-100Constant ILD0.0001, 0.0003, 0.0010.0003no1000matched
ViT · CIFAR-100Decreasing ILD0.0001, 0.0003, 0.0010.0001yes1000matched
ConvNeXt · CIFAR-100Dense0.0003, 0.001, 0.0030.0003yes1000matched
ConvNeXt · CIFAR-100Constant ILD0.0003, 0.001, 0.0030.0003yes1000matched
ConvNeXt · CIFAR-100Decreasing ILD0.0003, 0.001, 0.0030.0003yes1000matched
GPT · WikiText-103Dense0.0003, 0.0006, 0.00120.0012yes1000matched
GPT · WikiText-103Constant ILD0.0003, 0.0006, 0.00120.0012yes1000matched
GPT · WikiText-103Decreasing ILD0.0003, 0.0006, 0.00120.0012yes1000matched

Every candidate: CE and accuracy can favor different rates.

Full-depth validation scores below are means over the listed tuning seeds. The lowest validation CE determines the selected learning rate; top-1 accuracy is a diagnostic and never overrides that fixed rule. Accuracy can improve while CE worsens. These are tuning scores, not final test results.

FamilyRecipeCandidate LRValidation CE (nats)Validation top-1 accuracySelected by CE?
ViT · CIFAR-100Dense0.00013.012347.60%yes — lowest CE
ViT · CIFAR-100Dense0.00033.141748.18%no
ViT · CIFAR-100Dense0.0013.365518.80%no
ViT · CIFAR-100Constant ILD0.00012.668248.86%no
ViT · CIFAR-100Constant ILD0.00032.619550.36%yes — lowest CE
ViT · CIFAR-100Constant ILD0.0013.346719.30%no
ViT · CIFAR-100Decreasing ILD0.00012.517548.58%yes — lowest CE
ViT · CIFAR-100Decreasing ILD0.00032.814848.68%no
ViT · CIFAR-100Decreasing ILD0.0013.226321.84%no
ConvNeXt · CIFAR-100Dense0.00033.372353.22%yes — lowest CE
ConvNeXt · CIFAR-100Dense0.0013.593453.08%no
ConvNeXt · CIFAR-100Dense0.0033.958451.12%no
ConvNeXt · CIFAR-100Constant ILD0.00032.681856.54%yes — lowest CE
ConvNeXt · CIFAR-100Constant ILD0.0012.924057.82%no
ConvNeXt · CIFAR-100Constant ILD0.0032.958058.96%no
ConvNeXt · CIFAR-100Decreasing ILD0.00032.819656.64%yes — lowest CE
ConvNeXt · CIFAR-100Decreasing ILD0.0013.064457.40%no
ConvNeXt · CIFAR-100Decreasing ILD0.0033.193359.02%no
GPT · WikiText-103Dense0.00033.752236.26%no
GPT · WikiText-103Dense0.00063.598837.55%no
GPT · WikiText-103Dense0.00123.514438.33%yes — lowest CE
GPT · WikiText-103Constant ILD0.00033.971533.67%no
GPT · WikiText-103Constant ILD0.00063.713836.19%no
GPT · WikiText-103Constant ILD0.00123.679936.37%yes — lowest CE
GPT · WikiText-103Decreasing ILD0.00034.036533.09%no
GPT · WikiText-103Decreasing ILD0.00063.753135.90%no
GPT · WikiText-103Decreasing ILD0.00123.728535.98%yes — lowest CE

This larger experiment includes neither separately trained smaller dense models nor alternating-dropout training. It cannot establish that pruning beats training a smaller network directly, or settle the paper's distinct alternating-mask claim. Earlier toy controls do not replace these controls at A100 scale.

ConvNeXt: test the easy identity-path explanation.

ConvNeXt's residual LayerScale starts small. A network that barely uses its residual branches can appear robust to deleting them. Inspect the measured full-model learning and gamma magnitudes before attributing small pruning damage to useful learned redundancy. The residual contribution is γ × f(h), so gamma magnitude alone cannot identify how much the branch is used. Neither small gamma nor gamma growth establishes functional irrelevance or importance; direct residual-activation attribution is absent.

The original ConvNeXt implementation already uses depth-increasing stochastic depth. Applying layer dropout to a CNN is therefore an established idea. This experiment tests full/pruned behavior of these specified schedules, at this scale and geometry.

LayerScale before and after training, by stage
RecipeStageInitial mean |γ|Final mean |γ|Mean of per-seed stage maxima
Dense11e-060.08448340.260047
Dense21e-060.04024070.166386
Dense31e-060.08476160.265381
Dense41e-060.1436640.281691
Constant ILD11e-060.09733170.298273
Constant ILD21e-060.07755620.265521
Constant ILD31e-060.1203710.237378
Constant ILD41e-060.1589960.25658
Decreasing ILD11e-060.1043540.304274
Decreasing ILD21e-060.094910.275536
Decreasing ILD31e-060.1263960.240835
Decreasing ILD41e-060.1465790.217821

Values summarize absolute gamma within a stage and then across training seeds. The last column averages each seed's within-stage maximum; it is not the maximum across all runs.

Measured memory; an explicit cost ledger.

FamilyRecipeValidation CE
initial → final
Last minibatch CEPeak allocated GiBPeak reserved GiBTraining seconds
ViT · CIFAR-100Dense4.7318 → 3.06640.03516.666.96960.4
ViT · CIFAR-100Constant ILD4.7318 → 2.67450.15676.666.96976.0
ViT · CIFAR-100Decreasing ILD4.7318 → 2.53290.24976.6616.59976.1
ConvNeXt · CIFAR-100Dense4.7516 → 3.39760.00163.514.00432.4
ConvNeXt · CIFAR-100Constant ILD4.7516 → 2.67910.08293.514.00477.2
ConvNeXt · CIFAR-100Decreasing ILD4.7516 → 2.85850.00533.514.00443.9
GPT · WikiText-103Dense10.9857 → 3.53053.567230.3535.84941.6
GPT · WikiText-103Constant ILD10.9857 → 3.66563.760930.3535.84939.9
GPT · WikiText-103Decreasing ILD10.9857 → 3.75573.803030.3635.84943.2

Values above are means across the three final seeds. GPU peaks cover the recorded training window. Allocated memory measures live tensors, including resident data. Reserved memory can include allocator cache inherited from earlier calls or prevalidation; resetting peak counters does not clear that cache. Reserved values therefore do not establish a minimum VRAM requirement. Last-minibatch CE is noisy, not full-training-set loss. Full intervals and raw values are in the downloadable JSON.

These are execution and provenance measurements, not a controlled efficiency comparison. The masking implementation executes dense branches. The experiment establishes no reduction in FLOPs, wall time, GPU memory, energy, or dollars. No joules were measured.

$27.81Latest saved metered-usage snapshot
$39.67Conservative invocation reservations; not spend
$50.00User-authorized external-compute ceiling

Metered snapshot queried 2026-09-09 20:31 UTC; it may lag and is not a final invoice. The ledger records 60 invocation reservations, including the audit trail of preparation, pilots, tuning and final work. Reserved amounts are neither invoices nor measured charges.

Evidence you can inspect and reproduce.

All 27 declared final runs passed the analysis audit.

The generator refuses partial cells, failed audit flags, missing seed measurements, or inconsistent interval arithmetic. The upstream analysis checks paired initial states against exact matching or the separately disclosed vision numerical allowlists, matched data streams, completed training horizons, source and dataset hashes, fixed mask panels, GPU all-kept equivalence, equal learning-rate searches and reservation arithmetic.

Frozen source hashes and audit checks

train.py
50915ffcb0e0d0b40c55db02b31d3e91e9ad2013c127d187b8d2f96b033ebe4e

language.py
9aaa1230aa655fca3d18716f65885c67a545a51ba6991ab0310ab4c58215b42b

vision.py
edd014c05c041793025db1f64224de194ae0345c47b4e6720a631404f3848bf0

  • exact locked run specs and complete paired cells
  • no silently discarded successful final runs
  • bit-identical GPT initialization; separate exact allowlists and full numerical audits for ConvNeXt and ViT variants
  • identical paired data streams
  • executed core source matches published source
  • executed dataset manifests match preparation artifacts
  • identical heldout panels and masks
  • nominal20% mean omission for ILD
  • completed terminal steps
  • passed GPU full/allkeep equivalence
  • unique prescribed mask panel and fixed target counts
  • equal full-validation-only LR search
  • reservation ledger arithmetic/specs
What these results cannot establish
  • Few final training seeds; intervals are fragile and unadjusted across primary/secondary comparisons.
  • Learning-rate selection uncertainty and heldout dataset uncertainty are not included in seed intervals; tuning uses its recorded independent seed count.
  • ConvNeXt and ViT same-seed initial weights can differ by CPU host. Separate post hoc audits accept only two measured, numerically close hashes per architecture/seed with matching post-init RNG; identical training trajectories are not established. GPT initialization remains bit-identically paired.
  • Compute-then-mask executes dense branches; these runs do not establish FLOP, latency, memory, energy, or dollar savings.
  • Peak reserved GPU memory can include allocator cache from previous calls or prevalidation; it is not the model's required VRAM. Peak allocated memory describes live tensors in this implementation, including resident data.
  • Two-thirds retained residual blocks does not imply two-thirds FLOPs; ConvNeXt stage transitions always remain.
  • Fixed image resizing adds no observed information; CIFAR spatial geometry differs from ImageNet.
  • Small ConvNeXt LayerScale and undertraining can make deletion appear harmless; inspect full learning and recorded gamma magnitudes.
  • Language uses its recorded fixed document-contained context windows, not canonical full-corpus perplexity.
  • This is training from scratch under a short declared compute budget, not a reproduction of the target paper's LLM training scale.
Exportable three-panel primary effect figure with paired confidence intervals
Scientific figure export: paired treatment-minus-dense effects on the predefined intervention, with the three individual seed differences.
python experiments/a100_transfer/analyze.py
python scripts/build_a100_report.py
node scripts/check_a100_report.cjs

These commands analyze existing completed measurements and rebuild the report; they do not launch cloud training. All visualization code and assets are local, with no CDN dependency.

  1. Elhoushi et al., Don't Drop Dropout, September 2026 extended PDF. The empirical motivation, especially depth-elastic inference and layer/time schedules.
  2. Huang et al., Deep Networks with Stochastic Depth (2016) and Fan et al., Reducing Transformer Depth on Demand with Structured Dropout (2020). Existing CNN stochastic depth and transformer pruning without fine-tuning precede the reviewed paper.
  3. Interpretation audit written before final results. Scope, confounds and missing controls.
  4. Frozen experiment protocol. Exposures, schedule, tuning rules, primary endpoint and budget boundary.
  5. Locked final-run manifest. Every declared final seed and selected recipe configuration.
  6. Vision architecture and dataset provenance. ConvNeXt transitions, LayerScale, CIFAR preprocessing and ViT readout.
  7. Language architecture and dataset provenance. GPT-style initialization, WikiText103, tokenizer and document-contained windows.
  8. Critical paper review. Mathematical counterexamples, paper audit and energy/locality connections.
  9. Earlier local toy experiments. A different scale, dataset and mask enumeration; not pooled into this report's seed statistics.