Independent A100 experiment · verified final measurements
Does depth robustness survive a change of architecture?
Pruning lowers mean ViT CE for all recipes, most for dense training; ConvNeXt's primary CE intervals include zero.
We train GPT, a vision transformer, and a convolutional network with two layer-dropout recipes and a dense baseline. The test is whether removing a fixed fraction of residual blocks causes less damage—and whether the retained model is actually better.
3 architectures27 final runs3 paired seeds per comparison9 fixed interventions per model27 validation-only tuning runs
Scientific interpretation
Useful pruned models do transfer beyond GPT in these experiments, but the conclusion depends on the architecture, deletion pattern and metric. For both image architectures, the predefined two-thirds-depth models have higher mean accuracy under ILD than under dense training. The predefined relative-CE endpoint tells a more qualified story: GPT clearly favors ILD; ConvNeXt has large observed gains with wide intervals; ViT's early-exit comparison reverses because pruning improves the dense model's CE more. All intervals discussed here use three paired training seeds and are unadjusted for multiple comparisons; accuracy and additional masks are secondary outcomes.
ConvNeXt supplies the strongest practical image-classification contrast. Retaining 12 of 18 residual blocks reduces dense accuracy from 51.34% to 14.19%. Constant ILD retains 51.01% from a 54.62% full model, and decreasing ILD retains 50.62% from 54.12%. Nevertheless, the primary CE-damage effects are −4.795 [−10.705, 1.116] and −4.815 [−10.913, 1.283] nats: all three seed differences favor ILD, but both intervals include zero. The secondary paired pruned-accuracy gains are +36.82 [13.55, 60.10] and +36.43 [9.60, 63.26] percentage points. These are substantial observed benefits, with considerable uncertainty about their magnitude. ILD pruning slightly improves CE while reducing accuracy by about 3.5 points, illustrating why the two metrics should remain separate. Learned LayerScale stage means reach roughly 0.04–0.16 from an initialization of 10⁻⁶; residual scales remaining near their initial values cannot explain this result, although gamma alone does not measure branch importance.
ViT exposes a limitation of interpreting relative pruning damage as absolute predictor quality. Dense CE falls from 3.079 to 2.244 when retaining the first eight blocks; constant ILD goes from 2.713 to 2.155, and decreasing ILD from 2.561 to 2.209. Thus ILD receives less CE improvement from pruning, producing primary effects of +0.278 [0.226, 0.330] and +0.483 [0.389, 0.578]. Its pruned models still have lower mean CE and higher accuracy: 50.57% and 48.77%, versus dense 46.68%. Paired pruned-CE intervals include zero; secondary accuracy gains are +3.89 [2.04, 5.75] and +2.09 [1.52, 2.67] points. Dense ViT's large CE improvement accompanies only a 0.36-point aggregate accuracy decline. This is consistent with confidence or generalization problems, but calibration, logit magnitudes and individual prediction agreement were not measured. Low final minibatch losses alongside much larger held-out losses support concern about generalization, rather than an explanation that the vision models never learned.
The deletion pattern changes the ViT conclusion. With the three predefined random masks that retain the first block and eight blocks in total, dense accuracy averages only 17.52–24.15%; constant ILD gives 44.90–46.75%, and decreasing ILD 35.37–42.23%. These masks also favor both ILD recipes on relative CE damage, with all six secondary unadjusted intervals below zero. Prefix exit and intermediate deletion therefore probe different behavior at the same retained depth. Conversely, deleting the first block—a deletion never sampled by ILD training—leaves ViT at roughly 1% classification accuracy and GPT at only 1–3% token accuracy, and reduces ConvNeXt accuracy to 24.93%/17.85% under ILD, versus 37.56% for dense. The recipe does not create robustness to arbitrary missing blocks.
The 124M-parameter GPT arm reproduces the pruning effect with a clear full-model tradeoff. Prefix-eight CE damage falls from 4.199 nats for dense training to 0.093 with constant ILD and 0.116 with decreasing ILD; the paired effects are −4.106 [−5.390, −2.823] and −4.083 [−5.364, −2.802]. However, full-model CE worsens from 3.512 to 3.645/3.736, and token accuracy falls by 1.47/2.42 points. Robust subnetworks came with lower full-model quality at this budget. Each GPT run saw only 0.843 token presentations per parameter, so these measurements do not settle the paper's much longer training regime.
Decreasing-dropout training ends with useful subnetworks, but a universal advantage for that schedule does not appear. Constant ILD has better mean full and primary-pruned CE in GPT and ConvNeXt. Decreasing ILD has ViT's best full CE, while constant ILD has better prefix-eight CE and accuracy; at the more aggressive four-block exit, the ViT accuracy ranking changes again. These schedule comparisons are descriptive. Equal average masking does not isolate temporal order, because maximum probabilities and noise distributions also differ. The final step is dense, with no prolonged dense continuation. Stochastic depth in CNNs, LayerDrop and growing-subnetwork methods already provide precedents; the contribution here is a controlled, limited comparison of this recipe across the specified settings.
The next decisive checks would repeat vision training in a regime with a smaller train–held-out gap, add validation-only calibration diagnostics, and include directly trained smaller dense controls. Eight of nine learning-rate winners lie at a grid boundary, only one tuning seed was used, and both vision models share CIFAR-100. Separate initialization audits permit only explicitly reproduced numerical variants for the vision models; their tiny differences do not establish identical trajectories or negligible final-score effects. These limits belong to this experiment, not to the paper's reported runs. Training executed every branch before masking, so this study measures robustness and predictive quality, not saved training work, latency or energy.
Diamonds are paired means; small dots are individual training-seed differences. Bars are 95% Student-t intervals with 2 degrees of freedom. Six primary intervals are shown across this report, without multiplicity adjustment. Seed intervals omit dataset and learning-rate-selection uncertainty.
Keep the intact model and the pruned model in view.
Recipe
Full CE
Primary CE
Full accuracy
Primary accuracy
Full CE vs dense paired effect [95% CI]
Primary CE vs dense paired effect [95% CI]
A smaller pruning penalty can coexist with worse predictions. Absolute full and pruned losses are therefore reported alongside robustness. For a CE difference, negative favors the dropout-trained model. When both full-to-pruned CE changes are negative, a positive paired effect means a smaller CE improvement, not worse absolute pruned CE. We do not interpret an interval crossing zero as equivalence. Symmetric t intervals are untruncated; accuracy bounds outside 0–100% are not observed accuracies.
All six primary effects and their three paired seed values
Architecture / data
Recipe
Paired effect [95% CI]
Three seed differences
What the interval says
ViT · CIFAR-100
Constant ILD
+0.2777 [+0.2259, +0.3296]
+0.2805, +0.2557, +0.2971
The interval is entirely above zero: a higher CE change on pruning than with dense training.
ViT · CIFAR-100
Decreasing ILD
+0.4833 [+0.3887, +0.5780]
+0.5228, +0.4469, +0.4803
The interval is entirely above zero: a higher CE change on pruning than with dense training.
ConvNeXt · CIFAR-100
Constant ILD
-4.7947 [-10.7050, +1.1156]
-6.0409, -6.2919, -2.0513
The interval includes zero: the direction is unresolved by these three seeds; this is not an equivalence result.
ConvNeXt · CIFAR-100
Decreasing ILD
-4.8148 [-10.9129, +1.2833]
-6.2250, -6.2391, -1.9802
The interval includes zero: the direction is unresolved by these three seeds; this is not an equivalence result.
GPT · WikiText-103
Constant ILD
-4.1062 [-5.3895, -2.8228]
-4.6689, -3.9963, -3.6533
The interval is entirely below zero: a lower CE change on pruning than with dense training.
GPT · WikiText-103
Decreasing ILD
-4.0830 [-5.3643, -2.8016]
-4.6454, -3.9715, -3.6320
The interval is entirely below zero: a lower CE change on pruning than with dense training.
Nine interventions, fixed before final evaluation.
The same original readout is used for every retained network. No fine-tuning, adapters, output calibration, or inverse-survival rescaling is applied after pruning. These are nine specified masks, not an exhaustive subset search.
Interventions are categorical. Equal retained counts can remove different blocks; no interpolation between masks is implied. A negative change from full depth is allowed. Accuracy is token top-1 for GPT and image-class top-1 for vision; raw values should not be averaged across families.
Transfer across model families, under a declared budget.
ViT · CIFAR-100
A 12-block, width-768 ViT-style classifier, with 4×4 patches and mean patch pooling. There is no class token. The primary intervention keeps the first eight blocks and the original final readout. This changes architecture and modality relative to GPT, but shares CIFAR-100 with the CNN experiment.
A ConvNeXt-style CNN with 18 residual blocks in stages of 3, 3, 9, 3. The primary mask retains stage prefixes of 2, 2, 6, 2; downsampling transitions remain. This is stagewise thinning, not early exit. Native 32×32 CIFAR images are resized to 64×64; interpolation adds no observed information.
A randomly initialized, tied-head GPT-2-style decoder with 12 blocks and width 768, trained on WikiText-103 with GPT-2 tokens. 0.843 token presentations per parameter is far below the reviewed paper's 20-token-per-parameter principal regime. Fixed document-contained test windows are not canonical full-corpus perplexity.
Here, “transfer” means an effect replicating across settings; no pretrained weights transfer between families. The two image architectures use the same CIFAR-100 dataset and split. They test transfer across architecture, not replication across two independent vision datasets. All models are trained from scratch; GPT here is a short-run architectural baseline, not a reproduction of the paper's large-language-model pretraining.
Family
Parameters
Steps × batch
Training exposure per run
Input
Fixed test panel
ViT · CIFAR-100
85,219,684
10,240 × 256
2,621,440 image presentations 58.25 equivalent passes with replacement
32 × 32 pixels
10,000 images
ConvNeXt · CIFAR-100
27,897,028
10,000 × 512
5,120,000 image presentations 113.78 equivalent passes with replacement
Token and image counts are presentations, not unique examples. Training samples with replacement. “Equivalent passes” means presentations divided by training-set size, with replacement; it does not mean shuffled epochs or guaranteed complete passes. The 10,000 official CIFAR test images are kept out of tuning; GPT uses the actual recorded document-contained test windows and target count shown above. A requested maximum is only a cap; valid spans can be fewer.
Two recipes with the same expected masking rate
Increasing layer dropout (ILD) drops later blocks more often and always retains the first. Constant ILD uses maximum probability 0.4. Decreasing ILD starts at maximum 0.8 and linearly reaches zero. Both schedules imply a 20% expected masking rate for example-block residual contributions, averaged over depth and training steps; realized Bernoulli masks vary. Decreasing ILD reaches exactly one zero-dropout terminal step, with no extended dense continuation.
These schedules differ in maximum probability and variance as well as temporal order. Their comparison does not isolate time ordering alone. Training uses compute-then-mask execution, so this expected 20% masks residual contributions after execution; it is not saved execution or a measured 20% compute reduction.
Matched comparisons, independently tuned recipes
Within each family, recipes share training exposure, initialization seeds, data streams, optimizer family and evaluation panels. Each selects a learning rate using full-depth validation CE only, with equal grids and a separate tuning seed. No test or pruned metric chooses the learning rate.
The result compares the complete recipe plus its chosen learning rate. Final three-seed intervals condition on the selection made from one tuning seed; they do not include search uncertainty.
Learning-rate grids, selections and boundary diagnostics
Family
Recipe
Learning-rate grid
Selected LR
Grid boundary?
Tuning seeds
Training horizon
ViT · CIFAR-100
Dense
0.0001, 0.0003, 0.001
0.0001
yes
1000
matched
ViT · CIFAR-100
Constant ILD
0.0001, 0.0003, 0.001
0.0003
no
1000
matched
ViT · CIFAR-100
Decreasing ILD
0.0001, 0.0003, 0.001
0.0001
yes
1000
matched
ConvNeXt · CIFAR-100
Dense
0.0003, 0.001, 0.003
0.0003
yes
1000
matched
ConvNeXt · CIFAR-100
Constant ILD
0.0003, 0.001, 0.003
0.0003
yes
1000
matched
ConvNeXt · CIFAR-100
Decreasing ILD
0.0003, 0.001, 0.003
0.0003
yes
1000
matched
GPT · WikiText-103
Dense
0.0003, 0.0006, 0.0012
0.0012
yes
1000
matched
GPT · WikiText-103
Constant ILD
0.0003, 0.0006, 0.0012
0.0012
yes
1000
matched
GPT · WikiText-103
Decreasing ILD
0.0003, 0.0006, 0.0012
0.0012
yes
1000
matched
Every candidate: CE and accuracy can favor different rates.
Full-depth validation scores below are means over the listed tuning seeds. The lowest validation CE determines the selected learning rate; top-1 accuracy is a diagnostic and never overrides that fixed rule. Accuracy can improve while CE worsens. These are tuning scores, not final test results.
Family
Recipe
Candidate LR
Validation CE (nats)
Validation top-1 accuracy
Selected by CE?
ViT · CIFAR-100
Dense
0.0001
3.0123
47.60%
yes — lowest CE
ViT · CIFAR-100
Dense
0.0003
3.1417
48.18%
no
ViT · CIFAR-100
Dense
0.001
3.3655
18.80%
no
ViT · CIFAR-100
Constant ILD
0.0001
2.6682
48.86%
no
ViT · CIFAR-100
Constant ILD
0.0003
2.6195
50.36%
yes — lowest CE
ViT · CIFAR-100
Constant ILD
0.001
3.3467
19.30%
no
ViT · CIFAR-100
Decreasing ILD
0.0001
2.5175
48.58%
yes — lowest CE
ViT · CIFAR-100
Decreasing ILD
0.0003
2.8148
48.68%
no
ViT · CIFAR-100
Decreasing ILD
0.001
3.2263
21.84%
no
ConvNeXt · CIFAR-100
Dense
0.0003
3.3723
53.22%
yes — lowest CE
ConvNeXt · CIFAR-100
Dense
0.001
3.5934
53.08%
no
ConvNeXt · CIFAR-100
Dense
0.003
3.9584
51.12%
no
ConvNeXt · CIFAR-100
Constant ILD
0.0003
2.6818
56.54%
yes — lowest CE
ConvNeXt · CIFAR-100
Constant ILD
0.001
2.9240
57.82%
no
ConvNeXt · CIFAR-100
Constant ILD
0.003
2.9580
58.96%
no
ConvNeXt · CIFAR-100
Decreasing ILD
0.0003
2.8196
56.64%
yes — lowest CE
ConvNeXt · CIFAR-100
Decreasing ILD
0.001
3.0644
57.40%
no
ConvNeXt · CIFAR-100
Decreasing ILD
0.003
3.1933
59.02%
no
GPT · WikiText-103
Dense
0.0003
3.7522
36.26%
no
GPT · WikiText-103
Dense
0.0006
3.5988
37.55%
no
GPT · WikiText-103
Dense
0.0012
3.5144
38.33%
yes — lowest CE
GPT · WikiText-103
Constant ILD
0.0003
3.9715
33.67%
no
GPT · WikiText-103
Constant ILD
0.0006
3.7138
36.19%
no
GPT · WikiText-103
Constant ILD
0.0012
3.6799
36.37%
yes — lowest CE
GPT · WikiText-103
Decreasing ILD
0.0003
4.0365
33.09%
no
GPT · WikiText-103
Decreasing ILD
0.0006
3.7531
35.90%
no
GPT · WikiText-103
Decreasing ILD
0.0012
3.7285
35.98%
yes — lowest CE
This larger experiment includes neither separately trained smaller dense models nor alternating-dropout training. It cannot establish that pruning beats training a smaller network directly, or settle the paper's distinct alternating-mask claim. Earlier toy controls do not replace these controls at A100 scale.
ConvNeXt: test the easy identity-path explanation.
ConvNeXt's residual LayerScale starts small. A network that barely uses its residual branches can appear robust to deleting them. Inspect the measured full-model learning and gamma magnitudes before attributing small pruning damage to useful learned redundancy. The residual contribution is γ × f(h), so gamma magnitude alone cannot identify how much the branch is used. Neither small gamma nor gamma growth establishes functional irrelevance or importance; direct residual-activation attribution is absent.
The original ConvNeXt implementation already uses depth-increasing stochastic depth. Applying layer dropout to a CNN is therefore an established idea. This experiment tests full/pruned behavior of these specified schedules, at this scale and geometry.
LayerScale before and after training, by stage
Recipe
Stage
Initial mean |γ|
Final mean |γ|
Mean of per-seed stage maxima
Dense
1
1e-06
0.0844834
0.260047
Dense
2
1e-06
0.0402407
0.166386
Dense
3
1e-06
0.0847616
0.265381
Dense
4
1e-06
0.143664
0.281691
Constant ILD
1
1e-06
0.0973317
0.298273
Constant ILD
2
1e-06
0.0775562
0.265521
Constant ILD
3
1e-06
0.120371
0.237378
Constant ILD
4
1e-06
0.158996
0.25658
Decreasing ILD
1
1e-06
0.104354
0.304274
Decreasing ILD
2
1e-06
0.09491
0.275536
Decreasing ILD
3
1e-06
0.126396
0.240835
Decreasing ILD
4
1e-06
0.146579
0.217821
Values summarize absolute gamma within a stage and then across training seeds. The last column averages each seed's within-stage maximum; it is not the maximum across all runs.
Measured memory; an explicit cost ledger.
Family
Recipe
Validation CE initial → final
Last minibatch CE
Peak allocated GiB
Peak reserved GiB
Training seconds
ViT · CIFAR-100
Dense
4.7318 → 3.0664
0.0351
6.66
6.96
960.4
ViT · CIFAR-100
Constant ILD
4.7318 → 2.6745
0.1567
6.66
6.96
976.0
ViT · CIFAR-100
Decreasing ILD
4.7318 → 2.5329
0.2497
6.66
16.59
976.1
ConvNeXt · CIFAR-100
Dense
4.7516 → 3.3976
0.0016
3.51
4.00
432.4
ConvNeXt · CIFAR-100
Constant ILD
4.7516 → 2.6791
0.0829
3.51
4.00
477.2
ConvNeXt · CIFAR-100
Decreasing ILD
4.7516 → 2.8585
0.0053
3.51
4.00
443.9
GPT · WikiText-103
Dense
10.9857 → 3.5305
3.5672
30.35
35.84
941.6
GPT · WikiText-103
Constant ILD
10.9857 → 3.6656
3.7609
30.35
35.84
939.9
GPT · WikiText-103
Decreasing ILD
10.9857 → 3.7557
3.8030
30.36
35.84
943.2
Values above are means across the three final seeds. GPU peaks cover the recorded training window. Allocated memory measures live tensors, including resident data. Reserved memory can include allocator cache inherited from earlier calls or prevalidation; resetting peak counters does not clear that cache. Reserved values therefore do not establish a minimum VRAM requirement. Last-minibatch CE is noisy, not full-training-set loss. Full intervals and raw values are in the downloadable JSON.
These are execution and provenance measurements, not a controlled efficiency comparison. The masking implementation executes dense branches. The experiment establishes no reduction in FLOPs, wall time, GPU memory, energy, or dollars. No joules were measured.
$27.81Latest saved metered-usage snapshot
$39.67Conservative invocation reservations; not spend
$50.00User-authorized external-compute ceiling
Metered snapshot queried 2026-09-09 20:31 UTC; it may lag and is not a final invoice. The ledger records 60 invocation reservations, including the audit trail of preparation, pilots, tuning and final work. Reserved amounts are neither invoices nor measured charges.
All 27 declared final runs passed the analysis audit.
The generator refuses partial cells, failed audit flags, missing seed measurements, or inconsistent interval arithmetic. The upstream analysis checks paired initial states against exact matching or the separately disclosed vision numerical allowlists, matched data streams, completed training horizons, source and dataset hashes, fixed mask panels, GPU all-kept equivalence, equal learning-rate searches and reservation arithmetic.
bit-identical GPT initialization; separate exact allowlists and full numerical audits for ConvNeXt and ViT variants
identical paired data streams
executed core source matches published source
executed dataset manifests match preparation artifacts
identical heldout panels and masks
nominal20% mean omission for ILD
completed terminal steps
passed GPU full/allkeep equivalence
unique prescribed mask panel and fixed target counts
equal full-validation-only LR search
reservation ledger arithmetic/specs
What these results cannot establish
Few final training seeds; intervals are fragile and unadjusted across primary/secondary comparisons.
Learning-rate selection uncertainty and heldout dataset uncertainty are not included in seed intervals; tuning uses its recorded independent seed count.
ConvNeXt and ViT same-seed initial weights can differ by CPU host. Separate post hoc audits accept only two measured, numerically close hashes per architecture/seed with matching post-init RNG; identical training trajectories are not established. GPT initialization remains bit-identically paired.
Compute-then-mask executes dense branches; these runs do not establish FLOP, latency, memory, energy, or dollar savings.
Peak reserved GPU memory can include allocator cache from previous calls or prevalidation; it is not the model's required VRAM. Peak allocated memory describes live tensors in this implementation, including resident data.
Two-thirds retained residual blocks does not imply two-thirds FLOPs; ConvNeXt stage transitions always remain.
Fixed image resizing adds no observed information; CIFAR spatial geometry differs from ImageNet.
Small ConvNeXt LayerScale and undertraining can make deletion appear harmless; inspect full learning and recorded gamma magnitudes.
Language uses its recorded fixed document-contained context windows, not canonical full-corpus perplexity.
This is training from scratch under a short declared compute budget, not a reproduction of the target paper's LLM training scale.
Scientific figure export: paired treatment-minus-dense effects on the predefined intervention, with the three individual seed differences.
These commands analyze existing completed measurements and rebuild the report; they do not launch cloud training. All visualization code and assets are local, with no CDN dependency.