{
  "scientific_input_sha256": "5bad33d925dd3eb17e43351a69239f8a661df59c0a61a2916a1079dbd3c7d64f",
  "paragraphs": [
    "Useful pruned models do transfer beyond GPT in these experiments, but the conclusion depends on the architecture, deletion pattern and metric. For both image architectures, the predefined two-thirds-depth models have higher mean accuracy under ILD than under dense training. The predefined relative-CE endpoint tells a more qualified story: GPT clearly favors ILD; ConvNeXt has large observed gains with wide intervals; ViT's early-exit comparison reverses because pruning improves the dense model's CE more. All intervals discussed here use three paired training seeds and are unadjusted for multiple comparisons; accuracy and additional masks are secondary outcomes.",
    "ConvNeXt supplies the strongest practical image-classification contrast. Retaining 12 of 18 residual blocks reduces dense accuracy from 51.34% to 14.19%. Constant ILD retains 51.01% from a 54.62% full model, and decreasing ILD retains 50.62% from 54.12%. Nevertheless, the primary CE-damage effects are −4.795 [−10.705, 1.116] and −4.815 [−10.913, 1.283] nats: all three seed differences favor ILD, but both intervals include zero. The secondary paired pruned-accuracy gains are +36.82 [13.55, 60.10] and +36.43 [9.60, 63.26] percentage points. These are substantial observed benefits, with considerable uncertainty about their magnitude. ILD pruning slightly improves CE while reducing accuracy by about 3.5 points, illustrating why the two metrics should remain separate. Learned LayerScale stage means reach roughly 0.04–0.16 from an initialization of 10⁻⁶; residual scales remaining near their initial values cannot explain this result, although gamma alone does not measure branch importance.",
    "ViT exposes a limitation of interpreting relative pruning damage as absolute predictor quality. Dense CE falls from 3.079 to 2.244 when retaining the first eight blocks; constant ILD goes from 2.713 to 2.155, and decreasing ILD from 2.561 to 2.209. Thus ILD receives less CE improvement from pruning, producing primary effects of +0.278 [0.226, 0.330] and +0.483 [0.389, 0.578]. Its pruned models still have lower mean CE and higher accuracy: 50.57% and 48.77%, versus dense 46.68%. Paired pruned-CE intervals include zero; secondary accuracy gains are +3.89 [2.04, 5.75] and +2.09 [1.52, 2.67] points. Dense ViT's large CE improvement accompanies only a 0.36-point aggregate accuracy decline. This is consistent with confidence or generalization problems, but calibration, logit magnitudes and individual prediction agreement were not measured. Low final minibatch losses alongside much larger held-out losses support concern about generalization, rather than an explanation that the vision models never learned.",
    "The deletion pattern changes the ViT conclusion. With the three predefined random masks that retain the first block and eight blocks in total, dense accuracy averages only 17.52–24.15%; constant ILD gives 44.90–46.75%, and decreasing ILD 35.37–42.23%. These masks also favor both ILD recipes on relative CE damage, with all six secondary unadjusted intervals below zero. Prefix exit and intermediate deletion therefore probe different behavior at the same retained depth. Conversely, deleting the first block—a deletion never sampled by ILD training—leaves ViT at roughly 1% classification accuracy and GPT at only 1–3% token accuracy, and reduces ConvNeXt accuracy to 24.93%/17.85% under ILD, versus 37.56% for dense. The recipe does not create robustness to arbitrary missing blocks.",
    "The 124M-parameter GPT arm reproduces the pruning effect with a clear full-model tradeoff. Prefix-eight CE damage falls from 4.199 nats for dense training to 0.093 with constant ILD and 0.116 with decreasing ILD; the paired effects are −4.106 [−5.390, −2.823] and −4.083 [−5.364, −2.802]. However, full-model CE worsens from 3.512 to 3.645/3.736, and token accuracy falls by 1.47/2.42 points. Robust subnetworks came with lower full-model quality at this budget. Each GPT run saw only 0.843 token presentations per parameter, so these measurements do not settle the paper's much longer training regime.",
    "Decreasing-dropout training ends with useful subnetworks, but a universal advantage for that schedule does not appear. Constant ILD has better mean full and primary-pruned CE in GPT and ConvNeXt. Decreasing ILD has ViT's best full CE, while constant ILD has better prefix-eight CE and accuracy; at the more aggressive four-block exit, the ViT accuracy ranking changes again. These schedule comparisons are descriptive. Equal average masking does not isolate temporal order, because maximum probabilities and noise distributions also differ. The final step is dense, with no prolonged dense continuation. Stochastic depth in CNNs, LayerDrop and growing-subnetwork methods already provide precedents; the contribution here is a controlled, limited comparison of this recipe across the specified settings.",
    "The next decisive checks would repeat vision training in a regime with a smaller train–held-out gap, add validation-only calibration diagnostics, and include directly trained smaller dense controls. Eight of nine learning-rate winners lie at a grid boundary, only one tuning seed was used, and both vision models share CIFAR-100. Separate initialization audits permit only explicitly reproduced numerical variants for the vision models; their tiny differences do not establish identical trajectories or negligible final-score effects. These limits belong to this experiment, not to the paper's reported runs. Training executed every branch before masking, so this study measures robustness and predictive quality, not saved training work, latency or energy."
  ]
}
