# Ciresan-width MNIST: optimization results

Generated 2026-09-09T22:52:25.126458+00:00. **Search finalized; outcomes remain test-monitored.**

The target is the first observed **98.63% official-test accuracy (at most 137 errors out of 10,000)**. These runs fit the full 60,000-example training pool and monitor the test set after scheduled epochs. Recipe choices respond to test results: this is explicitly a test-target speed search, separate from the interrupted validation-only stochastic-depth study.

The public [historical W&B run](https://wandb.ai/yaroslavvb/train_ciresan/runs/ts4k9n55) first and uniquely reached 98.63% at epoch 94 and **388.4867 seconds of logged W&B runtime**, among 91 verified evaluations. Its GPU is unknown; its clock includes logging/evaluation and is not isolated training time. Historical source/clock gaps are documented in the [history audit](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/history/README.md).

## Selected recipes: all three confirmation seeds

The two selected linear-head, unit-dropout recipes are modified training recipes. They use B256, BF16 training, dropout 0.2, initial shrinkage 0.00008, and a factor 0.1 applied to both LR and shrinkage at epoch 21. Raw pixels use LR 0.004; normalized pixels (÷255) use LR 0.12. Seeds 101–103 differ from exploratory seed 1 but were also used in earlier recipe checks; the test set was already known and repeatedly monitored.

- **raw-linear-dropout-step: 3/3 reached the target**; observed training time 7.414–9.677s and invocation-to-score 15.807–20.551s. These ranges span the hardware listed below; they are not hardware-normalized recipe comparisons.
- **normalized-dropout-step: 3/3 reached the target**; observed training time 5.320–5.819s and invocation-to-score 14.897–17.236s. These ranges span the hardware listed below; they are not hardware-normalized recipe comparisons.

| Selected recipe | Seed | Reported GPU | Target epoch / accuracy | Training s | Invocation-to-score s | Whole invocation s | Dispatch-to-result s |
|---|---|---|---|---|---|---|---|
| [opt-confirm-raw-linear-dropout-step-s101](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-raw-linear-dropout-step-s101.json) | 101 | NVIDIA A100-SXM4-80GB | 32 / 98.64% | 7.413962 | 18.293905 | 18.556291 | 27.632757 |
| [opt-confirm-raw-linear-dropout-step-s102](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-raw-linear-dropout-step-s102.json) | 102 | NVIDIA A100-SXM4-40GB | 30 / 98.64% | 7.598947 | 15.807211 | 16.040970 | 24.587625 |
| [opt-confirm-raw-linear-dropout-step-s103](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-raw-linear-dropout-step-s103.json) | 103 | NVIDIA A100-SXM4-80GB | 42 / 98.63% | 9.676896 | 20.550714 | 20.838683 | 30.598894 |
| [opt-confirm-normalized-dropout-step-s101](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-normalized-dropout-step-s101.json) | 101 | NVIDIA A100 80GB PCIe | 24 / 98.67% | 5.608824 | 17.235739 | 17.491769 | 28.590364 |
| [opt-confirm-normalized-dropout-step-s102](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-normalized-dropout-step-s102.json) | 102 | NVIDIA A100-SXM4-40GB | 23 / 98.63% | 5.818916 | 15.592140 | 15.828836 | 25.146443 |
| [opt-confirm-normalized-dropout-step-s103](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-normalized-dropout-step-s103.json) | 103 | NVIDIA A100-SXM4-40GB | 21 / 98.71% | 5.320114 | 14.897264 | 15.167187 | 22.777005 |

Only seed 102 has the same reported GPU (A100-SXM4-40GB) for both selected recipes; the other two pairs mix SXM4/PCIe or 40/80GB variants. Keep the per-run numbers rather than attributing a pooled difference entirely to input normalization or LR.

**Confirmation check for raw-linear-dropout-step: 3/3 reached the target** by 100 epochs. Seeds: [101, 102, 103]; outcomes: {'reached': 3}. Same-seed repeats are counted separately; stable-recipe seeds were reused from earlier raw/ReLU checks. This supports repeatability across these tested seeds; it does not undo test-target selection.

**Confirmation check for normalized-dropout-step: 3/3 reached the target** by 100 epochs. Seeds: [101, 102, 103]; outcomes: {'reached': 3}. Same-seed repeats are counted separately; stable-recipe seeds were reused from earlier raw/ReLU checks. This supports repeatability across these tested seeds; it does not undo test-target selection.

**Confirmation check for frozen-b256-bf16: 1/5 reached the target** by 200 epochs. Seeds: [101, 102, 103, 104, 105]; outcomes: {'reached': 1, 'censored': 4}. Same-seed repeats are counted separately; stable-recipe seeds were reused from earlier raw/ReLU checks. The observed successes do not establish a robust recipe across these tested seeds.

## Kernel effects that are actually isolated

- `opt-speed-eager-tf32-b64-s1` → `opt-speed-fused-tf32-b64-s1`: same seed, recipe, initialization, dataset and reported NVIDIA A100-SXM4-40GB; 150 epochs take 248.926 → 106.164s of training (**2.345×**). All scored test trajectories identical: **True**. Target reached: reference=False, treatment=False. This measures fixed-horizon training speed, not a successful time-to-target ratio.
- `opt-speed-eager-tf32-b64-s1` → `opt-speed-graph-tf32-b64-s1`: same seed, recipe, initialization, dataset and reported NVIDIA A100-SXM4-40GB; 150 epochs take 248.926 → 122.238s of training (**2.036×**). All scored test trajectories identical: **True**. Target reached: reference=False, treatment=False. This measures fixed-horizon training speed, not a successful time-to-target ratio.

## Output-head diagnostic: digit 6

At seed 104, epoch 5, the paired [ReLU-head run](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-diagnostic-seed104-relu.json) predicts **zero sixes and correctly classifies 0/958 true sixes**. The [linear-head run](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-diagnostic-seed104-linear.json) gets **945/958 sixes correct (98.64%)**. Overall test accuracy changes **88.69% → 97.91%**. Initialization, data, implementation, recorded source, GPU and other recipe settings match; only the output ReLU differs. Pairing checks passed: **True**.

This directly supports an output-head failure in this diagnostic, not a claim that every stalled run has the same cause. The ReLU class 6 preactivation is positive on 1.98% of test inputs (maximum 0.691340), and its last-minibatch head-row gradient norm is 1.48919e-06; the linear head's corresponding norm is 0.170158. Thus “permanently dead output” or “identically zero gradient” would overstate the evidence. The original fixed-seed history could hide this initialization sensitivity.


CUDA graphs reduce replay launch overhead, while fused SGD reduces optimizer orchestration. The relevant evidence here is the matched measurements, not an assumption that these features always help. [PyTorch CUDA graphs](https://docs.pytorch.org/docs/2.14/notes/cuda.html#cuda-graphs), [SGD implementation choices](https://docs.pytorch.org/docs/2.14/generated/torch.optim.SGD.html).

## All accuracy attempts

Time columns for successful rows refer to the first target score. For censored/diverged rows they refer to the last evaluated checkpoint and are explicitly not successful target times. Dispatch is always the whole returned call. Diagnostic runs have deliberately shorter horizons; their lack of a target crossing is not directly comparable to a 150/200-epoch search.

| Attempt | Recipe | GPU | Status / cohort | Last / maximum epochs | First target epoch / accuracy | Best / last accuracy (%) | Training s at target or censor | Invocation s at target or censor | Setup s | Whole invocation s | Dispatch s |
|---|---|---|---|---|---|---|---|---|---|---|---|
| [opt-confirm-b256-bf16-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-b256-bf16-s1.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | reached / confirmation | 99 / 200 | 99 / 98.67% | 98.67 / 98.67 | 24.359 | 36.477 | 10.993 | 36.613 | 78.844 |
| [opt-confirm-b256-bf16-s101](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-b256-bf16-s101.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-80GB | reached / confirmation | 121 / 200 | 121 / 98.63% | 98.63 / 98.63 | 26.818 | 37.032 | 9.113 | 37.153 | 80.820 |
| [opt-confirm-b256-bf16-s102](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-b256-bf16-s102.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / confirmation | 200 / 200 | not observed | 77.97 / 77.72 | 48.698 | 58.768 | 8.310 | 58.773 | 98.331 |
| [opt-confirm-b256-bf16-s103](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-b256-bf16-s103.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / confirmation | 200 / 200 | not observed | 79.28 / 79.17 | 48.454 | 59.890 | 9.240 | 59.898 | 98.029 |
| [opt-confirm-b256-bf16-s104](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-b256-bf16-s104.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / confirmation | 200 / 200 | not observed | 89.27 / 89.11 | 48.653 | 59.419 | 7.563 | 59.423 | 68.117 |
| [opt-confirm-b256-bf16-s105](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-b256-bf16-s105.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / confirmation | 200 / 200 | not observed | 98.60 / 98.52 | 48.739 | 59.022 | 8.489 | 59.027 | 64.899 |
| [opt-confirm-normalized-dropout-step-s101](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-normalized-dropout-step-s101.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100 80GB PCIe | reached / confirmation | 24 / 100 | 24 / 98.67% | 98.67 / 98.67 | 5.609 | 17.236 | 11.257 | 17.492 | 28.590 |
| [opt-confirm-normalized-dropout-step-s102](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-normalized-dropout-step-s102.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | reached / confirmation | 23 / 100 | 23 / 98.63% | 98.63 / 98.63 | 5.819 | 15.592 | 9.461 | 15.829 | 25.146 |
| [opt-confirm-normalized-dropout-step-s103](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-normalized-dropout-step-s103.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | reached / confirmation | 21 / 100 | 21 / 98.71% | 98.71 / 98.71 | 5.320 | 14.897 | 9.251 | 15.167 | 22.777 |
| [opt-confirm-raw-linear-dropout-step-s101](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-raw-linear-dropout-step-s101.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-80GB | reached / confirmation | 32 / 100 | 32 / 98.64% | 98.64 / 98.64 | 7.414 | 18.294 | 10.406 | 18.556 | 27.633 |
| [opt-confirm-raw-linear-dropout-step-s102](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-raw-linear-dropout-step-s102.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | reached / confirmation | 30 / 100 | 30 / 98.64% | 98.64 / 98.64 | 7.599 | 15.807 | 7.801 | 16.041 | 24.588 |
| [opt-confirm-raw-linear-dropout-step-s103](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-confirm-raw-linear-dropout-step-s103.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-80GB | reached / confirmation | 42 / 100 | 42 / 98.63% | 98.63 / 98.63 | 9.677 | 20.551 | 10.315 | 20.839 | 30.599 |
| [opt-diagnostic-seed104-linear](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-diagnostic-seed104-linear.json) | cuda_graph+fused, BF16, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / diagnostic | 5 / 5 | not observed | 97.91 / 97.91 | 1.253 | 13.741 | 12.250 | 13.965 | 49.838 |
| [opt-diagnostic-seed104-relu](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-diagnostic-seed104-relu.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / diagnostic | 5 / 5 | not observed | 88.69 / 88.69 | 1.259 | 14.314 | 12.874 | 14.495 | 53.390 |
| [opt-recipe-b1024-linear-bf16-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-recipe-b1024-linear-bf16-s1.json) | cuda_graph+fused, BF16, B1024, linear head, plain dropout=0, scale=1, LR 0.016, shrink 0.00032, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.58 / 98.43 | 13.553 | 33.400 | 18.141 | 33.405 | 79.156 |
| [opt-recipe-b256-linear-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-recipe-b256-linear-s1.json) | cuda_graph+fused, BF16, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={} | NVIDIA A100-SXM4-80GB | reached / exploratory | 139 / 150 | 139 / 98.63% | 98.63 / 98.63 | 30.712 | 39.651 | 7.758 | 39.752 | 54.490 |
| [opt-recipe-b256-lower-lr-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-recipe-b256-lower-lr-s1.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.002, shrink 8e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.40 / 98.15 | 36.391 | 40.693 | 2.698 | 40.698 | 91.908 |
| [opt-recipe-b256-stronger-shrink-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-recipe-b256-stronger-shrink-s1.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 0.0002, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.60 / 98.33 | 36.521 | 50.087 | 11.864 | 50.092 | 58.807 |
| [opt-recipe-b512-linear-bf16-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-recipe-b512-linear-bf16-s1.json) | cuda_graph+fused, BF16, B512, linear head, plain dropout=0, scale=1, LR 0.008, shrink 0.00016, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.51 / 95.05 | 21.401 | 36.719 | 13.485 | 36.725 | 82.181 |
| [opt-recipe-b512-lower-lr-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-recipe-b512-lower-lr-s1.json) | cuda_graph+fused, TF32, B512, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 0.00016, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.37 / 98.21 | 19.003 | 36.662 | 16.049 | 36.668 | 48.713 |
| [opt-speed-eager-tf32-b64-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-eager-tf32-b64-s1.json) | eager, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.61 / 98.33 | 248.926 | 263.807 | 13.513 | 263.810 | 357.763 |
| [opt-speed-fused-bf16-b128-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-fused-bf16-b128-s1.json) | cuda_graph+fused, BF16, B128, ReLU head, plain dropout=0, scale=1, LR 0.002, shrink 3.99996e-05, schedule={} | NVIDIA A100-SXM4-40GB | reached / exploratory | 104 / 150 | 104 / 98.66% | 98.66 / 98.66 | 45.424 | 48.765 | 2.427 | 48.850 | 53.598 |
| [opt-speed-fused-bf16-b256-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-fused-bf16-b256-s1.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | reached / exploratory | 99 / 150 | 99 / 98.67% | 98.67 / 98.67 | 24.125 | 41.731 | 16.216 | 41.901 | 137.646 |
| [opt-speed-fused-bf16-b512-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-fused-bf16-b512-s1.json) | cuda_graph+fused, BF16, B512, ReLU head, plain dropout=0, scale=1, LR 0.008, shrink 0.0001599888, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 78.93 / 78.86 | 21.485 | 25.387 | 2.574 | 25.390 | 28.137 |
| [opt-speed-fused-bf16-b64-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-fused-bf16-b64-s1.json) | cuda_graph+fused, BF16, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-40GB | reached / exploratory | 114 / 150 | 114 / 98.63% | 98.63 / 98.63 | 92.074 | 110.481 | 17.171 | 110.575 | 216.136 |
| [opt-speed-fused-tf32-b256-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-fused-tf32-b256-s1.json) | cuda_graph+fused, TF32, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} | NVIDIA A100-SXM4-40GB | reached / exploratory | 113 / 150 | 113 / 98.66% | 98.66 / 98.66 | 24.738 | 43.160 | 16.942 | 43.331 | 139.528 |
| [opt-speed-fused-tf32-b64-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-fused-tf32-b64-s1.json) | cuda_graph+fused, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.61 / 98.33 | 106.164 | 122.210 | 14.372 | 122.214 | 216.411 |
| [opt-speed-graph-tf32-b64-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-speed-graph-tf32-b64-s1.json) | cuda_graph, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-40GB | censored / exploratory | 150 / 150 | not observed | 98.61 / 98.33 | 122.238 | 137.500 | 13.431 | 137.504 | 231.416 |
| [opt-stable-linear-dropout-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-stable-linear-dropout-s1.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={} | NVIDIA A100-SXM4-40GB | reached / exploratory | 32 / 100 | 32 / 98.63% | 98.63 / 98.63 | 8.078 | 16.270 | 7.935 | 16.374 | 23.520 |
| [opt-stable-linear-dropout-step-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-stable-linear-dropout-step-s1.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | reached / exploratory | 28 / 100 | 28 / 98.64% | 98.64 / 98.64 | 7.071 | 8.621 | 1.197 | 8.709 | 35.926 |
| [opt-stable-linear-step-fp32-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-stable-linear-step-fp32-s1.json) | cuda_graph+fused, TF32, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | censored / exploratory | 100 / 100 | not observed | 98.38 / 98.32 | 21.797 | 30.240 | 7.754 | 30.243 | 37.963 |
| [opt-stable-linear-step-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-stable-linear-step-s1.json) | cuda_graph+fused, BF16, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | censored / exploratory | 100 / 100 | not observed | 98.41 / 98.32 | 24.321 | 32.844 | 7.562 | 32.847 | 39.620 |
| [opt-stable-normalized-dropout-step-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-stable-normalized-dropout-step-s1.json) | cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | reached / exploratory | 24 / 100 | 24 / 98.67% | 98.67 / 98.67 | 6.069 | 14.990 | 8.698 | 15.105 | 25.197 |
| [opt-stable-relu-step-s1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-stable-relu-step-s1.json) | cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} | NVIDIA A100-SXM4-40GB | censored / exploratory | 100 / 100 | not observed | 98.39 / 98.27 | 25.210 | 32.925 | 4.861 | 32.928 | 59.773 |

All attempts retain source-style raw 0–255 pixels unless their recipe fields say otherwise. Larger batch changes update count and momentum dynamics even when LR and multiplicative shrinkage are scaled by batch size. BF16 changes training arithmetic; evaluation uses the FP32 stored parameters under TF32. Linear heads, unit dropout and normalization are explicit recipe changes. Listed schedule factors multiply both initial LR and initial direct shrinkage, requiring graph recapture; recapture time is excluded from training-loop time and included in invocation time. The six affine layers and 11,972,510 parameters are retained.

### First target and preceding observation

| Run | Previous epoch / accuracy | Previous training / invocation s | First target epoch / accuracy | Target training / invocation s |
|---|---|---|---|---|
| opt-confirm-b256-bf16-s1 | 98 / 98.58% | 24.113 / 36.221 | 99 / 98.67% | 24.359 / 36.477 |
| opt-confirm-b256-bf16-s101 | 120 / 98.62% | 26.597 / 36.802 | 121 / 98.63% | 26.818 / 37.032 |
| opt-confirm-normalized-dropout-step-s101 | 23 / 98.61% | 5.376 / 16.996 | 24 / 98.67% | 5.609 / 17.236 |
| opt-confirm-normalized-dropout-step-s102 | 22 / 98.61% | 5.568 / 15.334 | 23 / 98.63% | 5.819 / 15.592 |
| opt-confirm-normalized-dropout-step-s103 | 20 / 98.41% | 5.070 / 14.612 | 21 / 98.71% | 5.320 / 14.897 |
| opt-confirm-raw-linear-dropout-step-s101 | 31 / 98.58% | 7.184 / 18.056 | 32 / 98.64% | 7.414 / 18.294 |
| opt-confirm-raw-linear-dropout-step-s102 | 29 / 98.62% | 7.347 / 15.548 | 30 / 98.64% | 7.599 / 15.807 |
| opt-confirm-raw-linear-dropout-step-s103 | 41 / 98.58% | 9.447 / 20.312 | 42 / 98.63% | 9.677 / 20.551 |
| opt-recipe-b256-linear-s1 | 138 / 98.26% | 30.492 / 39.422 | 139 / 98.63% | 30.712 / 39.651 |
| opt-speed-fused-bf16-b128-s1 | 103 / 98.54% | 44.988 / 48.320 | 104 / 98.66% | 45.424 / 48.765 |
| opt-speed-fused-bf16-b256-s1 | 98 / 98.58% | 23.883 / 41.478 | 99 / 98.67% | 24.125 / 41.731 |
| opt-speed-fused-bf16-b64-s1 | 113 / 98.60% | 91.264 / 109.663 | 114 / 98.63% | 92.074 / 110.481 |
| opt-speed-fused-tf32-b256-s1 | 112 / 98.55% | 24.520 / 42.925 | 113 / 98.66% | 24.738 / 43.160 |
| opt-stable-linear-dropout-s1 | 31 / 98.50% | 7.827 / 16.011 | 32 / 98.63% | 8.078 / 16.270 |
| opt-stable-linear-dropout-step-s1 | 27 / 98.61% | 6.820 / 8.362 | 28 / 98.64% | 7.071 / 8.621 |
| opt-stable-normalized-dropout-step-s1 | 23 / 98.59% | 5.817 / 14.731 | 24 / 98.67% | 6.069 / 14.990 |

The preceding score describes observation cadence; accuracy need not change monotonically, and an unobserved transient earlier within an epoch cannot be ruled out.

## Kernel microbenchmarks and compiler cost

These repeatedly update one fixed real minibatch. Three timing repetitions are not three independent training seeds. Setup and steady-state timing are both shown; microbenchmark timings do not establish convergence.

| Run / variant | GPU | Status | Warm median ms/update | Variant setup s | Whole invocation / dispatch s |
|---|---|---|---|---|---|
| [opt-compile-214-v1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-compile-214-v1.json) | unknown | failed (missing_module) | — | — | — / 28.216 |
| [opt-compile-214-v2](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-compile-214-v2.json) | unknown | failed (precision_api_compatibility_failure) | — | — | — / 24.033 |
| [opt-compile-214-v3](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-compile-214-v3.json) / compile+fused, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} / max-autotune | NVIDIA A100-SXM4-40GB | completed | 1.1605 | 229.410 | 245.712 / 345.993 |
| [opt-kernels-214-v1](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-kernels-214-v1.json) | unknown | failed (missing_module) | — | — | — / 16.129 |
| [opt-kernels-214-v2](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-kernels-214-v2.json) / eager, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-80GB | completed | 1.8983 | 0.236 | 16.833 / 26.952 |
| [opt-kernels-214-v2](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-kernels-214-v2.json) / cuda_graph, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-80GB | completed | 0.7600 | 0.258 | 16.833 / 26.952 |
| [opt-kernels-214-v2](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-kernels-214-v2.json) / cuda_graph+fused, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-80GB | completed | 0.6637 | 0.307 | 16.833 / 26.952 |
| [opt-kernels-214-v2](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-kernels-214-v2.json) / cuda_graph+fused, BF16, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-80GB | completed | 0.7672 | 0.458 | 16.833 / 26.952 |
| [opt-kernels-214-v2](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/opt-kernels-214-v2.json) / cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} | NVIDIA A100-SXM4-80GB | completed | 0.9394 | 0.280 | 16.833 / 26.952 |

The compiler candidate compiles the model with eager optimizer orchestration; explicit CUDA graph captures the full update. Different capture scope, A100 memory variants and cache history limit cross-run attribution. A large compile/setup cost matters for this short workload even if later replays are faster. [PyTorch compile modes](https://docs.pytorch.org/docs/2.14/generated/torch.compile.html), [compiler caching](https://docs.pytorch.org/tutorials/recipes/torch_compile_caching_tutorial.html).

## Confirmation runs

Expected confirmation attempts: 12; all declared attempts complete: **True**. Missing: none declared/missing.

Target means ± sample SD below include successful seeds only. The attempted/reached denominator and censored/failed counts must accompany those means. Hardware-specific groups and seed-1 repeats versus other confirmation seeds are separate. The preserved raw label `fresh-seed` means different from exploratory seed 1; stable-recipe seeds 101–103 were already tested with the fragile raw/ReLU recipe.

| Recipe / GPU / confirmation type | Reached / attempted | Statuses | Target epochs | Target training seconds | Target invocation seconds |
|---|---|---|---|---|---|
| cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100 80GB PCIe / fresh-seed | 1 / 1 | {'reached': 1} | 24.000 ± — (n=1) | 5.609 ± — (n=1) | 17.236 ± — (n=1) |
| cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100-SXM4-40GB / fresh-seed | 2 / 2 | {'reached': 2} | 22.000 ± 1.414 (n=2) | 5.570 ± 0.353 (n=2) | 15.245 ± 0.491 (n=2) |
| cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} / NVIDIA A100-SXM4-40GB / fresh-seed | 0 / 4 | {'censored': 4} | — | — | — |
| cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100-SXM4-40GB / fresh-seed | 1 / 1 | {'reached': 1} | 30.000 ± — (n=1) | 7.599 ± — (n=1) | 15.807 ± — (n=1) |
| cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} / NVIDIA A100-SXM4-40GB / same-seed-repeat | 1 / 1 | {'reached': 1} | 99.000 ± — (n=1) | 24.359 ± — (n=1) | 36.477 ± — (n=1) |
| cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} / NVIDIA A100-SXM4-80GB / fresh-seed | 1 / 1 | {'reached': 1} | 121.000 ± — (n=1) | 26.818 ± — (n=1) | 37.032 ± — (n=1) |
| cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100-SXM4-80GB / fresh-seed | 2 / 2 | {'reached': 2} | 37.000 ± 7.071 (n=2) | 8.545 ± 1.600 (n=2) | 19.422 ± 1.596 (n=2) |

## Timing and interpretation

- **training_seconds:** Synchronized minibatch loops: gather, graph-input copies, forward/backward, SGD and direct shrinkage. Excludes epoch permutation, evaluation, checkpoint writes, imports, data loading and graph/compiler preparation or schedule-driven recapture.
- **run_wall_seconds:** Invocation start before benchmark import through availability of this score: includes invocation setup, loading/transfers, graph/compiler preparation, shuffling, training, prior progress writes and test evaluation. Excludes this score's subsequent checkpoint/progress writes, remote dispatch/container startup, image build and final volume commit.
- **total_run_seconds:** Invocation through post-score result/checkpoint work to its final timer; still excludes driver dispatch/container startup and final volume commit.
- **local_dispatch_elapsed_seconds:** Driver time from dispatch to returned result, including queue/startup/import/run/final volume commit/transfer. Whole-call time, not an exact timestamp of first target availability.
- **setup_seconds:** Per-invocation setup until training begins; does not prove a cold process/cache. Container reuse can change this value.
- **microbench_step_seconds:** Warm repeated-single-batch completed updates measured with synchronization; excludes compiler/setup cost and epoch gather/shuffle. Not time to accuracy.

Container/cache reuse was not recorded as a reliable cold/warm flag. Differences in setup or driver delay must not be credited to SGD, precision or batch size. No cold-process benchmark or energy measurement is claimed.

- Official test outcomes actively guided configuration choices and early stopping. These results are test-target optimization, not unbiased held-out generalization estimates.
- A first qualifying epoch is an observed checkpoint; the preceding evaluation gives temporal resolution but cannot exclude an earlier transient within-epoch crossing.
- Censored means target not observed by the evaluation budget; it does not prove the recipe can never reach it. Missing/failed jobs are not successful fast results.
- Larger batch, BF16, output-head changes, LR and shrinkage scaling alter the optimization recipe. Only matched-recipe comparisons identify kernel implementation effects.
- All rows use unchanged Ciresan widths, but this search has no augmentation and is not the record-setting historical Ciresan training system.
- Reported A100-SXM4-40GB, A100-SXM4-80GB and A100 80GB PCIe devices are distinguished. No hardware-normalized speed inference is made across those variants.
- Confirmation seeds 101–103 differ from exploratory seed 1 but were already used with the fragile raw/ReLU recipe before the stable recipes were selected. Raw metadata fresh-seed means distinct from seed 1, not untouched by this search. These confirmations do not undo test-set-driven selection or make the 10000 examples independent across runs.
- Training, invocation-to-score, whole invocation and driver dispatch clocks answer different questions. None is an energy measurement or invoice.

## Completeness and evidence

Accuracy statuses: {'reached': 16, 'censored': 18}. Kernel/compile statuses: {'failed': 3, 'completed': 2}.


The [machine-readable summary](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/results/optimization-analysis.json) includes result-byte hashes, effective recipes, initialization/data identity, source hashes when available, first/previous/last evaluations and all failures. It copies no checkpoints, arbitrary remote errors, private paths or credentials. The [implementation research and audit](https://github.com/yaroslavvb/gradient-dissent/blob/main/experiments/ciresan_stochastic_depth/optimization/research.md) documents graph-state restoration, precision-API compatibility and primary references.
