GRADIENT DISSENT

A100 · MNIST · 11.97 million parameters · measured September 9, 2026

98.63% in seconds.
Speed comes from kernels and training.

The original Ciresan widths reach the requested accuracy on all three confirmation seeds. The recommended recipe takes 15.59 seconds from invocation to score at the median, including setup and evaluation. Its synchronized training loop takes 5.61 seconds.

15.59 sMedian invocation → score
5.61 sMedian training loop
3 / 3Confirmation seeds reached 98.63%
2.345×Matched kernel-only speedup

Three assigned A100s include 40GB SXM4 and 80GB PCIe devices; these medians describe the observed service runs, not a hardware-controlled recipe comparison. The matched kernel comparison below uses the same reported 40GB SXM4 GPU model and unchanged learning curves.

Watch time turn into accuracy

Every point is a measured epoch

Accessible chart data

The actual times, without hiding setup

First qualifying epoch, three confirmation seeds. Seconds; all models use the same recommended recipe.
SeedAssigned A100EpochTest accuracyTraining sInvocation → score sDispatch → result s
101A100 80GB PCIe2498.67%5.6117.2428.59
102A100-SXM4-40GB2398.63%5.8215.5925.15
103A100-SXM4-40GB2198.71%5.3214.9022.78

Training excludes shuffling, initialization, evaluation and graph capture. Invocation-to-score includes those costs and imports; it starts inside the remote function, after container startup. Dispatch-to-result includes queueing, startup, output and final persistence. Cache/container reuse was not reliably recorded; none of these is claimed to be a controlled cold-start measurement.

The recommended baseline

784 → 2500 → 2000 → 1500 → 1000 → 500 → 10. Same six affine layers, 11,972,510 parameters. Hidden ReLUs, linear output logits, pixels divided by 255. Ordinary unit dropout at probability 0.2 after all five hidden ReLUs.

SGD momentum 0.9, batch 256, learning rate 0.12, direct parameter shrinkage 0.00008 per update. At the start of epoch 21, multiply both learning rate and shrinkage by 0.1. FP32 parameters and momentum, BF16 autocast, fused SGD, full-step CUDA Graphs. PyTorch 2.14.0, CUDA 13, Triton 3.8. Data stay on the GPU. Each epoch shuffles the 60,000-example pool and uses 59,904 examples in complete batches; the remainder changes with the shuffle.

This is an optimized Ciresan-width recipe with explicit input, head and regularization changes. It uses no augmentation. Test accuracy was monitored and used for tuning and stopping. Seeds 101–103 had already been examined with the earlier raw-input/ReLU recipe; these are confirmation runs, not untouched holdout seeds. This is a time-to-quality study, not an unbiased generalization estimate.

What made the difference?

Remove launch overhead

Full-step graph replay plus fused SGD cut identical 150-epoch work from 248.93 to 106.16 training seconds. Both traces peaked at 98.61%, so this comparison establishes speed, not time to the target.

Fix the output head

With the original output ReLU, only 1 of 5 confirmation seeds reached the target. A matched seed-104 diagnostic gave digit 6 recall of 0/958; changing only the head to linear raised it to 98.64% after five epochs.

Reach useful accuracy sooner

Dropout and the combined LR/shrinkage schedule reached the target in fewer epochs in exploratory runs. BF16 formed part of the selected batch-256 recipe; at batch 64, its matched microsteps were slower.

The tested compiler candidate needed 229.41 seconds of setup. For this short job, full-step CUDA Graphs were the practical choice. That compiler test compiled the model, with optimizer orchestration outside; it was not an exhaustive comparison of compiler strategies.

Read the complete audit: all attempts, failures, clocks, sources and caveats

Ciresan-width MNIST: optimization results

Generated 2026-09-09T22:52:25.126458+00:00. Search finalized; outcomes remain test-monitored.

The target is the first observed 98.63% official-test accuracy (at most 137 errors out of 10,000). These runs fit the full 60,000-example training pool and monitor the test set after scheduled epochs. Recipe choices respond to test results: this is explicitly a test-target speed search, separate from the interrupted validation-only stochastic-depth study.

The public historical W&B run first and uniquely reached 98.63% at epoch 94 and 388.4867 seconds of logged W&B runtime, among 91 verified evaluations. Its GPU is unknown; its clock includes logging/evaluation and is not isolated training time. Historical source/clock gaps are documented in the history audit.

Selected recipes: all three confirmation seeds

The two selected linear-head, unit-dropout recipes are modified training recipes. They use B256, BF16 training, dropout 0.2, initial shrinkage 0.00008, and a factor 0.1 applied to both LR and shrinkage at epoch 21. Raw pixels use LR 0.004; normalized pixels (÷255) use LR 0.12. Seeds 101–103 differ from exploratory seed 1 but were also used in earlier recipe checks; the test set was already known and repeatedly monitored.

  • raw-linear-dropout-step: 3/3 reached the target; observed training time 7.414–9.677s and invocation-to-score 15.807–20.551s. These ranges span the hardware listed below; they are not hardware-normalized recipe comparisons.
  • normalized-dropout-step: 3/3 reached the target; observed training time 5.320–5.819s and invocation-to-score 14.897–17.236s. These ranges span the hardware listed below; they are not hardware-normalized recipe comparisons.
Selected recipe Seed Reported GPU Target epoch / accuracy Training s Invocation-to-score s Whole invocation s Dispatch-to-result s
opt-confirm-raw-linear-dropout-step-s101 101 NVIDIA A100-SXM4-80GB 32 / 98.64% 7.413962 18.293905 18.556291 27.632757
opt-confirm-raw-linear-dropout-step-s102 102 NVIDIA A100-SXM4-40GB 30 / 98.64% 7.598947 15.807211 16.040970 24.587625
opt-confirm-raw-linear-dropout-step-s103 103 NVIDIA A100-SXM4-80GB 42 / 98.63% 9.676896 20.550714 20.838683 30.598894
opt-confirm-normalized-dropout-step-s101 101 NVIDIA A100 80GB PCIe 24 / 98.67% 5.608824 17.235739 17.491769 28.590364
opt-confirm-normalized-dropout-step-s102 102 NVIDIA A100-SXM4-40GB 23 / 98.63% 5.818916 15.592140 15.828836 25.146443
opt-confirm-normalized-dropout-step-s103 103 NVIDIA A100-SXM4-40GB 21 / 98.71% 5.320114 14.897264 15.167187 22.777005

Only seed 102 has the same reported GPU (A100-SXM4-40GB) for both selected recipes; the other two pairs mix SXM4/PCIe or 40/80GB variants. Keep the per-run numbers rather than attributing a pooled difference entirely to input normalization or LR.

Confirmation check for raw-linear-dropout-step: 3/3 reached the target by 100 epochs. Seeds: [101, 102, 103]; outcomes: {'reached': 3}. Same-seed repeats are counted separately; stable-recipe seeds were reused from earlier raw/ReLU checks. This supports repeatability across these tested seeds; it does not undo test-target selection.

Confirmation check for normalized-dropout-step: 3/3 reached the target by 100 epochs. Seeds: [101, 102, 103]; outcomes: {'reached': 3}. Same-seed repeats are counted separately; stable-recipe seeds were reused from earlier raw/ReLU checks. This supports repeatability across these tested seeds; it does not undo test-target selection.

Confirmation check for frozen-b256-bf16: 1/5 reached the target by 200 epochs. Seeds: [101, 102, 103, 104, 105]; outcomes: {'reached': 1, 'censored': 4}. Same-seed repeats are counted separately; stable-recipe seeds were reused from earlier raw/ReLU checks. The observed successes do not establish a robust recipe across these tested seeds.

Kernel effects that are actually isolated

  • opt-speed-eager-tf32-b64-s1 → opt-speed-fused-tf32-b64-s1: same seed, recipe, initialization, dataset and reported NVIDIA A100-SXM4-40GB; 150 epochs take 248.926 → 106.164s of training (2.345×). All scored test trajectories identical: True. Target reached: reference=False, treatment=False. This measures fixed-horizon training speed, not a successful time-to-target ratio.
  • opt-speed-eager-tf32-b64-s1 → opt-speed-graph-tf32-b64-s1: same seed, recipe, initialization, dataset and reported NVIDIA A100-SXM4-40GB; 150 epochs take 248.926 → 122.238s of training (2.036×). All scored test trajectories identical: True. Target reached: reference=False, treatment=False. This measures fixed-horizon training speed, not a successful time-to-target ratio.

Output-head diagnostic: digit 6

At seed 104, epoch 5, the paired ReLU-head run predicts zero sixes and correctly classifies 0/958 true sixes. The linear-head run gets 945/958 sixes correct (98.64%). Overall test accuracy changes 88.69% → 97.91%. Initialization, data, implementation, recorded source, GPU and other recipe settings match; only the output ReLU differs. Pairing checks passed: True.

This directly supports an output-head failure in this diagnostic, not a claim that every stalled run has the same cause. The ReLU class 6 preactivation is positive on 1.98% of test inputs (maximum 0.691340), and its last-minibatch head-row gradient norm is 1.48919e-06; the linear head's corresponding norm is 0.170158. Thus “permanently dead output” or “identically zero gradient” would overstate the evidence. The original fixed-seed history could hide this initialization sensitivity.

CUDA graphs reduce replay launch overhead, while fused SGD reduces optimizer orchestration. The relevant evidence here is the matched measurements, not an assumption that these features always help. PyTorch CUDA graphs, SGD implementation choices.

All accuracy attempts

Time columns for successful rows refer to the first target score. For censored/diverged rows they refer to the last evaluated checkpoint and are explicitly not successful target times. Dispatch is always the whole returned call. Diagnostic runs have deliberately shorter horizons; their lack of a target crossing is not directly comparable to a 150/200-epoch search.

Attempt Recipe GPU Status / cohort Last / maximum epochs First target epoch / accuracy Best / last accuracy (%) Training s at target or censor Invocation s at target or censor Setup s Whole invocation s Dispatch s
opt-confirm-b256-bf16-s1 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB reached / confirmation 99 / 200 99 / 98.67% 98.67 / 98.67 24.359 36.477 10.993 36.613 78.844
opt-confirm-b256-bf16-s101 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-80GB reached / confirmation 121 / 200 121 / 98.63% 98.63 / 98.63 26.818 37.032 9.113 37.153 80.820
opt-confirm-b256-bf16-s102 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB censored / confirmation 200 / 200 not observed 77.97 / 77.72 48.698 58.768 8.310 58.773 98.331
opt-confirm-b256-bf16-s103 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB censored / confirmation 200 / 200 not observed 79.28 / 79.17 48.454 59.890 9.240 59.898 98.029
opt-confirm-b256-bf16-s104 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB censored / confirmation 200 / 200 not observed 89.27 / 89.11 48.653 59.419 7.563 59.423 68.117
opt-confirm-b256-bf16-s105 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB censored / confirmation 200 / 200 not observed 98.60 / 98.52 48.739 59.022 8.489 59.027 64.899
opt-confirm-normalized-dropout-step-s101 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100 80GB PCIe reached / confirmation 24 / 100 24 / 98.67% 98.67 / 98.67 5.609 17.236 11.257 17.492 28.590
opt-confirm-normalized-dropout-step-s102 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB reached / confirmation 23 / 100 23 / 98.63% 98.63 / 98.63 5.819 15.592 9.461 15.829 25.146
opt-confirm-normalized-dropout-step-s103 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB reached / confirmation 21 / 100 21 / 98.71% 98.71 / 98.71 5.320 14.897 9.251 15.167 22.777
opt-confirm-raw-linear-dropout-step-s101 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-80GB reached / confirmation 32 / 100 32 / 98.64% 98.64 / 98.64 7.414 18.294 10.406 18.556 27.633
opt-confirm-raw-linear-dropout-step-s102 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB reached / confirmation 30 / 100 30 / 98.64% 98.64 / 98.64 7.599 15.807 7.801 16.041 24.588
opt-confirm-raw-linear-dropout-step-s103 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-80GB reached / confirmation 42 / 100 42 / 98.63% 98.63 / 98.63 9.677 20.551 10.315 20.839 30.599
opt-diagnostic-seed104-linear cuda_graph+fused, BF16, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB censored / diagnostic 5 / 5 not observed 97.91 / 97.91 1.253 13.741 12.250 13.965 49.838
opt-diagnostic-seed104-relu cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB censored / diagnostic 5 / 5 not observed 88.69 / 88.69 1.259 14.314 12.874 14.495 53.390
opt-recipe-b1024-linear-bf16-s1 cuda_graph+fused, BF16, B1024, linear head, plain dropout=0, scale=1, LR 0.016, shrink 0.00032, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.58 / 98.43 13.553 33.400 18.141 33.405 79.156
opt-recipe-b256-linear-s1 cuda_graph+fused, BF16, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={} NVIDIA A100-SXM4-80GB reached / exploratory 139 / 150 139 / 98.63% 98.63 / 98.63 30.712 39.651 7.758 39.752 54.490
opt-recipe-b256-lower-lr-s1 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.002, shrink 8e-05, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.40 / 98.15 36.391 40.693 2.698 40.698 91.908
opt-recipe-b256-stronger-shrink-s1 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 0.0002, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.60 / 98.33 36.521 50.087 11.864 50.092 58.807
opt-recipe-b512-linear-bf16-s1 cuda_graph+fused, BF16, B512, linear head, plain dropout=0, scale=1, LR 0.008, shrink 0.00016, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.51 / 95.05 21.401 36.719 13.485 36.725 82.181
opt-recipe-b512-lower-lr-s1 cuda_graph+fused, TF32, B512, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 0.00016, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.37 / 98.21 19.003 36.662 16.049 36.668 48.713
opt-speed-eager-tf32-b64-s1 eager, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.61 / 98.33 248.926 263.807 13.513 263.810 357.763
opt-speed-fused-bf16-b128-s1 cuda_graph+fused, BF16, B128, ReLU head, plain dropout=0, scale=1, LR 0.002, shrink 3.99996e-05, schedule={} NVIDIA A100-SXM4-40GB reached / exploratory 104 / 150 104 / 98.66% 98.66 / 98.66 45.424 48.765 2.427 48.850 53.598
opt-speed-fused-bf16-b256-s1 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB reached / exploratory 99 / 150 99 / 98.67% 98.67 / 98.67 24.125 41.731 16.216 41.901 137.646
opt-speed-fused-bf16-b512-s1 cuda_graph+fused, BF16, B512, ReLU head, plain dropout=0, scale=1, LR 0.008, shrink 0.0001599888, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 78.93 / 78.86 21.485 25.387 2.574 25.390 28.137
opt-speed-fused-bf16-b64-s1 cuda_graph+fused, BF16, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-40GB reached / exploratory 114 / 150 114 / 98.63% 98.63 / 98.63 92.074 110.481 17.171 110.575 216.136
opt-speed-fused-tf32-b256-s1 cuda_graph+fused, TF32, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} NVIDIA A100-SXM4-40GB reached / exploratory 113 / 150 113 / 98.66% 98.66 / 98.66 24.738 43.160 16.942 43.331 139.528
opt-speed-fused-tf32-b64-s1 cuda_graph+fused, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.61 / 98.33 106.164 122.210 14.372 122.214 216.411
opt-speed-graph-tf32-b64-s1 cuda_graph, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-40GB censored / exploratory 150 / 150 not observed 98.61 / 98.33 122.238 137.500 13.431 137.504 231.416
opt-stable-linear-dropout-s1 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={} NVIDIA A100-SXM4-40GB reached / exploratory 32 / 100 32 / 98.63% 98.63 / 98.63 8.078 16.270 7.935 16.374 23.520
opt-stable-linear-dropout-step-s1 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB reached / exploratory 28 / 100 28 / 98.64% 98.64 / 98.64 7.071 8.621 1.197 8.709 35.926
opt-stable-linear-step-fp32-s1 cuda_graph+fused, TF32, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB censored / exploratory 100 / 100 not observed 98.38 / 98.32 21.797 30.240 7.754 30.243 37.963
opt-stable-linear-step-s1 cuda_graph+fused, BF16, B256, linear head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB censored / exploratory 100 / 100 not observed 98.41 / 98.32 24.321 32.844 7.562 32.847 39.620
opt-stable-normalized-dropout-step-s1 cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB reached / exploratory 24 / 100 24 / 98.67% 98.67 / 98.67 6.069 14.990 8.698 15.105 25.197
opt-stable-relu-step-s1 cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} NVIDIA A100-SXM4-40GB censored / exploratory 100 / 100 not observed 98.39 / 98.27 25.210 32.925 4.861 32.928 59.773

All attempts retain source-style raw 0–255 pixels unless their recipe fields say otherwise. Larger batch changes update count and momentum dynamics even when LR and multiplicative shrinkage are scaled by batch size. BF16 changes training arithmetic; evaluation uses the FP32 stored parameters under TF32. Linear heads, unit dropout and normalization are explicit recipe changes. Listed schedule factors multiply both initial LR and initial direct shrinkage, requiring graph recapture; recapture time is excluded from training-loop time and included in invocation time. The six affine layers and 11,972,510 parameters are retained.

First target and preceding observation

Run Previous epoch / accuracy Previous training / invocation s First target epoch / accuracy Target training / invocation s
opt-confirm-b256-bf16-s1 98 / 98.58% 24.113 / 36.221 99 / 98.67% 24.359 / 36.477
opt-confirm-b256-bf16-s101 120 / 98.62% 26.597 / 36.802 121 / 98.63% 26.818 / 37.032
opt-confirm-normalized-dropout-step-s101 23 / 98.61% 5.376 / 16.996 24 / 98.67% 5.609 / 17.236
opt-confirm-normalized-dropout-step-s102 22 / 98.61% 5.568 / 15.334 23 / 98.63% 5.819 / 15.592
opt-confirm-normalized-dropout-step-s103 20 / 98.41% 5.070 / 14.612 21 / 98.71% 5.320 / 14.897
opt-confirm-raw-linear-dropout-step-s101 31 / 98.58% 7.184 / 18.056 32 / 98.64% 7.414 / 18.294
opt-confirm-raw-linear-dropout-step-s102 29 / 98.62% 7.347 / 15.548 30 / 98.64% 7.599 / 15.807
opt-confirm-raw-linear-dropout-step-s103 41 / 98.58% 9.447 / 20.312 42 / 98.63% 9.677 / 20.551
opt-recipe-b256-linear-s1 138 / 98.26% 30.492 / 39.422 139 / 98.63% 30.712 / 39.651
opt-speed-fused-bf16-b128-s1 103 / 98.54% 44.988 / 48.320 104 / 98.66% 45.424 / 48.765
opt-speed-fused-bf16-b256-s1 98 / 98.58% 23.883 / 41.478 99 / 98.67% 24.125 / 41.731
opt-speed-fused-bf16-b64-s1 113 / 98.60% 91.264 / 109.663 114 / 98.63% 92.074 / 110.481
opt-speed-fused-tf32-b256-s1 112 / 98.55% 24.520 / 42.925 113 / 98.66% 24.738 / 43.160
opt-stable-linear-dropout-s1 31 / 98.50% 7.827 / 16.011 32 / 98.63% 8.078 / 16.270
opt-stable-linear-dropout-step-s1 27 / 98.61% 6.820 / 8.362 28 / 98.64% 7.071 / 8.621
opt-stable-normalized-dropout-step-s1 23 / 98.59% 5.817 / 14.731 24 / 98.67% 6.069 / 14.990

The preceding score describes observation cadence; accuracy need not change monotonically, and an unobserved transient earlier within an epoch cannot be ruled out.

Kernel microbenchmarks and compiler cost

These repeatedly update one fixed real minibatch. Three timing repetitions are not three independent training seeds. Setup and steady-state timing are both shown; microbenchmark timings do not establish convergence.

Run / variant GPU Status Warm median ms/update Variant setup s Whole invocation / dispatch s
opt-compile-214-v1 unknown failed (missing_module) — — — / 28.216
opt-compile-214-v2 unknown failed (precision_api_compatibility_failure) — — — / 24.033
opt-compile-214-v3 / compile+fused, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} / max-autotune NVIDIA A100-SXM4-40GB completed 1.1605 229.410 245.712 / 345.993
opt-kernels-214-v1 unknown failed (missing_module) — — — / 16.129
opt-kernels-214-v2 / eager, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-80GB completed 1.8983 0.236 16.833 / 26.952
opt-kernels-214-v2 / cuda_graph, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-80GB completed 0.7600 0.258 16.833 / 26.952
opt-kernels-214-v2 / cuda_graph+fused, TF32, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-80GB completed 0.6637 0.307 16.833 / 26.952
opt-kernels-214-v2 / cuda_graph+fused, BF16, B64, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-80GB completed 0.7672 0.458 16.833 / 26.952
opt-kernels-214-v2 / cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.001, shrink 2e-05, schedule={} NVIDIA A100-SXM4-80GB completed 0.9394 0.280 16.833 / 26.952

The compiler candidate compiles the model with eager optimizer orchestration; explicit CUDA graph captures the full update. Different capture scope, A100 memory variants and cache history limit cross-run attribution. A large compile/setup cost matters for this short workload even if later replays are faster. PyTorch compile modes, compiler caching.

Confirmation runs

Expected confirmation attempts: 12; all declared attempts complete: True. Missing: none declared/missing.

Target means ± sample SD below include successful seeds only. The attempted/reached denominator and censored/failed counts must accompany those means. Hardware-specific groups and seed-1 repeats versus other confirmation seeds are separate. The preserved raw label fresh-seed means different from exploratory seed 1; stable-recipe seeds 101–103 were already tested with the fragile raw/ReLU recipe.

Recipe / GPU / confirmation type Reached / attempted Statuses Target epochs Target training seconds Target invocation seconds
cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100 80GB PCIe / fresh-seed 1 / 1 {'reached': 1} 24.000 ± — (n=1) 5.609 ± — (n=1) 17.236 ± — (n=1)
cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=0.0039216, LR 0.12, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100-SXM4-40GB / fresh-seed 2 / 2 {'reached': 2} 22.000 ± 1.414 (n=2) 5.570 ± 0.353 (n=2) 15.245 ± 0.491 (n=2)
cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} / NVIDIA A100-SXM4-40GB / fresh-seed 0 / 4 {'censored': 4} — — —
cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100-SXM4-40GB / fresh-seed 1 / 1 {'reached': 1} 30.000 ± — (n=1) 7.599 ± — (n=1) 15.807 ± — (n=1)
cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} / NVIDIA A100-SXM4-40GB / same-seed-repeat 1 / 1 {'reached': 1} 99.000 ± — (n=1) 24.359 ± — (n=1) 36.477 ± — (n=1)
cuda_graph+fused, BF16, B256, ReLU head, plain dropout=0, scale=1, LR 0.004, shrink 7.99976e-05, schedule={} / NVIDIA A100-SXM4-80GB / fresh-seed 1 / 1 {'reached': 1} 121.000 ± — (n=1) 26.818 ± — (n=1) 37.032 ± — (n=1)
cuda_graph+fused, BF16, B256, linear head, unit_dropout dropout=0.2, scale=1, LR 0.004, shrink 8e-05, schedule={'21': 0.1} / NVIDIA A100-SXM4-80GB / fresh-seed 2 / 2 {'reached': 2} 37.000 ± 7.071 (n=2) 8.545 ± 1.600 (n=2) 19.422 ± 1.596 (n=2)

Timing and interpretation

  • training_seconds: Synchronized minibatch loops: gather, graph-input copies, forward/backward, SGD and direct shrinkage. Excludes epoch permutation, evaluation, checkpoint writes, imports, data loading and graph/compiler preparation or schedule-driven recapture.
  • run_wall_seconds: Invocation start before benchmark import through availability of this score: includes invocation setup, loading/transfers, graph/compiler preparation, shuffling, training, prior progress writes and test evaluation. Excludes this score's subsequent checkpoint/progress writes, remote dispatch/container startup, image build and final volume commit.
  • total_run_seconds: Invocation through post-score result/checkpoint work to its final timer; still excludes driver dispatch/container startup and final volume commit.
  • local_dispatch_elapsed_seconds: Driver time from dispatch to returned result, including queue/startup/import/run/final volume commit/transfer. Whole-call time, not an exact timestamp of first target availability.
  • setup_seconds: Per-invocation setup until training begins; does not prove a cold process/cache. Container reuse can change this value.
  • microbench_step_seconds: Warm repeated-single-batch completed updates measured with synchronization; excludes compiler/setup cost and epoch gather/shuffle. Not time to accuracy.

Container/cache reuse was not recorded as a reliable cold/warm flag. Differences in setup or driver delay must not be credited to SGD, precision or batch size. No cold-process benchmark or energy measurement is claimed.

  • Official test outcomes actively guided configuration choices and early stopping. These results are test-target optimization, not unbiased held-out generalization estimates.
  • A first qualifying epoch is an observed checkpoint; the preceding evaluation gives temporal resolution but cannot exclude an earlier transient within-epoch crossing.
  • Censored means target not observed by the evaluation budget; it does not prove the recipe can never reach it. Missing/failed jobs are not successful fast results.
  • Larger batch, BF16, output-head changes, LR and shrinkage scaling alter the optimization recipe. Only matched-recipe comparisons identify kernel implementation effects.
  • All rows use unchanged Ciresan widths, but this search has no augmentation and is not the record-setting historical Ciresan training system.
  • Reported A100-SXM4-40GB, A100-SXM4-80GB and A100 80GB PCIe devices are distinguished. No hardware-normalized speed inference is made across those variants.
  • Confirmation seeds 101–103 differ from exploratory seed 1 but were already used with the fragile raw/ReLU recipe before the stable recipes were selected. Raw metadata fresh-seed means distinct from seed 1, not untouched by this search. These confirmations do not undo test-set-driven selection or make the 10000 examples independent across runs.
  • Training, invocation-to-score, whole invocation and driver dispatch clocks answer different questions. None is an energy measurement or invoice.

Completeness and evidence

Accuracy statuses: {'reached': 16, 'censored': 18}. Kernel/compile statuses: {'failed': 3, 'completed': 2}.

The machine-readable summary includes result-byte hashes, effective recipes, initialization/data identity, source hashes when available, first/previous/last evaluations and all failures. It copies no checkpoints, arbitrary remote errors, private paths or credentials. The implementation research and audit documents graph-state restoration, precision-API compatibility and primary references.