TRAIN · COMPARE · INSPECT
MNIST browser lab
Choose a network, make branches optional, then train it here.
The active network
Four optional branchesA skipped branch follows its bypass. At 50% dropout, a kept branch contributes twice during training and once at testing.
Measured learning curves
No prerecorded results: the graph fills from your training runs. The first run also qualifies the backend and compiles its training operations.
Experiment results
| Run | Dropout branches | Test accuracy | Test loss | Training |
|---|---|---|---|---|
| Your completed runs will appear here. | ||||
Each configuration starts from fresh, seed-matched weights. In a sweep, dropout eligibility changes between separately trained models. Confidence intervals require repeated seeds; these local runs are not the 48-run A100 result.
REAL MNIST IMAGES
Meet the training data
28 × 28 grayscale images. Training, validation and official test images occupy disjoint splits.
Select a digit
UntrainedTry removing branches at inference
Uses the latest completed model and the checkpoint selected above. This is a separate evaluation: it does not change the training results. Unchecked branches follow their bypass at testing; retained branches have gain one.
Train a model to enable this experiment. Feature layer and classifier remain active.
How this reproduces the experiment
These are fully connected residual networks with Ciresan-style widths, not convolutional ResNets. Every middle branch has a prefix-crop bypass. Each selected branch independently skips a whole minibatch with the chosen probability. A retained training branch is scaled by 1 / (1 − drop probability); evaluation uses gain one. Skipped weights and momentum remain unchanged.
Each plan uses the same seeded initialization, image order, and underlying gate draws across its configurations. GPU arithmetic can differ across devices and backends. The browser validates forward, backward and update support before training; Auto can fall back if a backend fails that check.
The minimum validation-loss checkpoint is selected without using test labels. The official test set is evaluated after training, at both selected and final checkpoints. Repeated interactive exploration can still overfit this test set. Smaller networks, dataset subsets and browser kernels make this an exploratory reproduction, not a reproduction of the A100 timing or accuracy numbers.
Dataset provenance and license · Dataset hashes and split indices · TensorFlow.js backends · WebGPU implementation · Source code