GRADIENT DISSENTRuns on your device

TRAIN · COMPARE · INSPECT

MNIST browser lab

Choose a network, make branches optional, then train it here.

The active network

Four optional branches

A skipped branch follows its bypass. At 50% dropout, a kept branch contributes twice during training and once at testing.

Validation accuracy—
Cross-entropy · lower is better—
Training time—
Epoch—

Measured learning curves

No prerecorded results: the graph fills from your training runs. The first run also qualifies the backend and compiles its training operations.

Experiment results

RunDropout branchesTest accuracyTest lossTraining
Your completed runs will appear here.

Each configuration starts from fresh, seed-matched weights. In a sweep, dropout eligibility changes between separately trained models. Confidence intervals require repeated seeds; these local runs are not the 48-run A100 result.

REAL MNIST IMAGES

Select a digit

Untrained

Try removing branches at inference

Uses the latest completed model and the checkpoint selected above. This is a separate evaluation: it does not change the training results. Unchecked branches follow their bypass at testing; retained branches have gain one.

Train a model to enable this experiment. Feature layer and classifier remain active.

How this reproduces the experiment

These are fully connected residual networks with Ciresan-style widths, not convolutional ResNets. Every middle branch has a prefix-crop bypass. Each selected branch independently skips a whole minibatch with the chosen probability. A retained training branch is scaled by 1 / (1 − drop probability); evaluation uses gain one. Skipped weights and momentum remain unchanged.

Each plan uses the same seeded initialization, image order, and underlying gate draws across its configurations. GPU arithmetic can differ across devices and backends. The browser validates forward, backward and update support before training; Auto can fall back if a backend fails that check.

The minimum validation-loss checkpoint is selected without using test labels. The official test set is evaluated after training, at both selected and final checkpoints. Repeated interactive exploration can still overfit this test set. Smaller networks, dataset subsets and browser kernels make this an exploratory reproduction, not a reproduction of the A100 timing or accuracy numbers.

Dataset provenance and license · Dataset hashes and split indices · TensorFlow.js backends · WebGPU implementation · Source code