Train a network
Weights start random. Press Train and watch the loss fall, then draw a digit to test it. Switch models, or turn on reduced precision to see accuracy degrade.
Loss is measured on the training images, accuracy on unseen test images. When loss keeps falling but accuracy levels off, the network is memorizing its 2,000 training images rather than learning more general features (overfitting).
How it learns
The default model is a multilayer perceptron trained on MNIST. Your drawing is downsampled to 28 × 28, flattened into 784 numbers and passed through three layers of weights. Training nudges every weight to reduce the prediction error.
No framework
Forward pass, backpropagation and SGD are written in plain JavaScript. No TensorFlow, ONNX or WebAssembly.
In-browser training uses 2,000 MNIST images (200 per class) with mini-batch SGD. "Load pre-trained" swaps in weights trained on the full 60,000-image set in PyTorch.
Fewer bits
Weights are usually 32-bit floats. Quantization stores each one with fewer bits, saving memory and compute, which matters on phones and microcontrollers.
Turn on Precision in the lab and drag the bit slider. At 8 bits accuracy holds. At 4 bits the model is 8× smaller but starts to degrade. At 2 bits each weight is one of four values and predictions fall apart.
Softmax's exponentials resist quantization, so most "integer-only" models keep it in floating point. This demo tries I-BERT (polynomial approximation), Softermax (base-2) and the shift trick.
For the Vision Transformer and GPT, attention scores are quantized before softmax. The
Shift toggle uses shift invariance, softmax(z + c) = softmax(z): subtracting
the mean centers values around zero for a tighter fit. The difference is clearest in the
GPT's text at low bit-widths.
Softmax in integer arithmetic
A semester project under Prof. Dr. Luca Benini, optimizing softmax in MobileBERT for integer-only inference while holding accuracy down to 4 bits.
Exponentials are expensive in fixed-point hardware, so we approximate them with second-order polynomials.
Fewer bits turn the signal into a staircase. Drag to watch the L1 error grow.
Low-bit quantization distorts attention patterns. The shift trick (toggle) centers logits around zero, allowing tighter clipping and higher accuracy at 4 and 5 bits.
Full precision (FP32)
Quantized
I-BERT
Approximates exponentials using integer-only polynomials and power-of-2 shifts.
Softermax
Base-2 with online normalization, avoiding multiple passes over the data.
ITAmax
Picks optimal scaling factors by focusing on narrow input ranges.