Initialization & Signal Scale

How models actually learn, from vanilla gradient descent to Adam

Optimization can fail before it starts if the initial scale is wrong. If weights are too small, signals and gradients can vanish. If weights are too large, activations and gradients can explode or saturate.

🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.

▶ Initialization & Signal Scale
← Gradient AccumulationHyperparameter Search →