How models actually learn, from vanilla gradient descent to Adam
A fixed learning rate is rarely best for a whole training run. Early training can handle larger moves because the parameters are far from useful settings. Later training often needs smaller moves to settle.
🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.
▶ Schedules & Warmup