Schedules & Warmup

How models actually learn, from vanilla gradient descent to Adam

A fixed learning rate is rarely best for a whole training run. Early training can handle larger moves because the parameters are far from useful settings. Later training often needs smaller moves to settle.

🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.

▶ Schedules & Warmup
← The Learning RateConditioning & Zig-Zag →