How a network computes, why gradients vanish, and what makes depth trainable
optim.22 gives you two rules of thumb. Gradient noise falls like 1/√B, and if you multiply the batch size by k, try multiplying the learning rate by k as well. Both are stated there as things that usually work.
🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.
▶ Gradient Noise Scale: Why η and B Are One Knob