Gradient Noise Scale: Why η and B Are One Knob

How a network computes, why gradients vanish, and what makes depth trainable

optim.22 gives you two rules of thumb. Gradient noise falls like 1/√B, and if you multiply the batch size by k, try multiplying the learning rate by k as well. Both are stated there as things that usually work.

🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.

▶ Gradient Noise Scale: Why η and B Are One Knob
← Why Warmup Is Not Optional With AdamWhy Depth Works: Folding Space and Reusing Features →