Initialization I: Keeping the Forward Signal Alive

How a network computes, why gradients vanish, and what makes depth trainable

Everything dl.20 measured was settled before the first training step, by the numbers the weights happened to be drawn from. This lesson picks that distribution, and it turns out there is only one free parameter worth arguing about: how wide the draw should be.

๐Ÿ”’ This is a Pro lesson โ€” the interactive figure, worked examples, quiz and practice open with Pro access.

โ–ถ Initialization I: Keeping the Forward Signal Alive
โ† Vanishing and Exploding GradientsInitialization II: He, Gain and Orthogonal โ†’