How a network computes, why gradients vanish, and what makes depth trainable
Everything dl.20 measured was settled before the first training step, by the numbers the weights happened to be drawn from. This lesson picks that distribution, and it turns out there is only one free parameter worth arguing about: how wide the draw should be.
๐ This is a Pro lesson โ the interactive figure, worked examples, quiz and practice open with Pro access.
โถ Initialization I: Keeping the Forward Signal Alive