How a network computes, why gradients vanish, and what makes depth trainable
dl.5 already told you that one hidden layer is enough. Given enough units, a single layer can approximate any continuous function on an interval as closely as you like. So the question this whole course has been putting off: why stack layers at all?
๐ This is a Pro lesson โ the interactive figure, worked examples, quiz and practice open with Pro access.
โถ Why Depth Works: Folding Space and Reusing Features