Reverse Mode vs Forward Mode

How a network computes, why gradients vanish, and what makes depth trainable

calc2.15 established that the derivative of a composition is a product of Jacobians, J = Jᴸ ⋯ J² J¹. Matrix multiplication is associative, so you may bracket that product however you like and the answer is identical. The cost is not identical, and the gap is not small.

🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.

▶ Reverse Mode vs Forward Mode
← The Embedding Layer's GradientThe Tape: Memory, Recomputation and Checkpointing →