模型究竟是如何學習的,從原始梯度下降法到Adam
梯度裁剪限制了一次更新能變得多大。如果某個批次產生了一個巨大的梯度,裁剪會在最佳化器執行更新之前把它縮小。
🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.