模型究竟是如何學習的,從原始梯度下降法到Adam
固定不變的學習率很少能貫穿整個訓練過程都是最優的。訓練早期通常可以承受較大的移動,因為參數離有用的設定還很遠。訓練後期則往往需要較小的移動才能穩定下來。
🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.