模型究竟是如何學習的,從原始梯度下降法到Adam
Adam 和 AdamW 的區別在於如何處理權重衰減。Adam 把 L2 懲罰項混入自適應梯度更新之中;AdamW 則把權重衰減作為一個單獨的收縮步驟來執行。
🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.