Adam 與 AdamW

模型究竟是如何學習的,從原始梯度下降法到Adam

Adam 和 AdamW 的區別在於如何處理權重衰減。Adam 把 L2 懲罰項混入自適應梯度更新之中;AdamW 則把權重衰減作為一個單獨的收縮步驟來執行。

🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.

▶ Adam 與 AdamW
← 梯度裁剪學習率查找器 →