Bagaimana model sebenarnya belajar, dari gradient descent dasar hingga Adam
Batch size, ditulis B, mengubah noise dalam estimasi gradien. Batch kecil memberi estimasi yang berisik tetapi murah. Batch besar memberi estimasi yang lebih stabil, tetapi setiap update lebih mahal.
🔒 This is a Pro lesson — the interactive figure, worked examples, quiz and practice open with Pro access.
▶ Penskalaan Batch Size