Tag: SGD
-

Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
What a gradient is, what an optimizer decides, and why Adam’s two running statistics per weight are worth 8 bytes each on…

What a gradient is, what an optimizer decides, and why Adam’s two running statistics per weight are worth 8 bytes each on…