Tag: GPU memory
-

Activation Memory: Why the Forward Pass Costs More Than the Weights
Activation memory scales with batch, sequence, hidden size and depth, not parameters. The lifecycle, the quadratic attention trap, and every lever that…
-

Training Memory: The Four Tenants and the 16 Bytes Per Parameter
Why a 3 GB model needs 25 GB to train. The four memory tenants derived from first principles, where the 16 bytes…