Exercise Solutions: Training Memory: The Four Tenants and the 16 Bytes Per Parameter

Exercise solutions: Training Memory: The Four Tenants and the 16 Bytes Per Parameter

These are the worked solutions for the exercises in Part 2, Training Memory: The Four Tenants and the 16 Bytes Per Parameter. Read the exercise first; coming here before you have tried it defeats the point.

Exercise 1

Weights: 500,000,000 times 2 bytes is 1,000,000,000 bytes, 1 GB. Gradients: the same shape, one number per weight, also 1 GB. Optimizer: 500,000,000 times 12 bytes is 6,000,000,000 bytes, 6 GB.

Sum: 1 plus 1 plus 6 is 8 GB, which matches 500,000,000 times 16 bytes computed directly. The optimizer alone is 6 of the 8 GB, three quarters of the bill, exactly the proportion this part states for any parameter count under standard mixed-precision AdamW. Note this is static state only; activations still have to be added separately for a real fit-or-not decision.

Exercise 2

Frozen 90 percent: 1,800,000,000 parameters. Frozen weights receive no gradient and need no optimizer state, so they only pay for their bf16 compute copy, 2 bytes each: 1,800,000,000 times 2 is 3,600,000,000 bytes, 3.6 GB.

Trainable 10 percent: 200,000,000 parameters at the full 16 bytes: 200,000,000 times 16 is 3,200,000,000 bytes, 3.2 GB.

Total: 3.6 plus 3.2 is 6.8 GB, against 2,000,000,000 times 16 bytes, 32 GB, for fully fine-tuning the whole model. That is a saving of about 25.2 GB, roughly 79 percent, from freezing alone, with no low-rank parameterisation involved at all. This is the tenant-level truth underneath LoRA: freezing removes the gradient tenant and the optimizer tenant for whatever it touches, 14 of 16 bytes per frozen parameter, leaving only the 2-byte weight copy. LoRA is a particular, parameter-efficient way of choosing what stays trainable, a small low-rank patch rather than a literal 10 percent of the original layers, but the byte accounting for what gets removed is identical either way. What this trick does not remove is activations: every one of those frozen layers still runs a full forward pass, and the backward pass still has to flow through them to reach whichever 10 percent is updating, so the activation tenant is untouched.

Exercise 3

The gap between representable bf16 values near 0.5 is about 0.0039. The per-step update of 0.00005 is roughly 78 times smaller than that gap, far below one representable step. So 0.5000 plus 0.00005, computed and stored directly in bf16, rounds straight back to 0.5000. Every single one of the 200 steps hits this same wall, because there is no separate higher-precision accumulator remembering the sub-threshold remainder between steps.

The logged value after 200 steps is exactly 0.5000, bit for bit identical to the starting value. It does not slowly creep toward the mathematically intended 0.5100 and stall partway; it never moves at all, because each individual step is invisible to bf16’s resolution and nothing carries the invisible part forward. This is exactly why standard mixed precision keeps a separate fp32 master copy of the weight: fp32’s gap near 0.5 is about 6e-8, easily capturing a 0.00005 update, so the running total in fp32 correctly reaches about 0.5100 after 200 steps, and only that accumulated fp32 value gets rounded down to bf16 for the next forward and backward pass’s compute copy.

Exercise 4

8-bit Adam quantises the two running moments, m and v, from 4 bytes each in fp32 down to about 1 byte each, while leaving the fp32 master weight copy at 4 bytes untouched. So the optimizer tenant drops from 12 bytes per parameter, 4 master plus 4 plus 4, to about 6 bytes per parameter, 4 master plus roughly 1 plus 1, a saving of about 6 bytes per parameter.

Full fine-tune, Qwen2.5-7B, about 7.6 billion parameters (16 times 7.6 billion gives the roughly 122 GB this series already quotes): 6 bytes times 7.6 billion is about 45.6 GB saved, taking the run from about 122 GB down to about 76.4 GB. That is a real, meaningful reduction, about 37 percent, even though it still will not fit one 32 GB card; it is decisive for something like an 80 GB data-centre card.

LoRA, same model: only about 1 percent of the parameters are trainable and therefore carry any optimizer state at all, about 76 million parameters. 8-bit Adam’s saving applies only to those: 6 bytes times 76 million is about 456 MB. Against an 18 GB total static state, that is roughly 2.5 percent, close to noise, and probably not worth the added dependency on a quantised-optimizer implementation for a run this size.

The general rule underneath both numbers: 8-bit Adam’s payoff scales with how many parameters actually carry live optimizer state. Full fine-tuning gives it 100 percent of the model to work on, so the saving is tens of gigabytes. LoRA already froze 99 percent of the model before 8-bit Adam ever got involved, so there is almost nothing left for it to shrink.