Exercise Solutions: Activation Memory: Why the Forward Pass Costs More Than the Weights

Exercise solutions: Activation Memory: Why the Forward Pass Costs More Than the Weights

These are the worked solutions for the exercises in Part 3, Activation Memory: Why the Forward Pass Costs More Than the Weights. Read the exercise first; coming here before you have tried it defeats the point.

Exercise 1

Feed-forward intermediate, shape batch by sequence by intermediate width, bf16: 8 times 1024 times 8960 times 2 bytes is 146,800,640 bytes, about 146.8 MB.

The roughly ten hidden-width tensors, shape 10 by batch by sequence by hidden: 10 times 8 times 1024 times 1536 times 2 bytes is 251,658,240 bytes, about 251.7 MB.

Attention scores, shape batch by heads by sequence by sequence: 8 times 12 times 1024 times 1024 times 2 bytes is 201,326,592 bytes, about 201.3 MB.

Per-layer total: about 146.8 plus 251.7 plus 201.3 is about 599.8 MB, essentially exactly double the roughly 300 MB baseline, because batch appears linearly in all three terms. Across 28 layers: about 16.8 GB, again essentially double the roughly 8 GB baseline. Doubling batch doubles activation memory, cleanly and predictably, because none of these three pieces have a batch-squared term.

Exercise 2

Feed-forward intermediate: 4 times 2048 times 8960 times 2 bytes is 146,800,640 bytes, about 146.8 MB, doubled from the baseline because it scales linearly in sequence.

The ten hidden-width tensors: 10 times 4 times 2048 times 1536 times 2 bytes is 251,658,240 bytes, about 251.7 MB, also doubled, same reason.

Attention scores: 4 times 12 times 2048 times 2048 times 2 bytes is 402,653,184 bytes, about 402.7 MB, roughly four times the roughly 100 MB baseline, not two.

Per-layer total: about 146.8 plus 251.7 plus 402.7 is about 801.2 MB. Across 28 layers: about 22.4 GB, against the roughly 8 GB baseline, an increase of about 2.8 times from doubling one dimension.

Compare the two exercises directly. Doubling batch took the total from about 8 GB to about 16.8 GB, a clean 2x, because batch is a linear multiplier everywhere it appears, including inside the attention score term. Doubling sequence took the total from about 8 GB to about 22.4 GB, nearly 2.8x, because the attention score term has sequence twice over, batch by heads by sequence by sequence, so it quadruples while the other two terms merely double. The attention score term is the one responsible for the extra growth, and it is exactly the quadratic trap this part names: sequence length is far more expensive to increase than batch size, byte for byte.

Exercise 3

gradient_checkpointing_enable() is a persistent setting on the model object, not a one-shot flag scoped to a single call. Once it is turned on, it stays on until something explicitly turns it off. In this scenario, the first cell calls peak_gb(model, batch, True), which calls model.gradient_checkpointing_enable() and leaves it enabled on that live model. The second cell then calls peak_gb(model, batch, False) on the same object. The function’s checkpointing argument being False only means it skips calling gradient_checkpointing_enable() again; it never calls the corresponding disable. So the model is still checkpointed when the supposedly “without checkpointing” measurement runs, and both readings reflect the checkpointed, lower peak.

The fix is to make each measurement start from a known, symmetric state, either by disabling explicitly or by using a fresh model per condition:

# torch 2.x
def peak_gb(model, batch, checkpointing):
    torch.cuda.reset_peak_memory_stats()
    if checkpointing:
        model.gradient_checkpointing_enable()
    else:
        model.gradient_checkpointing_disable()
    model(**batch).loss.backward()
    return torch.cuda.max_memory_allocated() / 1e9

With the explicit else branch, calling the function twice in either order gives each condition its own honest state, and the two numbers should now differ by roughly the amount this part describes.

Exercise 4

Feed-forward intermediate: 1 times 2048 times 18944 times 2 bytes is 77,594,624 bytes, about 77.6 MB.

Roughly ten hidden-width tensors: 10 times 1 times 2048 times 3584 times 2 bytes is 146,800,640 bytes, about 146.8 MB.

Attention scores, using the 28 query heads: 1 times 28 times 2048 times 2048 times 2 bytes is 234,881,024 bytes, about 234.9 MB.

Per-layer total: about 77.6 plus 146.8 plus 234.9 is about 459.3 MB. Across 28 layers: about 12.9 GB, call it about 13 GB, remembering this is an order-of-magnitude estimate and, for this exercise, the worst case with attention scores fully materialised rather than streamed by a flash or SDPA backend.

Now combine with static state. LoRA static state for Qwen2.5-7B is about 18 GB. Add about 13 GB of activations: about 31 GB. Against a 32 GB card’s roughly 29.8 GiB of usable memory, and keeping the 2 to 3 GB safety margin this series insists on, that does not fit; it is over budget before you even account for the margin. QLoRA static state is about 8 GB. Add the same about 13 GB of activations: about 21 GB, which fits with several gigabytes of genuine headroom.

So at batch 1, sequence 2048, with attention scores materialised in full, QLoRA is the method that actually fits this card, and LoRA is not, even though LoRA’s own static-state number looks perfectly comfortable in isolation. The right next move before abandoning LoRA, though, is to attack the activation term directly: enabling a flash or SDPA attention backend removes the roughly 235 MB per layer the attention scores are currently costing, which would pull LoRA’s total back under 32 GB without switching methods at all. That is the same lesson this part keeps returning to, that a static-state number and an activation number have to be added before you can say anything fits.