Training memory is the reason a model that takes 3 GB of space on disk can need 25 GB or more of graphics card memory to fine-tune. That gap is not waste, and it is not a bug in anyone’s code. Every byte of it has a name and a job. Once you know the four names, you can predict the bill before you start the run.
Part 1 arrived at a figure of 16 bytes per parameter by counting what one turn of the training loop has to hold. This part takes those sixteen bytes apart. It also answers the two questions Part 1 left standing: why the model’s weights are stored twice on purpose, and why the obvious saving (finish with one thing, free it, then start the next) cannot be made to work.
By the end you will be able to itemise every one of the sixteen bytes and say who owns it, work out how much a memory-saving optimizer would actually save you before you install anything, and tell apart two ideas that are constantly confused: being a tenant of the card, and being part of the 16 bytes.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter (you are here)
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
Explain why a 3 GB model needs 25 GB to train
Download Qwen2.5-1.5B, the model this series uses for every calculation, and you get a file of roughly 3 GB. That file holds 1.54 billion parameters, which are simply the numbers inside the model, stored at 2 bytes each. The Qwen team’s technical report is where that count and the layer and width figures used below come from. Load it for serving, which is the job also called inference, and it occupies about that much on the card. Load it to fine-tune and the card fills to around 25 GB before you have processed a single example. Nothing has been added to the model. The number of parameters has not changed.
Training memory is the total amount of graphics card memory a fine-tuning run occupies while it is running. The card’s own memory is called VRAM, and it is a separate and much smaller pool than the system RAM in your machine. Everything a training step touches has to fit inside it.
The reason training memory dwarfs the model file is that a training step is not just reading the model. It is asking a question about every number in the model at once: how should this one change? Answering that question takes bookkeeping, and the bookkeeping is bigger than the model.
This is also why sizing a card by parameter count is a habit worth breaking. Parameter count tells you the size of exactly one of the four things you have to hold. The useful question is never “is 7 billion times 2 bytes small enough”. It is “what is the per-parameter rate for the job I am actually running”, and that rate swings by a factor of sixteen depending on the answer.
Name the four tenants and say which one is measured differently
Four kinds of data share the card during training. Calling them tenants is the frame this series uses, because they behave like tenants: they all hold their space at the same time, and getting one of them to leave is the only thing that reliably makes room.
Here they are, in the order the training loop creates them. The 2 plus 2 plus 12 split of the first three is Rajbhandari et al.’s accounting, from the ZeRO paper.
- The weights, at 2 bytes per parameter. One number per parameter, held in a 16-bit format called
bf16. This is the copy the model actually computes with. Part 1 introduced the format. For now, 16 bits is 2 bytes, and that is all you need. - The gradients, at 2 bytes per parameter. First the model is scored, giving the loss, one number saying how wrong its predictions were on this batch of text. Then a sweep runs backwards through the model and works out, for every weight, which way to nudge that weight to make the loss smaller and how steeply the loss responds. That number is the weight’s gradient. The sweep is called the backward pass, and the procedure it runs is called backpropagation. There is exactly one gradient per weight, in the same
bf16format, which is why this tenant is exactly the same size as the first one. If gradients are new to you, the walkthrough of how a neural network learns derives them from a single neuron up. - The optimizer state, at 12 bytes per parameter. The optimizer is the component that turns a gradient into an actual weight change. The one nearly every fine-tune uses, AdamW, keeps three extra numbers per parameter between steps, each in the more precise 32-bit format called
fp32, which is 4 bytes. Three times 4 is 12. This tenant is three quarters of the whole per-parameter bill, and most of this article is about it. - The activations, which are not measured per parameter at all. Data flows forward through the model’s 28 layers, and each layer produces intermediate values and hands them on. That trip is the forward pass. Those intermediate values are called activations, and they are stored as tensors, which is simply the word for a grid of numbers of any shape. They cannot be thrown away when the forward pass finishes, because the backward pass needs them.
That fourth tenant is the one that breaks the pattern, so it is worth being precise about what sets its size. Four numbers do:
- the batch, meaning how many training examples you push through at once,
- the sequence length, meaning how many tokens are in each example (a token is a chunk of text, roughly word-sized),
- the hidden size, meaning how wide the list of numbers flowing between layers is, which is 1,536 in the running example,
- and the layer count, which is 28.
Not one of those four is the parameter count. So activation memory cannot be written as bytes per parameter, and it never appears inside the 16. This is why the series always calls the 16-byte figure the static state, meaning weights plus gradients plus optimizer state, the part that stays the same size from the first step to the last. Activations are counted separately, every single time. Part 3 follows one activation from birth to death and puts real numbers on it.
Why the four cannot take turns
The obvious saving looks like this. Compute the gradients, free the activations, then run the optimizer, reusing the same memory for each stage. Peak memory would then be the largest tenant instead of the sum of all of them. It would roughly halve the bill.
It does not work, and the reason is a chain of dependencies you can walk in one direction only.
- The optimizer step needs the gradients. It has nothing to do until they exist, so it cannot run first and it cannot run early.
- The gradients were computed from the activations. Working out how a layer’s weights affected the loss requires knowing what went into that layer, so the activations had to still be present while the backward pass ran.
- The weights are read by the forward pass, read again by the backward pass, and then written by the optimizer. They are never idle at any point in the step.
Put those together and all four tenants are live at the same instant, somewhere between the end of the forward pass and the end of the update. One full turn of that loop is called a training step, and peak memory is what the card holds at the worst moment inside it. So peak memory is the sum, and the sum is what has to fit.
Put the running example through the sum
Qwen2.5-1.5B has about 1.54 billion parameters. Here is the whole static-state bill, itemised, in code you can run with nothing installed.
PARAMS = 1.54e9 # Qwen2.5-1.5B
tenants = {
"weights (bf16)": 2,
"gradients (bf16)": 2,
"optimizer (fp32)": 12,
}
total = 0.0
for name, bytes_each in tenants.items():
gb = PARAMS * bytes_each / 1e9
total += gb
print(f"{name} {bytes_each:2d} B/param {gb:6.2f} GB")
print(f"static state 16 B/param {total:6.2f} GB")
print(f"the same figure in gibibytes {total / 1.074:6.2f} GiB")
print("activations are NOT in this number. Add them separately.")
It prints 3.08 GB of weights, 3.08 GB of gradients and 18.48 GB of optimizer state, for 24.64 GB of static state. The optimizer alone is more than the model file by a factor of six.
Now add the fourth tenant. At a batch of 4 sequences of 1,024 tokens each, the activations come to roughly 8 GB. Treat that as an order of magnitude and not a measurement, because the exact bytes depend on which intermediate values the framework decides to keep. So the run’s peak is about 33 GB in total: 24.6 GB of static state plus roughly 8 GB of activations.
The target hardware for every calculation in this series is one 32 GB consumer card. Two unit systems collide here, and the collision is bigger than the margin. Card memory is reported in gibibytes, written GiB, which are 2 to the power 30 bytes and about 7 percent larger than the decimal gigabytes the multiplication above produces. Divide by 1.074 to compare like with like. The static state is 22.9 GiB, the activations are roughly 7.5 GiB, and the card offers about 29.8 GiB once the driver has taken its share. That is 30.4 against 29.8, and it does not fit. Leave 2 to 3 GB of headroom on top of any estimate you make, because one unusually long batch lives in that last gigabyte.
Explain why the weights are stored twice
Look back at the tenant list and something odd stands out. The weights appear twice: once as the 2-byte bf16 copy, and again inside the optimizer’s 12 bytes as a 4-byte fp32 copy. That is 6 bytes to store 1.54 billion numbers that are supposed to be the same numbers.
It is deliberate, it is called mixed precision, and essentially every modern training run works this way. The arrangement, including the high-precision copy, comes from Micikevicius et al. Here is why, built from a decimal example first.
What precision means, in ordinary decimal
Imagine a cheap till that keeps four significant digits and stores exactly what it displays. It shows 500.0 for five hundred pounds. Now ring up two pence.
- The true answer is 500.02.
- The till has four digits to work with, so the nearest values it can store are 500.0 and 500.1.
- 500.02 is much nearer to 500.0, so 500.0 is what gets stored.
- The addition was computed correctly and then vanished on the way into storage.
Now ring up two pence a thousand times over. The true total should be 520. The till still reads 500.0. Every single addition was individually too small to survive being written down, and nothing anywhere remembered the leftovers. The till is not broken. It simply has fewer digits than this job needs.
The same idea in binary, where the digits are called the mantissa
A computer stores a decimal number as a floating-point number, which splits a fixed budget of bits into two parts.
- The exponent bits say how big or small the number is, in powers of two. More exponent bits means the format reaches further before a value is too large or too small for it to hold at all.
- The mantissa bits hold the significant digits. These are the till’s four digits. More mantissa bits means finer steps between the values the format can actually represent.
The two formats in the tenant list split their budgets like this. bf16 spends 16 bits as 8 exponent bits and 7 mantissa bits. fp32 spends 32 bits as 8 exponent bits and 23 mantissa bits. Same reach, very different fineness.
You can work out what 7 mantissa bits buys. Take the range from 0.5 to 1.0, where the exponent is fixed and only the mantissa is moving. Seven bits give 2 to the power 7, which is 128 evenly spaced values across that range. The range is 0.5 wide, so the gap between one representable value and the next is 0.5 divided by 128, which is about 0.0039. There is nothing in between. A bf16 number near 0.5 can be 0.5000 or 0.5039, and no value between them exists.
Do the same for fp32. Twenty-three bits give 8,388,608 values across the same range, so the gap is 0.5 divided by 8,388,608, which is about 0.00000006, usually written 6e-8.
Now the consequence, which is the till again with different digits. A weight sitting at 0.5 receives an update of 0.0001. That update is about forty times smaller than bf16‘s gap, so 0.5 plus 0.0001 rounds straight back to 0.5. Store the weights in bf16 only and the update is not merely inaccurate. It is gone. Repeat for three thousand steps and the weight has never moved, while the loss curve sits flat and tells you nothing about why.
Hence two copies with two jobs. The bf16 copy does the arithmetic, because 16-bit arithmetic runs faster on the hardware and 16-bit storage is half the size. The fp32 copy is where the optimizer applies each update, because 6e-8 captures a 0.0001 step with room to spare. After the update, the fp32 copy is rounded down to bf16 to refresh the compute copy for the next step. The high-precision copy is the one that actually remembers, and it is known as the fp32 master weights.
The older fp16 arrangement, and the config line it explains
Mixed precision predates bf16. The original version used a different 16-bit format called fp16, which splits its budget as 5 exponent bits and 10 mantissa bits. More mantissa than bf16, which sounds better, and fewer exponent bits, which turns out to matter far more.
Fewer exponent bits means less reach. Gradients are often very small numbers, and with only 5 exponent bits a lot of them fall off the bottom of what fp16 can hold and become exactly zero. That is called underflow, and the word to notice is exactly. An underflowed gradient has not been made noisy. It has been deleted, and averaging over a batch cannot bring it back.
The workaround is loss scaling: multiply the loss by a large constant before the backward pass, so every gradient it produces is lifted into a range fp16 can hold, then divide them all back down before the optimizer step. It works, and it is one more thing to get wrong. Scale too high and values run off the top of the format instead, which produces NaN, short for “not a number”, the value floating-point arithmetic returns when an operation has no defined answer. One NaN reaching the weights ends the run.
bf16 has the same 8 exponent bits as fp32, so it reaches just as far and gradients do not underflow. Loss scaling becomes unnecessary. That is the whole reason modern recipes set bf16=True rather than fp16=True on hardware that supports it. The memory accounting is identical either way, at 2 plus 2 plus 12. What bf16 removes is a tuning knob and a silent failure mode. Part 4 works through every number format in the field, including the 8-bit and 4-bit ones.
Justify the optimizer’s 12 bytes
Twelve of the sixteen bytes belong to one component, so it deserves an argument rather than an assertion.
Start from what an optimizer is for. The gradient tells you which direction to push a weight. It says nothing about how far. Somebody has to decide the distance, and that somebody is the optimizer.
The simplest possible answer is SGD, or stochastic gradient descent: multiply the gradient by a single number called the learning rate, and subtract. The learning rate scales every weight change in the model, and a schedule can raise or lower it over the course of a run. SGD stores nothing between steps at all, so it would cost zero extra bytes.
The problem is that one global learning rate has to suit every weight in the model, and in a transformer the gradients vary enormously from one part of the model to another. A transformer is a stack of identical blocks, each containing attention, the step where every position in the text looks at the other positions. Some weights in that stack routinely see gradients hundreds of times larger than others. Pick a learning rate small enough to keep the loud weights stable and the quiet ones barely move. Pick one large enough to move the quiet ones and the loud ones blow up. Part 5 works through why that spread makes a single global step size unworkable.
The two running averages, on numbers you can check
AdamW’s answer is to give every weight its own step size, worked out from that weight’s own history. It keeps two numbers per parameter to do it.
- m, a running average of the weight’s recent gradients.
- v, a running average of the squares of those same gradients.
A running average is a number updated a little at each step to track recent history, so old values fade out instead of being stored. Kingma and Ba, who introduced both, call them the first moment and the second moment. The names mean nothing more than the average of the gradient and the average of its square.
Watch what those two do, using a plain mean over five steps in place of the real fading average, which changes nothing about the effect. Take a wild weight whose last five gradients were 0.4, -0.3, 0.5, -0.2 and 0.4.
- Add them: 0.8. Divide by 5, so m is 0.16. The swings partly cancelled, leaving a modest signal pointing the consistent way.
- Square each one: 0.16, 0.09, 0.25, 0.04, 0.16. Squaring throws away the sign, so nothing cancels.
- Add those and divide by 5, so v is 0.14. Take the square root: about 0.374.
- Divide m by that root: 0.16 divided by 0.374 is about 0.43.
Now a steady weight whose last five gradients were 0.05, 0.05, 0.06, 0.05 and 0.04. Its gradients are roughly eight times smaller.
- m is 0.05.
- The squares are 0.0025, 0.0025, 0.0036, 0.0025 and 0.0016, so v is 0.00254, and its square root is about 0.050.
- Divide: 0.05 divided by 0.050 is about 0.99.
That is the entire idea. The wild weight’s raw gradients were eight times bigger, and after the division both weights get a step in the same band, roughly between minus one and one. The quiet weight is allowed to move nearly a full unit. The loud, self-contradicting one is throttled back to less than half. The learning rate then scales both, and it finally means something, because it is being applied to numbers that already sit on a common scale.
Written as one line, with the two averages already computed:
new weight = old weight - lr * m / (sqrt(v) + eps) - lr * decay * old weight
Every symbol in words:
lris the learning rate, the single number scaling every change.mis the running average of this weight’s gradients, the direction.vis the running average of their squares, andsqrt(v)is the brake.epsis a tiny constant, typically 1e-8, present only so that a weight whose gradients have all been zero does not cause a division by zero.decayis the weight decay setting, which is the subject of the next subsection.
Part 6 walks that line through in full, including the small correction needed on the earliest steps. The point here is what it costs to keep.
Where twelve comes from, and why it is not eight
Both running averages are stored in fp32, at 4 bytes each. Storing them in a 16-bit format would put them straight back into the till problem from the last section, because v in particular can be extremely small. So m and v are 8 bytes per parameter.
The remaining 4 bytes are the fp32 master copy of the weight. It is booked to the optimizer tenant rather than to the weights tenant, because the optimizer is the only thing that ever writes to it. Eight plus four is twelve.
On the running example that is 18.48 GB of the 24.64 GB of static state, which is exactly three quarters. Activations are not in either of those two figures. Every memory technique in the rest of this series is aimed at that three quarters, for the obvious reason.
The W in AdamW, and what it costs
Three plain definitions first, because the W stands for an idea nobody has needed yet in this series.
- Overfitting is when a model starts memorising its training examples instead of learning the pattern in them. Training loss keeps falling while performance on anything unseen stalls or gets worse.
- Regularisation is any deliberate constraint added to training to make that less likely.
- Weight decay is one such constraint, and a very simple one. On every step, every weight is pulled slightly toward zero, so no weight grows larger than it needs to be. It is applied as a fixed small fraction, typically 0.01.
Plain Adam applied that pull by folding it into the gradient, before the adaptive machinery ran. You will see that older arrangement called an L2 penalty, which just means the size of the weights was added into the loss so the training loop treated a large weight as a cost in its own right.
Folding it in has a side effect nobody wanted. Adam divides by the square root of v, so anything folded into the gradient gets divided too. A weight with large, noisy gradients has a large v, so its pull toward zero came out shrunk. A quiet weight kept the full pull. The constraint landed unevenly across the model, and how unevenly depended on gradient statistics rather than on anything you chose.
AdamW applies the pull directly to the weight instead, outside the adaptive step, which is the last term on the update line above. Every weight then decays by the same fraction, which is what everybody meant in the first place. The name for that change is decoupled weight decay, and Loshchilov and Hutter’s paper is where it comes from.
The fix costs nothing in memory. It changes which line the decay is applied on, and stores no new numbers. That is why every current framework default names adamw rather than adam, and why setting weight_decay above zero on plain Adam is now treated as a mild bug rather than a choice.
Shrink the biggest tenant when a run will not fit
Because the optimizer is three quarters of the per-parameter bill, it is also where the savings are. Three techniques attack it, and they attack it in three different ways.
Squeeze the two averages. Quantisation means storing numbers with fewer bits by mapping them onto a small set of allowed values. An 8-bit optimizer quantises m and v from 4 bytes each down to about 1 byte each, in small blocks so that local variation is preserved. The algorithm is unchanged and so is its behaviour, which is the finding in Dettmers et al.’s paper. The fp32 master copy is left alone at 4 bytes, because that is the one thing precision cannot be taken from.
Work out what that is worth on the running example. The optimizer tenant falls from 12 bytes per parameter to about 6: four for the master copy, plus roughly one each for m and v. Six bytes saved across 1.54 billion parameters is about 9.2 GB. Static state drops from about 24.6 GB to about 15.4 GB, and the roughly 8 GB of activations is still owed on top of both of those numbers.
Move the state out of the way during a spike. A paged optimizer keeps its state where it normally lives, but can move it out to ordinary system RAM when GPU memory spikes and pull it back when the pressure passes. It saves nothing in the steady state. What it does is convert an out-of-memory crash three hours into a run into a brief slowdown, which is worth having on a card with no margin left.
Remove the tenant entirely. A weight that is never going to be updated needs no optimizer state and no gradient. Freezing most of the model is the largest saving of the three by a wide margin, and it is the next section.
One tempting move is worth ruling out explicitly. You could switch to SGD, which keeps no running averages, and delete 8 of the 12 bytes at a stroke. In practice, transformer training with SGD is notoriously fragile, needs careful hand-tuning of the learning rate over the run, and often still ends up worse. When memory binds, the better trades are to shrink AdamW’s state, to freeze most of the model, or to split the run across more than one card, which is what Part 14 covers. Giving up per-weight adaptivity is close to the bottom of the list.
Say exactly which tenants LoRA removes and which it leaves
LoRA freezes the pretrained model and trains a small pair of skinny matrices beside each weight matrix it targets. A weight matrix is just a rectangular grid of the model’s numbers, and a typical one in the running example is 1,536 by 1,536. Methods of this shape are collectively called parameter-efficient fine-tuning, and the small trained piece you ship at the end is called an adapter. Part 11 is the full treatment. What matters here is what freezing does to the tenant list, because that is the entire memory story and it is simpler than most summaries suggest.
A frozen weight is one the training loop has been told never to change. Follow it through the four tenants.
- It still takes part in the forward pass, so it still needs its 2-byte
bf16copy. That tenant stays. - It will never be updated, so nobody needs to know which way to nudge it. No gradient. Two bytes gone.
- With no updates to apply, there is nothing for the optimizer to track. No running averages and no master copy. Twelve bytes gone.
Fourteen of sixteen bytes disappear for every frozen weight, which is the LoRA memory win told from the tenant side rather than from the low-rank side. For Qwen2.5-1.5B it takes static state from about 24.6 GB down to about 4 GB. QLoRA then attacks the one tenant freezing cannot touch, by storing the frozen weights in 4 bits instead of 16, which takes the first row of the table below from 2 bytes per parameter to about 0.5.
| Tenant | What it is | Bytes per parameter | Sized by | Removed or shrunk by |
|---|---|---|---|---|
| Weights | the bf16 copy the model computes with | 2 | parameter count | quantisation, which is how QLoRA takes it to about 0.5 |
| Gradients | one bf16 number per weight, saying which way to nudge it | 2 | parameter count | freezing the weight |
| Optimizer state | the fp32 master copy plus the two fp32 running averages | 12 | parameter count | freezing, or 8-bit and paged variants |
| Activations | forward values kept until the backward pass consumes them | not per-parameter | batch, sequence length, hidden size, layers | gradient checkpointing, a flash attention backend, a smaller batch or sequence |
The last row is the one to read twice, because it is what most summaries of LoRA leave out. Freezing a weight does nothing whatsoever to its activations. The forward pass still runs through every frozen layer, and the backward pass still has to flow all the way back through them to reach the adapters at the far end. Every activation is still produced and still stored.
So the same roughly 8 GB of activations that a full fine-tune of the running example owes is still owed by a LoRA run of it, on top of the 4 GB of static state. That is why the practical advice in Part 11 still tells you to switch gradient checkpointing on. Gradient checkpointing throws away most stored activations and recomputes them during the backward pass, trading roughly 20 to 30 percent more time for a large drop in peak memory. Despite the name it has nothing to do with a checkpoint, which is a saved copy of the model. The two names collide and the ideas do not. A flash attention backend is the other lever, and it computes attention in small tiles without ever building the full grid of scores, which removes the largest single activation.
Every static-state figure in this series excludes activations. The 24.6 GB, the 15.4 GB and the 4 GB above are weights plus gradients plus optimizer state and nothing else. So are the figures the series quotes for the larger Qwen2.5-7B: about 122 GB for a full fine-tune, about 18 GB for LoRA and about 8 GB for QLoRA. None of them is a total. Add activations back before comparing any of them against a real card, and at 7B scale do not expect a single activation number to be quotable at all, because it moves with batch size and with which of the levers above are switched on.
Tell a memory tenant apart from the 16 bytes
Two questions get run together constantly, and keeping them apart is most of what this article is for.
- Is this thing a training memory tenant? Meaning: does it occupy space on the card during a training run?
- Is this thing part of the 16 bytes per parameter? Meaning: does its size scale with the parameter count?
They are different questions, and the interesting answers are the ones where the two disagree.
Activations are yes to the first and no to the second. They are a genuine tenant, often several gigabytes in size, and they are sized by batch and sequence rather than by parameters, so they are outside the 16 entirely. Quoting the 16-byte figure as though it were the run’s total is the single most common sizing mistake in the field.
The KV cache is no to both. During generation, a model saves the keys and values of the tokens it has already produced so it does not recompute them for every new token, and that saved pile is the KV cache. Training has no generation loop. The model is shown real text and grades every position in one pass, so there is nothing to cache. The KV cache is an inference tenant and never a training one.
One last look-alike worth separating, because it produces confident wrong answers. The optimizer’s 12 bytes are three fp32 numbers per parameter and nothing else: the master copy, m, and v. They do not include the bf16 compute weights, which are their own separate 2-byte tenant. They do not include the gradients, which are another separate 2-byte tenant. Three numbers, twelve bytes, no more and no less.
Key takeaways
- Training memory is the VRAM a run holds while training. It dwarfs the model file because a training step has to work out how every weight should change, and that bookkeeping is larger than the model.
- The four tenants are weights, gradients, optimizer state and activations. A dependency chain keeps all four live at the same instant, so peak memory is their sum and never the largest of them.
- Three tenants scale with parameters and total 16 bytes each under standard mixed-precision AdamW: weights 2, gradients 2, optimizer 12. Activations are the fourth and are sized by batch, sequence length, hidden size and layer count instead.
- The weights are stored twice on purpose. bf16 keeps only 7 mantissa bits, which leaves about 0.0039 between representable values near 0.5, so a smaller update rounds away to nothing. The fp32 master copy is what accumulates those updates instead.
- The optimizer’s 12 bytes are the fp32 master copy plus two running averages, m and v. They buy a per-weight step size, which is what makes transformer training work with a single learning rate setting.
- An 8-bit optimizer takes that tenant from 12 bytes to about 6, which is about 9.2 GB of static state on a 1.5B model. Freezing a weight removes 14 of its 16 bytes outright.
- Freezing never removes activations. Every static-state figure in this series, including 24.6 GB and 4 GB, is weights plus gradients plus optimizer state only, and is never a total.
You can now
- Explain to somebody else why a model that is 3 GB on disk needs 25 GB of card memory to fine-tune, from “Explain why a 3 GB model needs 25 GB to train”.
- Name the four tenants, say which one is not measured per parameter, and give the three-step reason they cannot take turns, from “Name the four tenants and say which one is measured differently”.
- Work out the gap between representable values in a 16-bit format and say why an update smaller than that gap disappears, from “Explain why the weights are stored twice”.
- Compute an adaptive step size by hand from five gradients and explain why a noisy weight ends up moving less than a steady one, from “Justify the optimizer’s 12 bytes”.
- Estimate in gigabytes what an 8-bit optimizer would save on a given model before installing anything, from “Shrink the biggest tenant when a run will not fit”.
- State which tenants freezing removes and which it leaves, and correct a memory estimate that quotes static state as a total, from “Say exactly which tenants LoRA removes and which it leaves”.
Glossary
- 8-bit optimizer
- AdamW with its two running averages stored in 1 byte each instead of 4, using block-wise quantisation. The algorithm and its behaviour are unchanged, and it reclaims most of 8 bytes per parameter. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
- Activation
- Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
- Activation memory
- The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
- Adam
- The optimizer that gives every weight its own step size by tracking two running averages of that weight’s own gradients. AdamW is the corrected version everyone actually uses. Kingma and Ba, Adam
- AdamW
- Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Attention
- The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
- Backpropagation
- The procedure that computes a gradient for every weight in one sweep backwards through the model, from the loss at the end to the first layer. It works by applying the chain rule one step at a time. The calculus behind backpropagation
- Backward pass
- Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- bf16
- A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
- Checkpoint
- A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
- Decoupled weight decay
- Applying weight decay straight to the weight rather than folding it into the gradient. It is the W in AdamW, it costs no extra memory, and it makes the decay land evenly across the model. Loshchilov and Hutter, Decoupled Weight Decay Regularization
- Exponent
- The part of a floating-point number that says how big or small it is, in powers of two. More exponent bits means the format reaches further before values overflow or underflow. Number formats for training
- Flash attention
- An attention implementation that computes the answer in small tiles inside fast on-chip memory, never building the full score grid. Same result, far less memory, and the term that grew with the square of sequence length becomes linear. Dao et al., FlashAttention
- Forward pass
- Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
- fp16
- A 16-bit number format with 5 exponent bits and 10 mantissa bits. Precise for its size but short of reach, so small gradients underflow to zero and loss scaling becomes mandatory. Number formats for training
- fp32
- 32-bit floating point, with 8 exponent bits and 23 mantissa bits. It is the precise reference format, used for the master copy of the weights and for the optimizer’s running averages. Number formats for training
- fp32 master weights
- The full-precision copy of the weights that the optimizer actually updates in a mixed-precision run. It exists because a 16-bit weight cannot record an update far smaller than itself, so without it the updates round away and training stalls. Micikevicius et al., Mixed Precision Training
- Frozen weights
- Weights marked as not trainable, so they never receive an update. A frozen weight needs no gradient and no optimizer state, which removes 14 of its 16 bytes, though its activations are still stored. LoRA explained
- Full fine-tuning
- Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most.
- GiB
- A gibibyte, 2 to the power 30 bytes, which is how drivers and vendors report GPU memory. Byte arithmetic such as 16 times the parameter count lands in decimal GB instead, so 24.6 GB is about 22.9 GiB and a 32 GB card gives roughly 29.8 GiB.
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Gradient checkpointing
- Throwing away most stored activations and recomputing them during the backward pass. Peak activation memory drops a long way in exchange for roughly 20 to 30 percent more time. Chen et al., Training Deep Nets with Sublinear Memory Cost
- Hidden size
- The width of the list of numbers that flows between layers, written d_model. The example model’s is 1,536, and it multiplies straight into activation memory. Inside one transformer block
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- KV cache
- The keys and values of past tokens, kept during generation so they are not recomputed for every new token. It exists only at inference; training has no generation loop and therefore no KV cache. Continuous batching and paged attention
- Layer
- One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
- Learning rate
- A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
- Learning-rate schedule
- A rule that changes the learning rate over the course of a run, typically warming up and then decaying. The schedule is a separate thing from the optimizer, and both act on every step. Warmup and schedule
- LoRA
- Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Loss scaling
- Multiplying the loss by a large constant before the backward pass so small gradients do not underflow fp16, then dividing them back before the optimizer step. bf16 does not need it at all. Micikevicius et al., Mixed Precision Training
- Mantissa
- The part of a floating-point number that holds its significant digits. More mantissa bits means finer steps between the values the format can represent. Number formats for training
- Memory tenant
- One of the four things sharing the card during training: weights, gradients, optimizer state and activations. All four are live at the same moment, so peak memory is their sum rather than the largest of them.
- Mixed precision
- Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
- NaN
- Not a number, the value floating-point arithmetic produces from an undefined operation such as infinity minus infinity. Once one reaches the weights the run is dead, and it is the usual end state of a badly scaled fp16 run. Number formats for training
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Optimizer state
- The numbers an optimizer keeps between steps, such as running averages of past gradients. Under standard mixed-precision AdamW it is 12 of the 16 bytes per parameter, which makes it the largest memory tenant. Rajbhandari et al., ZeRO
- Overfitting
- When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse. The six silent failure modes
- Paged optimizer
- An optimizer that can move its state out to ordinary system RAM when GPU memory spikes, then bring it back when the pressure passes. It turns a hard crash into a brief slowdown. Dettmers et al., QLoRA
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them.
- Parameter-efficient fine-tuning
- Any method that trains a small number of new parameters and leaves the pretrained ones frozen. LoRA and QLoRA are the ones in practical use, and PEFT is the library that implements them. Hugging Face PEFT, LoRA guide
- QLoRA
- LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
- Quantisation
- Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
- Regularisation
- Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
- Running average
- A number updated a little at each step to track recent history, where old values fade away instead of being stored. Adam keeps two of them per weight, which is where its memory cost comes from. Gradients and optimizers explained
- Second moment
- Adam’s running average of the squared gradient, written v. It measures how large and erratic a weight’s gradients have been, and dividing the step by its square root is what gives each weight its own step size. Gradients and optimizers explained
- Sequence length
- How many tokens are in one training example after tokenisation. Activation memory grows in step with it, and the attention part grows with its square. Activation memory and gradient checkpointing
- SGD
- Stochastic gradient descent, the simplest optimizer: multiply the gradient by the learning rate and subtract. It stores nothing extra, and it is unreliable on transformers because one global learning rate has to suit every weight. Gradients and optimizers explained
- Static state
- Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total.
- Tensor
- A grid of numbers with any number of dimensions. One number is a scalar, a row of them is a vector, a table is a matrix, and anything past that is still a tensor with more dimensions. How a neural network learns
- Token
- The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
- Training step
- One cycle of the loop: forward pass, loss, backward pass, optimizer step, then clear the gradients. Everything else in a training script is arrangements around those five moves. Supervised fine-tuning end to end
- Transformer
- The architecture behind every model in this series: a stack of blocks that alternate attention with a feed-forward network. Vaswani et al., Attention Is All You Need
- Underflow
- What happens when a number is too small for the format to represent, so it becomes zero. A gradient that underflows has not been made noisy, it has been deleted, and averaging cannot bring it back. Number formats for training
- VRAM
- The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about.
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
- Weight decay
- A small pull on every weight toward zero on every step, so no weight grows larger than it needs to be. It is a form of regularisation, and it is separate from gradient clipping, which is a safety limit. Loshchilov and Hutter, Decoupled Weight Decay Regularization
Practical exercises
Itemise the bill for a 500M-parameter full fine-tune
A 500 million parameter model is going through standard mixed-precision AdamW full fine-tuning. Using the four-tenant breakdown from this part, weights 2 bytes, gradients 2 bytes, optimizer 12 bytes, compute each tenant’s contribution in GB and confirm they sum to the 16-bytes-per-parameter total.
See the worked solution (opens in a new tab)
Freeze 90 percent of a 2B model without touching LoRA
Instead of LoRA’s low-rank patch, imagine literally freezing 90 percent of a 2 billion parameter model’s weights and fully fine-tuning the remaining 10 percent at standard mixed-precision AdamW. Compute the static state this way, compare it to the 32 GB a full fine-tune of the whole model would need, and say which tenant this trick removes and which one it does not.
See the worked solution (opens in a new tab)
Predict what happens to a bf16-only weight update
Suppose a weight sits at 0.5000 and receives a per-step update of 0.00005 for 200 consecutive steps, all in the same direction, with no fp32 master copy anywhere, updates are applied directly to a bf16 weight. Using the gap figures this part gives for bf16 near 0.5, about 0.0039, predict what the weight’s logged value looks like after all 200 steps, and say whether it slowly approaches the mathematically intended change or something else entirely.
See the worked solution (opens in a new tab)
Decide whether 8-bit Adam is worth it, twice
Take the canonical Qwen2.5-7B figures this series already quotes: full fine-tuning needs about 122 GB of static state, LoRA needs about 18 GB. Roughly 1 percent of the model’s parameters are trainable under LoRA. Work out, in GB, how much 8-bit Adam would save on the full fine-tune, and how much it would save on the LoRA run, then explain why the same technique is a big win in one case and close to noise in the other.
Frequently asked questions
Why does a 3 GB model need 25 GB of VRAM to train?
Because the weights are only one of four things resident during a training step. A gradient is stored for every weight, the optimizer stores three more numbers for every weight, and every intermediate value the forward pass produced is held until the backward pass consumes it. Under standard mixed-precision AdamW the parameter-scaled part alone is 16 bytes per parameter, which is 24.6 GB of static state for a 1.5B model, and activations sit on top of that.
What exactly are the 16 bytes per parameter?
Two bytes for the bf16 weight copy the model computes with, two bytes for the bf16 gradient, and twelve bytes of optimizer state. Those twelve are three fp32 numbers at 4 bytes each: a master copy of the weight, a running average of its gradient, and a running average of its squared gradient. Activations are a real cost and are not part of this per-parameter figure.
Is the KV cache part of training memory?
No. The KV cache holds the keys and values of past tokens so a model does not recompute them for every new token it generates. Training has no generation loop, because the real text is already known and every position is graded in one pass. The KV cache is an inference tenant only.
Can I just use bf16 everywhere and skip the fp32 master copy?
No, because bf16 keeps only 7 mantissa bits, which puts about 0.0039 between representable values near 0.5. A typical update is far smaller than that, so it rounds away and the weight never moves. The fp32 master copy exists to accumulate updates that bf16 cannot record, and dropping it stalls the run rather than speeding it up.
Should I switch from AdamW to SGD to save memory?
Usually no. SGD would delete 8 of the optimizer’s 12 bytes, but it applies one global learning rate to weights whose gradients differ by orders of magnitude across a transformer, so it is fragile and often trains worse. Shrinking AdamW’s state with an 8-bit optimizer, or freezing most of the model, gets you similar memory for far less risk.
Does LoRA reduce activation memory?
No. Freezing a weight removes its gradient and its optimizer state, which is 14 of its 16 bytes, and leaves its forward activation exactly where it was. The forward pass still runs through every frozen layer and the backward pass still flows back through them, so the same activation bill is owed. Gradient checkpointing is the lever that attacks it.
Sources and further reading
- Rajbhandari et al., ZeRO: Memory Optimizations Toward Training Trillion Parameter Models, the source of the 2 plus 2 plus 12 accounting used throughout this series.
- Micikevicius et al., Mixed Precision Training, which introduced the fp32 master weights and loss scaling described above.
- Kingma and Ba, Adam: A Method for Stochastic Optimization, the origin of the two running averages m and v.
- Loshchilov and Hutter, Decoupled Weight Decay Regularization, the paper that put the W in AdamW.
- Dettmers et al., 8-bit Optimizers via Block-wise Quantization, which takes the largest tenant from 12 bytes to about 6.
- Qwen Team, Qwen2.5 Technical Report, the source for the running example’s 1.54 billion parameters, 28 layers and hidden size of 1,536.
- Qwen2.5-1.5B model card, for the on-disk file size and the released weight format.
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, the method behind the frozen-weight tenant arithmetic in the LoRA section.
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, which attacks the one tenant freezing cannot remove.
- Chen et al., Training Deep Nets with Sublinear Memory Cost, the gradient checkpointing result named in the last row of the tenant table.
- Dao et al., FlashAttention, the tiled attention backend that removes the largest single activation.
- Hugging Face, Methods and tools for efficient training on a single GPU, the practical checklist for every lever named here.
