AdamW is the optimizer behind almost every language model you are ever likely to fine-tune. In a training script it is one line, usually optim="adamw_torch", and most people never look inside it. Inside are four short lines of arithmetic. They run once for every weight, on every step.
This part opens those four lines and names every symbol in them in plain words. There is no calculus here. Every number is one you could check on a phone.
By the end you will be able to write the AdamW update from memory, walk one weight through three steps by hand, say what each setting in your optimizer config actually does, and tell weight decay apart from gradient clipping. Those last two get confused constantly, and they do completely different jobs.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line (you are here)
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
See where AdamW runs inside the training loop
Start with placement, because it makes everything after it concrete. A training step has four stages, and they repeat for every batch of examples.
- Forward. Push a batch of text through the model and get predictions out.
- Loss. Compare those predictions against the text that actually came next. The result is one number saying how wrong the model was.
- Backward. Work out, for every single weight, which direction that weight should move to make the loss smaller. That is one number per weight, and it is called the gradient.
- Optimizer step. Actually move the weights.
AdamW is stage four. Nothing else in that list changes a single weight. In PyTorch the last two stages are literally loss.backward(), then optimizer.step(), then optimizer.zero_grad() to clear the gradients before the next batch.
It is worth being clear about how little AdamW knows. It never sees your text. It does not know what a token is, how many layers the model has, or which weight belongs to which layer. It receives a long list of numbers, one gradient per weight, and returns a decision about how far each of those weights moves. Which positions in your text were graded in the first place is decided long before this, and the next part, on masking, is about that choice.
One more thing about stage four matters for the rest of the series. AdamW keeps a small record for every weight, and that record survives from one step to the next. It has to, because the whole design rests on remembering. That stored record is the largest single consumer of memory in a training run, which is why the part on the four memory tenants gives it 12 of its 16 bytes per parameter.
Read the name: adaptive moment estimation plus weight decay
The name is not a brand. It is a description, and reading it letter by letter tells you most of the design.
Take “moment” first, because it is the only genuinely unfamiliar word. It is a term from statistics. The first moment of a set of numbers is their plain average. The second moment is the average of their squares. Adam keeps both, for every weight, over that weight’s own gradients:
- The first moment, written
m, is a running average of the gradient. It says which way this weight has been heading lately. - The second moment, written
v, is a running average of the gradient squared. Squaring removes the minus signs, sovonly tracks size. It says how big this weight’s pushes have been lately.
“Estimation” is in the name because Adam never has the true averages. It has cheap running approximations of them, updated a little at a time. “Adaptive” is there because those two numbers let every weight get its own step size instead of sharing one.
The W was added later. It stands for weight decay, and it names a decision about where one small subtraction goes. The last section of this article is entirely about it.
What each ancestor fixed
AdamW is not one invention. It is the end of a chain, where each link fixed one specific complaint about the link before it.
- Plain gradient descent. Move every weight a fixed fraction of its gradient. Simple, and it remembers nothing at all.
- Momentum. Keep a running average of recent gradients and move along that instead. The path stops zig-zagging on noisy batches.
- AdaGrad. Give each weight its own step size, by dividing its step by the running total of its past gradient sizes. Good idea, one flaw: a total only ever grows, so every step size shrank toward zero and long runs ground to a halt.
- RMSProp. Replace that growing total with a running average that forgets old gradients. Now a weight’s step size can recover instead of starving.
- Adam. Use both at once: momentum for the direction, RMSProp’s average for the size, plus a small fix for the first few steps.
- AdamW. Move the weight decay out of the gradient and apply it to the weight directly.
Every one of those patches is still visible in the four lines below. Nothing was thrown away.
Read the AdamW update, one line at a time
Here is the whole algorithm. It is four lines, and AdamW runs them separately for every weight in the model. Before the lines, the cast.
thetais one weight. One number, such as 0.5. Part 1 used theta for all 1.54 billion parameters at once. Here it means a single one of them, because AdamW does this same arithmetic on each weight on its own.gis this weight’s gradient on this step, handed over by the backward pass. Also one number.tis the step number. It is 1 on the first step, 2 on the second, and so on.lris the learning rate, the single overall step size you set in the config.mandvare the two running averages the previous part introduced. Both start at 0 before the first step, because there is no history yet.
In all four lines the equals sign means “replace the old value with this”. It is an instruction, not a statement of fact.
Line 1: the smoothed direction
m = beta1 * m + (1 - beta1) * g
Every symbol, in words:
mon the right is the old running average, carried over from the last step.mon the left is the new one.beta1is a setting, 0.9 by default. It is the share of the old average you keep.(1 - beta1)is therefore 0.1, the share of the new gradient you mix in.gis this weight’s gradient right now.
Read as a sentence: keep 90 percent of what you already had, and stir in 10 percent of the latest gradient. One odd batch can only move m by a tenth of the way, so a single bad example cannot swing the direction. m is the direction, with the noise taken out.
Line 2: the record of how big this weight’s gradients have been
v = beta2 * v + (1 - beta2) * g * g
Every symbol, in words:
von the right is the old value,von the left the new one. Same shape as line 1.beta2is a setting, 0.999 by default. So you keep 99.9 percent of the old value and mix in 0.1 percent of the new one. Much slower to move thanm.g * gis the gradient multiplied by itself.
That multiplication is the only real difference from line 1, and it does one useful thing: it throws the minus sign away. A gradient of -0.2 and a gradient of +0.2 both contribute 0.04. So v knows nothing about direction. It only knows size, averaged over a long stretch of recent steps.
Line 3: the startup fix
m_hat = m / (1 - beta1 ** t)
v_hat = v / (1 - beta2 ** t)
Every symbol, in words:
m_hatis said “m hat”. It is the corrected version ofm. Same forv_hat.beta1 ** tmeans beta1 raised to the power t, which is beta1 multiplied by itself t times. On step 3 with beta1 at 0.9 that is 0.9 times 0.9 times 0.9, which is 0.729.tis the step number again.
Both m and v started at 0. That zero is a lie about the model, and it drags both averages down for the first stretch of the run. These two divisions cancel that lie. The next section but one shows exactly what they are worth at step 1, step 2 and step 100, and how quickly they stop mattering.
Line 4: the step that actually moves the weight
theta = theta - lr * m_hat / (sqrt(v_hat) + epsilon) - lr * lambda * theta
Every symbol, in words:
thetaon the right is the weight’s current value, and on the left its new value.lris the learning rate. It scales the whole move.m_hatis the corrected smoothed direction from line 1. This is the “which way” half.sqrt(v_hat)is the square root of the corrected size record. The square root undoes the squaring in line 2, so this number is back in the same units as a gradient. This is the “how big have this weight’s pushes been” half.epsilonis a tiny number, 1e-8 by default, added so the division can never be a division by zero.lambdais the weight decay strength, typically 0.01. It is the only symbol in the line that has nothing to do with the gradient.
The line has two subtractions, and they are separate ideas bolted together. The first subtraction is the Adam part: direction divided by size. The second, lr * lambda * theta, is the W. You can set lambda to 0 and the second subtraction disappears entirely.
Now read that first fraction on its own, because it is the whole point of Adam. If this weight’s gradients have been large, sqrt(v_hat) is large, so the fraction is small and the weight barely moves. If they have been small and consistent, the fraction is close to 1 and the weight moves by roughly the full learning rate. Each weight is measured against its own history, never against the rest of the model.
Follow one weight through three steps by hand
Now watch it run. One weight, small numbers, nothing hidden.
Set the weight theta to 0.500. Learning rate 0.001. The defaults for the rest: beta1 0.9, beta2 0.999, epsilon 1e-8. Weight decay is 0 for now, so the last term drops out. Say the backward pass hands this weight the same gradient, 0.10, on every step.
Step 1, one line at a time:
m = 0.9 * 0 + 0.1 * 0.10 = 0.010v = 0.999 * 0 + 0.001 * 0.10 * 0.10 = 0.00001m_hat = 0.010 / (1 - 0.9) = 0.010 / 0.1 = 0.10v_hat = 0.00001 / (1 - 0.999) = 0.00001 / 0.001 = 0.01- The square root of 0.01 is 0.1, so the fraction is
0.10 / 0.1 = 1.0 theta = 0.500 - 0.001 * 1.0 = 0.499
Two things in that arithmetic are worth stopping on. First, m_hat came out at 0.10, which is exactly the gradient, and v_hat came out at 0.01, which is exactly the gradient squared. From a single sample, the correction handed back the true values. Second, the weight moved by 0.001, which is exactly the learning rate.
Keep going with the same gradient and the pattern holds.
| Step | m | v | m_hat | v_hat | the fraction | weight after the step |
|---|---|---|---|---|---|---|
| 1 | 0.0100 | 0.00001000 | 0.10 | 0.01 | 1.0 | 0.499 |
| 2 | 0.0190 | 0.00001999 | 0.10 | 0.01 | 1.0 | 0.498 |
| 3 | 0.0271 | 0.00002997 | 0.10 | 0.01 | 1.0 | 0.497 |
Look at the raw m and v columns crawling upward, and then at the corrected columns sitting perfectly still. That gap is the entire job of line 3.
The result in the last column is the single most useful fact about AdamW. When a weight’s gradients are steady, it moves by about the learning rate every step. Not by the gradient. By the learning rate. Feed the same weight a steady gradient of 100 instead of 0.10 and it still moves by 0.001 per step, because 100 divided by 100 is also 1. Adam divides the size out.
That is why the learning rate is the dial you tune and the gradient scale is mostly not your problem. It is also why a learning rate that is ten times too large is such a reliable way to wreck a run: every weight in the model then travels ten times too far, every step.
Here is the same arithmetic as code. It needs nothing installed.
beta1, beta2, eps = 0.9, 0.999, 1e-8
lr, decay = 0.001, 0.0
theta = 0.5 # the weight
m = v = 0.0 # both running averages start at zero
for t in (1, 2, 3):
g = 0.10 # the same gradient every step
m = beta1 * m + (1 - beta1) * g # line 1
v = beta2 * v + (1 - beta2) * g * g # line 2
m_hat = m / (1 - beta1 ** t) # line 3
v_hat = v / (1 - beta2 ** t) # line 3
theta = theta - lr * m_hat / (v_hat ** 0.5 + eps) # line 4, the Adam part
theta = theta - lr * decay * theta # line 4, the W part
print(f"step {t} m={m:.4f} v={v:.8f} theta={theta:.4f}")
Change g to 100 and watch the weight land on the same three values. Change decay to 0.01 and watch the last line start shaving a little off the weight as well.
Watch bias correction fade to nothing
Line 3 is the line people skip, so here it is on its own, in numbers.
The problem is the starting value. Both m and v have to begin somewhere before any gradient exists, and zero is the only honest choice. But zero is also a claim: it says this weight has had no gradients and no movement. On step 1 that false claim is 90 percent of m and 99.9 percent of v.
Work out how far off they are. On step 1, m holds 0.1 times the gradient, so it reads at a tenth of the truth. v holds 0.001 times the gradient squared, so its square root reads at about a thirty-second of the gradient’s size. Both are too low, and they are too low by very different amounts.
Now divide them, uncorrected, as line 4 would: 0.1 divided by 0.0316 is about 3.2. Without the correction, the first step would be about three times the learning rate rather than one times it. That is the real damage. The two estimates start cold, they warm up at different speeds, and the ratio between them comes out wrong until they do.
The fix is one division per average, by 1 - beta ** t. Here is what that division is worth as the run goes on. Each column is the number the average gets multiplied by, once you turn the division around.
| Step t | m is scaled up by | v is scaled up by | What is happening |
|---|---|---|---|
| 1 | 10x | 1,000x | both averages are nearly empty, so both get scaled up hard |
| 2 | 5.3x | 500x | one more sample in each, and the correction has already halved |
| 10 | 1.5x | 100x | m is nearly warm; v has seen 10 of the 1,000 steps it wants |
| 100 | 1.00003x | 10.5x | m’s correction has switched itself off entirely |
| 1,000 | 1x | 1.6x | v is finally most of the way there |
| 5,000 | 1x | 1.007x | both corrections are now doing nothing at all |
Read the last column down and the mechanism is obvious. The correction is huge when the averages are empty, and it puts itself out of business as they fill. Nobody switches it off, because beta ** t shrinks toward zero on its own and the division quietly becomes a division by 1.
Two practical notes follow from that table. The v column fades far more slowly than the m column, because beta2 is set to remember far more history. And there is nothing here for you to configure. Bias correction is arithmetic inside the optimizer, not a setting.
It is also a different thing from learning-rate warmup, which is a setting, and one you will meet in every training config. Warmup ramps the learning rate up from zero over the first few percent of steps, to protect the model from big early moves. Bias correction repairs the optimizer’s own two estimates. You want both, and neither replaces the other.
Explain why a jumpy weight takes smaller steps than a steady one
The last two sections used a weight whose gradient never changed. Real gradients change constantly, so put two weights side by side.
Weight A gets a gradient of +0.10 on every step. It is being pushed steadily in one direction.
Weight B gets +0.10, then -0.10, then +0.10, then -0.10. Same size every time, opposite direction every time. Successive batches disagree about which way it should go.
Start with v, because it is the boring one. Line 2 squares the gradient, and squaring kills the sign, so both weights build exactly the same v and exactly the same brake. All the difference lands in m.
| Step | A’s gradient | A’s m | B’s gradient | B’s m |
|---|---|---|---|---|
| 1 | +0.10 | 0.0100 | +0.10 | 0.0100 |
| 2 | +0.10 | 0.0190 | -0.10 | -0.0010 |
| 3 | +0.10 | 0.0271 | +0.10 | 0.0091 |
| 4 | +0.10 | 0.0344 | -0.10 | -0.0018 |
Check row 2 by hand if you like: 0.9 * 0.010 + 0.1 * (-0.10) = 0.009 - 0.010 = -0.001. The new gradient cancelled almost everything the first one built.
After step 4, correct both averages by the same factor and divide by the same brake. Weight A’s fraction is 1.0, so it moves by a full learning rate. Weight B’s is about 0.05, so it moves by a twentieth of one, and in the opposite direction to its last move. Keep going and B keeps hovering there, taking tiny steps in alternating directions.
That is the behaviour worth remembering. A weight that keeps changing its mind moves slowly. A weight that gets pushed the same way every step moves at the full rate. Neither one needed a setting for that, and neither weight knows anything about the other. The previous part covers why this matters so much for a transformer, where gradients differ in size by orders of magnitude between the embeddings, the attention layers and the normalisation layers.
Set the learning rate, the two betas and epsilon with intent
Every symbol in line 4 that is not a gradient is a line in your config. Here they all are.
| Symbol | Config name | Typical value | What it does |
|---|---|---|---|
| lr | learning_rate |
1e-5 to 3e-4 | scales every step, so it sets roughly how far any weight can travel per step |
| beta one | adam_beta1 |
0.9 | how much gradient history goes into the direction |
| beta two | adam_beta2 |
0.999, or 0.95 when pretraining | how much history goes into the per-weight brake |
| epsilon | adam_epsilon |
1e-8 | keeps the division in line 4 away from zero |
| lambda | weight_decay |
0.0 to 0.1 | how hard every weight is pulled toward zero, the W, covered in the next section |
The two betas have a rule of thumb that makes them intuitive. A running average with decay rate beta remembers roughly the last 1 / (1 - beta) steps. That is all there is to it.
- beta1 at 0.9 gives
1 / 0.1, so the direction remembers about the last 10 gradients. Lower it and the optimizer reacts faster and gets noisier. - beta2 at 0.999 gives
1 / 0.001, so the brake remembers about the last 1,000. That is deliberately much longer. You want the brake to reflect a settled view of this weight’s usual gradient size, not whatever the last batch happened to contain.
The one place the default moves is very large pretraining runs, where 0.95 is common. That is a memory of about 20 steps rather than 1,000. Early in a pretraining run the loss landscape is still changing shape quickly, and a brake built from 1,000 stale steps can be badly wrong about the present. Fine-tuning starts from a model that has already settled and moves it a short distance, so 0.999 is right and there is no reason to touch it.
Epsilon looks like a rounding detail and mostly is one. It has two jobs. The first is protection: if a weight’s gradients have all been zero, v is zero, its square root is zero, and dividing by zero produces infinity or a NaN that kills the run. With epsilon the denominator is 1e-8 instead, and since m is zero too, the step is zero, which is exactly right for a weight with no signal. The second job is damping. Any weight whose sqrt(v_hat) is not much larger than epsilon gets a smaller step than it otherwise would. At 1e-8 that affects almost nothing. Raise it to 1e-6 and it quietly steadies the weights with the tiniest gradients, which is why you sometimes see 1e-6 in a heavily quantised or low-precision recipe that otherwise looks standard.
Tell weight decay apart from an L2 penalty and from gradient clipping
Now the W, which is the last term in line 4 and the reason the letter is there at all.
What weight decay is
Weight decay is a small pull on every weight, toward zero, applied on every step. That is the whole idea. Look again at the term: theta = theta - lr * lambda * theta. With a learning rate of 0.001 and lambda of 0.01, that shrinks every weight by 0.001 percent per step. A weight receiving no gradient at all would end a 3,000-step run about 3 percent smaller than it started.
Why bother shrinking weights on purpose? Because of a failure called overfitting: the model fits your training examples so closely that it does worse on anything it has not seen. Large weights are one way that happens, since they let the model lean very hard on a few specific patterns. A constant gentle pull toward zero means a weight only stays large if the gradients keep insisting on it.
Any deliberate constraint of that kind, added to make a model behave better on unseen data rather than to make the training loss smaller, is called regularisation. Weight decay is the mildest and most common one. Nobody tunes it much. 0.01 is a fine default, 0.0 is a legitimate choice, and above 0.1 is unusual.
Why folding the decay into the gradient goes wrong
There is an older way of asking for the same thing, called an L2 penalty. You add a term to the loss that is proportional to the square of each weight, so a large weight literally costs you loss. Work out its gradient and it comes to lambda times the weight, which means the whole scheme reduces to one instruction: before the optimizer sees g, add lambda * theta to it.
With plain gradient descent those two routes are identical. With Adam they are not, and that is the entire finding of the paper that added the W.
Follow the decay once it is inside g. It goes into m, into v, and then through the division by sqrt(v_hat). So each weight’s decay gets divided by that weight’s own brake. Take two weights, both sitting at 0.5, with a learning rate of 0.001 and lambda of 0.01. One is busy, with a large brake. One is quiet, with a small one. Once the running averages have settled, the decay part of their steps looks like this.
| Weight | Its sqrt(v_hat) | Decay part of the step, L2 inside the gradient | Decay part of the step, AdamW |
|---|---|---|---|
| a busy weight, large gradients | 1.0 | about 0.000005 | 0.000005 |
| a quiet weight, small gradients | 0.1 | about 0.00005 | 0.000005 |
The right-hand column is what you asked for: the same percentage shrink for every weight. The column beside it is what the old route delivers. The quiet weight gets pulled ten times harder than the busy one, for no reason anybody chose. Set weight_decay above zero on plain Adam and this is what you get.
AdamW takes the decay out of the gradient and applies it to the weight directly, after the adaptive part is finished. It costs no extra memory, because it changes where a subtraction happens and stores nothing new. When weight_decay is exactly 0 the two are the same optimizer. That is the whole of the W, and it is why every current recipe says adamw rather than adam.
Weight decay is not gradient clipping
These two get mixed up constantly, probably because they sit next to each other in every config file. They are unrelated.
Gradient clipping works like this. After the backward pass and before the optimizer runs, add up the size of every gradient in the model into one number, called the gradient norm. If that number is above a limit, scale all the gradients down together until it sits exactly at the limit. The standard limit is max_grad_norm = 1.0, and it needs no tuning.
The technique is Pascanu et al.’s. They were working on recurrent networks, where a single batch could produce a gradient large enough to throw the weights somewhere useless in one step, and rescaling the whole gradient back to a fixed norm was the remedy that kept the direction while capping the distance. Transformers are far better behaved, which is exactly why clipping at 1.0 almost never fires and costs almost nothing to leave on.
| Weight decay | Gradient clipping | |
|---|---|---|
| What it touches | the weights | the gradients |
| When it runs | inside the optimizer step, as part of AdamW | before the optimizer step, separately from AdamW |
| What it is for | regularisation: better behaviour on data the model has not seen | safety: stopping one strange batch from throwing the weights into nonsense |
| Typical setting | 0.01 | 1.0 |
| Effect when nothing is wrong | a tiny shrink, on every step, forever | none at all, it does not trigger |
Clipping costs nothing and prevents the classic mid-training loss explosion, so leave it on. It also gives you a free diagnostic: if the gradient norm is being clipped on almost every step, your learning rate is too hot. The part on running a supervised fine-tune end to end covers how to read that log while a run is in progress.
Choose between AdamW and its alternatives
AdamW has one well-known cost, and it is memory. Two extra numbers per weight, stored at 4 bytes each, and they exist for the whole run. Almost every alternative in practical use is an attempt to pay less of that.
| Optimizer | What it changes | Stored per weight | Reach for it when |
|---|---|---|---|
| AdamW | the four lines above | 2 numbers, 8 bytes | always, unless you have a specific reason not to |
| 8-bit AdamW | nothing about the algorithm. It squeezes m and v into 1 byte each instead of 4 | about a quarter of AdamW’s | the run does not fit, and you do not want to change how it trains |
| Paged 8-bit AdamW | 8-bit, and it can park its state in ordinary system RAM during a memory spike | same as 8-bit | as above, and the run dies at spikes rather than steadily |
| Lion | keeps one running average instead of two, and uses only its sign to pick each step | 1 number, 4 bytes | memory is still tight after 8-bit, and you can afford to retune the learning rate and the decay |
| Adafactor | stores a compressed summary of v instead of the full record | far less than 2 numbers | very large models, and it is fussier about the learning rate |
| SGD with momentum | drops the per-weight brake entirely, so one learning rate serves the whole model | 1 number, 4 bytes | rarely for language models. Part 5 takes that argument apart |
For a fine-tune the decision is nearly always made for you. Start on AdamW. If the run does not fit, move to the 8-bit version, which is the same algorithm with smaller storage and behaviour close enough that you do not have to re-tune anything. Only reach past that once you have hit a wall those two cannot clear, and on a fine-tune you rarely will.
One name is missing from that table on purpose, because it is the alternative that is not about memory at all. You et al.’s LAMB stores as much as AdamW does. It adds a per-layer rescaling on top of the per-weight one, so that training stays stable at batch sizes in the tens of thousands. That is a pretraining problem. A fine-tune on one card never gets near those batch sizes, so it is background rather than a choice.
Three other names turn up in 2026 papers: Muon, Shampoo and SOAP. They aim at frontier-scale pretraining runs rather than at fine-tunes, and they are further reading rather than a decision anyone owes you.
Configure AdamW for a real fine-tuning run
In a training config, all of the above is one string and a handful of numbers.
# TRL SFTConfig or Transformers TrainingArguments. Names current as of August 2026.
optim = "adamw_torch" # TrainingArguments default; SFTConfig defaults to adamw_torch_fused
# optim = "adamw_torch_fused" # identical maths, one faster kernel, no memory cost
# optim = "paged_adamw_8bit" # 8-bit state plus paging, for QLoRA or a tight card
learning_rate = 2e-4 # LoRA regime; a full fine-tune wants 1e-5 to 2e-5
adam_beta1 = 0.9 # direction memory, about 10 steps
adam_beta2 = 0.999 # brake memory, about 1,000 steps
adam_epsilon = 1e-8 # floor under the divide in line 4
weight_decay = 0.01 # the W, applied straight to the weights
max_grad_norm = 1.0 # gradient clipping, a separate safety limit
lr_scheduler_type = "cosine" # schedules the learning rate, separate from AdamW
warmup_ratio = 0.03 # also separate from bias correction
Library APIs drift. Pin your versions and check the current trainer documentation before a real run. The names above are current as of August 2026.
Two picks cover most cases. For a small full fine-tune on one card, the fused variant is the fastest and costs no extra memory, because it runs the same arithmetic in fewer trips to memory. For LoRA or QLoRA on a card that is already close to full, paged_adamw_8bit keeps AdamW’s behaviour, cuts the stored state, and survives a memory spike instead of crashing on it.
What those four lines cost on the running example
Qwen2.5-1.5B has 1.54 billion weights, so AdamW runs its four lines 1.54 billion times per step. Count the storage.
mandvare one 32-bit number each per weight, so 4 plus 4 bytes. That is1.54e9 x 8 = 12.3 GBfor the two running averages alone.- A mixed-precision run also keeps a 32-bit master copy of each weight, which the optimizer is the component that updates. Another 4 bytes, so the optimizer tenant comes to 12 bytes per parameter, about 18.5 GB.
- Add the weights and the gradients at 2 bytes each and you have the 16 bytes per parameter this series quotes everywhere: about 24.6 GB of static state for a full fine-tune of this model.
That 24.6 GB is static state only, meaning weights, gradients and optimizer state. Activations sit on top and are counted separately. At a batch of 4 sequences of 1,024 tokens they come to roughly 8 GB, which is an order of magnitude rather than a measurement. So the run wants about 33 GB, against roughly 29.8 GiB usable on a 32 GB card once the driver has taken its share. It does not fit, and that is before leaving the 2 to 3 GB of safety margin every estimate in this series keeps.
Switching to 8-bit takes the m and v part from 12.3 GB to about 3.1 GB, so the optimizer tenant drops from 12 bytes per parameter to about 6. LoRA attacks the same bill from the other end: it freezes almost every weight, and a frozen weight gets no gradient, no m and no v. Static state for this model at rank 16 falls to about 4 GB. The same roughly 8 GB of activations is still owed on top of that, because freezing weights does not change the forward pass.
Key takeaways
- AdamW is stage four of the training step. The forward pass, the loss and the backward pass work out which way each weight should go, and AdamW is the only stage that moves it.
- The update is four lines per weight: smooth the gradient into m, smooth the squared gradient into v, correct both for their cold start, then step by m over the square root of v, and shrink the weight slightly.
- Dividing by the square root of v is the whole adaptive idea. A weight whose gradients have been large or contradictory takes a small step; a weight with steady gradients moves by about the full learning rate.
- Bias correction exists because m and v both start at zero and warm up at different speeds, which would make the first step about three times too large. It scales m by 10 at step 1 and by 1.00003 at step 100, so it retires on its own.
- Weight decay is regularisation: a constant gentle pull toward zero. The W means applying it to the weight directly, so every weight shrinks by the same percentage, instead of routing it through the brake where quiet weights get pulled ten times harder.
- Gradient clipping is not weight decay. It caps the gradients before the step, does nothing at all when the run is healthy, and belongs at 1.0.
- The alternatives almost all attack AdamW’s memory. For a fine-tune, use AdamW, drop to its 8-bit form if the run does not fit, and treat the rest as reading.
You can now
- Point at the exact line of a training script where weights change, and say what AdamW does and does not know about your model, from “See where AdamW runs inside the training loop”.
- Write the four lines of the AdamW update and name every symbol in them without looking, from “Read the AdamW update, one line at a time”.
- Compute a weight’s next value by hand from its gradient and the two running averages, from “Follow one weight through three steps by hand”.
- Say what bias correction is worth at any step number, and why nobody has to switch it off, from “Watch bias correction fade to nothing”.
- Predict which of two weights moves further given their gradient histories, from “Explain why a jumpy weight takes smaller steps than a steady one”.
- Set every optimizer line in a training config on purpose, and explain why weight decay and gradient clipping are separate settings doing separate jobs, from “Tell weight decay apart from an L2 penalty and from gradient clipping”.
Glossary
- 8-bit optimizer
- AdamW with its two running averages stored in 1 byte each instead of 4, using block-wise quantisation. The algorithm and its behaviour are unchanged, and it reclaims most of 8 bytes per parameter. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
- Adafactor
- An optimizer that stores far less than Adam by approximating the second moment with one row and one column of numbers per weight matrix. Used at very large scale, and fussier about the learning rate. Shazeer and Stern, Adafactor
- Adam
- The optimizer that gives every weight its own step size by tracking two running averages of that weight’s own gradients. AdamW is the corrected version everyone actually uses. Kingma and Ba, Adam
- AdamW
- Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter.
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Attention
- The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
- Backpropagation
- The procedure that computes a gradient for every weight in one sweep backwards through the model, from the loss at the end to the first layer. It works by applying the chain rule one step at a time. The calculus behind backpropagation
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- Beta one and beta two
- The two decay rates in Adam that set how much history goes into each running average. Beta one at 0.9 gives a memory of roughly the last ten gradients, and beta two at 0.999 roughly the last thousand.
- Bias correction
- A small division inside Adam that undoes the fact that its two running averages start at zero and are therefore far too small on the first steps. It fades to no effect as the run goes on. Kingma and Ba, Adam
- Block-wise quantisation
- Quantising numbers in small blocks, each with its own scale factor, instead of using one scale for a whole tensor. It handles local variation, and it is how both 8-bit optimizer state and NF4 weights work. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
- Cosine schedule
- A learning-rate schedule that decays smoothly from the peak down toward zero along a cosine curve. It is the reliable default for fine-tuning. Hugging Face TRL, SFT Trainer
- Decoupled weight decay
- Applying weight decay straight to the weight rather than folding it into the gradient. It is the W in AdamW, it costs no extra memory, and it makes the decay land evenly across the model. Loshchilov and Hutter, Decoupled Weight Decay Regularization
- Embedding
- The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
- First moment
- Adam’s running average of the gradient, written m. It gives a smoothed direction to move in, so the path stops zig-zagging on noisy batches. Gradients and optimizers explained
- Generalisation
- How well a model does on inputs it never saw during training. It is the thing you actually want, and training loss does not measure it. Evaluating a fine-tune
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Gradient clipping
- Capping the overall size of the gradient before the update, so one strange batch cannot throw the weights into nonsense. A limit of 1.0 is the standard setting and needs no tuning.
- Gradient norm
- The overall size of the gradient across all weights, logged every step. Constant clipping means the learning rate is too high, and a norm near zero from the first step means it is too low. What to watch while it runs
- Hyperparameter
- A setting you choose before training rather than something the model learns, such as the learning rate, the batch size or the number of epochs. The knobs that decide whether it learns
- Layer
- One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
- Layer norm
- A step that rescales the numbers flowing through a layer so they stay in a sensible range, which keeps training stable. Modern LLMs use a cheaper version of it called RMSNorm. Inside one transformer block
- Learning rate
- A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
- Learning-rate schedule
- A rule that changes the learning rate over the course of a run, typically warming up and then decaying. The schedule is a separate thing from the optimizer, and both act on every step. Warmup and schedule
- Lion
- An optimizer that keeps one running average per weight and uses only its sign to decide each step. Less memory than AdamW, and it needs its own learning rate and weight decay settings. Chen et al., Symbolic Discovery of Optimization Algorithms
- LoRA
- Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Low-rank
- A matrix is low-rank when it can be rebuilt exactly from two much smaller matrices multiplied together. LoRA assumes the update a fine-tune needs is low-rank, which is why two skinny matrices can stand in for a full one. Aghajanyan et al., Intrinsic Dimensionality of Fine-Tuning
- Mixed precision
- Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
- Momentum
- Keeping a running average of recent gradients so the path stops zig-zagging and builds speed in a consistent direction. It costs one extra number per weight. Gradients and optimizers explained
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Optimizer state
- The numbers an optimizer keeps between steps, such as running averages of past gradients. Under standard mixed-precision AdamW it is 12 of the 16 bytes per parameter, which makes it the largest memory tenant. Rajbhandari et al., ZeRO
- Paged optimizer
- An optimizer that can move its state out to ordinary system RAM when GPU memory spikes, then bring it back when the pressure passes. It turns a hard crash into a brief slowdown. Dettmers et al., QLoRA
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- QLoRA
- LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
- Quantisation
- Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
- Regularisation
- Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
- Running average
- A number updated a little at each step to track recent history, where old values fade away instead of being stored. Adam keeps two of them per weight, which is where its memory cost comes from. Gradients and optimizers explained
- Second moment
- Adam’s running average of the squared gradient, written v. It measures how large and erratic a weight’s gradients have been, and dividing the step by its square root is what gives each weight its own step size. Gradients and optimizers explained
- SGD
- Stochastic gradient descent, the simplest optimizer: multiply the gradient by the learning rate and subtract. It stores nothing extra, and it is unreliable on transformers because one global learning rate has to suit every weight. Gradients and optimizers explained
- Training step
- One cycle of the loop: forward pass, loss, backward pass, optimizer step, then clear the gradients. Everything else in a training script is arrangements around those five moves. Supervised fine-tuning end to end
- Transformer
- The architecture behind every model in this series: a stack of blocks that alternate attention with a feed-forward network. Vaswani et al., Attention Is All You Need
- VRAM
- The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
- Warmup
- Ramping the learning rate up from zero over the first few percent of steps, so the earliest updates cannot damage the model while the optimizer still has no history. Set it as a ratio rather than a step count. Warmup and schedule
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
- Weight decay
- A small pull on every weight toward zero on every step, so no weight grows larger than it needs to be. It is a form of regularisation, and it is separate from gradient clipping, which is a safety limit. Loshchilov and Hutter, Decoupled Weight Decay Regularization
Practical exercises
Derive the memory window rule behind beta one and beta two
This part states that beta one at 0.9 gives the smoothed gradient a memory of roughly the last ten gradients, and beta two at 0.999 gives the per-weight brake a memory of roughly the last thousand. Work out the simple formula, built from the decay rate alone, that produces both of those numbers, then use it to compute the effective memory window for beta two at 0.95, the value this part names for large scale pretraining, and for beta one at 0.99. State both new windows in steps.
See the worked solution (opens in a new tab)
Work the bias correction arithmetic by hand for three steps
Bias correction divides the running average by one minus the decay rate raised to the step number. Assume a constant gradient of 1.0 on every step and beta one at 0.9. By hand, compute the raw running average and the bias corrected estimate at steps one, two and three, then state what you notice about the corrected estimate across all three steps and why the mechanism produces that result for a constant gradient specifically.
See the worked solution (opens in a new tab)
Compute the memory saved by switching to an 8-bit optimizer
Qwen2.5-1.5B has 1.54 billion parameters. Using the two states, four bytes each, this part gives for adamw_torch’s m and v, compute how many gigabytes those two running averages alone occupy for a full fine tune, leaving the fp32 master weight copy out of the count. Then use the about a quarter figure this part gives for paged_adamw_8bit to compute the same two states’ memory under that variant, and state the gigabytes saved.
See the worked solution (opens in a new tab)
Diagnose a beta two typo using the memory window formula
A colleague meant to set adam_beta2 to 0.95 for a large pretraining style run but typed 0.5 instead. Using the memory window idea from the first exercise, compute the resulting window in steps, compare it against the intended 0.95 setting’s window, and explain in mechanism terms, not just with a word like unstable, why a per-weight brake with a two-step memory makes a run prone to blowing up.
See the worked solution (opens in a new tab)
Decide whether paged_adamw_8bit is solving a real problem for a LoRA run
You are running LoRA at rank 16 on Qwen2.5-1.5B on one 32 GB card, batch 4 and sequence 1024. Using this series’ static state figure for that LoRA setup and its order-of-magnitude activation figure for this batch and sequence length, add up the total VRAM this run actually needs, compare it against the roughly 29.8 GiB a 32 GB card has usable, and decide whether paged_adamw_8bit is solving a real problem here or is a memory trick reached for out of habit.
Frequently asked questions
What does the W in AdamW stand for?
Weight decay, and specifically decoupled weight decay. Original Adam folded the decay into the gradient, where the optimizer’s own per-weight division then distorted it: quiet weights got pulled toward zero far harder than busy ones. AdamW subtracts the decay from the weight directly, after the adaptive part, so every weight shrinks by the same percentage each step. It costs no extra memory.
Why does Adam need bias correction?
Because both running averages start at zero, so both read too low at the beginning, and they warm up at different speeds. On step one the first average holds a tenth of the true gradient while the square root of the second holds about a thirty-second of its size, so the uncorrected step would come out roughly three times the learning rate. Dividing each average by one minus its decay rate raised to the step number cancels that exactly, and the correction fades to nothing as the run goes on.
What do adam_beta1 and adam_beta2 actually control?
How much history goes into each running average. A decay rate of beta remembers roughly the last one over one minus beta steps, so beta1 at 0.9 gives the direction a memory of about 10 gradients and beta2 at 0.999 gives the per-weight brake a memory of about 1,000. The brake is deliberately slower, because it should reflect a settled estimate of a weight’s usual gradient size. Leave both alone for fine-tuning; 0.95 for beta2 shows up in large pretraining runs, where the landscape changes faster.
What is the difference between weight decay and gradient clipping?
They touch different things at different moments. Weight decay pulls every weight slightly toward zero inside the optimizer step, every step, as a form of regularisation. Gradient clipping runs before the optimizer, measures the overall size of the gradients, and scales them down only if that size crosses a limit. Clipping does nothing at all on a healthy run, which is exactly why it is safe to leave on at 1.0.
Which optim string should I use?
Use adamw_torch as the default and the fused variant for extra speed on a recent GPU at no memory cost. Move to paged_adamw_8bit when VRAM is tight or you are running QLoRA. The 8-bit variant runs the same algorithm with its two running averages stored in one byte each, so behaviour is close to unchanged and the stored state drops to about a quarter.
Does weight_decay hurt a LoRA run?
Not inherently, but its effective strength changes. LoRA trains a small adapter on top of a frozen model, so the same weight_decay value now pulls that whole small adapter toward zero every step, and it has far less capacity to absorb the pull than a full weight matrix would. Most LoRA recipes leave it at 0.0 or keep it small, around 0.01, because the low-rank bottleneck already constrains the model structurally.
Sources and further reading
- Kingma and Ba, Adam: A Method for Stochastic Optimization, the source of the two moment estimates, the bias correction and the name.
- Loshchilov and Hutter, Decoupled Weight Decay Regularization, the paper that showed the L2 route and the decay route are not the same under Adam, and added the W.
- Dettmers et al., 8-bit Optimizers via Block-wise Quantization, the method behind every 8-bit AdamW variant.
- Shazeer and Stern, Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.
- Chen et al., Symbolic Discovery of Optimization Algorithms, which introduced Lion.
- Pascanu et al., On the difficulty of training recurrent neural networks, the origin of gradient clipping and of the norm-based form used at
max_grad_norm = 1.0. - You et al., Large Batch Optimization for Deep Learning, LAMB, the per-layer adaptive rule aimed at very large batch pretraining rather than at memory.
- Micikevicius et al., Mixed Precision Training, for the fp32 master copy that makes the optimizer tenant 12 bytes rather than 8.
- Rajbhandari et al., ZeRO, where the 16 bytes per parameter accounting quoted in the last section is laid out.
- Qwen Team, Qwen2.5 Technical Report, the source of the 1.54 billion parameter count behind the 12.3 GB and 18.5 GB figures.
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, the method that removes m and v for almost every weight instead of shrinking them.
- Hugging Face TRL, SFT Trainer documentation, for the current optimizer and scheduler configuration surface.
