Supervised fine-tuning is the plainest idea in this series once you see what it actually does. You take a model somebody else already trained, show it a few thousand examples of the answers you want, and let a loop nudge its numbers toward producing more of those. There is no new architecture and no new maths. The scoring function is the one the model was pretrained with. The only real decisions are which text you train on and which parts of that text you grade.
This part runs one complete fine-tune of Qwen2.5-1.5B on a single 32 GB card, twice on purpose. First as a plain script with nothing hidden, printed in one piece so you can copy it and run it. Then through a trainer, which is the form you would actually use, where every configuration line maps onto a mechanic you have already watched work.
By the end you will be able to say which parts of a model a fine-tune can move, take one training step apart and read the single number that summarises it, tell the two masks apart and recognise the failure that happens when one goes missing, read a training script line by line, and run the checks that say whether a falling loss produced a better model. A falling loss is not evidence of anything on its own.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run (you are here)
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
Say what pretraining left behind, and what a fine-tune installs
One idea decides how you read every decision in this part, and it is usually stated backwards.
The common story says a base model has not learned to behave and a fine-tune teaches it. That is not what happens. A model that has read trillions of tokens has already absorbed almost every behaviour you could ask for: polite support replies, terse commit messages, JSON payloads, forum arguments, exam answers. Ask it anything and something fluent comes out. Brown et al. showed the strong version of this with GPT-3, a model given no instruction tuning at all: put a few worked examples in the prompt and it performs tasks nobody trained it on. The capability was already in the weights, and the prompt only selected it. What it lacks is any opinion about which of those behaviours this moment calls for.
So the honest summary is this. Pretraining does not fail to absorb behaviour. It absorbs every behaviour and holds no default policy over them. Supervised fine-tuning does not install a behaviour that was missing. It installs the default: it decides which already-absorbed behaviour surfaces when the prompt does not force one. Wei et al. demonstrated that at scale in FLAN: fine-tune a pretrained model on a large collection of tasks phrased as instructions, and it begins following instructions on task types held out of that collection entirely.
That reframing pays for itself twice. It explains why a thousand carefully written examples can turn a base model into a usable assistant, which is what Zhou and colleagues reported in the LIMA work. And it explains why ten thousand cannot teach it your product catalogue. Allen-Zhu and Li reach the same wall from the other side: they find that whether a fact can be pulled back out of a model is largely settled during pretraining, by how that fact was presented there, and that later training does not rescue one that went in badly. Selecting among absorbed behaviours is cheap. Adding a fact that was never absorbed is not.
What “absorbed” actually means
The word carries a lot of weight, so pin it down. Absorbed knowledge has four properties, and each one has a consequence.
- It is frequency-weighted. A pattern that appeared constantly presses harder on the weights than one that appeared twice. Nothing is stored once. Things are stored more or less strongly, in proportion to how often the data showed them.
- It is lossy compression. Trillions of tokens went into a few billion numbers, so nothing is kept word for word. What survives is a tendency: given this context, these continuations are likely. That is how a model can be fluent about a topic and wrong about its details in one sentence.
- It is not addressable. There is no row to look up and no field to edit. You cannot point at the parameters where “refuse politely” lives, because it does not live in one place. It is smeared across the pile.
- It depends on conditioning. Which absorbed behaviour comes out is decided by the context in front of the weights. Same model, same numbers, different prompt, genuinely different behaviour. That is why prompting works at all, and why it is unreliable. You are steering rather than setting.
Read those together and fine-tuning stops being mysterious. You cannot edit an unaddressable, lossy, frequency-weighted store directly. You can add a few thousand examples that all press in one direction, and shift which behaviour wins when nothing else is pushing.
The twelve layers, and the six a fine-tune actually moves
“Behaviour” is too coarse a word to plan a project with. Split what a model holds into twelve layers, from the mechanics of words up to what it believes about itself. This is a working taxonomy, not something the architecture enforces: none of the twelve corresponds to one of the model’s 28 stacked layers. It is a checklist for deciding whether a fine-tune is the right tool at all.
Be plain about where the list comes from. The twelve layers below are this series’ own way of organising the question, not a taxonomy you will find in a paper, and nobody else numbers them this way. The organising is ours. The claims each row makes are not, and the paragraphs after the table give the published evidence for them.
| Layer | What it holds | Does a supervised fine-tune move it? |
|---|---|---|
| 1. Lexical | which tokens exist and how they spell words | No. The tokenizer and pretraining own it. |
| 2. Syntactic | grammar, agreement, word order | No. |
| 3. Semantic | what words mean and how meanings relate | No. |
| 4. Factual | specific claims about the world | Not reliably. Use retrieval instead. |
| 5. Procedural | how to do a thing: arithmetic, valid SQL, summarising | No new skill. It can only pick among skills already there. |
| 6. Reasoning form | whether working is shown, and in what shape | Yes. |
| 7. Discourse and format | the shape of an answer: sections, keys, length | Yes, more reliably than anything else. |
| 8. Register and style | tone, diction, formality | Yes. |
| 9. Pragmatic | reading intent, when to ask rather than guess | Yes, partly. |
| 10. Persona and role | who the model is being while it answers | Yes. |
| 11. Normative | what it will and will not do | Yes, partly. Fine calibration is preference tuning. |
| 12. Meta | knowing what it does not know | Barely, and a careless fine-tune makes it worse. |
Read the last column as a boundary. Layers 1 to 5 are the base model’s competence, bought with millions of dollars of pretraining, and a well-run fine-tune barely disturbs them. Choose your base model for those five, as Part 8 does, because supervised fine-tuning will not repair them. A badly run fine-tune does damage them, and that damage has a name and a section further down.
Two lines of published work stand behind that boundary. Kaplan et al. and then Hoffmann et al. measured what pretraining scale actually buys, as smooth curves in parameters, data and compute, which is the sense in which layers 1 to 5 are paid for rather than taught. On layer 4 specifically, Allen-Zhu and Li put a capacity figure on stored facts, around 2 bits of knowledge per parameter. That is a property of the pretrained model, and no run of a few thousand examples adds to it.
Layers 6 to 11 are what you are buying: format, the shape of the reasoning, tone, role, what to refuse. Each is a selection among behaviours the model already produces, which is why a few thousand examples is enough. That is the LIMA result read row by row: Zhou et al. got a usable assistant out of 1,000 curated examples, and what those examples supplied was format, style and the habit of answering. None of it was knowledge.
Layer 12 needs a warning, because a fine-tune can make it worse while every other layer improves. Supervised fine-tuning shows one gold answer per prompt, and in a normal dataset every gold answer is confident. Train on ten thousand of those and you have taught the model that confidence is the house style, including where it should have admitted ignorance. Nothing in the loss reports it, and it is part of what preference tuning in Part 13 exists to fix.
See why one forward pass can grade a whole answer
Now the mechanics, starting with what surprises people who have used a chatbot but never trained one. You might expect training to work the way generation does, one token at a time, each guess fed back in to produce the next. It does not. Training shows the model the whole correct answer at once.
Picture a sentence on a card: “The cat sat on the mat.” Now imagine a reader who has to guess each word from the words before it, and who is told the true word immediately after every guess, right or wrong. They guess word two from word one. Then word three from the real words one and two, never from whatever they guessed a moment ago. On to the end of the card.
That is teacher forcing, and it is how every language model in this series is trained. The name and the concern both predate transformers: Lamb et al. in Professor Forcing set out the gap it leaves, since a model trained only on true prefixes has never had its own output as input, which is the mismatch that shows up at generation time. The model is never fed its own guesses during training. It is fed the true text and asked, at every position at the same time, what comes next.
Two consequences follow, and both shape every training script you will read.
- One pass grades the entire example. A 400-token answer is not 400 separate runs. It is one forward pass, meaning one trip through the model from input to output, and it produces a prediction at all 400 positions at once. Nothing lets position 10 peek at position 11 while it does this, because the architecture forbids it. That rule is the causal mask, and a section below covers what happens when it goes missing.
- The correct answers look exactly like the input. What you hand over as targets is the same run of tokens you fed in, because the token sitting at position 5 is the right answer for the prediction made at position 4.
That second point confuses everybody the first time. You pass in a list called labels that is a copy of the input. The model shifts it by one internally, scores every position with cross-entropy, and hands back a single number. You never write the shift and you never write the scoring.
You do write one thing: which positions get graded. In a chat fine-tune you want the model to learn how to answer, not how to ask, so the user’s words are switched off and only the assistant’s are scored. The switch is a label value of -100, and Part 7 explains the convention and why that particular number. A later section builds the mask by hand.
Take one training step apart, and read its gradient norm
The word “step” is overloaded here, and the overload causes one specific misreading worth killing before it starts.
A training step is not a model layer
Qwen2.5-1.5B has 28 layers stacked one on top of another. Data enters the bottom, passes through all 28, and a prediction comes out of the top. A layer is a place in the model.
A training step is not a place. It is one full turn of the training loop, and the entire model takes part in it: one group of examples going forward through all 28 layers, one loss computed at the end, one backward pass running back through all 28, and one update applied to every parameter at the same moment. Layers are not visited one per step. Every layer is visited twice in every step.
So training for 500 steps does not train some layers more than others, and a log line reading “step 400” says the loop has gone round 400 times. Throughout this part, one step means one optimizer update, which may be assembled from several smaller passes for reasons a later section explains.
The five moves of one training step
With teacher forcing in place, one step is five moves. They are the same five for every model trained by gradient descent, and everything else in a training script is arrangements around them.
- Forward pass. Push a batch of examples through the model. Out come predictions at every position.
- Loss. Compare those predictions against the true next tokens, at the positions you chose to grade. One number comes out, and lower is better.
- Backward pass. For every one of the model’s 1.54 billion numbers, work out which way it would have to move to make that one number smaller. Those directions are the gradients, one per number.
- Optimizer step. Move every number a little in the direction its gradient points. How far is a separate decision, and a later section is about it.
- Zero the gradients. Wipe the slate. This matters more than it looks: the framework adds each new gradient to whatever is already there rather than replacing it. Skip the wipe and the next step is polluted by this one. That same adding behaviour turns out to be useful on purpose, and a later section uses it.
One gradient per parameter, and one number that summarises them all
Move three deserves to be stated precisely, because the rest of this section is arithmetic on it.
The model has N parameters, and here N is about 1,540,000,000. Call them w_1 up to w_N. The loss is one number that depends on all of them. The backward pass produces one number per parameter, answering a single question: if I nudge this one parameter and hold every other still, at what rate does the loss change? That quantity is a partial derivative.
g_iis the gradient of parameter number i. One ordinary number.Lis the loss for this step, the single how-wrong-were-we score.w_iis parameter number i, one of the 1.54 billion.- The curly d is the partial derivative sign. It reads as “the rate at which the thing on top changes when only the thing underneath moves”. Part 5 builds that from a ball rolling into a valley.
One backward pass produces 1.54 billion of these, one per parameter. Together they are the gradient, a collection exactly the size of the model itself, which is why Part 2 charges 2 bytes per parameter for it.
Nobody can read 1.54 billion numbers per step, so training logs one summary: the length of that collection treated as a single very long arrow. Length from components is the Pythagoras rule, extended from two numbers to 1.54 billion. Square every component, add them up, take the square root.
- The double bars mean length, and the small 2 says which kind: the ordinary straight-line one, the same measure Pythagoras gives for two numbers.
- The tall sign is a sum. Add up what follows it, for i running from 1 to N.
g_isquared is that parameter’s gradient multiplied by itself, which makes it positive so gradients pointing opposite ways cannot cancel out.- N is the number of parameters, 1.54 billion here.
- The square root undoes the squaring, so the answer returns in the same units as one gradient.
Check it on two numbers. If the only gradients were 3 and 4, the sum of squares is 9 plus 16, which is 25, and the square root of 25 is 5. That is the whole calculation, done 1.54 billion terms wide.
Now the point people miss. That result is one scalar for the entire step. Not one per layer. Not one per parameter. One number standing for all 1.54 billion gradients at once, logged as grad_norm. It cannot tell you which layer moved or which parameter was large. It tells you how big the update wanted to be, and that diagnoses a surprising number of failures.
Where the gradient goes: two running averages and one cap
Two things consume the backward pass, and they consume different halves of what it produced.
The first is the optimizer, which works on the individual gradients rather than the summary. AdamW keeps two running averages for every parameter, updated each step.
tis the step number, som_tis this step’s value andm_(t-1)is last step’s.mis the first moment: a smoothed average of recent gradients, which gives a steadier direction than any single noisy batch.vis the second moment: an average of the squared gradients, measuring how large and erratic this parameter’s history has been.beta1andbeta2are the decay rates, normally 0.9 and 0.999. They set how much of the old average survives.g_tis this step’s gradient for this parameter. Both lines run once per parameter, so there is anmand avfor each of the 1.54 billion, which is 8 of the 16 bytes Part 2 counts. Part 6 walks the full update rule line by line.
The second consumer is the clip, and this one does use the norm. Clipping by the norm comes from Pascanu et al., who introduced it to stop one exploding update from destroying a network that had been training perfectly well until that batch. If the norm exceeds a limit you set, called max_norm, the whole gradient is scaled down until its norm equals that limit exactly.
gis the whole collection of gradients andg_clippedis what the optimizer is allowed to see.max_normis the limit, normally 1.0, set once.- The min picks the smaller of the two, so when the norm is under the limit the multiplier is exactly 1 and nothing happens.
- When the norm is over, every one of the 1.54 billion gradients is multiplied by the same fraction. A norm of 4.0 against a limit of 1.0 quarters all of them. Directions survive and only the overall size is capped.
Order matters. Clipping runs after the backward pass and before the optimizer step, exactly where the script below puts it, so the clipped gradient is what those running averages absorb. The norm is both a number you read and a control that acts.
Turn a conversation into numbers the model can train on
The model has never seen a conversation. It has no idea what a role called “user” is. It reads a flat run of integers. So something has to sit between your data and the model, and that something is the biggest source of bugs in the pipeline.
Here is one training example from the dataset this part uses, and the text it has to become.
# what one row of the dataset looks like
[{"role": "user", "content": "Name three coffee shops."},
{"role": "assistant", "content": "Uptime Coffee."}]
# what the chat template renders it into, as plain text, before tokenizing
<|im_start|>system
(a default system line)<|im_end|>
<|im_start|>user
Name three coffee shops.<|im_end|>
<|im_start|>assistant
Uptime Coffee.<|im_end|>
The tokenizer and the chat template, in that order
Two components do this work, and they are easy to confuse.
The tokenizer splits text into tokens and maps each one to an integer. It ships with the model and is specific to it. Token number 9707 means one thing in Qwen’s vocabulary and something else in another model’s, so a tokenizer and a model always have to come from the same checkpoint.
The chat template is a small text template that also ships with the tokenizer. It decides what a list of messages turns into: which markers open and close each turn, where the role name goes, and whether a default system line is inserted at the top. The markers above, the ones that look like <|im_start|>, are special tokens. A special token stands for structure rather than words, and it has its own vocabulary entry like any other token.
This is why you cannot find the answer’s starting point by counting characters in your own text. You did not write the system line or the markers. The template did, and only the template knows how long they came out.
Find the boundary by rendering the conversation twice
The fix is a trick rather than a parser. Make the template render the same conversation two ways and compare the lengths. Four steps.
- Render the whole thing. Take the full conversation, user turn and assistant turn together, run it through the template and tokenize it. Call the list of integers
full. This is what the model reads. - Render it again without the answer. Take the same conversation with the assistant’s message removed, and ask the template to add the opening scaffolding for an assistant turn on the end. The flag is
add_generation_prompt, and it exists because that is what you do at generation time: set the model up to speak and let it continue. Tokenize that too, and call itprompt. - Compare the lengths. Every token in
promptsits at the same index infull, because the second render is a prefix of the first. So the length ofpromptis the index where the answer begins. That is the boundary, in token space, with no guessing. - Build the labels. Copy
full, then overwrite everything before the boundary with -100. What is left graded is exactly the assistant’s answer, including the token that closes the turn.
Grading that closing token is not a detail. It is how the model learns to stop talking, and a mask that trims it off produces a model that answers correctly and then keeps going forever.
Tell the causal mask apart from the label mask, and catch a model that cheats
Two different masks act inside one training step, and they are constantly confused. Part 7 builds both from nothing. Here is the distinction in the form that matters when you are debugging.
One blocks attention, the other blocks scoring
The causal attention mask lives inside self-attention, the step where each position looks at other positions and mixes in what it finds useful. Its rule is one line: the position at index i may attend to positions 1 through i, and may not attend to positions i+1 through N. It is applied inside every attention layer on every pass, and you never set it. It comes with the architecture.
The label mask lives at the very end of the forward pass, at the loss. Its rule is also one line: any position whose label is -100 is skipped when cross-entropy adds up the penalties. You set this one, by hand in Version A below and through a flag in Version B.
| Causal attention mask | Label mask, the -100 convention | |
|---|---|---|
| Acts inside | self-attention, in all 28 layers | cross-entropy, once, at the end |
| Question it answers | which positions may this one look at | which positions do we score |
| Set by | the architecture, always on | you, per example |
| If it is missing | the model cheats and cannot generate | the run trains on the prompt, silently |
A label mask does not block attention flow
Here is the sentence to keep. Marking a prompt position with -100 removes it from the loss and removes nothing from attention. Those prompt tokens are still read, and later positions still attend to them, which is what you want: an answer has to be conditioned on the question.
Now run the same fact the other way, because that is where the bug hides. Labels have no say over what any position may look at. If the causal mask goes missing from the attention path, a response token at position i still attends forward to positions i+1 and beyond, whatever the labels say. Setting a label to -100 does not close an attention window, and setting a label to a real token does not open one.
Carry it as a rule. The label mask controls what is graded. The causal mask controls what is visible. Neither substitutes for the other, and confusing them produces a bug whose every symptom is reassuring.
The cheating model, step by step
Nothing in this part’s scripts can produce this failure, because AutoModelForCausalLM applies the causal mask for you. It appears when someone writes a custom attention block, ports a model by hand, or wires a bidirectional encoder-style path into a next-token objective. Know it on sight, because the loss curve actively lies about it.
- Training loss collapses toward zero, unnaturally fast. Position i is being asked to predict the token at i+1, and that token is now visible in its own input. Copying beats learning, so the model copies. A loss that should settle near 1.0 after thousands of steps instead falls to a fraction of that within a few hundred.
- Held-out loss looks perfect too. This is the check that normally saves you, and here it does not. The held-out set runs through the same buggy bidirectional path, so it can copy just as easily. Both curves agree, and both measure the same cheat.
- The gradient norm stays calm. It does not spike. Copying the next token is an easy optimisation with a smooth surface: the model finds it quickly and the gradients shrink and stay small. So the norm sits stable or drifts toward zero, which reads as a beautifully converged run. The usual advice to watch for a gradient spike will not catch this, because there is no spike to catch.
- Then inference collapses. At generation time position i+1 does not exist yet, because that token has not been produced. The model has learned a function of information that is no longer there, so it emits gibberish, or loops one phrase forever, or stops immediately. The distance between a flawless training log and a useless model is the whole signature.
Confirming it takes one forward pass. Run a single example, note the logits at some position i, then change the token at position i+1 and run it again. In a correct causal model the logits at position i cannot move. If they do, attention is looking forward and every symptom above follows.
Choose the few numbers that decide whether it learns
The five moves are fixed. What you choose is a short list of settings, called hyperparameters because you set them before training rather than the model learning them. Two of them decide whether the model learns anything useful. The rest are housekeeping.
The learning rate is the master dial
The gradient tells each weight which direction to move. It says nothing about how far. The learning rate decides that, and it multiplies every weight change in the run.
Picture walking down a hillside in thick fog. You can feel which way the ground slopes, so you know the direction. The learning rate is your stride length. Stride too far and you cross the bottom of the valley and end up higher on the other side, then overshoot again, worse each time. Stride too short and you are still on the hillside at nightfall.
Both failures are common and they look different in the logs. Too high shows up as a loss that spikes, oscillates, or turns into NaN, which is what floating-point arithmetic produces when asked something undefined and which kills a run permanently. Too low shows up as a loss that barely moves, which leads people to conclude fine-tuning does not work when really they underpowered it.
Two regimes matter, and mixing them up is one of the three most common beginner mistakes in this field.
- Full fine-tuning of a 1B to 2B model wants roughly 1e-5 to 2e-5. You are moving every pretrained weight, and large moves wreck what pretraining spent millions of dollars building.
- LoRA wants roughly 1e-4 to 3e-4, about ten times more, because only a small added patch is training and it starts from near zero. Part 11 covers LoRA.
Copy a LoRA rate into a full fine-tune and you will destroy the model. Copy a full fine-tune rate into LoRA and it will barely move. For this run, start at 2e-5.
Warm the rate up, then let it decay
Two more settings shape the learning rate over time rather than setting it once.
At step zero the optimizer has no history. AdamW sizes each weight’s step from the two running averages above, and on the first step those averages are empty, so the sizing is a guess. A full-size step on a guess can damage the model before training has begun. Warmup ramps the rate from zero to the target over the first few percent of steps.
Then a schedule brings it down. Cosine decay is the reliable default: a smooth curve from the peak toward zero across the rest of the run, so late training polishes instead of shoving. Set warmup as a fraction of the run rather than a step count, so it still makes sense when the dataset or epoch count changes.
Micro-batches, accumulation, and the batch you actually train with
A batch is a group of examples processed in one go. Bigger batches give a steadier gradient, because averaging over 32 examples cancels the quirks of any single one. So bigger is better, up to the point where memory runs out: every example in a batch has its activations stored until the backward pass consumes them, and Part 3 is about that cost.
Gradient accumulation breaks that ceiling. Run four small batches one after another without stepping. Their gradients add up in place, which is the framework’s default, as move five pointed out. After the fourth, take one optimizer step and wipe.
The result is an update built from 16 examples while only 4 examples of activations were ever in memory. The small batch that fits is the micro-batch. The number that matters for training behaviour is the effective batch:
effective batch = per-device batch x accumulation steps x number of GPUs
Here that is 4 times 4 times 1, so 16. For supervised runs on small models, aim for 16 to 64. The effective batch is the number worth quoting because it is the one the optimizer sees, and it is the number Goyal et al. tie the learning rate to in their linear scaling rule for SGD: multiply the batch by k and multiply the rate by k with it.
Divide the loss by the accumulation count
Accumulation comes with one trap that has quietly ruined a lot of runs, and it is invisible.
Gradients are summed across the four micro-batches, not averaged. So the total handed to the optimizer is four times what one batch would have produced, and the optimizer multiplies that by the learning rate. You have taken a step four times too large while the config still says 2e-5.
The fix is to divide each micro-batch’s loss by the accumulation count before its backward pass. Four quarters add up to one whole, and the sum becomes an average again. Without it, the optimizer sees a gradient four times larger than the one your configuration describes. It is tempting to say that this acts like a learning rate four times higher, and with plain SGD it would. With AdamW it does not, and the reason is worth knowing. AdamW divides the step by the square root of the second moment, so multiplying every gradient by four multiplies the first moment by four and the second moment by sixteen, and the square root of sixteen is four. The two fours cancel. AdamW is very nearly blind to a constant rescaling of the gradient.
What is not blind to it is gradient clipping. max_grad_norm acts on the raw norm, before any of that normalisation, so a gradient four times too large trips the clip on nearly every step. Once the clip is firing constantly, the update is set by the clip rather than by your learning rate, and the run is no longer the one you configured. That is the observable to watch for: not a divergent loss, but clipping on almost every step and a loss that will not settle. A trainer does this division for you. A hand-written loop does not.
Starting values that work
| Setting | Start here | What it does, and how it fails |
|---|---|---|
learning_rate |
2e-5 | stride length; higher diverges, lower barely moves the model |
lr_scheduler_type |
cosine | the shape of the decay; smooth and reliable |
warmup_steps |
0.03 | a value below 1 is read as a fraction of the run, so this is 3 percent |
num_train_epochs |
1 to 3 | one epoch is one pass over the data; more invites memorisation |
per_device_train_batch_size |
4 | as high as memory allows; drop to 2 if the run runs out of memory |
gradient_accumulation_steps |
4 | gives an effective batch of 16 on one card |
bf16 |
True | 16-bit arithmetic; halves activation memory and runs faster |
gradient_checkpointing |
True | the out-of-memory to fits switch, at roughly 20 to 30 percent slower |
max_grad_norm |
1.0 | caps the size of one update so a strange batch cannot blow up the model |
weight_decay |
0.01 | a small pull on every weight toward zero, which curbs overfitting |
max_length |
1024 | tokens per example; longer costs more activation memory |
seed |
pin it | fixes the shuffling, so two runs are actually comparable |
Change one dial at a time and watch the loss curve. The order of impact is the learning rate first, then the effective batch, then the memory settings and only if you ran out of memory. Three clean single-variable runs teach you more than thirty tangled ones.
Make a 1.5B full fine-tune fit on one 32 GB card
Before running anything, check that it fits. Qwen2.5-1.5B has 1.54 billion parameters, a figure from Qwen’s own technical report, and standard mixed-precision AdamW costs 16 bytes per parameter of static state, meaning weights plus gradients plus optimizer state. Part 2 derives all sixteen of those bytes.
1.54e9 x 16 bytes = 24.6 GB of static state
Compare like with like, because card memory is quoted in the binary unit. 24.6 GB is 22.9 GiB, and a 32 GB card offers roughly 29.8 GiB once the driver has taken its share. That leaves about 6.9 GiB, and 2 to 3 GB of it should stay free as a safety margin.
Activations have to fit in what is left. At a batch of 4 and a sequence length of 1024 they run to roughly 8 GB, about 7.5 GiB. Treat that as an order of magnitude rather than a measurement, because the exact figure depends on which intermediate values the framework keeps. Either way it does not fit into 6.9 GiB, so the run needs help. Three settings provide it.
bf16 mixed precision
Do the arithmetic in a 16-bit format for speed and memory, while keeping a 32-bit copy of the weights that the optimizer actually updates. The 32-bit copy exists because a 16-bit number cannot record an update far smaller than itself, so without it thousands of tiny changes would round away. Use bf16 rather than fp16, because fp16’s smaller range makes small gradients vanish and forces a loss-scaling workaround. Part 4 covers the formats.
Gradient checkpointing
Do not store every layer’s activations. Store a few, throw the rest away, and recompute what is missing during the backward pass. You buy a large cut in activation memory for roughly 20 to 30 percent more time, about 20 percent by Hugging Face’s own documentation and about 30 percent in the paper that introduced the technique. This is usually the setting that turns an out-of-memory crash into a working run.
An 8-bit optimizer, only if you are still short
AdamW’s two running averages are normally 4 bytes each. An 8-bit optimizer stores them in 1 byte each, taking the optimizer’s share from 12 bytes per parameter to about 6, since the 32-bit master copy is untouched, and the total from 16 to about 10. Reach for it only when the first two settings have not left enough room.
Run a complete fine-tune from one script (Version A, the bare loop)
Here is the whole thing in one piece. It downloads a model, prepares about 10,000 examples, masks the labels by hand, trains for one epoch and saves the result. Nothing is hidden inside a framework.
Read it once without worrying about individual lines. The next section takes it apart block by block, including the lines that look like decoration and are not.
import torch
from torch.utils.data import DataLoader
from transformers import (AutoModelForCausalLM, AutoTokenizer,
get_cosine_schedule_with_warmup)
from datasets import load_dataset
dev, mid = "cuda", "Qwen/Qwen2.5-1.5B" # the BASE checkpoint, not -Instruct
# 1. the tokenizer and the model
tok = AutoTokenizer.from_pretrained(mid)
model = AutoModelForCausalLM.from_pretrained(mid, dtype=torch.float32) # explicit fp32
model = model.to(dev)
model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
model.config.use_cache = False # caching and recomputation cannot both be on
model.train()
# 2. one example in, input ids and labels out
def build(example, tok, max_len=1024):
msgs = example["messages"]
# grade the LAST assistant turn, mask everything before it
last = max((i for i, m in enumerate(msgs) if m["role"] == "assistant"), default=-1)
# return_dict=False is required: this call returns a dict of tensors by
# default, and what this function wants back is a plain list of token ids
full = tok.apply_chat_template(msgs, tokenize=True, return_dict=False,
add_generation_prompt=False)
prompt = tok.apply_chat_template(msgs[:last], tokenize=True, return_dict=False,
add_generation_prompt=True)
full = full[:max_len]
labels = list(full)
for i in range(min(len(prompt), len(full))):
labels[i] = -100 # switch off the prompt region
return {"input_ids": full, "labels": labels}
# 3. the dataset and the batches
ds = load_dataset("HuggingFaceH4/no_robots", split="train").map(
lambda ex: build(ex, tok),
remove_columns=["messages", "prompt", "prompt_id", "category"])
ds = ds.filter(lambda r: any(l != -100 for l in r["labels"])) # drop what truncation ate
def collate(rows): # pad to the longest row in the batch
m = max(len(r["input_ids"]) for r in rows)
ii = [r["input_ids"] + [tok.pad_token_id] * (m - len(r["input_ids"])) for r in rows]
am = [[1] * len(r["input_ids"]) + [0] * (m - len(r["input_ids"])) for r in rows]
lb = [r["labels"] + [-100] * (m - len(r["labels"])) for r in rows]
return torch.tensor(ii), torch.tensor(am), torch.tensor(lb)
loader = DataLoader(ds, batch_size=4, shuffle=True, collate_fn=collate)
# 4. the optimizer and the schedule
accum, epochs = 4, 1 # effective batch = 4 x 4 = 16
opt = torch.optim.AdamW(model.parameters(), lr=2e-5, weight_decay=0.01)
total = (len(loader) // accum) * epochs
sched = get_cosine_schedule_with_warmup(opt, int(0.03 * total), total)
# 5. the loop
for ep in range(epochs):
for step, (ii, am, lb) in enumerate(loader):
ii, am, lb = ii.to(dev), am.to(dev), lb.to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
out = model(input_ids=ii, attention_mask=am, labels=lb) # 1 FORWARD, 2 LOSS
(out.loss / accum).backward() # 3 BACKWARD
if (step + 1) % accum == 0:
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step(); sched.step(); opt.zero_grad() # 4 STEP, 5 ZERO
if step % 80 == 0:
print(f"ep{ep} step{step} loss {out.loss.item():.3f}")
model.save_pretrained("qwen15b-sft/final"); tok.save_pretrained("qwen15b-sft/final")
One caution before you run it. These library APIs move between releases, so pin your versions and check the current documentation before a real run. Everything here is current as of August 2026.
Read Version A one block at a time
Five blocks, in the order the script runs them. Every excerpt is lifted verbatim from above.
Block 1: load the tokenizer and the model
tok = AutoTokenizer.from_pretrained(mid)
model = AutoModelForCausalLM.from_pretrained(mid, dtype=torch.float32)
model = model.to(dev)
model.train()
The two from_pretrained calls download a checkpoint from the model hub, cache it on disk, and build the object. Both are given the same identifier, which is the point: a tokenizer and a model must come from the same checkpoint or the integers mean different things at each end.
The word “CausalLM” in the class name means a model that only ever looks left, predicting each token from the ones before it. That class applies the causal mask discussed above, and attaches the output head that scores every token in the vocabulary, which is what makes a loss computable.
The next line moves the model onto the graphics card. Weights are built in system memory first, and model.to("cuda") copies them across. Everything in one computation has to live on the same device, which is why the loop later moves each batch across too.
Then model.train(), which trains nothing. It flips the model into training mode, changing the behaviour of a few layers that act differently while learning. Its counterpart is model.eval(), and forgetting to switch is a classic source of results that do not reproduce.
Why the number format is written out in full
The argument dtype=torch.float32 says which number format the weights are stored in, and it is the least decorative line in the script.
Leave it off and the loader does not quietly fall back to 32-bit. It reads the format declared in the checkpoint’s own configuration file, and Qwen2.5 declares bf16. So omitting the argument loads 16-bit weights, and nothing warns you.
That single omission changes what you are running. With 16-bit weights there is no 32-bit master copy for the optimizer to update, so the tiny updates a low learning rate produces round away against a weight that cannot represent them. The torch.autocast block further down becomes a no-op, because everything is already 16-bit. You get half the memory and none of the protection, and the run still looks fine while it happens.
With the weights genuinely in 32-bit, autocast casts values to bf16 for each forward pass while the stored weights, the gradients and the optimizer’s two averages stay 32-bit. Count the bytes for one parameter: 4 for the weight, 4 for its gradient, 4 and 4 for the averages. That is the same 16 bytes per parameter as the canonical 2 plus 2 plus 12 split, arranged differently.
Why gradient checkpointing takes a keyword argument
model.gradient_checkpointing_enable(gradient_checkpointing_kwargs={"use_reentrant": False})
model.config.use_cache = False
PyTorch has two implementations of checkpointing. The older one, called reentrant, needs at least one input to each checkpointed block to be carrying a gradient, which stops working the moment part of the model is frozen. The newer one has no such limitation and PyTorch’s documentation recommends it. It is not the default in every code path, so pass it explicitly.
The second line switches off the cache. During generation a model saves intermediate work for every token it has produced so it does not redo it for the next one. Training has no generation loop, so nothing is reused, and the cache conflicts with recomputing activations anyway. Set it yourself so the intent is written down.
Block 2: turn one example into input ids and labels
This is the build function, the boundary trick in code. Three excerpts.
last = max((i for i, m in enumerate(msgs) if m["role"] == "assistant"), default=-1)
Find the position of the last assistant message. The default=-1 stops the line crashing on a malformed row with no assistant turn at all.
full = tok.apply_chat_template(msgs, tokenize=True, return_dict=False,
add_generation_prompt=False)
prompt = tok.apply_chat_template(msgs[:last], tokenize=True, return_dict=False,
add_generation_prompt=True)
The two renders. The first takes every message and adds no generation scaffolding, because the answer is already there. The second stops before the last assistant message and does add the scaffolding, so it ends exactly where the answer would start.
The argument return_dict=False is not optional. This call now defaults to returning a dictionary of tensors, and the rest of the function wants a plain Python list of integers so it can slice and copy it. Leave the argument off and the following lines will not do what they appear to do.
full = full[:max_len]
labels = list(full)
for i in range(min(len(prompt), len(full))):
labels[i] = -100
Cut the sequence at the maximum length. Copy it to make the labels, since the targets are the inputs. Then walk from the start to the boundary and switch every position off. The min() guard covers the case where truncation cut into the prompt itself, so the loop never runs past the end of the list.
That is the job the trainer’s assistant_only_loss flag does automatically. Doing it by hand once shows there is no magic in it, and it leaves you a method that works when a template will not cooperate.
Block 3: build the dataset and the batches
ds = load_dataset("HuggingFaceH4/no_robots", split="train").map(
lambda ex: build(ex, tok),
remove_columns=["messages", "prompt", "prompt_id", "category"])
ds = ds.filter(lambda r: any(l != -100 for l in r["labels"]))
The first call downloads about 10,000 human-written instruction and response pairs. The .map() runs build over every row. The remove_columns argument throws away the original text columns, which matters: a batch has to end up a rectangle of numbers, and leftover strings break that.
The filter is the interesting line. Some answers sit past the truncation point, so cutting at 1024 tokens removes the graded region and leaves an example whose labels are all -100. It produces no loss and no gradient, still costs a forward pass, and is invisible in the logs. A whole batch of them is worse: there is nothing to average, and the loss comes back as NaN.
def collate(rows): # pad to the longest row in the batch
m = max(len(r["input_ids"]) for r in rows)
ii = [r["input_ids"] + [tok.pad_token_id] * (m - len(r["input_ids"])) for r in rows]
am = [[1] * len(r["input_ids"]) + [0] * (m - len(r["input_ids"])) for r in rows]
lb = [r["labels"] + [-100] * (m - len(r["labels"])) for r in rows]
return torch.tensor(ii), torch.tensor(am), torch.tensor(lb)
Examples have different lengths and a batch is a rectangle, so short rows get filler tokens until every row matches the longest. That filler is called padding, and this function builds three parallel lists out of it.
iiis the token ids, padded with the tokenizer’s dedicated padding token.amis the attention mask: a 1 for every real token and a 0 for every pad. This is a third mask, the one that stops real tokens attending to filler.lbis the labels, padded with -100, so the filler is never graded.
Then torch.tensor turns each nested Python list into a tensor, which is a grid of numbers with a fixed shape. The DataLoader line hands out groups of four rows, reshuffles between epochs, and calls this function on every group.
Block 4: the optimizer and the schedule
opt = torch.optim.AdamW(model.parameters(), lr=2e-5, weight_decay=0.01)
total = (len(loader) // accum) * epochs
sched = get_cosine_schedule_with_warmup(opt, int(0.03 * total), total)
The call to model.parameters() hands the optimizer every trainable number in the model. That one argument is what makes this a full fine-tune rather than a LoRA run. The weight decay is a small pull on every weight toward zero at each step.
The middle line is easy to get wrong. The schedule counts optimizer steps, not batches, and only every fourth batch takes a step. So the total is batches divided by the accumulation count, times the epochs. Give the scheduler the wrong total and the cosine curve finishes in the wrong place, either bottoming out early or never reaching the bottom. The int(0.03 * total) is the 3 percent warmup from the table.
Block 5: the loop itself
with torch.autocast("cuda", dtype=torch.bfloat16):
out = model(input_ids=ii, attention_mask=am, labels=lb) # 1 FORWARD, 2 LOSS
(out.loss / accum).backward() # 3 BACKWARD
The with line opens a block inside which the heavy arithmetic runs in bf16 while the stored weights stay 32-bit. That is mixed precision, switched on for exactly the region that benefits.
Passing labels into the model call makes the model compute the loss itself. It shifts the labels by one, runs cross-entropy at every position, skips every position marked -100, and returns an object whose .loss attribute is a single number. Moves one and two, in one call.
Then .backward() is move three. It works out a gradient for every parameter and adds it into that parameter’s own gradient slot. Note the division in front of it: four quarter-sized contributions summing to one whole.
if (step + 1) % accum == 0:
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step(); sched.step(); opt.zero_grad() # 4 STEP, 5 ZERO
Everything in this block happens once every four micro-batches. First the clipping, which is the formula from the gradient-norm section, with max_norm set to 1.0. The limit needs no tuning.
Then the three calls that end a step. opt.step() applies the update, move four. sched.step() advances the learning rate one position along the warmup and cosine curve. opt.zero_grad() is move five, wiping the slots so the next four micro-batches start from nothing.
The last two lines log and save. out.loss.item() pulls that number off the card into ordinary Python. save_pretrained writes the weights and config into a folder, and the tokenizer is saved beside them because a model without its tokenizer is unusable.
Get the same run from a trainer (Version B), with every config line explained
Nobody writes that loop in production. A trainer already contains it, plus evaluation, checkpointing, logging, resuming after a crash and running across several cards. What Version A buys you is the ability to read Version B with none of it a black box.
import torch
from datasets import load_dataset
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTConfig, SFTTrainer
mid = "Qwen/Qwen2.5-1.5B"
tok = AutoTokenizer.from_pretrained(mid)
# borrow the Instruct checkpoint's chat template so assistant_only_loss below can
# find the markers it needs. Same tokenizer, no new tokens, no resize.
tok.chat_template = AutoTokenizer.from_pretrained(mid + "-Instruct").chat_template
model = AutoModelForCausalLM.from_pretrained(mid, dtype=torch.float32) # explicit fp32 again
ds = load_dataset("HuggingFaceH4/no_robots") # this one has train and test splits
cfg = SFTConfig(
output_dir="qwen15b-sft",
# the mechanic
assistant_only_loss=True, # mask everything but assistant turns
eos_token="<|im_end|>", # what Qwen2.5's template closes turns with
max_length=1024,
# the knobs
learning_rate=2e-5, # FULL fine-tune regime
lr_scheduler_type="cosine", warmup_steps=0.03, # a float below 1 is read as a ratio
num_train_epochs=1,
per_device_train_batch_size=4, gradient_accumulation_steps=4,
bf16=True, gradient_checkpointing=True,
max_grad_norm=1.0, weight_decay=0.01,
# watch it
logging_steps=10, eval_strategy="steps", eval_steps=100,
save_strategy="steps", save_steps=100, seed=42,
load_best_model_at_end=True, metric_for_best_model="eval_loss", greater_is_better=False,
)
trainer = SFTTrainer(model=model, args=cfg,
train_dataset=ds["train"], eval_dataset=ds["test"],
processing_class=tok) # NOT tokenizer=, which is the old API
trainer.train()
trainer.save_model("qwen15b-sft/final")
Every line in that config stands for something you have watched run.
| Config line | What it replaces from Version A |
|---|---|
assistant_only_loss=True |
the whole build function, including the double render and the -100 loop |
eos_token |
nothing; this one is new, and the next subsection is about it |
max_length=1024 |
the truncation line, full = full[:max_len] |
learning_rate, weight_decay |
the same two arguments passed to torch.optim.AdamW |
lr_scheduler_type, warmup_steps |
the call to get_cosine_schedule_with_warmup, and its step arithmetic |
per_device_train_batch_size, gradient_accumulation_steps |
the loader’s batch size, the every-fourth-step condition, and the division by the accumulation count |
bf16=True |
the torch.autocast block |
gradient_checkpointing=True |
the enable call and the cache switch, both already the default here |
max_grad_norm=1.0 |
the call to clip_grad_norm_ |
| the eval and save lines | nothing; Version A never measured itself on held-out data |
seed=42 |
nothing; it pins the shuffling so two runs can be compared |
The model is loaded with an explicit 32-bit dtype here too, for the reason given above, and it is the bf16=True config line rather than the loading call that switches on mixed precision. Drop the explicit dtype and this becomes a pure 16-bit run with no master copy, the same trap as before.
The chat template patch, in four steps
One line looks like an odd hack, and it is worth understanding rather than copying.
- What the flag needs. Setting
assistant_only_loss=Trueasks the library to grade only the assistant’s tokens. For that it needs the chat template to carry markers, written{% generation %}, wrapping the assistant’s part so the tokenizer can report which positions those are. - What happens without them. The trainer will patch in a known-good template, but it recognises templates by comparing the whole template string against ones it already knows. It does not read the template’s logic.
- Why the base checkpoint misses. The base Qwen2.5-1.5B template differs from the Instruct one in exactly one string, its default system prompt. That is enough for the comparison to fail, so the trainer raises an error even though the base template is perfectly valid.
- The fix. Copy the Instruct checkpoint’s template string onto the base tokenizer, the line in the script above. The two checkpoints share a tokenizer and a vocabulary, so this introduces no new tokens and needs no resize. It is a string swap.
Two fallbacks if you would rather not. Drop to Version A, which masks explicitly and always works. Or point chat_template_path at a training-ready template from another model, such as chat_template_path="HuggingFaceTB/SmolLM3-3B", which the library’s own documentation suggests, bearing in mind that this one does add new special tokens and resize the embedding matrix before you have trained a step.
Two end tokens, and why you name both
The eos_token line is the other one that looks like decoration and is not.
A model stops talking when it produces its end token. Qwen2.5 has a complication: its tokenizer’s default end token is <|endoftext|>, while its chat template closes every assistant turn with a different token, <|im_end|>. Those are two separate vocabulary entries.
So the training data teaches the model to finish an answer by emitting the second, and the default configuration is watching for the first. Unless you say otherwise, generation never sees the stop signal the model was trained to produce, and the model keeps writing. Naming the template’s own token in the config keeps training and stopping consistent.
What drifts in this API, and what to pin
This surface changes between releases, and four details above are the current spellings rather than the ones in older tutorials.
- Data arguments such as the maximum length live on the config object, not on the trainer.
- The tokenizer is passed as
processing_class=. The oldertokenizer=is not the current name. - The length knob is
max_length=, notmax_seq_length=. - There is no
warmup_ratioargument any more. Usewarmup_steps, where a float below 1 is read as a ratio of the total steps.
Pin your library versions and check the current documentation before a real run. The names above are current as of August 2026.
Watch three numbers while the run is going
Three things are worth a glance, and only one of them is the loss.
Training loss. Watch the shape more than any value. A smooth descent that flattens out is healthy. A jagged or rising curve means the learning rate is too high, and a nearly flat one means it is too low. For this model and dataset a run typically starts around 1.5 to 2.5 and flattens approaching 1.0, illustrative for this setup rather than a target. A loss heading for zero is not a triumph, and the cheating-model section says why.
Gradient norm. The single scalar derived earlier, logged every step and read against the clip limit of 1.0. It moves before the loss curve does. Comfortably under the limit with clipping on the odd batch is healthy: an unusual example is being caught, which is the point. Clipping on nearly every step means the raw gradient is consistently larger than the model can absorb, so lower the learning rate rather than raising the limit, which only hides it. A norm near zero from the first step means the steps are too small to move the model off the base. The full symptom table is the diagnostic matrix further down.
GPU memory. Near full on the training card is expected. What is a warning sign is a second card starting to move when you did not intend to use it, which means your device isolation did not take effect.
On timing: for 10,000 examples at an effective batch of 16 and a sequence length of 1024, one epoch of full fine-tuning on a single modern card takes on the order of tens of minutes. Your first run’s job is not a great model. It is a correct run you understand end to end.
Prove the model actually changed
Here is the trap that separates people who have run a fine-tune from people who understand one. The training loss went down, so it worked. Not necessarily. A falling training loss proves one thing: the model fitted the signal on the tokens you graded. Whether that produced the behaviour you wanted is a separate question, with two answers.
The first is a second loss, measured on data the model never trained on. That set is the holdout, and Version B computes it every 100 steps.
While both curves fall, the model is generalising, meaning it is learning the pattern rather than the examples. When the training loss keeps falling and the held-out loss turns upward, it has started memorising, and further epochs make it worse on anything new. On a 10,000-example set that turn can come quickly, which is why the recommendation starts at one epoch.
So your model is the checkpoint at the held-out minimum, not the one from the last step. Version B’s load_best_model_at_end line does that bookkeeping for you.
The before-and-after check
Loss curves are necessary and not sufficient. The real proof is generating from the base model and from yours, on prompts neither trained on, and reading the answers side by side.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B")
msgs = [{"role": "user", "content": "Give me three names for a coffee shop near a data centre."}]
enc = tok.apply_chat_template(msgs, add_generation_prompt=True,
return_tensors="pt", return_dict=True) # dict form carries the mask
prompt_len = enc["input_ids"].shape[1]
eos_id = tok.convert_tokens_to_ids("<|im_end|>") # Qwen2.5's real end-of-turn token
for path in ["Qwen/Qwen2.5-1.5B", "qwen15b-sft/final"]: # base first, then yours
m = AutoModelForCausalLM.from_pretrained(path, dtype=torch.bfloat16).to("cuda")
out = m.generate(**enc.to("cuda"), max_new_tokens=120, do_sample=False,
eos_token_id=[tok.eos_token_id, eos_id]) # the default AND the template's
print(f"\n=== {path} ===\n", tok.decode(out[0][prompt_len:], skip_special_tokens=True))
Three details are worth naming. This call asks for return_dict=True, the opposite of the training script, because here the attention mask that comes with the dictionary is wanted. The do_sample=False makes generation deterministic, so the comparison is between two models rather than two dice rolls. And eos_token_id names both stop tokens from the subsection above.
The shape of the contrast, illustrative rather than a captured run, looks like this.
=== Qwen/Qwen2.5-1.5B (base) ===
Coffee shop names can come from many places: your neighbourhood, a favourite
memory, or a pun on coffee itself. Think about the feeling you want customers
to have, then brainstorm broadly before narrowing down to a final choice.
=== qwen15b-sft/final (fine-tuned) ===
1. Uptime Coffee
2. Redundant Grounds
3. The Failover Cafe
Read that against the twelve layers. Nothing in the second answer needed a fact the base model lacked. What changed is layer 7, the shape of the answer, and layer 9, taking the request as a request rather than as a topic.
If your fine-tune answers correctly and then keeps going, check the end token setup first: it is the more common cause and the quicker to rule out. Only then check whether the closing token was inside the graded region during training, because a model that got no gradient signal on it has no reason to stop.
The failure the loss cannot show you
Full fine-tuning moves every weight, so teaching it your format can erode abilities it already had: arithmetic, code, general knowledge, which is to say layers 4 and 5. This is catastrophic forgetting, and your held-out loss will not reveal it, because that holdout is made of text that looks like your training data.
Catching it needs a separate probe on tasks that look nothing like your data, even a handful you check by hand. You do not have to fix it here. You need to see that it exists, because it is the problem LoRA reduces by freezing the base model. Part 15 turns that probe into a procedure: score the same general tasks on the base and on your candidate, and read the difference rather than the bare number.
The other half of the answer is upstream. Every choice here assumed the data was good, and that assumption does more work than any hyperparameter in the table. Part 10 is about the data itself.
Run the whole project, from Stage 0 to the decision
Everything above is one run. A project is a sequence, and the order is most of what keeps it honest. The stages are numbered from zero, because the first happens before you have anything to train.
| Stage | What you do | Done when |
|---|---|---|
| 0 | Write down the target behaviour and build the evaluation for it, before anything else | you can score the base model and believe the number |
| 1 | Choose the base model and the bench (Part 8) | the model loads and one forward pass runs |
| 2 | Collect and curate the data (Part 10) | you have read a random sample of it yourself |
| 3 | Split into train, validation and a test set you touch once | no example appears in two splits |
| 4 | Build the transformation and verify one batch by hand | the decoded graded tokens are exactly the assistant’s words |
| 5 | Size the run: static state plus activations against the card | a single step runs without an out-of-memory error |
| 6 | Smoke run on 50 examples, deliberately overfitting them | the loss goes almost to zero, proving the loop can learn at all |
| 7 | The real run, watching loss, gradient norm and memory | the held-out curve has reached its minimum |
| 8 | Evaluate in three parts, below | all three have a number or a written verdict |
| 9 | Decide: ship, iterate, or stop | you can name which one and why |
Stage 6 is the one people skip and the one that saves days. If a deliberate overfit on 50 examples cannot drive the loss near zero, the problem is the loop or the mask, and no amount of real training will fix it.
Stage 8 has three parts because one measurement cannot see all three ways a fine-tune fails.
- In-distribution. Held-out loss and a task score on data that looks like your training data. Says whether it learned the job.
- Out-of-distribution. A probe on tasks that look nothing like your data. Says whether it forgot something, and nothing else in the pipeline can see that.
- Behavioural. Real generations from the base and the candidate side by side, on prompts neither trained on. Says whether a person would call it better.
None of the three substitutes for the others. Stage 9 compares all three against the baseline written at Stage 0. If you cannot say whether to ship, the missing piece is almost always Stage 0 rather than the model.
A diagnostic matrix for a run that has gone wrong
Read the loss, the gradient norm and one sample of output together. Individually they are ambiguous. Together they identify most first-run failures.
| What the loss does | What the gradient norm does | What generation looks like | Root cause | First move |
|---|---|---|---|---|
| spikes, oscillates, or turns to NaN | pinned at the clip limit, or jumps by orders of magnitude | degraded, repetitive, or nonsense | learning rate too high | halve the rate and restart from the last good checkpoint |
| almost flat from the first step | near zero from the first step, clipping never fires | indistinguishable from the base model | learning rate too cold | raise the rate before adding epochs |
| collapses toward zero fast, and the held-out curve agrees | calm and stable, or drifting to zero, with no spike | gibberish, or one phrase repeating | the causal mask is missing from the attention path | check whether logits at position i respond to the token at i+1 |
| falls to a high plateau and stays there | normal | continues or rephrases the question instead of answering | prompt tokens were never masked to -100 | decode one batch’s graded tokens and read them |
| jagged and will not settle, at a rate that should be safe | clipping fires on nearly every step | degraded | micro-batch losses were not divided by the accumulation count | divide by the accumulation count, or let the trainer do it |
| NaN on the very first step | NaN | nothing to inspect | a batch with no graded positions, or fp16 without loss scaling | filter all-masked examples, and use bf16 |
Two habits make most of this table unnecessary. Change one setting at a time, and read the gradient norm alongside the loss rather than after it.
Key takeaways
- Pretraining absorbs every behaviour and holds no default policy over them. A fine-tune installs the default, which is why it moves layers 6 to 11 reliably, leaves 1 to 5 alone, and can quietly damage layer 12.
- Teacher forcing feeds the model the whole true sequence at once, so one forward pass grades every position. That is why the labels look like the input, and why the prompt boundary comes from rendering the conversation twice.
- A training step is one turn of the loop across every layer at once. Its backward pass makes one partial derivative per parameter, and the gradient norm is one scalar summarising all 1.54 billion of them.
- The causal mask gates attention and the label mask gates scoring, and neither substitutes for the other. A missing causal mask gives near-zero training and held-out loss with a calm gradient norm, then gibberish at inference.
- Full fine-tuning wants a learning rate near 2e-5 and LoRA roughly ten times more. Accumulation buys a large effective batch cheaply, and forgetting to divide by the accumulation count silently multiplies your learning rate.
- Two lines that look like decoration are not: the explicit 32-bit dtype at load time, and the end-of-turn token, which decides whether the model stops or rambles.
- A falling training loss proves nothing alone. Keep the checkpoint at the held-out minimum, and evaluate in three parts: in-distribution, out-of-distribution, and by reading real output beside the base.
You can now
- Say what pretraining left behind and which of the twelve layers a fine-tune can actually move, from “Say what pretraining left behind, and what a fine-tune installs”.
- Take one training step apart, compute a gradient norm by hand from two numbers, and say what that scalar can and cannot tell you, from “Take one training step apart, and read its gradient norm”.
- Locate the boundary between prompt and answer in token space without parsing anything, from “Turn a conversation into numbers the model can train on”.
- Tell the causal mask apart from the label mask, and recognise a cheating model from its suspiciously perfect loss curve, from “Tell the causal mask apart from the label mask, and catch a model that cheats”.
- Read a complete training script line by line, then map every line of a trainer config onto the mechanic it replaces, from “Read Version A one block at a time” and “Get the same run from a trainer”.
- Run a project from Stage 0 to the decision, and diagnose a bad run from its loss curve, its gradient norm and its output together, from “Run the whole project, from Stage 0 to the decision”.
Glossary
- 8-bit optimizer
- AdamW with its two running averages stored in 1 byte each instead of 4, using block-wise quantisation. The algorithm and its behaviour are unchanged, and it reclaims most of 8 bytes per parameter. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
- Activation
- Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
- Activation memory
- The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
- AdamW
- Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Attention
- The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
- Backward pass
- Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
- Base model
- The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory.
- bf16
- A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
- Catastrophic forgetting
- When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
- Causal language model
- A model that only ever looks left, predicting each token from the ones before it. Every model in this series is one, as opposed to a model such as BERT that reads in both directions. The complete inference path
- Causal mask
- The rule inside attention that stops each position seeing anything later in the sequence. It is built into the architecture, you never set it, and it is what lets one pass train next-token prediction at every position without cheating. Masking in LLM training
- Chat template
- The rule that turns a list of role-and-content messages into the exact text and special tokens the model expects to see. It ships with the tokenizer, and it is where the boundary between prompt and response lives.
- Checkpoint
- A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name.
- Collator
- The small piece of code that takes several examples and assembles one batch out of them: padding them to a common length, building the attention mask, and setting padded labels to -100. Masking in LLM training
- Cosine schedule
- A learning-rate schedule that decays smoothly from the peak down toward zero along a cosine curve. It is the reliable default for fine-tuning. Hugging Face TRL, SFT Trainer
- Cross-entropy
- A score for how wrong a prediction was. It is small when the model gave high probability to the token that actually came next, and large when it did not. How a neural network learns
- Early stopping
- Keeping the checkpoint from the point where held-out loss was lowest, instead of the one from the last step. Past that point the model is memorising rather than learning.
- Effective batch
- The number of examples that actually go into one weight update: the per-device batch times the accumulation steps times the number of GPUs. This is the number that matters for training behaviour, rather than the batch that happens to fit.
- Embedding
- The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
- End-of-turn token
- The special token that marks where an answer stops. If it is left out of the graded region during training, or the wrong one is passed at generation time, the model rambles past the end of its answer. Masking in LLM training
- Epoch
- One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one.
- Exponent
- The part of a floating-point number that says how big or small it is, in powers of two. More exponent bits means the format reaches further before values overflow or underflow. Number formats for training
- First moment
- Adam’s running average of the gradient, written m. It gives a smoothed direction to move in, so the path stops zig-zagging on noisy batches. Gradients and optimizers explained
- Forward pass
- Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
- fp16
- A 16-bit number format with 5 exponent bits and 10 mantissa bits. Precise for its size but short of reach, so small gradients underflow to zero and loss scaling becomes mandatory. Number formats for training
- fp32
- 32-bit floating point, with 8 exponent bits and 23 mantissa bits. It is the precise reference format, used for the master copy of the weights and for the optimizer’s running averages. Number formats for training
- fp32 master weights
- The full-precision copy of the weights that the optimizer actually updates in a mixed-precision run. It exists because a 16-bit weight cannot record an update far smaller than itself, so without it the updates round away and training stalls. Micikevicius et al., Mixed Precision Training
- Full fine-tuning
- Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
- Generalisation
- How well a model does on inputs it never saw during training. It is the thing you actually want, and training loss does not measure it. Evaluating a fine-tune
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Gradient accumulation
- Running several small batches, adding their gradients together, and only updating the weights once at the end. You get the steadier signal of a big batch while holding just one small batch in memory.
- Gradient checkpointing
- Throwing away most stored activations and recomputing them during the backward pass. Peak activation memory drops a long way in exchange for roughly 20 to 30 percent more time. Chen et al., Training Deep Nets with Sublinear Memory Cost
- Gradient clipping
- Capping the overall size of the gradient before the update, so one strange batch cannot throw the weights into nonsense. A limit of 1.0 is the standard setting and needs no tuning. AdamW explained, line by line
- Gradient norm
- The overall size of the gradient across all weights, logged every step. Constant clipping means the learning rate is too high, and a norm near zero from the first step means it is too low.
- Holdout
- Data deliberately kept out of training so you can measure the model on something it has never seen. A number measured on data the model trained on is not evidence of anything. Building a holdout you can trust
- Hyperparameter
- A setting you choose before training rather than something the model learns, such as the learning rate, the batch size or the number of epochs.
- ignore_index
- The setting on PyTorch’s cross-entropy loss that names one label value to skip entirely. It defaults to -100, which is why -100 is the convention for marking positions you do not want graded. PyTorch, CrossEntropyLoss
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- Instruction tuning
- Supervised fine-tuning on instruction-and-answer pairs, so a base model learns to follow instructions and answer in a consistent shape. Zhou et al., LIMA
- Label
- The correct answer a training example is graded against. For a language model the labels are the token IDs the model was supposed to produce, and a label of -100 means do not grade this position at all. Masking in LLM training
- Layer
- One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
- Learning rate
- A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base.
- Learning-rate schedule
- A rule that changes the learning rate over the course of a run, typically warming up and then decaying. The schedule is a separate thing from the optimizer, and both act on every step.
- Logits
- The raw scores a model produces for every possible next token, before they are turned into probabilities. One number per token in the vocabulary, and higher means the model favours that token. The complete inference path
- LoRA
- Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Loss masking
- Deciding which positions in an example count toward the loss. For instruction tuning you grade only the assistant’s response and mark everything else with -100, so it contributes no loss and no gradient. Masking in LLM training
- Loss scaling
- Multiplying the loss by a large constant before the backward pass so small gradients do not underflow fp16, then dividing them back before the optimizer step. bf16 does not need it at all. Micikevicius et al., Mixed Precision Training
- Micro-batch
- The batch that actually fits in memory in one go. Several micro-batches are combined by gradient accumulation so the update behaves like one much larger batch.
- Mixed precision
- Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
- NaN
- Not a number, the value floating-point arithmetic produces from an undefined operation such as infinity minus infinity. Once one reaches the weights the run is dead, and it is the usual end state of a badly scaled fp16 run. Number formats for training
- Next-token prediction
- The one thing a language model is trained to do: given the text so far, put a probability on every possible next token. Fine-tuning uses exactly the same objective as pretraining. What actually changes inside the model
- OOM
- Out of memory, the error you get when a run needs more VRAM than the card has. In training it almost always strikes where the forward pass ends and the backward pass begins, which points straight at activations. Activation memory and gradient checkpointing
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Optimizer state
- The numbers an optimizer keeps between steps, such as running averages of past gradients. Under standard mixed-precision AdamW it is 12 of the 16 bytes per parameter, which makes it the largest memory tenant. Rajbhandari et al., ZeRO
- Out-of-distribution probe
- A small evaluation on tasks that look nothing like your training data, kept specifically to catch abilities you have lost. In-distribution evaluation cannot see catastrophic forgetting at all. Evaluating a fine-tune
- Overfitting
- When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse. The six silent failure modes
- Padding
- Filler tokens added to short sequences so every sequence in a batch has the same length. Padding is bookkeeping for the hardware, and it has to be masked out of both attention and the loss so it cannot change the answer. Masking in LLM training
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- Partial derivative
- The rate at which one output changes when you nudge one input and hold everything else still. A gradient is the partial derivative of the loss with respect to a single weight. The calculus behind backpropagation
- Preference tuning
- Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- Regularisation
- Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
- Second moment
- Adam’s running average of the squared gradient, written v. It measures how large and erratic a weight’s gradients have been, and dividing the step by its square root is what gives each weight its own step size. Gradients and optimizers explained
- Sequence length
- How many tokens are in one training example after tokenisation. Activation memory grows in step with it, and the attention part grows with its square. Activation memory and gradient checkpointing
- Softmax
- The step that turns a list of raw scores into probabilities that add up to 1. Larger scores get larger probabilities, and the gaps between the scores decide how confident the result looks. The complete inference path
- Special token
- A token that stands for structure rather than ordinary text, such as the marker that opens or closes an assistant turn. They come from the tokenizer and are placed by the chat template.
- Static state
- Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
- Supervised fine-tuning
- Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT.
- System prompt
- A message at the start of a conversation that sets the model’s role and rules. During training it is masked out of the loss along with the user’s turns. Masking in LLM training
- Teacher forcing
- During training the model is shown the real sequence and asked to predict the next token at every position at once, instead of being fed its own guesses. It is why one forward pass can grade a whole example.
- Tensor
- A grid of numbers with any number of dimensions. One number is a scalar, a row of them is a vector, a table is a matrix, and anything past that is still a tensor with more dimensions. How a neural network learns
- Tensor core
- The part of an NVIDIA GPU built to do matrix multiplies in low precision very fast. It is why bf16 training beats fp32, and why a format tensor cores cannot multiply directly, such as NF4, costs throughput. NVIDIA, Accelerating AI training with TF32 tensor cores
- Token
- The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
- Tokenizer
- The component that splits text into tokens and maps them to integer IDs, and back again. Every model has its own, and it has to match the model you are training. Choosing a base model and building a bench
- Training step
- One cycle of the loop: forward pass, loss, backward pass, optimizer step, then clear the gradients. Everything else in a training script is arrangements around those five moves.
- Truncation
- Cutting an example off at the maximum length. Done carelessly it can remove the entire response, leaving an example whose labels are all -100 and which teaches nothing. Masking in LLM training
- Underflow
- What happens when a number is too small for the format to represent, so it becomes zero. A gradient that underflows has not been made noisy, it has been deleted, and averaging cannot bring it back. Number formats for training
- Validation set
- The split you check during training to pick settings and choose the best checkpoint. Because you make decisions from it over and over, it slowly stops being an honest estimate. Splits, done honestly
- VRAM
- The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
- Warmup
- Ramping the learning rate up from zero over the first few percent of steps, so the earliest updates cannot damage the model while the optimizer still has no history. Set it as a ratio rather than a step count.
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
- Weight decay
- A small pull on every weight toward zero on every step, so no weight grows larger than it needs to be. It is a form of regularisation, and it is separate from gradient clipping, which is a safety limit. Loshchilov and Hutter, Decoupled Weight Decay Regularization
Practical exercises
Compute the effective batch size and predict what forgetting the division does
Version A of this part’s script sets per_device_train_batch_size to 4 and gradient_accumulation_steps to 4. Compute the effective batch size that trains on. Then suppose the line (out.loss / accum).backward() is changed to out.loss.backward(), leaving everything else the same. Using this part’s own explanation of what gradient accumulation actually sums, state what happens to the size of the gradient the optimizer sees at each step relative to the correctly divided version. Then answer the part that catches most people: work through what AdamW’s own normalisation does to a gradient that has been multiplied by a constant, say whether the run therefore behaves like one at a higher learning rate, and name the one thing in the configuration that is not scale-invariant and so gives the bug away in the logs.
See the worked solution (opens in a new tab)
Diagnose a fine-tune that will not stop generating
A fine-tuned Qwen2.5-1.5B answers a prompt correctly, then keeps generating unrelated text until it hits max_new_tokens. Using this part’s two named causes of exactly this symptom, describe which one to check first, what confirms it in code, and what you would only check second if the first comes back clean.
See the worked solution (opens in a new tab)
Predict what removing the attention mask actually breaks
Version A’s collate function builds input_ids, attention_mask and labels and passes all three into the forward call. Suppose you simplify the call to model(input_ids=ii, labels=lb), dropping attention_mask only from that call, while the collate function still pads labels with -100 exactly as before. For a batch of right-padded sequences, work out precisely what does and does not get corrupted, and name a concrete reason to keep passing it anyway even where the damage turns out to be smaller than expected.
See the worked solution (opens in a new tab)
Explain why assistant_only_loss needs the chat template patch
Version B copies the Instruct checkpoint’s chat_template onto the base tokenizer before training. Explain what specifically causes assistant_only_loss=True to fail on the base checkpoint’s own, valid chat template without that line, why borrowing the Instruct template fixes it at no cost to the model’s weights, and why that fix is preferable to the chat_template_path fallback TRL also documents.
See the worked solution (opens in a new tab)
Trace what a missing truncation filter does to one batch, and to a whole run
Suppose the line ds.filter(lambda r: any(l != -100 for l in r["labels"])) is deleted from Version A’s script, and one training example’s response falls entirely past the max_length truncation point. Describe what happens to that one example during an ordinary training step where it shares a batch with healthy examples. Then describe a rarer, more severe scenario where the same missing filter can corrupt the entire run rather than just wasting one example’s compute.
Frequently asked questions
What learning rate should I use for supervised fine-tuning?
For a full fine-tune of a 1B to 2B model, start at 2e-5, within a range of roughly 1e-5 to 2e-5. Higher risks divergence, because every pretrained weight is moving at once. LoRA is a different regime entirely at roughly 1e-4 to 3e-4, and using one regime’s value in the other is a very common failure.
How many epochs should I train for?
Start with one, and rarely go beyond three on a small instruction set. Watch the held-out loss rather than the training loss: while both fall you are generalising, and when the held-out curve turns up while training keeps falling you are memorising. Keep the checkpoint at the held-out minimum, not the one from the final step.
My training loss is falling but the model is not better. Why?
A falling training loss only means the model fitted the signal on the tokens you graded. Four common causes of the gap: the loss mask is wrong so you graded the wrong tokens, the model is overfitting and the held-out curve has already turned up, the behaviour you wanted was never expressible as a single gold answer per prompt and needs preference tuning, or, if the loss fell almost to zero, the causal mask is missing and the model is copying the next token out of its own input.
Do I need gradient accumulation?
You need it whenever the batch size you want is larger than activation memory allows, which for a full fine-tune on one card is almost always. It sums gradients across several small batches before stepping, so the update behaves like a large batch while memory stays at one small batch. Remember to divide each loss by the accumulation count.
Why does my fine-tuned model keep talking after it has answered?
Check the end token first. Qwen2.5’s tokenizer defaults its end token to the end-of-text token, while its chat template closes assistant turns with a different end-of-turn token, so unless you name that second token in both the training config and the generation call, nothing is watching for the stop signal the model was trained to produce. If that is already correct, check next whether the closing token was inside the graded region during training.
How long does a supervised fine-tune take?
For a 1.5B model on roughly 10,000 short examples at an effective batch of 16 and a sequence length of 1024, one epoch of full fine-tuning on a single modern GPU takes on the order of tens of minutes. LoRA on the same job is faster and uses far less memory, which is why most practical iteration happens there.
Sources and further reading
- Hugging Face TRL, SFT Trainer documentation, the current configuration surface, including the masking flags and the chat-template fallbacks.
- Qwen2.5-1.5B model card, the base checkpoint used throughout, including its tokenizer and chat template, and the Qwen2.5 Technical Report for the family’s sizes and training scale.
- PyTorch, CrossEntropyLoss, for the ignore index that makes label masking work.
- Hugging Face, Methods and tools for efficient training on a single GPU, for the memory settings in this part, together with Chen et al., Training Deep Nets with Sublinear Memory Cost, the paper behind gradient checkpointing and the 30 percent end of its cost.
- Zhou et al., LIMA: Less Is More for Alignment, the clearest published evidence that a small curated set selects behaviour rather than installing knowledge, which is the claim layers 6 to 11 rest on.
- Allen-Zhu and Li, Physics of Language Models 3.1, Knowledge Storage and Extraction and 3.3, Knowledge Capacity Scaling Laws, for how factual knowledge is stored during pretraining and why later training does not reliably rewrite it. These are the evidence for row 4 of the table.
- Kaplan et al., Scaling Laws for Neural Language Models and Hoffmann et al., Training Compute-Optimal Large Language Models, for what pretraining scale buys, which is what layers 1 to 5 are made of.
- Brown et al., Language Models are Few-Shot Learners, for a pretrained model performing tasks it was never trained on when the prompt shows it examples.
- Wei et al., Finetuned Language Models Are Zero-Shot Learners, the FLAN paper, for instruction tuning generalising to held-out task types.
- Lamb et al., Professor Forcing, for teacher forcing and the training-to-generation gap it leaves.
- Pascanu et al., On the difficulty of training recurrent neural networks, the origin of clipping gradients by their norm, which is what
max_grad_normdoes, and Goyal et al., Accurate, Large Minibatch SGD, for the linear scaling rule that makes the effective batch the number worth quoting. - HuggingFaceH4/no_robots, the dataset used in both scripts above.
