Loss masking decides which tokens a fine-tuning run is graded on. It is the highest-leverage setting in the pipeline and the quietest-failing one. Get it wrong and nothing crashes. The loss still falls, the run still finishes, and the model you ship is worse for a reason no log file will name.
Part of the trouble is the word itself. Three different things inside one training step are all called masking. They share one idea, marking some positions to ignore. They gate three different operations for three different reasons, and mixing them up is where the bugs live.
This part builds all three from nothing, using one short chat example you can check on paper. It assumes you have never written a training script.
By the end you will be able to write out the labels for a training example by hand, say which mask you own and which the model owns, explain why a blocked attention cell ends up with a weight of exactly zero, and run a short check that catches the most expensive silent bug in a fine-tuning pipeline.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding (you are here)
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
Read any mask as one keep-or-ignore flag per position
Strip away the variants and a mask is a list of flags, one per position, riding alongside your data. Each flag says keep or ignore. Some operation reads that list and honours it. That is the whole idea.
A position is just a slot in the sequence, counted from the left. Text goes into a model as tokens, which are chunks of text of roughly word size, and each token occupies one position. Six tokens means six positions, so a mask over them is six flags. Here is one:
| Position | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Token | Name | two | primes | Two | and | three |
| Flag | ignore | ignore | ignore | keep | keep | keep |
The first three tokens are a question somebody asked. The last three are the answer. The flags say that only the answer counts. Nothing about that table is specific to machine learning, and nothing about it is hard. Every mask in this part is this picture pointed at a different operation.
So when you meet the word, two questions decode it. Which operation does this mask gate? And what does hiding a position buy you there? Answer both and the confusion goes away, because the three maskings give three different answers.
- Loss masking gates the scoring. Hidden positions are not graded, so the model learns nothing from them. You set this one.
- Causal masking gates attention. Each position is stopped from looking at anything later in the sequence. The model does this on its own, always, and you never touch it.
- Padding masking gates both. Filler tokens added to square off a batch are hidden from attention and from the loss. You set it, though the tools mostly set it for you.
Say what a label is, and why -100 removes a position from training
Start with the mask you configure directly, because it is the one that decides what your fine-tune actually learns.
A label is the right answer at one position
Training compares what the model predicted against what was correct. The correct answer at a position is called its label. For a language model the correct answer is always the same thing: the token that actually came next. So you never invent labels for a language model. The labels are a copy of the tokens themselves.
The copy is lined up so that the guess made at one position is scored against the token at the next one. Hugging Face’s models perform that shift for you inside the model, which is why the array of labels you hand over is exactly the same length as the array of tokens, holding exactly the same values. That sounds pointless until you realise what it buys: a second array you are free to edit.
Editing it is loss masking. Take the copy and overwrite every position you do not want graded with the value -100.
Here is the running example of this part: the same six words as above, now with the turn markers a real chat format inserts around them. It is one short instruction-and-response pair of the kind that fills HuggingFaceH4/no_robots, the roughly 10,000 human-written pairs this series fine-tunes on, tokenized the way Qwen2.5-1.5B expects, a model whose family Qwen’s own technical report describes in full. Other models use different markers for the same job. The system turn and the whitespace tokens are dropped here so the table fits on a screen.
| Position | Token | Belongs to | Label | Attention mask |
|---|---|---|---|---|
| 1 | <|im_start|> |
turn marker | -100 | 1 |
| 2 | user |
role marker | -100 | 1 |
| 3 | Name | the user’s words | -100 | 1 |
| 4 | two | the user’s words | -100 | 1 |
| 5 | primes | the user’s words | -100 | 1 |
| 6 | <|im_end|> |
end of the user turn | -100 | 1 |
| 7 | <|im_start|> |
turn marker | -100 | 1 |
| 8 | assistant |
role marker | -100 | 1 |
| 9 | Two | the answer | the token ID for Two | 1 |
| 10 | and | the answer | the token ID for and | 1 |
| 11 | three | the answer | the token ID for three | 1 |
| 12 | <|im_end|> |
end of the answer | the token ID for that marker | 1 |
Three of those columns become three arrays that travel together into the model, and they are the three names you will see in code. The tokens are stored as integer IDs in an array called input_ids. The label column is an array called labels. The last column is the attention_mask, which the section on padding below is about. Every position in this example holds a 1 there, because nothing has been padded yet.
Read the label column top to bottom and you have the mask. Eight positions marked ignore, four graded. Position 8 is worth a glance: the model is not graded on producing the word assistant, because at inference time, meaning whenever the model is actually answering somebody, that marker is handed to it rather than generated. Position 12 is graded, and the next section explains why that single row matters more than it looks.
You rarely write labels one position at a time. A collator does it. A collator is the small piece of code sitting between your dataset and the model: it takes several examples, pads them to a common length, stacks them into one batch, and builds the label array and the attention mask that travel with them. Libraries ship collators, and you can write your own. The ones people write by hand are where masking bugs are born.
What “ignore” actually does to a weight
Being told a position is ignored is not the same as knowing what happens. Follow the chain in four steps. It is short, and it is the entire point of the -100 convention.
Step 1. The loss is a sum over graded positions. Part 1 built cross-entropy from scratch: at each position you graded, look up the probability the model gave to the correct token and take the negative logarithm of it. High probability gives a small penalty, low probability gives a large one. Add those penalties over the graded positions and divide by how many there were. One number for the example.
Step 2. -100 is a value no real label can hold. Token IDs are positions in the model’s vocabulary list, so they run from zero upwards and are never negative. That makes any negative number safe to use as a marker. PyTorch’s cross-entropy has a setting called ignore_index, which names one label value meaning skip this position entirely, and its default value is -100. So a position holding -100 is dropped from the sum in step 1, and also from the count you divide by.
Step 3. Dropped from the sum means it cannot move the total. This is the step people skip. A gradient answers one question about one weight: if I nudge this weight, does the loss go up or down, and how fast? If a position never enters the sum, then nothing the model predicted at that position can change the total, however wrong it was. Its contribution to every weight’s gradient is exactly zero.
Step 4. Zero gradient means zero weight change. Weights only ever move through the optimizer, and the optimizer only ever acts on gradients, as the previous part on AdamW walks through line by line. So a masked position contributes no loss, therefore no gradient, therefore no update. The model is not being gently discouraged from learning that position. It is never taught by it at all.
One honest caveat on that chain. A masked token is still read. It still runs through the forward pass, and it still acts as context that later positions attend to, which is exactly what you want for a prompt. Masking removes a position from the scoring, and it removes nothing from the input. That also means masking is free: the compute is identical either way, because those tokens were going through the model regardless.
Choose which positions your run should grade
The default for instruction tuning is to grade the assistant’s tokens and nothing else. That convention comes from the papers that established instruction tuning: Wei et al. in FLAN and Sanh et al. in T0 both rewrote large collections of tasks as an instruction plus a response, and in both the response is the part the model is scored on. It is worth seeing why, because there is a real case where the opposite is correct.
You could grade every position. For raw pretraining, and for continued pretraining on a pile of plain domain text, that is exactly what you do, because there is no prompt and no response, just text the model should learn to continue. That is the objective Brown et al. trained GPT-3 with: every token in the corpus is graded, because every token is one the model should learn to predict. Instruction tuning is a different shape. An example is a conversation, and only one participant in it is you.
Three concrete costs come with grading the prompt.
It spends capacity on text you do not control. Every graded position teaches the model to produce that text. Prompts come from your users, in their phrasing, and you never want the model generating them.
It drowns the signal you care about. Prompts are usually much longer than responses, and much harder to predict. Take a plausible example: a 200-token prompt on which the model averages a penalty of 2.0, and a 40-token response on which it averages 1.0. Grading everything gives a total of 200 times 2.0 plus 40 times 1.0, which is 440, spread over 240 positions. The response accounts for 40 of those 440 points, about 9 percent. Roughly nine tenths of the pull on the weights that step is going into predicting the user’s words. Mask the prompt and it is 100 percent.
It can teach the wrong reflex. A model graded on continuing prompts learns to continue prompts. With a diverse instruction set that shows up as a model which extends your question instead of answering it.
Grade the closing token, or the model never learns to stop
An end-of-turn token is a special token the chat format puts at the end of each turn, marking where an answer finishes. In the table above it is the marker at position 12. It is a token like any other, so the model has to learn to produce it, and the only way it learns is by being graded on it.
Mask that closing token and you get a symptom that looks exactly like a quality problem. The fine-tuned model answers correctly, then keeps going: another sentence, an invented follow-up question, a second answer, on until it hits whatever length limit the caller set. Nobody suspects the collator, because the loss curve was fine and the answers are good. The graded region has to include the closing token. Make it the first thing you check when a tuned model rambles.
Grade every assistant turn, and keep the mask honest under packing
A single-turn example has one boundary and one graded span. Find where the answer starts and you are done. A multi-turn conversation has several, and the tempting shortcut is to grade only the final assistant turn, because the end of the sequence is the easiest place to find.
That shortcut throws away supervision you already paid for. In a conversation that runs user, assistant, user, assistant, with replies of 40 and 60 tokens, grading only the last one grades 60 tokens and discards 40. The first reply is just as valid a lesson. Grading all of them also teaches the model to hold its role across turns instead of drifting into narrating both sides of the conversation.
The mechanism does not change. Mask every system turn and every user turn, keep every assistant turn, and keep each turn’s closing token. If you build labels by hand, walk through the turns in a loop. Do not compute a single boundary.
Where the boundary comes from
Two pieces of tooling decide how hard this is, and both are worth naming.
A tokenizer is the component that splits text into tokens and maps them to integer IDs, and back again. Every model ships its own. A chat template is a rule that ships alongside it, turning a list of role-and-content messages into the exact text and special tokens that model expects, turn markers included. The template is where the boundary between prompt and response physically lives, so anything that finds the boundary is really reading the template.
Some templates go further and emit generation markers. These are small pieces of bookkeeping wrapped around each assistant turn, so that a tokenizer asked for them hands back an extra row of flags saying precisely which positions are assistant tokens. That row is the loss mask, worked out by the template author instead of by you.
Modern trainers use this. The flag completion_only_loss covers datasets stored as a prompt and a completion. The flag assistant_only_loss covers conversational datasets and finds every assistant span rather than the last one, which is why it needs those generation markers to exist in the first place. Whether your base model’s template has them is a real selection criterion, and one that the next part, on choosing a base model and building a bench, covers alongside the rest. Pin your library versions before you rely on any of these flag names, because this corner of the API has moved more than once.
The one setting that quietly invalidates your mask
Sequence packing concatenates several short examples end to end into one full-length sequence, so the card stops spending compute on padding. It is a genuine speed win and it changes the geometry every mask in this article assumes.
Two things have to be reset at each example boundary inside a packed sequence. The first is attention: without a reset, tokens from the second example can attend back into the first, because the causal rule only says a position may look leftwards, and after packing there is more to the left than there used to be. The second is position indices, the numbers telling the model where each token sits counted from the start of the sequence. Left alone, the second example is told it begins at position 900 rather than position 1.
Libraries that support packing handle both, and the ones that do not will happily train nonsense with no error. The part on data covers packing properly. The rule to carry away is narrow and useful: after you turn packing on, verify the mask again from scratch.
Verify the mask on one batch before you trust a run
This is the single cheapest habit in the pipeline, and it takes about thirty seconds. Build one batch, then look at it. Decode the graded positions back into text and confirm they are exactly the assistant’s words, nothing more and nothing less.
# sanity-check ONE example's mask. Do this every time you change data code.
labels = batch["labels"][0]
graded = labels[labels != -100] # the tokens that actually count
print(tokenizer.decode(graded)) # should be ONLY the assistant reply
assert (labels != -100).any(), "mask is all -100, this batch teaches nothing"
# and the reverse check: what did we throw away?
masked_positions = (labels == -100).sum().item()
print(f"{masked_positions} of {len(labels)} positions masked")
If you have never used PyTorch, here is every line.
batchis what the collator produced: a dictionary whose keys are names such asinput_ids,attention_maskandlabels. Each value is a tensor, which is simply a grid of numbers. This one has a row per example in the batch and a column per position.batch["labels"][0]pulls out the labels tensor and then takes row zero, the first example. Counting starts at zero, so row zero is the first row.labels != -100does not return a single true or false. It compares every position at once and returns a tensor of the same shape holding true or false per position. That is the mask, made visible.- Putting that true-or-false tensor inside the square brackets of
labelskeeps only the positions where it is true. Sogradedis the list of label IDs that will actually be scored. tokenizer.decodeturns a list of token IDs back into readable text. This is the line that matters. Read the output. It should be the assistant’s reply, and it should end with the closing marker.assertstops the program with your message if the condition is false..any()is true when at least one position is true, so this fires when every single label is -100.(labels == -100).sum()counts the ignored positions, because summing true-or-false values counts the trues..item()takes that one-element tensor and hands back a plain Python number you can print.
Two silent failure modes make the assertion worth keeping forever.
The first is truncation. Truncation is cutting an example off at the maximum length you configured. If a prompt is long enough that the cut lands before the response begins, that example’s labels are all -100 and it teaches nothing at all. No error is raised, because an example with nothing to learn from is not an error to any library.
The second is an off-by-a-few boundary. Shift the split by three tokens and you quietly grade the tail of the prompt along with the answer. The loss curve is indistinguishable.
The assertion above catches the first mode only when a whole batch is empty. That is not enough on its own. Filter examples whose labels are entirely -100 out of the dataset before training starts, so a handful of them cannot hide inside otherwise healthy batches. The part on running a full supervised fine-tune shows the boundary calculation that gets the split right in token space rather than in characters.
Explain the causal mask, and why negative infinity gives exactly zero attention
The second masking is built into the model. To see what it does, you need one paragraph on attention.
Attention is the step where each position looks at the other positions and mixes in whatever it finds useful. It works in three moves. First, for every pair of positions, the model computes a raw score saying how relevant one is to the other. Second, it runs softmax across each row of those scores, which turns them into weights that are all positive and add up to 1. Third, each position takes a blend of the other positions’ contents using those weights. The part on multi-head attention covers the machinery in full.
Now the problem the causal mask solves. During training the whole sequence goes through the model in one pass, for speed, and every position is simultaneously being trained to predict the token after it. Feeding the model the true text rather than its own guesses is called teacher forcing, and Lamb et al. named its cost in Professor Forcing: a model trained only on true prefixes has never had its own output as input. If position 5 were allowed to attend to position 6, it would be reading the very answer it is being graded on. It would score beautifully and learn nothing.
So the rule is imposed: a position may attend only to positions at or before its own. Draw the grid with queries down the side and keys across the top and the allowed region is the lower triangle.
Why the blocked cells are set to negative infinity
The blocked cells are not deleted. They are set to negative infinity, and then the softmax runs as usual. That sounds like a strange way to say no. Work it through and it is the neatest possible way.
Softmax, built from raw scores in Part 1, does two things in order. It raises e, about 2.718, to the power of each score, which makes every value positive. Then it divides each result by the total, which makes them add up to 1. Take three positions with scores 2.0, 1.0 and 0.1 and no mask at all:
- Raise e to each score: 7.39, 2.72, 1.11.
- Add them: 11.22.
- Divide each by the total: 0.66, 0.24, 0.10.
Now block the third position by replacing its score with negative infinity, and redo the same arithmetic. Raising e to a large negative power collapses fast: e to the power of minus 1 is 0.37, e to the power of minus 10 is 0.000045, e to the power of minus 100 is a decimal point followed by forty-three zeros. Push it all the way and the answer is exactly 0.
- Raise e to each score: 7.39, 2.72, 0.
- Add them: 10.11.
- Divide each by the total: 0.73, 0.27, 0.00.
The blocked position gets a weight of exactly zero, so none of its content reaches the output. The two survivors still add to 1, and they have shared out the blocked position’s weight between them in proportion to their own scores. Nothing downstream needs a special case, and nothing needs to know a mask was applied.
That is why negative infinity and not, say, zero. A score of zero would give e to the power of zero, which is 1, a perfectly respectable weight. Negative infinity is the one value softmax turns into no contribution at all, which lets the mask be applied by plain addition before the softmax and then forgotten. In real code libraries add the largest negative number the format can hold rather than true infinity, because arithmetic on actual infinities can produce a not-a-number value that poisons the run. The effect is the same to more decimal places than anything cares about.
The practical takeaway is short. You never set this mask. It is part of the architecture of any causal language model, applied inside every attention layer, on every pass, whether you are training or serving. The masks you own are the ones that depend on your data: where the prompt ends, and how long each sequence is.
Pad a batch without changing what it teaches
A batch is several examples processed together, and a GPU wants that batch as a rectangle: the same number of columns in every row. Real examples are different lengths. So you pad, which means appending a filler token to the short ones until every row reaches the length of the longest.
Padding is bookkeeping for the hardware. It carries no meaning, so you have to tell the model to disregard it, and the thing that does the telling is the attention mask: a 1 for each real token and a 0 for each pad.
The zeros work through the same machinery as the causal mask. A column marked zero is turned into negative infinity before the softmax, so every real token gives it a weight of exactly zero. Without that, tokens would blend meaningless filler into their own representations, and the amount of nonsense mixed in would depend on how long the longest unrelated example in the batch happened to be.
Padding also overlaps with loss masking, and this is the part people forget. A pad position needs two markings, not one: an attention mask of 0 so nothing attends to it, and a label of -100 so it is not graded. In a collator that is two lines, extend the attention mask with zeros and extend the labels with -100. Miss the second and your model spends part of every step being graded on its ability to predict filler.
Which side you pad on, and why it costs people a day
Padding has a side, and the right side differs between training and generation. This is the highest-value paragraph in the article for anyone who is about to generate text from their own fine-tune.
Generation works by appending. You hand the model a prompt, it produces one token, that token is written into the next column of the row, and the loop runs again from there. The model always continues from the end of the row. Now put a four-token prompt into a six-column batch two ways:
| Layout | 1 | 2 | 3 | 4 | 5 | 6 | The next token is written after slot 6, so it continues from |
|---|---|---|---|---|---|---|---|
| Right-padded | Name | two | primes | marker | PAD | PAD | two pad tokens, which is not where the prompt ended |
| Left-padded | PAD | PAD | Name | two | primes | marker | the last real token of the prompt, which is correct |
Read the last column. With right padding, the sequence the model is asked to continue reads prompt, filler, filler, and then its own new tokens land on the far side of that filler. The attention mask does stop the pad tokens being attended to, and it still does not help, because the model is being asked to write the continuation in the wrong place. With left padding every row’s real final token sits flush against the right-hand edge, which is exactly the position generation continues from, and every row in the batch lines up.
So: left-pad for generation, right-pad for training. Training has no appending loop, every position is scored where it sits, and right padding is the ordinary convention. Generation libraries will usually warn you about right padding on a decoder-only model, which is a model that only ever looks leftwards. Heed the warning. A wrong padding side produces garbage output from a perfectly good model, and there is nothing in the model or the training log to point at.
See all three masks fire in one forward pass
In a single training step all three run, and they never interfere, because they gate different operations at different moments.
Inside every attention layer, the causal mask and the padding mask are combined into one grid of allowed and blocked cells, and applied together in the same addition before the softmax. A cell survives only if both agree: the key position is at or before the query, and the key position is a real token. Then, at the very end of the forward pass, once the model has produced a guess at every position, the loss mask decides which of those guesses are graded at all.
They compose without conflict because they answer different questions. Causal and padding answer which positions may this token look at. Loss answers which positions do we score. Here is the whole thing in one table.
| Mask | Hides | Gates | Mechanism | Who sets it | If you forget it |
|---|---|---|---|---|---|
| Loss | prompt and pad tokens | the loss | label set to -100 | you, in the collator | trains on the wrong tokens, silently |
| Causal | future tokens | attention | upper triangle set to negative infinity | the model, automatically | the model cheats and cannot generate |
| Padding | pad tokens | attention and the loss | attention mask of 1 or 0, plus a label of -100 | you, via tokenizer and collator | attends to filler, grades filler, corrupts the batch |
Optional: recognise a fourth meaning of masking in other models
This section does not apply to fine-tuning a causal language model. Skip it if you like. It earns its place only because the word turns up in one more sense and the confusion is common.
Masked language modelling is the training objective behind BERT-style models. About 15 percent of the input tokens are replaced with a placeholder, and the model is trained to reconstruct them using context from both sides. There is no causal mask involved, because attention is deliberately allowed to look in both directions. Nothing is being hidden from the loss. Something is being hidden from the input, which is the opposite arrangement to everything above.
It is a pretraining objective for understanding models, the kind used for classification, search and embeddings, rather than for generating text. You will not use it to fine-tune a chat model. It does matter when you build a retrieval pipeline with an embedding model, because that embedding model was very likely trained this way and your decoder model was not. When someone says masking and means “predict the hidden word”, this is the sense they mean.
Key takeaways
- A mask is a keep-or-ignore flag per position. Three of them show up in LLM training, and the way to tell them apart is to ask which operation each one gates.
- Labels for a language model are a copy of the tokens. Loss masking is editing that copy: overwrite every position you do not want graded with -100.
- PyTorch’s cross-entropy skips any label equal to -100, so that position leaves the sum, contributes zero gradient, and changes no weight. It is still read as context, and masking it costs no compute.
- Include the assistant turn’s closing token in the graded region. Mask it and the model never learns to stop, which reads as a quality problem and is a collator bug.
- Causal masking is the model’s job, not yours. Blocked cells are set to negative infinity, and softmax turns that into a weight of exactly zero while the surviving weights still add to 1.
- A pad position needs both markings: an attention mask of 0 and a label of -100. Left-pad for generation so the row ends on a real token, right-pad for training.
- Decode the graded tokens of one batch and read them. It is thirty seconds, and it is the cheapest bug prevention in the pipeline.
You can now
- Write out the labels and attention mask for a chat example by hand, and say how many positions end up graded, from “Say what a label is, and why -100 removes a position from training”.
- Trace an ignored position through to the weights: no loss, no gradient, no update, from the same section.
- Decide whether a given run should grade the prompt, and put a number on what grading it would cost, from “Choose which positions your run should grade”.
- Run a five-line check on one batch and read its output, without prior PyTorch, from “Verify the mask on one batch before you trust a run”.
- Explain why a blocked attention cell ends up with a weight of exactly zero, using the softmax arithmetic, from “Explain the causal mask, and why negative infinity gives exactly zero attention”.
- Pick the correct padding side for training and for generation, and say what breaks with the wrong one, from “Pad a batch without changing what it teaches”.
Glossary
- Attention
- The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
- Backward pass
- Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- Causal language model
- A model that only ever looks left, predicting each token from the ones before it. Every model in this series is one, as opposed to a model such as BERT that reads in both directions. The complete inference path
- Causal mask
- The rule inside attention that stops each position seeing anything later in the sequence. It is built into the architecture, you never set it, and it is what lets one pass train next-token prediction at every position without cheating.
- Chat template
- The rule that turns a list of role-and-content messages into the exact text and special tokens the model expects to see. It ships with the tokenizer, and it is where the boundary between prompt and response lives. From messages to tensors
- Collator
- The small piece of code that takes several examples and assembles one batch out of them: padding them to a common length, building the attention mask, and setting padded labels to -100.
- Cross-entropy
- A score for how wrong a prediction was. It is small when the model gave high probability to the token that actually came next, and large when it did not. How a neural network learns
- Decoder-only
- The standard LLM architecture: one stack of identical transformer blocks, with a causal mask so each position sees only what came before it. Llama, Qwen and most open models are all this shape. Inside one transformer block
- Embedding
- The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
- End-of-turn token
- The special token that marks where an answer stops. If it is left out of the graded region during training, or the wrong one is passed at generation time, the model rambles past the end of its answer.
- Forward pass
- Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- ignore_index
- The setting on PyTorch’s cross-entropy loss that names one label value to skip entirely. It defaults to -100, which is why -100 is the convention for marking positions you do not want graded. PyTorch, CrossEntropyLoss
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- Instruction tuning
- Supervised fine-tuning on instruction-and-answer pairs, so a base model learns to follow instructions and answer in a consistent shape. Zhou et al., LIMA
- Label
- The correct answer a training example is graded against. For a language model the labels are the token IDs the model was supposed to produce, and a label of -100 means do not grade this position at all.
- Layer
- One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
- Logits
- The raw scores a model produces for every possible next token, before they are turned into probabilities. One number per token in the vocabulary, and higher means the model favours that token. The complete inference path
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Loss masking
- Deciding which positions in an example count toward the loss. For instruction tuning you grade only the assistant’s response and mark everything else with -100, so it contributes no loss and no gradient.
- Masked language modelling
- A different training objective, used by BERT-style models: hide about 15 percent of the words and train the model to fill them back in using context from both sides. Search and embedding models are trained this way; a causal LLM is not. Devlin et al., BERT
- Next-token prediction
- The one thing a language model is trained to do: given the text so far, put a probability on every possible next token. Fine-tuning uses exactly the same objective as pretraining. What actually changes inside the model
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Padding
- Filler tokens added to short sequences so every sequence in a batch has the same length. Padding is bookkeeping for the hardware, and it has to be masked out of both attention and the loss so it cannot change the answer.
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- Sequence packing
- Concatenating several short examples into one full-length sequence so you stop wasting compute on padding. It needs boundary resets, or one example can attend back into another and the model learns nonsense continuations. Hugging Face TRL, SFT Trainer
- Softmax
- The step that turns a list of raw scores into probabilities that add up to 1. Larger scores get larger probabilities, and the gaps between the scores decide how confident the result looks. The complete inference path
- Supervised fine-tuning
- Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
- System prompt
- A message at the start of a conversation that sets the model’s role and rules. During training it is masked out of the loss along with the user’s turns.
- Tensor
- A grid of numbers with any number of dimensions. One number is a scalar, a row of them is a vector, a table is a matrix, and anything past that is still a tensor with more dimensions. How a neural network learns
- Token
- The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
- Tokenizer
- The component that splits text into tokens and maps them to integer IDs, and back again. Every model has its own, and it has to match the model you are training. Choosing a base model and building a bench
- Training step
- One cycle of the loop: forward pass, loss, backward pass, optimizer step, then clear the gradients. Everything else in a training script is arrangements around those five moves. Supervised fine-tuning end to end
- Truncation
- Cutting an example off at the maximum length. Done carelessly it can remove the entire response, leaving an example whose labels are all -100 and which teaches nothing.
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
Practical exercises
Build a label array by hand for a toy example
Take a toy tokenized example with prompt token ids 11, 12, 13, response token ids 21, 22, 23, an end-of-turn token id of 99, and a pad token id of 0, padded to a batch length of 8. Following this part’s rules, write out the input_ids, the attention_mask, and the labels, then state how many of the eight positions end up graded.
See the worked solution (opens in a new tab)
Diagnose a model that will not stop generating
After a fine tuning run with a normal-looking loss curve, the model answers a question correctly at inference, then keeps generating unrelated sentences until it hits the token limit. Using only this part’s ideas about the mask you own, describe the specific check you would run on a training batch to test whether the masking bug named in this part is the cause, and what a passing versus a failing result looks like once you decode the graded span.
See the worked solution (opens in a new tab)
Quantify the supervision lost by grading only the last turn
A training example is a two-turn conversation: user, then an assistant reply of 12 tokens including its closing token, then another user turn, then an assistant reply of 18 tokens including its closing token. Compute the fraction of the assistant tokens’ supervision that a mask grading only the final assistant turn throws away, compared with a mask that grades every assistant turn. State both the fraction kept and the fraction lost.
See the worked solution (opens in a new tab)
Explain the packing bleed using only the causal mask’s own rule
Sequence packing concatenates several short training examples into one full-length sequence with no boundary reset. Using only the causal mask’s actual rule, a position may attend to any position at or before it, explain exactly why a token from the second packed example can still attend to every token of the first example even though the causal mask itself has not changed at all, and state the fix this part names for it.
See the worked solution (opens in a new tab)
Write the dataset-level check the batch-level assertion cannot replace
This part’s sanity check asserts that at least one label in a batch is not -100. Explain why that assertion can still pass on a run where one specific training example’s entire response was removed by truncation, then write the one line of dataset-level code that removes such examples before training starts, rather than only catching them if they happen to dominate a batch.
Frequently asked questions
What is loss masking in fine-tuning?
It is choosing which positions in a training example are graded. Positions you mask contribute no loss and no gradient, so the model learns nothing from them. In instruction tuning you mask the system and user turns and any padding, leaving only the assistant’s response. Mechanically you set those label positions to -100.
Why is -100 the number everything uses?
Because PyTorch’s cross-entropy loss takes an ignore_index setting that names one label value to skip, and its default is -100. Token IDs are never negative, so the value can never collide with a real label. Any position holding it leaves the sum that makes up the loss, which means it contributes zero gradient and moves no weight at all.
Should I mask the prompt when fine-tuning?
For instruction tuning, yes. Grading prompt tokens spends capacity on reproducing user phrasing you never want generated, and because prompts are usually longer than responses it can send most of the training signal to the wrong place. For continued pretraining on plain domain text there is no prompt, so grading everything is correct.
Do I need to set the causal mask myself?
No. It is part of the architecture, applied inside every attention layer of any causal language model, on every pass. The masks you own are the ones that depend on your data: the loss mask, which needs to know where the prompt ends, and the padding mask, which needs to know how long each sequence is.
Why does the padding side matter?
Generation continues from the end of the row, appending each new token after the last column. Right-pad for generation and the model is asked to continue from filler rather than from the end of the prompt, which produces garbage output from a perfectly good model. Left-pad for generation, right-pad for training.
What happens if I get the mask boundary wrong?
Nothing visible, which is the danger. The loss still descends and the run still completes, but you may have trained partly on the prompt, or on nothing at all if truncation removed the response region before the graded span began. Decoding the graded tokens of one batch and reading them is what catches it.
Sources and further reading
- PyTorch, CrossEntropyLoss, where the
ignore_indexdefault of -100 is documented. - Vaswani et al., Attention Is All You Need, for masked self-attention and the negative-infinity trick before the softmax.
- Wei et al., Finetuned Language Models Are Zero-Shot Learners, the FLAN paper, for instruction tuning as a collection of tasks written as an instruction and a graded response.
- Sanh et al., Multitask Prompted Training Enables Zero-Shot Task Generalization, the T0 paper, the same instruction-and-response framing across a large multitask mixture.
- Brown et al., Language Models are Few-Shot Learners, for the pretraining objective that grades every token, which is what response-only masking departs from.
- Lamb et al., Professor Forcing, for teacher forcing and the gap it leaves between training and generation.
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers, the masked language modelling objective in the optional section.
- Qwen2.5 Technical Report, for the model family whose chat markers this part’s running example uses.
- Hugging Face TRL, SFT Trainer documentation, for
completion_only_loss,assistant_only_lossand packing. - HuggingFaceH4/no_robots, the chat-structured dataset this series fine-tunes on, where the mask boundary is unambiguous.
