Exercise solutions: Supervised Fine-Tuning End to End: The First Real Run
These are the worked solutions for the exercises in Part 9, Supervised Fine-Tuning End to End: The First Real Run. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
Effective batch equals 4 * 4 = 16, meaning one optimizer step happens after the model has seen 16 examples worth of gradient. Gradient accumulation works by summing, not averaging, each micro-batch’s gradient into .grad across the four backward calls; dividing the loss by accum before each backward call is what turns that sum back into the same-scale average a single batch of 16 would have produced. Removing the division leaves the four micro-batch gradients summed at full scale, so the gradient the optimizer step reacts to is four times larger than the correctly scaled version, for the same learning_rate = 2e-5 line sitting untouched in the config.
Now the part that catches most people, and the reason this exercise exists. The intuitive next step is to say the run therefore behaves like one at 8e-5. That is true for plain SGD, where the update is the learning rate times the gradient, so scaling the gradient by four is identical to scaling the rate by four. It is not true for AdamW. AdamW divides the step by the square root of the second moment. Multiply every gradient by four and the first moment grows by four, while the second moment, which averages the squared gradient, grows by sixteen. The square root of sixteen is four, so the four in the numerator and the four in the denominator cancel. AdamW is very nearly invariant to a constant rescaling of the gradient, and the effective learning rate barely moves. Only the epsilon term in the denominator breaks that invariance, and at 1e-8 it is negligible for gradients of ordinary size.
So what actually goes wrong? Gradient clipping, which is the one thing in this configuration that is not scale-invariant. max_grad_norm = 1.0 acts on the raw gradient norm, before AdamW normalises anything. At four times the intended size the clip fires on nearly every step, and once it is firing constantly the update is being set by the clip threshold rather than by your learning rate or your data. That is the tell in the logs: sustained clipping on almost every step, and a loss that will not settle. Read it as the accumulation bug rather than as a too-hot learning rate, because lowering learning_rate will not fix it. Restoring the division will.
Exercise 2
Check the EOS setup first. Qwen2.5’s tokenizer defaults eos_token to <|endoftext|>, while its chat template actually closes assistant turns with <|im_end|>, so unless both the training config’s eos_token and the eos_token_id list passed to generate() explicitly include <|im_end|>, the model can emit that token forever without it ever being recognised as a stop condition. Confirming it in code means checking that SFTConfig set eos_token="<|im_end|>" during training, and that the generation call passes eos_token_id=[tok.eos_token_id, eos_id] where eos_id is tok.convert_tokens_to_ids("<|im_end|>"), not just the tokenizer’s own default id. Only once both of those check out clean should you go back to the loss mask itself: decode the graded span from a training batch and confirm the assistant turn’s closing token sits inside it, since a mask that excluded that one token never gave the model gradient signal to want to stop in the first place, a training-time cause that produces the identical symptom at inference.
Exercise 3
For a right-padded batch under ordinary causal attention, position i can only attend to positions at or before i, and every pad token sits strictly after every real token in its own row, since padding is added on the right. That means no real token’s query can ever reach a pad token’s key, attention_mask present or not, so real tokens’ own representations do not actually get corrupted by attending into padding in this specific, right-padded setup, contrary to what a blanket reading of the general warning might suggest. The labels stay correctly masked too, since the -100 padding is built earlier in the collate function and does not depend on what gets passed into the forward call, so the loss also stays clean. What is genuinely lost is something else: relying on this protection at all makes the run correct only by accident of how this particular script pads, and only for causal, one-sequence-per-row batches. The moment padding side changes to the left, as generation requires, or sequences get packed and concatenated rather than padded per row, that same causal argument no longer holds, and real tokens can attend straight into padding or into an unrelated example. Passing attention_mask explicitly costs nothing and keeps the code correct independent of those choices, rather than correct because of one specific, easy-to-change assumption about how this collate function happens to pad.
Exercise 4
assistant_only_loss=True needs the chat template to carry generation markers, and the trainer only patches that capability into templates it recognises through an exact string match against its own known-good templates, not by parsing the template’s logic. The base Qwen2.5-1.5B checkpoint ships a template that differs from the trainer’s known Qwen2.5-Instruct template in exactly one string, its default system prompt, so the match fails even though the base template is a completely valid, working chat template on its own terms. Copying the Instruct checkpoint’s chat_template string onto the base tokenizer costs nothing in weights, because the base and Instruct checkpoints of the same model share the same tokenizer and vocabulary, so no new tokens are introduced and no embedding matrix needs resizing; it is purely a string swap on the tokenizer object. The chat_template_path=”HuggingFaceTB/SmolLM3-3B” fallback comes from an entirely different model family, and per TRL’s own documentation it adds new special tokens and resizes the embedding matrix, meaning it changes the actual shape and content of the model before a single training step runs, adding untrained embedding rows that then have to be learned from nothing. That is a real architectural change, where the Instruct-template swap used here is not.
Exercise 5
With the filter removed, that one example’s labels list is entirely -100 after truncation removed its whole response region. Placed inside an otherwise healthy batch, cross-entropy’s default mean reduction is computed across the flattened non-ignored positions of the whole batch, not per example, so this example contributes zero valid positions and zero gradient of its own: its forward pass runs, its compute is spent, and it teaches nothing, quietly shrinking how much real signal that step actually contained below what per_device_train_batch_size implies, with nothing in the logs to say so. The rarer and more severe case is a step where every example in that micro-batch happens to be fully truncated this way, which becomes more likely at small batch sizes or on a dataset with an unusually long tail of over-length prompts. Then the count of non-ignored label positions across the whole batch is exactly zero, and cross-entropy’s mean reduction divides a total loss of zero by zero valid positions, producing NaN. Backward() propagates that NaN into every parameter’s gradient, and the very next optimizer step writes NaN into the model’s weights permanently, ending the run in a way no amount of later training can undo. This is precisely the failure a batch-level assertion alone cannot catch, since it only ever fires once such a batch has already arrived, which is why filtering at the dataset level before training starts is the actual fix, not a defensive extra.