Exercise solutions: Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
These are the worked solutions for the exercises in Part 10, Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
The four train strings hold one exact duplicate, “a b c” appearing twice, so train examples equals 4, and the count of unique strings is 3, giving exact duplicates in train equal to 4 - 3 = 1. Test examples equals 2. The set intersection of train and test is the single string “d e f”, shared by both, so train and eval exact overlap equals 1. Printed in the shape of this part’s script: train examples 4, test examples 2, exact duplicates in train 1, train and eval exact overlap 1.
Exercise 2
The dataset shrank by a factor of 52,000 / 9,000 = 5.78, rounded to two decimal places. The training time shrank by a factor of 80 / 14 = 5.71, rounded the same way. The two factors land within about one percent of each other because, holding sequence length, batch size and hardware fixed, wall-clock time per epoch scales close to linearly with how many examples the loop processes, so cutting the dataset by roughly 5.8 times cuts the time by very nearly the same amount. What makes the result notable is not that speed tracked size, which is expected, but that quality went up at the same time, which only happens if the roughly 43,000 removed examples were net harmful rather than neutral filler.
Exercise 3
This points to contamination, evaluation examples leaking into the training set. It is the row that fails dishonestly because the model memorised the leaked material rather than learning the underlying skill, so the eval number reads as good as, or better than, a genuinely well-trained model instead of showing the damage the way every other corruption in the table does. The pipeline step that should have caught it is decontamination, scanning training text for overlap against the evaluation set, and it has to run before the split is finalised, not after: decontamination’s entire job is guaranteeing that nothing which ends up in evaluation was ever seen in training, so splitting first and only decontaminating afterward, or not at all, leaves near-duplicate pairs free to land on both sides of the wall the whole exercise relies on to keep the numbers honest.
Exercise 4
Expected count is the test fraction times the category’s total size: 0.05 * 50 = 2.5 examples. Treating this as a rare event drawn from a large pool, the Poisson approximation with mean 2.5 gives the probability of exactly zero as e^-2.5 = 0.082, about an eight percent chance. That is not a one-in-a-million fluke, it is a real, easily reachable outcome of an ordinary random shuffle, which means an unstratified split can silently produce a test set that says nothing at all about that subcategory roughly one run in twelve. There would be no way to tell from the resulting eval number alone, since the metric would simply be silent on a category it never happened to sample, rather than visibly wrong, which is exactly the case for stratifying the split rather than trusting a shuffle.
Exercise 5
The injection: truncate every training response to its first sentence, or its first handful of tokens, regardless of what the task actually called for, leaving prompts, roles and masking otherwise untouched. Training loss falls normally, likely faster and lower than a clean run, because predicting a short, low-entropy continuation before the model’s own learned stop point is a strictly easier objective per example, a property of the corrupted task, not evidence of a better model. A held-out eval loss computed the ordinary way, on an evaluation set that mirrors the same truncated shape, looks fine for the identical reason: it is measuring how well the model predicts short answers, which the corrupted model now does very well, so the corruption and the metric are built from the same assumption and the number cannot see past it. What actually reveals the damage is a before-and-after generation check, this series’ proving-the-shift idea, run on prompts that specifically call for a long, multi-step answer such as a how-to guide: the corrupted model will visibly cut a multi-step answer down to one sentence regardless of what the prompt needed, a failure that only shows up once you look at what the model actually produces against a task the training distribution’s shape cannot quietly excuse, rather than at a loss number computed on data shaped the same way the corruption shaped the training set.