A fine-tuning dataset is just a file of examples. That file decides more about your finished model than any setting in the training script. The script is the same for everyone. The data is yours, and it is where both the gains and the silent failures live.
This part is about the file, not the loop. What shape the examples take. How to tell good ones from bad. How to split them so your numbers mean something. And where examples written by another model help, and where they rot.
One fact sits under all of it, and it is uncomfortable. Train on badly damaged data and the training score still improves, smoothly, every time. The curve looks healthy while the model gets worse. Only a fair test on data the model has never seen shows the damage. And one specific mistake makes even that test lie to you.
By the end you will be able to choose a format that grades the right tokens, clean a dataset in the right order, split it three ways so a reported number is honest, and audit any dataset in about a minute before you spend a GPU on it.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data (you are here)
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
Pick the file format that grades the tokens you care about
Before quality comes shape. Start with one example, written out in plain English.
The instruction is “Rewrite this in plain English: The applicant is required to submit the form.” The answer is “You need to send us the form.”
That single example can be stored three different ways. The three are not interchangeable, because they hand the training loop three different jobs. One word has to come first.
What “graded” means
Training walks through an example one token at a time. A token is a chunk of text, usually a word or a piece of one. At each position the model guesses what comes next. That guess is scored against the token that really came next. The score is called the loss, and lower means the guess was closer.
Not every position has to be scored. You choose. A position that counts toward the loss is graded. A position that does not count is masked out, and the model learns nothing from it. Deciding which positions count is called loss masking, and Part 7 in the prerequisites above is entirely about it.
Hold on to that, because your file format picks the masking for you.
The three formats
Raw text. One field, usually called text, holding the instruction and the answer glued together as one string. Every token is graded, including the instruction.
Prompt and completion. Two fields, prompt and completion. The instruction goes in one, the answer in the other. Trainers grade the completion and skip the prompt.
Conversational. A list of messages, each with a role and some content. The roles are system, user and assistant. A system message sets the rules, a user message asks, an assistant message answers. You grade the assistant messages and skip the rest.
Now the consequence. Raw text grades the instruction as well as the answer. So the model is being taught to write instructions too. Ask it a question later and it will answer, then cheerfully write the next user question by itself, because that is a thing it saw and was rewarded for. Use raw text only when you want the model to soak up a body of domain writing. Never use it for instruction tuning.
Instruction tuning means training a model on instruction and answer pairs so it learns to follow instructions. That is what the rest of this part assumes you are doing.
Two details that break real runs
First, multi-turn conversations. If an example has three assistant replies, grade all three. Grading only the last one throws away two thirds of the teaching. It also lets the model drift into narrating both sides of a conversation, because it never saw the earlier replies rewarded.
Second, the stop marker. Every chat model has a special end-of-turn token that means “the answer stops here”. It has to sit inside the graded region. Leave it outside and the model never learns to stop talking. It will answer your question and then keep going, forever, which looks like a decoding bug and is not one.
The thing that lays those markers down is the chat template. It is a rule that ships with the model’s tokenizer, and it turns a list of messages into the exact text and special markers the model expects. Use the template that came with your model. Writing your own string format is the fastest way to teach a model a structure it will never see again at serving time.
Pack short examples without letting them leak into each other
Most instruction examples are short. Most training runs use a fixed maximum length. Those two facts cost you money, and the fix has a trap in it.
Say your maximum length is 1,024 tokens and your typical example is 200. Every sequence needs to be the same length for the hardware, so the other 824 positions get filled with padding: meaningless filler tokens that exist only to make the shapes line up. They are masked out of the loss, so they teach nothing. You are paying for 1,024 positions of compute and getting 200 positions of teaching. That is about 80 percent waste.
Packing fixes it by gluing several short examples end to end into one full-length sequence. Five 200-token examples fill the 1,024 almost exactly. Now nearly every position is real text.
Here is the trap. Inside the model, each token looks back at the tokens before it and pulls in whatever seems useful. That step is called attention, and it is what lets a model use context. Glue two examples together naively and the second one can look back into the first.
Picture it concretely. Example A is a recipe. Example B is a Python question. Packed naively, the tokens of the Python answer can see the recipe sitting behind them. The model quietly learns that Python answers sometimes follow recipes. Nothing errors. The loss looks fine. You have just taught the model a relationship that does not exist.
There is a second half to the trap. Each token also carries a number saying where it sits in the sequence. If example B starts at position 500, the model treats B’s first word as the 500th word of one long document. Correct packing resets that count at every boundary, so each example starts at position zero again.
In current TRL you get correct packing by setting packing=True in SFTConfig. The default packing_strategy="bfd" also switches on padding-free training, which flattens the batch into one continuous run of tokens. Per TRL’s own documentation that path needs a FlashAttention 2 or 3 backend, because that backend is what knows how to keep the flattened pieces apart. FlashAttention is a faster way of computing attention, covered in Part 3 of this series.
Two practical rules. Pack when you have many short examples and the padding waste is costing real time. Do not pack on your first learning runs, because it is one more thing that can corrupt the signal silently while you are still checking that your masks are right. And pin your library versions before you build anything on these flags, because packing and masking options have moved between releases more than once.
Judge a dataset by quality before you judge it by size
Here is a decision you will actually face. You have written 200 examples yourself, and they are good. A colleague offers 2,000 more scraped from a support forum. Maybe half of those are decent. Do you take them?
Instinct says yes, because more data is better. For supervised fine-tuning, usually it is not. The same instinct fails at pretraining scale too. Gunasekar and colleagues trained a 1.3B-parameter code model on a small curated corpus of textbook-quality text, and reported it beating much larger models trained on far more scraped code. Work out why from what training does.
The loss is averaged over the graded positions in the batch. So every example gets a vote on what the model becomes, and the votes are equal. Add 1,000 sloppy examples to 200 good ones and the sloppy ones now cast 83 percent of the votes. You have not added a little noise on top of your signal. You have made noise the majority position.
The reason this bites harder for fine-tuning than you might expect is that a fine-tune is not teaching facts. The model learned its facts during pretraining, over trillions of tokens. To picture the gap, Gao and colleagues assembled The Pile out of 22 separate sources into a corpus of roughly 825 GiB of text, and Brown and colleagues trained GPT-3 on a few hundred billion tokens. A few thousand examples cannot compete with that and are not trying to. What they install is a habit: the shape of the answer, its length, its tone, when to stop. Habits are learned from a few clean demonstrations and wrecked by many dirty ones.
The two published results everyone cites
Two studies made this concrete, and they are worth stating carefully rather than as folklore.
LIMA. Zhou and colleagues fine-tuned a 65B model on exactly 1,000 carefully written examples. Human raters then compared its answers against three other systems, one comparison at a time. LIMA’s answer was rated equal or better in 43 percent of comparisons against GPT-4, 58 percent against Bard, and 65 percent against DaVinci-003. Read those numbers as “how often the 1,000-example model held its own”. Against the strongest opponent in the set it still tied or won more than four times in ten.
AlpaGasus. Chen and colleagues started from a noisy 52,000-example instruction set. They asked a strong model to score every example, kept about 9,000, and threw the rest away. The model trained on the 9,000 was rated better than the model trained on all 52,000. It also trained about 5.7 times faster, cutting a 7B fine-tune from roughly 80 minutes to roughly 14.
The rule that falls out: adding data helps only while the new data is at least as clean as what you already have. Past that line, every extra example teaches the model your noise. So the honest answer to your colleague is “yes, if I can filter them first”, and the rest of this part is about how.
What “quality” decomposes into
Quality is not a feeling. It is five properties you can check, and each one has a failure it prevents.
| Property | What good looks like | The failure it prevents |
|---|---|---|
| Correctness | the answer is right, and it answers the question asked | the model learns confident wrong answers |
| Format consistency | every example uses the same template and the same roles | broken stop markers, drifting structure, rambling |
| Diversity | a wide spread of tasks, phrasings and lengths | mode collapse, where every answer comes out in one shape |
| Difficulty match | examples span the difficulty you will see in real use | a model that handles only easy cases, or only hard ones |
| Length sanity | answer length fits the task, not a fixed habit | learned padding, or learned terseness |
Two of those names need unpacking. Mode collapse is when tuning crushes the variety out of a model, so it produces the same phrasings over and over. You only catch it by reading many outputs, never a single sample. Length bias is a systematic pull toward one answer length that has nothing to do with the task. It usually arrives from raters preferring longer answers, and you can also install it directly by making all your own answers the same size.
Run the three cleaning passes, and run them in the right order
Real pipelines spend most of their effort here. Three operations do nearly all the work. The order they run in matters as much as the operations themselves.
Pass one: deduplication
Duplicates are not harmless filler. Go back to the voting picture. If one answer appears 50 times in a 1,000-example set, it casts 5 percent of the votes rather than 0.1 percent. It gets 50 times the pull on the weights that a single example would. The model memorises that answer and over-weights its style.
Deduplication runs in two passes.
Exact duplicates are easy. Turn each example into a fingerprint with a hash function, which maps any text to a short fixed-length value, and identical text always produces the same value. Group by fingerprint, keep one from each group. This is fast and catches copy-paste.
Near duplicates are the real problem. Compare these two:
“write a python function that reverses a string”
“write a function in python that reverses a string”
Those are the same example. Their hashes are completely different, because one word moved. To catch this you need a similarity score rather than an equality test. The standard one is Jaccard similarity, and you can compute it by hand.
- Take the set of words in the first text: write, a, python, function, that, reverses, string. That is 7 words.
- Take the set of words in the second: write, a, function, in, python, that, reverses, string. That is 8.
- Count the words in both: 7 of them.
- Count the words in either: 8, the 7 shared plus “in”.
- Divide. 7 divided by 8 is 0.875.
A score of 0.875 out of a possible 1.0 says these are near-copies. Set a threshold, say 0.8, and drop anything above it. Real tools chop the text into short runs of consecutive words rather than single words, which is stricter, but the arithmetic has exactly this shape.
One problem remains. Comparing every example against every other example means comparing every pair. For 10,000 examples that is about 50 million pairs, which is fine. For a million examples it is about 500 billion pairs, which is not. The standard escape is called min-hashing with locality-sensitive hashing. It builds a short fingerprint for each example whose chance of matching another fingerprint is roughly the Jaccard score you would have computed, then sorts examples into buckets so that only plausible matches ever get compared directly. You do not need to implement it. You need to know that a library such as datasketch exists for exactly this, and that near-duplicate removal without it does not scale.
The payoff is measured. Lee and colleagues found that more than 1 percent of the tokens a model produces unprompted are copied verbatim from its training set when that set was not deduplicated. Deduplicating cut the rate of memorised output by roughly a factor of ten.
Pass two: decontamination
Decontamination means checking that nothing in your training data also appears in the data you will measure on, and removing whatever does. Skip it and your reported number stops measuring skill and starts measuring memory.
Work the arithmetic once and you will never skip it again. Suppose you have 200 test questions. Your model genuinely answers 60 percent of unseen questions correctly. Now suppose some test questions accidentally got copied into the training set, so the model has simply memorised those answers.
- Nothing leaked. The score is 60 percent of 200, so 120 correct. Reported: 60 percent. Honest.
- 20 questions leaked, which is 10 percent. Those 20 are all correct. Of the remaining 180, 60 percent is 108. Total 128. Reported: 64 percent.
- 50 questions leaked, which is 25 percent. Those 50 are correct, plus 60 percent of the remaining 150, which is 90. Total 140. Reported: 70 percent.
The model never got better. In the last row you would report 70 percent for a model that is worth 60. That is why one leaked test item invalidates a benchmark: a benchmark is only a claim about unseen questions, and a leaked question is not unseen. A benchmark, here, means a fixed test set with a published score, the kind of thing people put on slides.
Decontamination in practice means scanning training text for overlap with your held-out set, and with any public benchmark you intend to quote. The usual check is n-gram overlap, which means looking for runs of the same consecutive words, typically eight or thirteen of them, appearing in both. Exact string matching is not enough, for the same reason exact deduplication is not enough. Reworded copies leak just as effectively.
This is not a rare hazard. The same deduplication study found train-test overlap affecting more than 4 percent of the validation set in standard benchmarks. Public benchmark questions also leak into pretraining data, which is why a public leaderboard score is a coarse filter and never a decision.
Pass three: quality scoring and filtering
You are not going to hand-read ten thousand examples. Score them instead, then keep a slice off the top. There is a ladder of rigour, and cheap comes first.
- Heuristics. Length bounds, a language check, regular expressions for boilerplate such as “As an assistant” or leftover markup. Costs nothing and removes a surprising amount.
- A scoring model. A reward model is a small model trained on human comparisons to give any answer a single number. Run it over the set and threshold.
- A strong model as judge. Ask a capable model to rate each example against a written rubric. This is the technique AlpaGasus used to get from 52,000 to 9,000.
Judging with a model works better than it has any right to. Zheng and colleagues found a strong judge agrees with human raters more than 80 percent of the time, which is about as often as two humans agree with each other. It also has known habits: it favours the answer shown first, it favours longer answers, and it favours its own writing style. Randomise the order, and be suspicious when the winners are all the long ones.
Finally, protect diversity on the way out. Cluster the survivors and cap how many near-copies of any one cluster you keep. Otherwise a top-scoring pattern quietly takes over the set, and you have optimised your way into mode collapse.
Both orderings in that figure are load-bearing. Deduplicate first so you do not spend judge calls scoring fifty copies of the same example. Decontaminate before you split, because once the data is split, a leak has already crossed the wall and no later step will find it.
One more leak deserves its own name, because a duplicate check cannot see it. Label leakage is when information from the expected answer sneaks into the input field of an example. Imagine a support dataset where the ticket text quietly includes the resolution code, because the export ran after the ticket was closed. The model reads the answer off the prompt. Your metrics soar. In production the field is empty and the model is useless. Part 15 covers how to hunt for it.
Every one of these passes is also an auditable control. Being able to say you deduplicated, decontaminated against the evaluation set, and filtered with documented thresholds is exactly the evidence a model-risk review will ask for. Build it as a logged, repeatable script rather than a notebook you ran once.
Split the data three ways and keep the wall between the parts
You need three sets, not two. The reason is about how each one gets used up.
- Train. The model learns on it. Nothing measured here means anything, because the model has seen every answer.
- Validation. You check this during training to pick settings and to choose which saved copy of the model to keep. Keeping the copy from the point where validation loss was lowest is called early stopping.
- Test. You look at this once, at the very end, for one honest number.
Why validation cannot also be test is worth building slowly, because it is the part people talk themselves out of.
Say you try twelve learning rates and keep whichever scores best on validation. You have now made twelve decisions using that set. The winner is partly genuinely better and partly just lucky on those particular examples. Repeat that over a project and validation slowly turns into a set you have fitted to, one decision at a time. It gets used up. That is not a moral failing, it is arithmetic, and it is why the test set exists and stays sealed.
Tune against the test set even once and it becomes a second validation set. There is no ceremony that restores it.
Split by category, not at random
A random split can betray you, and here is the shape of it.
The word you need first is distribution, which just means the tally of how much of each kind of thing your data contains. The no_robots dataset is a good example. Its 10,000 examples carry a category label, and the tally is lopsided. Generation covers about 4,560 of them. Extract covers about 190.
Now take a random 5 percent test split, so 500 examples. On average about 9 or 10 of those should be Extract examples. On average. A random draw can easily hand you two, or none. If it hands you none, your test score says precisely nothing about extraction, and you will not notice, because the number will look fine.
A stratified split fixes this. Split within each category separately, then combine. Take 5 percent of Generation, 5 percent of Extract, 5 percent of every other category, and glue the pieces together. Every category then appears in every split in the same proportion it appears overall. Do the same by difficulty if you have difficulty labels, and by source if your data came from several places.
Keep one probe that looks nothing like your training data
Your validation and test sets are drawn from the same pile as your training set. So they can only tell you how the model does on data that looks like what you trained on. They are blind to everything else.
Everything else includes the thing most likely to hurt you. Fine-tuning hard on one task can quietly erode abilities the model already had, such as arithmetic or following a formatting instruction. That is called catastrophic forgetting. Your in-distribution test set will report success the whole time it is happening.
The fix is a small out-of-distribution probe: a handful of tasks deliberately unlike your training data, scored before and after the fine-tune. A few maths questions, a few general-knowledge questions, a couple of instruction-following checks. It costs almost nothing and it is the only instrument that sees this failure. Part 15 turns it into a full procedure.
The rule to carry away: any number you show someone is worth exactly as much as the wall between the set it was measured on and the set the model trained on. Write the split method down, version it, and never move the wall to make a number look better.
Generate synthetic data without collapsing the model
Human-written examples are slow and expensive. Synthetic data means training examples written by a model instead of a person. It is how many teams reach scale, and it is a genuine capability with a genuine failure mode attached.
Three reasons to generate at all. Scale, because a hundred thousand examples costs API calls rather than annotator hours. Targeted coverage, because if you need 500 examples of one rare edge case you can simply ask for exactly those. And distillation, where a strong model writes high-quality answers and a small model trains to copy them.
| Method | The idea | Landmark work |
|---|---|---|
| Self-Instruct | seed with a few human-written tasks, then let the model write thousands more instructions and answers from those | Wang et al., 2022 |
| Evol-Instruct | take an existing prompt and rewrite it to be deeper, harder or more constrained, so you manufacture difficulty on demand | Xu et al., 2023 |
| Distillation | a stronger teacher model writes the answers and a smaller student trains to match them | the prune-and-distil lineage behind Llama-3.2-1B, Part 8 |
The failure that defines the topic: model collapse
Build this one up from marbles, because the mechanism is simpler than the name suggests.
Put 1,000 marbles in a bag. 990 are red and 10 are blue. Draw 100 at random. On average you expect 1 blue marble. Sometimes you get 2. Often you get 0.
Now throw the original bag away and refill it from what you drew, scaled back up to 1,000. If your 100 draws contained no blue, there is no blue in the new bag. Blue is gone. Not rare, gone. Repeat the whole procedure and the next-rarest colour goes the same way.
A model generating training data behaves like that draw. It produces common patterns often and rare ones rarely. The rare ones live in what statisticians call the tails of the distribution: the edge cases that are individually unlikely and collectively important. Train the next model on that output and the tails thin. Do it again and they vanish.
Shumailov and colleagues published exactly this experiment in Nature in 2024, and named the effect model collapse. They fine-tuned a small language model, OPT-125m, on the wikitext2 dataset. Then they used it to generate a dataset the same size, trained the next generation on that, and repeated for nine generations, five separate times.
Their headline measure was perplexity, which is one score for how surprised a model is by a piece of real text. Lower is better. Before fine-tuning, the model scored about 115 on the real test text. After fine-tuning on the real wikitext2 data it scored about 34, which is the model doing its job well. When each generation trained only on the previous generation’s output, later generations gave back 20 to 28 points of that gain.
The examples are more vivid than the numbers. Given a prompt about fourteenth-century English church architecture, the first model wrote a sensible paragraph about Perpendicular Revival architecture and St. John’s Cathedral in London. By generation nine, the same prompt produced text about the world’s largest populations of black-tailed jackrabbits, white-tailed jackrabbits, blue-tailed jackrabbits, red-tailed jackrabbits and yellow-tailed jackrabbits. The model has not become random. It has fallen into a groove and is repeating variations of one pattern.
The paper separates two stages. In early collapse the model starts losing the tails, so rare cases stop appearing. In late collapse the model converges on something with little resemblance to the original data, with much less variety left in it. Early collapse is the dangerous one, because a model in early collapse still looks fine on ordinary examples.
The primary cause they name is plain sampling. Any finite sample of a distribution loses some of the rare events, and information lost at one step can never come back at the next. That is the marble bag. Two secondary causes sit on top: the model cannot represent the original distribution perfectly, and the training procedure itself has its own biases.
One detail worth knowing, because it kills the obvious fix. The collapsed models repeat phrases constantly, so the natural response is to force variety with a repetition penalty. The researchers tried it. The models produced worse continuations, and perplexity doubled compared with the original. Suppressing the symptom made the underlying problem worse.
The defence is mixing, not abstinence
The same paper points at the fix. In a second setting they kept 10 percent of the original real data in every generation’s training set. The degradation dropped to minor. Real data anchors the tails that generated data drops.
So the working recipe for synthetic data is a pipeline, not a prompt.
- Keep real data in the mix. Never train a generation purely on the previous generation’s output.
- Deduplicate the generated set first. Models repeat themselves far more than people do, so near-duplicate rates in generated data are high.
- Verify anything checkable. For maths, check the answer. For code, run the tests. A teacher model’s mistakes get copied faithfully and then amplified.
- Then score and filter, and cap per cluster so one dominant pattern cannot swamp the set.
The useful way to hold it: generation sets the ceiling on possible quality, and curation decides how much of that ceiling you actually reach. Producing a million examples is a weekend of API calls. Making them usable is the job.
The licence follows the weights
One compliance point that is not a footnote. Data generated by a commercial model is governed by that model’s terms of service, and several of those explicitly forbid using the outputs to train a competing model. So “we distilled our instruction set from a frontier model” can carry legal exposure that travels with the resulting weights, long after anyone remembers where the data came from.
For a defensible pipeline, use permissively licensed teacher models or your own. Record the provenance of every synthetic batch: which model, which version, which prompt, which date. Keep that record with the dataset. It is precisely what a model-risk review asks for, and it is why a fully documented training-data story has real commercial value.
Poison a copy of your data on purpose and see what each fault breaks
Everything above is easier to believe once you have watched it happen. The protocol is small. Take a clean baseline dataset. Make several copies, and damage each copy in exactly one way, so only one variable moves at a time. Train on each. Compare all of them against the same clean held-out set and the same handful of generation checks.
The damage column below is reasoned from the mechanism rather than measured from one particular run. Treat it as what to expect, and confirm it against your own numbers.
| Corruption | What you inject | Expected damage |
|---|---|---|
| Duplication | copy about 30 percent of the examples several times over | memorisation and verbatim regurgitation, and one style over-weighted |
| Format corruption | break the role markers, or drop the end-of-turn token on some examples | the model rambles, will not stop, and its structure drifts |
| Length bias | truncate every answer to a single sentence | the model becomes uselessly terse no matter what the task needs |
| Refusal injection | replace about 10 percent of answers with a refusal | the model refuses harmless requests, a self-inflicted alignment tax |
| Contamination | copy test examples into the training set | the evaluation looks excellent and is fiction |
The refusal row introduces one term. An alignment tax is capability lost in exchange for making a model better behaved. Usually it arrives from a deliberate safety decision. This row shows you can install it entirely by accident, through data hygiene alone.
Now the reason this protocol teaches more than any amount of reading. Every training curve in that experiment goes down. All of them. The training loss measures how well the model predicts the text it was given, and it has no opinion about whether that text was any good. Truncate every answer to one sentence and the model learns to produce one sentence, and gets steadily better at it, and the curve looks textbook.
Four of the five corruptions at least fail honestly: run a clean held-out evaluation and the damage shows up as a worse score. Contamination is the exception. It fails dishonestly, because the model memorised the leaked answers, so the reported number looks as good as baseline while the model is worse. That is why decontamination is the control you never skip, and why it has to run before the split.
Audit any dataset in one minute before you train on it
You do not need a purpose-built tool to catch most of this. Three questions cover the bulk of it. How many examples are actually here? How many are exact copies? And did anything in train leak into the evaluation split?
from collections import Counter
from datasets import load_dataset
ds = load_dataset("HuggingFaceH4/no_robots")
train, test = ds["train"], ds["test"]
def norm(example):
joined = " ".join(m["content"] for m in example["messages"])
return " ".join(joined.split()).lower()
train_texts = [norm(r) for r in train]
test_texts = [norm(r) for r in test]
print("train examples: ", len(train_texts))
print("test examples: ", len(test_texts))
print("exact duplicates:", len(train_texts) - len(set(train_texts)))
print("train/test leak: ", len(set(train_texts) & set(test_texts)))
print("categories: ", Counter(train["category"]).most_common())
Line by line, assuming you have never used any of these libraries.
from collections import Counterpulls in a standard Python helper that tallies how many times each value appears in a list.from datasets import load_datasetpulls in the Hugging Face datasets library, which downloads a named dataset and caches it locally.load_dataset("HuggingFaceH4/no_robots")fetches the dataset. It returns a dictionary-like object with one entry per split.ds["train"], ds["test"]takes the two splits out. Here that is 9,500 and 500 examples.def norm(example)defines a function that flattens one example into a single plain string, so that two examples can be compared.example["messages"]is the list of turns. Each turn is a small dictionary with a role and a content. The join glues every turn’s text together with spaces." ".join(joined.split())collapses every run of spaces, tabs and newlines into one space..lower()then lowercases. Both exist so that two examples differing only in spacing or capitals are treated as the same example.- The two list comprehensions run
normover every row and collect the results, giving one string per example. set(train_texts)builds a Python set, which silently discards repeats. So the list length minus the set length is the number of extra copies.set(train_texts) & set(test_texts)is set intersection: the strings that appear in both. Anything above zero is a leak, and it is the one line here you must never ignore.Counter(train["category"]).most_common()reads the whole category column as a list, tallies it, and sorts it by count. This is the lopsided distribution from the splitting section, printed out.
Extend the same shape in three directions when you need more. Add a near-duplicate rate by running min-hash over the normalised text and thresholding on Jaccard similarity, using a library rather than writing it. Add a format check by confirming every example survives the tokenizer’s chat template without an error. Add a length distribution by counting tokens per example with the exact tokenizer you are about to train with, because that is the only count that matches what training will see.
Read the numbers your own run prints rather than trusting any figure quoted here. They depend on the dataset, on how you chose to normalise the text, and on your similarity threshold.
Run it before every training job, not once. It costs a minute, and it catches the class of problem that otherwise appears as an unexplained quality regression three weeks later. This discipline pays off whatever you train next, including the low-rank adapters in the next part on LoRA.
Key takeaways
- The file format picks your masking, and masking decides what the model learns. Raw text grades the instruction too, which teaches the model to write instructions. Use prompt-and-completion or conversational data for instruction tuning.
- Packing removes most of the padding waste on short examples. Done naively it lets one example read the one before it, so it needs boundary resets, and it is not a first-run feature.
- Quality beats quantity because every example gets an equal vote. Adding data helps only while the new data is at least as clean as what you already have.
- Deduplicate before scoring, and decontaminate before splitting. A duplicate multiplies an example’s pull on the weights. A leak turns your reported score into fiction.
- Use three sets. Validation is used up by the decisions you make from it, test is opened once, and a separate out-of-distribution probe is the only thing that sees forgetting.
- Synthetic data scales cheaply and collapses if you train on it recursively, because each generation drops the rare cases. Mix real data back in, deduplicate, verify what is checkable, and record the licence.
- Corrupted data still produces a falling training curve. Only a clean held-out evaluation shows the damage, and contamination makes even that lie.
You can now
- Choose between raw text, prompt-and-completion and conversational data by asking which tokens you want graded, from “Pick the file format that grades the tokens you care about”.
- Decide whether to switch packing on, and name the two boundary resets correct packing has to apply, from “Pack short examples without letting them leak into each other”.
- Compute a Jaccard similarity by hand and set a near-duplicate threshold from it, from “Run the three cleaning passes, and run them in the right order”.
- Work out how much a given leak rate inflates a reported benchmark score, from “Run the three cleaning passes, and run them in the right order”.
- Build a stratified three-way split plus an out-of-distribution probe, and say what each part is for, from “Split the data three ways and keep the wall between the parts”.
- Design a synthetic-data pipeline that resists model collapse, and justify each step, from “Generate synthetic data without collapsing the model”.
- Audit an unfamiliar dataset for size, duplicates, leaks and category balance in about a minute, from “Audit any dataset in one minute before you train on it”.
Glossary
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Alignment tax
- The capability a model loses in exchange for being made more helpful or better behaved. Running supervised fine-tuning and then preference tuning is enough on its own to cause it. Lin et al., Mitigating the Alignment Tax of RLHF
- Attention
- The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
- Backward pass
- Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
- Base model
- The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- Benchmark
- A fixed public test set with a published score, such as a maths or general-knowledge exam for models. Useful as a coarse filter and untrustworthy as a decision, because public questions leak into pretraining data. Xu et al., Benchmark Data Contamination survey
- Catastrophic forgetting
- When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
- Chat template
- The rule that turns a list of role-and-content messages into the exact text and special tokens the model expects to see. It ships with the tokenizer, and it is where the boundary between prompt and response lives. From messages to tensors
- Checkpoint
- A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
- Data contamination
- Overlap between your training data and your evaluation data, including reworded near-copies. It is the one data problem that fails dishonestly: the score looks excellent because the model memorised the answers. Xu et al., Benchmark Data Contamination survey
- Decontamination
- Checking that nothing in the training data also appears in the evaluation data, and removing whatever does. Skip it and your reported score measures memorisation rather than skill.
- Deduplication
- Removing exact and near-duplicate examples from a dataset. Every copy of an example multiplies its pull on the weights, so duplicates cause memorisation and over-weight one style. Lee et al., Deduplicating Training Data Makes Language Models Better
- Early stopping
- Keeping the checkpoint from the point where held-out loss was lowest, instead of the one from the last step. Past that point the model is memorising rather than learning. Proving the shift
- Embedding
- The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
- End-of-turn token
- The special token that marks where an answer stops. If it is left out of the graded region during training, or the wrong one is passed at generation time, the model rambles past the end of its answer. Masking in LLM training
- Flash attention
- An attention implementation that computes the answer in small tiles inside fast on-chip memory, never building the full score grid. Same result, far less memory, and the term that grew with the square of sequence length becomes linear. Dao et al., FlashAttention
- Forward pass
- Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Holdout
- Data deliberately kept out of training so you can measure the model on something it has never seen. A number measured on data the model trained on is not evidence of anything. Building a holdout you can trust
- Hyperparameter
- A setting you choose before training rather than something the model learns, such as the learning rate, the batch size or the number of epochs. The knobs that decide whether it learns
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- Instruction tuning
- Supervised fine-tuning on instruction-and-answer pairs, so a base model learns to follow instructions and answer in a consistent shape. Zhou et al., LIMA
- Knowledge distillation
- Training a small model to copy a larger model’s outputs rather than learning from raw data alone. Llama-3.2-1B was built this way, after a larger model had been pruned down. Llama-3.2-1B model card
- Label
- The correct answer a training example is graded against. For a language model the labels are the token IDs the model was supposed to produce, and a label of -100 means do not grade this position at all. Masking in LLM training
- Length bias
- The tendency of both human and model annotators to rate longer answers as better, which quietly teaches a preference-tuned model to pad. If a tuned model starts waffling, suspect the data before the algorithm. Meng et al., SimPO
- LLM as a judge
- Using a strong model to score or compare answers against a written rubric. It agrees with human raters over 80 percent of the time, about as often as humans agree with each other, and it favours the first answer shown, longer answers, and its own style. Zheng et al., Judging LLM-as-a-Judge
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Loss masking
- Deciding which positions in an example count toward the loss. For instruction tuning you grade only the assistant’s response and mark everything else with -100, so it contributes no loss and no gradient. Masking in LLM training
- Mode collapse
- When tuning crushes the variety out of a model’s answers, so it produces the same shapes and phrasings over and over. You catch it by reading many outputs, never a single sample. The six silent failure modes
- Model collapse
- What happens when models are trained on their own generated output over and over: the rare cases at the edges disappear first, then diversity, then quality. The defence is mixing real data back in, not abstinence. Shumailov et al., AI models collapse on recursively generated data
- Out-of-distribution probe
- A small evaluation on tasks that look nothing like your training data, kept specifically to catch abilities you have lost. In-distribution evaluation cannot see catastrophic forgetting at all. Evaluating a fine-tune
- Overfitting
- When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse. The six silent failure modes
- Padding
- Filler tokens added to short sequences so every sequence in a batch has the same length. Padding is bookkeeping for the hardware, and it has to be masked out of both attention and the loss so it cannot change the answer. Masking in LLM training
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- Reward model
- A model trained on human comparisons to give any answer a single score. RLHF optimises against it because you cannot put a human in the loop for millions of steps. Ouyang et al., InstructGPT
- Sequence packing
- Concatenating several short examples into one full-length sequence so you stop wasting compute on padding. It needs boundary resets, or one example can attend back into another and the model learns nonsense continuations. Hugging Face TRL, SFT Trainer
- Stratified split
- Splitting a dataset so every category appears in the same proportion in each split, instead of splitting at random. Without it a whole category can land entirely in the training set and your test score says nothing about it.
- Structural pruning
- Permanently removing whole pieces of a trained model, such as layers or attention heads, to make it smaller. It is normally followed by more training to recover the quality that was lost. Llama-3.2-1B model card
- Supervised fine-tuning
- Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
- Synthetic data
- Training examples written by a model rather than by a person. Cheap and endlessly scalable, and it needs curation, a mix of real data alongside it, and a check on the licence of whatever model produced it. Wang et al., Self-Instruct
- Test set
- The split you look at once, at the very end, for one honest number. Tune against it and it quietly becomes a second validation set and stops being trustworthy.
- Token
- The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
- Tokenizer
- The component that splits text into tokens and maps them to integer IDs, and back again. Every model has its own, and it has to match the model you are training. Choosing a base model and building a bench
- Validation set
- The split you check during training to pick settings and choose the best checkpoint. Because you make decisions from it over and over, it slowly stops being an honest estimate.
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
Practical exercises
Hand-compute the forensics script’s output on a toy dataset
Using the shape of this part’s forensics script, not the script itself, work out what it would print for a train set of the four normalised strings “a b c”, “a b c”, “d e f”, “g h i”, and a test set of the two strings “d e f”, “j k l”. State the train example count, the test example count, the exact duplicate count in train, and the exact train and eval overlap count.
See the worked solution (opens in a new tab)
Reconcile AlpaGasus’s data reduction against its speedup
AlpaGasus filtered 52,000 examples down to about 9,000, and this part states that cut a 7B fine tune from about 80 minutes to about 14. Compute the dataset size reduction factor and the training time reduction factor separately, compare the two, and explain in one sentence why they land so close to each other.
See the worked solution (opens in a new tab)
Diagnose an eval score that looks too good to be true
After fine tuning, your held-out eval loss looks as good as baseline. Spot checking generations on prompts that closely resemble your test set, several answers are near word-for-word identical to the test set’s own gold answers, phrasing included. Using this part’s corruption table, name the corruption this points to, explain why it is the one row that fails dishonestly rather than showing up as a worse score, and name the pipeline step, and its correct order relative to the split, that should have caught it.
See the worked solution (opens in a new tab)
Quantify the risk an unstratified split poses to a minority category
Assume a rare subcategory makes up about 0.5 percent of a 10,000-example dataset, roughly 50 examples total, and you take a pure random 5 percent test split, 500 examples, with no stratification. Compute the expected number of that subcategory’s examples landing in the test set, then use the Poisson approximation to estimate the probability that the test set ends up with zero examples of it purely by chance. State what that probability tells you about relying on an unstratified split for a minority category this size.
See the worked solution (opens in a new tab)
Design a length-bias corruption experiment and find what actually reveals it
Design the length-bias row from this part’s corruption table as an actual experiment on a copy of no_robots. Describe exactly what you would inject, then reason through what happens to the training loss curve, the held-out eval loss computed the ordinary way, and a before-and-after generation check on a prompt that calls for a genuinely long, multi-step answer. Explain why the first two would not tell you this run went wrong, and what would.
Frequently asked questions
How much data do I need to fine-tune an LLM?
Far less than most people expect, if it is clean. LIMA reached strong instruction following on exactly 1,000 curated examples, and AlpaGasus filtered a noisy 52,000-example set down to about 9,000 and produced a better model that trained about 5.7 times faster. Start with hundreds to a few thousand high-quality examples rather than tens of thousands of scraped ones.
Is more data always better for fine-tuning?
No. The loss is averaged over examples, so every example gets an equal vote on what the model becomes. Adding data helps only while the new data is at least as clean as what you already have, and past that point each noisy example teaches the model your noise.
What is data contamination and why does it matter so much?
Contamination is any overlap between your training data and the data you evaluate on, including reworded near-copies rather than exact ones. It matters because it is the one data problem that fails dishonestly: the model memorises the leaked answers, so the reported number looks excellent while the model is worse. Leak 10 percent of a test set into training and a model worth 60 percent reports 64.
Can I train on synthetic data generated by another model?
Yes, and many teams blend it into their instruction sets, but two constraints apply. Training each generation only on the last one’s output causes model collapse as the rare cases disappear, so keep real data in the mix and curate hard. And commercial model outputs are governed by that provider’s terms, several of which forbid training a competing model.
Do I need train, validation and test, or is two enough?
Three, if you will ever report a number to anyone. Validation is used up by the repeated decisions you make from it, such as picking settings and choosing a checkpoint, so it stops being an honest estimate. Test is opened once at the end, and a fourth out-of-distribution probe catches forgetting that in-distribution sets cannot see.
Should I use packing for my fine-tuning dataset?
Only once your masks are verified and padding waste is measurably costing you time. Packing concatenates short examples to fill the sequence, which removes most of the padding waste, but it needs boundary resets so that attention and position counting restart at each example. Without them the model learns continuations that run across two unrelated examples.
Sources and further reading
- Zhou et al., LIMA: Less Is More for Alignment, the thousand-example result and the comparison percentages quoted above.
- Chen et al., AlpaGasus: Training a Better Alpaca with Fewer Data, the 52,000 down to 9,000 filtering result.
- Lee et al., Deduplicating Training Data Makes Language Models Better, for the memorisation rate and the train-test overlap figure.
- Shumailov et al., AI models collapse when trained on recursively generated data, Nature 2024, the source of every model-collapse number and example above.
- Wang et al., Self-Instruct: Aligning Language Models with Self-Generated Instructions.
- Xu et al., WizardLM: Empowering Large Language Models to Follow Complex Instructions, which introduced Evol-Instruct.
- Gunasekar et al., Textbooks Are All You Need, the strongest published statement of the argument that data quality beats data volume.
- Gao et al., The Pile: An 800GB Dataset of Diverse Text for Language Modeling, for what a large pretraining corpus actually contains.
- Brown et al., Language Models are Few-Shot Learners, for the pretraining data scale a fine-tuning set is being compared against.
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, for the judge agreement rate and its biases.
- Hugging Face TRL, SFT Trainer documentation, for the packing, padding-free and masking flags.
- HuggingFaceH4/no_robots dataset card, the source of the split sizes and category counts.
