Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data

Inside LLM Fine-Tuning, part 10 of 15: Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data

A fine-tuning dataset is just a file of examples. That file decides more about your finished model than any setting in the training script. The script is the same for everyone. The data is yours, and it is where both the gains and the silent failures live.

This part is about the file, not the loop. What shape the examples take. How to tell good ones from bad. How to split them so your numbers mean something. And where examples written by another model help, and where they rot.

One fact sits under all of it, and it is uncomfortable. Train on badly damaged data and the training score still improves, smoothly, every time. The curve looks healthy while the model gets worse. Only a fair test on data the model has never seen shows the damage. And one specific mistake makes even that test lie to you.

By the end you will be able to choose a format that grades the right tokens, clean a dataset in the right order, split it three ways so a reported number is honest, and audit any dataset in about a minute before you spend a GPU on it.

Pick the file format that grades the tokens you care about

Before quality comes shape. Start with one example, written out in plain English.

The instruction is “Rewrite this in plain English: The applicant is required to submit the form.” The answer is “You need to send us the form.”

That single example can be stored three different ways. The three are not interchangeable, because they hand the training loop three different jobs. One word has to come first.

What “graded” means

Training walks through an example one token at a time. A token is a chunk of text, usually a word or a piece of one. At each position the model guesses what comes next. That guess is scored against the token that really came next. The score is called the loss, and lower means the guess was closer.

Not every position has to be scored. You choose. A position that counts toward the loss is graded. A position that does not count is masked out, and the model learns nothing from it. Deciding which positions count is called loss masking, and Part 7 in the prerequisites above is entirely about it.

Hold on to that, because your file format picks the masking for you.

The three formats

Raw text. One field, usually called text, holding the instruction and the answer glued together as one string. Every token is graded, including the instruction.

Prompt and completion. Two fields, prompt and completion. The instruction goes in one, the answer in the other. Trainers grade the completion and skip the prompt.

Conversational. A list of messages, each with a role and some content. The roles are system, user and assistant. A system message sets the rules, a user message asks, an assistant message answers. You grade the assistant messages and skip the rest.

① raw text (continuation / pretraining-style) every token graded — no mask use for: domain adaptation on plain corpora. Not instruction-following. ② prompt / completion (two fields) prompt — masked (-100) completion — graded flag: completion_only_loss=True (TRL default for this shape) ③ conversational (messages: list of role turns) system+user assistant ✓ user assistant ✓ flag: assistant_only_loss=True — grades EVERY assistant turn (multi-turn) no_robots is this shape. Note: both assistant turns are graded, not just the last.

Same instruction, same answer, three files. The shaded band on each row is the part that counts toward the loss, and it moves.

Now the consequence. Raw text grades the instruction as well as the answer. So the model is being taught to write instructions too. Ask it a question later and it will answer, then cheerfully write the next user question by itself, because that is a thing it saw and was rewarded for. Use raw text only when you want the model to soak up a body of domain writing. Never use it for instruction tuning.

Instruction tuning means training a model on instruction and answer pairs so it learns to follow instructions. That is what the rest of this part assumes you are doing.

Two details that break real runs

First, multi-turn conversations. If an example has three assistant replies, grade all three. Grading only the last one throws away two thirds of the teaching. It also lets the model drift into narrating both sides of a conversation, because it never saw the earlier replies rewarded.

Second, the stop marker. Every chat model has a special end-of-turn token that means “the answer stops here”. It has to sit inside the graded region. Leave it outside and the model never learns to stop talking. It will answer your question and then keep going, forever, which looks like a decoding bug and is not one.

The thing that lays those markers down is the chat template. It is a rule that ships with the model’s tokenizer, and it turns a list of messages into the exact text and special markers the model expects. Use the template that came with your model. Writing your own string format is the fastest way to teach a model a structure it will never see again at serving time.

Pack short examples without letting them leak into each other

Most instruction examples are short. Most training runs use a fixed maximum length. Those two facts cost you money, and the fix has a trap in it.

Say your maximum length is 1,024 tokens and your typical example is 200. Every sequence needs to be the same length for the hardware, so the other 824 positions get filled with padding: meaningless filler tokens that exist only to make the shapes line up. They are masked out of the loss, so they teach nothing. You are paying for 1,024 positions of compute and getting 200 positions of teaching. That is about 80 percent waste.

Packing fixes it by gluing several short examples end to end into one full-length sequence. Five 200-token examples fill the 1,024 almost exactly. Now nearly every position is real text.

Here is the trap. Inside the model, each token looks back at the tokens before it and pulls in whatever seems useful. That step is called attention, and it is what lets a model use context. Glue two examples together naively and the second one can look back into the first.

Picture it concretely. Example A is a recipe. Example B is a Python question. Packed naively, the tokens of the Python answer can see the recipe sitting behind them. The model quietly learns that Python answers sometimes follow recipes. Nothing errors. The loss looks fine. You have just taught the model a relationship that does not exist.

UNPACKED — three examples, padded to max_length (wasted compute in grey): ~40% padding = wasted FLOPs PACKED — same three, concatenated into one full sequence: boundaries = attention resets

The top row wastes most of its width on filler. The bottom row is full of real text, and the vertical bars are the boundary resets that stop one example from reading the one before it.

There is a second half to the trap. Each token also carries a number saying where it sits in the sequence. If example B starts at position 500, the model treats B’s first word as the 500th word of one long document. Correct packing resets that count at every boundary, so each example starts at position zero again.

In current TRL you get correct packing by setting packing=True in SFTConfig. The default packing_strategy="bfd" also switches on padding-free training, which flattens the batch into one continuous run of tokens. Per TRL’s own documentation that path needs a FlashAttention 2 or 3 backend, because that backend is what knows how to keep the flattened pieces apart. FlashAttention is a faster way of computing attention, covered in Part 3 of this series.

Two practical rules. Pack when you have many short examples and the padding waste is costing real time. Do not pack on your first learning runs, because it is one more thing that can corrupt the signal silently while you are still checking that your masks are right. And pin your library versions before you build anything on these flags, because packing and masking options have moved between releases more than once.

Judge a dataset by quality before you judge it by size

Here is a decision you will actually face. You have written 200 examples yourself, and they are good. A colleague offers 2,000 more scraped from a support forum. Maybe half of those are decent. Do you take them?

Instinct says yes, because more data is better. For supervised fine-tuning, usually it is not. The same instinct fails at pretraining scale too. Gunasekar and colleagues trained a 1.3B-parameter code model on a small curated corpus of textbook-quality text, and reported it beating much larger models trained on far more scraped code. Work out why from what training does.

The loss is averaged over the graded positions in the batch. So every example gets a vote on what the model becomes, and the votes are equal. Add 1,000 sloppy examples to 200 good ones and the sloppy ones now cast 83 percent of the votes. You have not added a little noise on top of your signal. You have made noise the majority position.

The reason this bites harder for fine-tuning than you might expect is that a fine-tune is not teaching facts. The model learned its facts during pretraining, over trillions of tokens. To picture the gap, Gao and colleagues assembled The Pile out of 22 separate sources into a corpus of roughly 825 GiB of text, and Brown and colleagues trained GPT-3 on a few hundred billion tokens. A few thousand examples cannot compete with that and are not trying to. What they install is a habit: the shape of the answer, its length, its tone, when to stop. Habits are learned from a few clean demonstrations and wrecked by many dirty ones.

The two published results everyone cites

Two studies made this concrete, and they are worth stating carefully rather than as folklore.

LIMA. Zhou and colleagues fine-tuned a 65B model on exactly 1,000 carefully written examples. Human raters then compared its answers against three other systems, one comparison at a time. LIMA’s answer was rated equal or better in 43 percent of comparisons against GPT-4, 58 percent against Bard, and 65 percent against DaVinci-003. Read those numbers as “how often the 1,000-example model held its own”. Against the strongest opponent in the set it still tied or won more than four times in ten.

AlpaGasus. Chen and colleagues started from a noisy 52,000-example instruction set. They asked a strong model to score every example, kept about 9,000, and threw the rest away. The model trained on the 9,000 was rated better than the model trained on all 52,000. It also trained about 5.7 times faster, cutting a 7B fine-tune from roughly 80 minutes to roughly 14.

dataset size → quality out 52k Alpaca (noisy) 9k filtered ✓ beats 52k LIMA 1,000 curated more examples ≠ better if the added examples are worse than what you have

Both bars point the same way. Fewer, cleaner examples beat many more noisy ones, and the smaller set finishes training in a fraction of the time.

The rule that falls out: adding data helps only while the new data is at least as clean as what you already have. Past that line, every extra example teaches the model your noise. So the honest answer to your colleague is “yes, if I can filter them first”, and the rest of this part is about how.

What “quality” decomposes into

Quality is not a feeling. It is five properties you can check, and each one has a failure it prevents.

Property What good looks like The failure it prevents
Correctness the answer is right, and it answers the question asked the model learns confident wrong answers
Format consistency every example uses the same template and the same roles broken stop markers, drifting structure, rambling
Diversity a wide spread of tasks, phrasings and lengths mode collapse, where every answer comes out in one shape
Difficulty match examples span the difficulty you will see in real use a model that handles only easy cases, or only hard ones
Length sanity answer length fits the task, not a fixed habit learned padding, or learned terseness

Two of those names need unpacking. Mode collapse is when tuning crushes the variety out of a model, so it produces the same phrasings over and over. You only catch it by reading many outputs, never a single sample. Length bias is a systematic pull toward one answer length that has nothing to do with the task. It usually arrives from raters preferring longer answers, and you can also install it directly by making all your own answers the same size.

Run the three cleaning passes, and run them in the right order

Real pipelines spend most of their effort here. Three operations do nearly all the work. The order they run in matters as much as the operations themselves.

Pass one: deduplication

Duplicates are not harmless filler. Go back to the voting picture. If one answer appears 50 times in a 1,000-example set, it casts 5 percent of the votes rather than 0.1 percent. It gets 50 times the pull on the weights that a single example would. The model memorises that answer and over-weights its style.

Deduplication runs in two passes.

Exact duplicates are easy. Turn each example into a fingerprint with a hash function, which maps any text to a short fixed-length value, and identical text always produces the same value. Group by fingerprint, keep one from each group. This is fast and catches copy-paste.

Near duplicates are the real problem. Compare these two:

“write a python function that reverses a string”

“write a function in python that reverses a string”

Those are the same example. Their hashes are completely different, because one word moved. To catch this you need a similarity score rather than an equality test. The standard one is Jaccard similarity, and you can compute it by hand.

  1. Take the set of words in the first text: write, a, python, function, that, reverses, string. That is 7 words.
  2. Take the set of words in the second: write, a, function, in, python, that, reverses, string. That is 8.
  3. Count the words in both: 7 of them.
  4. Count the words in either: 8, the 7 shared plus “in”.
  5. Divide. 7 divided by 8 is 0.875.

A score of 0.875 out of a possible 1.0 says these are near-copies. Set a threshold, say 0.8, and drop anything above it. Real tools chop the text into short runs of consecutive words rather than single words, which is stricter, but the arithmetic has exactly this shape.

One problem remains. Comparing every example against every other example means comparing every pair. For 10,000 examples that is about 50 million pairs, which is fine. For a million examples it is about 500 billion pairs, which is not. The standard escape is called min-hashing with locality-sensitive hashing. It builds a short fingerprint for each example whose chance of matching another fingerprint is roughly the Jaccard score you would have computed, then sorts examples into buckets so that only plausible matches ever get compared directly. You do not need to implement it. You need to know that a library such as datasketch exists for exactly this, and that near-duplicate removal without it does not scale.

The payoff is measured. Lee and colleagues found that more than 1 percent of the tokens a model produces unprompted are copied verbatim from its training set when that set was not deduplicated. Deduplicating cut the rate of memorised output by roughly a factor of ten.

Pass two: decontamination

Decontamination means checking that nothing in your training data also appears in the data you will measure on, and removing whatever does. Skip it and your reported number stops measuring skill and starts measuring memory.

Work the arithmetic once and you will never skip it again. Suppose you have 200 test questions. Your model genuinely answers 60 percent of unseen questions correctly. Now suppose some test questions accidentally got copied into the training set, so the model has simply memorised those answers.

  • Nothing leaked. The score is 60 percent of 200, so 120 correct. Reported: 60 percent. Honest.
  • 20 questions leaked, which is 10 percent. Those 20 are all correct. Of the remaining 180, 60 percent is 108. Total 128. Reported: 64 percent.
  • 50 questions leaked, which is 25 percent. Those 50 are correct, plus 60 percent of the remaining 150, which is 90. Total 140. Reported: 70 percent.

The model never got better. In the last row you would report 70 percent for a model that is worth 60. That is why one leaked test item invalidates a benchmark: a benchmark is only a claim about unseen questions, and a leaked question is not unseen. A benchmark, here, means a fixed test set with a published score, the kind of thing people put on slides.

Decontamination in practice means scanning training text for overlap with your held-out set, and with any public benchmark you intend to quote. The usual check is n-gram overlap, which means looking for runs of the same consecutive words, typically eight or thirteen of them, appearing in both. Exact string matching is not enough, for the same reason exact deduplication is not enough. Reworded copies leak just as effectively.

This is not a rare hazard. The same deduplication study found train-test overlap affecting more than 4 percent of the validation set in standard benchmarks. Public benchmark questions also leak into pretraining data, which is why a public leaderboard score is a coarse filter and never a decision.

Pass three: quality scoring and filtering

You are not going to hand-read ten thousand examples. Score them instead, then keep a slice off the top. There is a ladder of rigour, and cheap comes first.

  1. Heuristics. Length bounds, a language check, regular expressions for boilerplate such as “As an assistant” or leftover markup. Costs nothing and removes a surprising amount.
  2. A scoring model. A reward model is a small model trained on human comparisons to give any answer a single number. Run it over the set and threshold.
  3. A strong model as judge. Ask a capable model to rate each example against a written rubric. This is the technique AlpaGasus used to get from 52,000 to 9,000.

Judging with a model works better than it has any right to. Zheng and colleagues found a strong judge agrees with human raters more than 80 percent of the time, which is about as often as two humans agree with each other. It also has known habits: it favours the answer shown first, it favours longer answers, and it favours its own writing style. Randomise the order, and be suspicious when the winners are all the long ones.

Finally, protect diversity on the way out. Cluster the survivors and cap how many near-copies of any one cluster you keep. Otherwise a top-scoring pattern quietly takes over the set, and you have optimised your way into mode collapse.

raw dump 100% exact dedup hash near-dup MinHash decontaminate vs eval/benchmarks score + filter judge, keep top trainable clean

Read this left to right and note two orderings. Deduplication comes before scoring, and decontamination comes before the split.

Both orderings in that figure are load-bearing. Deduplicate first so you do not spend judge calls scoring fifty copies of the same example. Decontaminate before you split, because once the data is split, a leak has already crossed the wall and no later step will find it.

One more leak deserves its own name, because a duplicate check cannot see it. Label leakage is when information from the expected answer sneaks into the input field of an example. Imagine a support dataset where the ticket text quietly includes the resolution code, because the export ran after the ticket was closed. The model reads the answer off the prompt. Your metrics soar. In production the field is empty and the model is useless. Part 15 covers how to hunt for it.

Every one of these passes is also an auditable control. Being able to say you deduplicated, decontaminated against the evaluation set, and filtered with documented thresholds is exactly the evidence a model-risk review will ask for. Build it as a logged, repeatable script rather than a notebook you ran once.

Split the data three ways and keep the wall between the parts

You need three sets, not two. The reason is about how each one gets used up.

  • Train. The model learns on it. Nothing measured here means anything, because the model has seen every answer.
  • Validation. You check this during training to pick settings and to choose which saved copy of the model to keep. Keeping the copy from the point where validation loss was lowest is called early stopping.
  • Test. You look at this once, at the very end, for one honest number.

Why validation cannot also be test is worth building slowly, because it is the part people talk themselves out of.

Say you try twelve learning rates and keep whichever scores best on validation. You have now made twelve decisions using that set. The winner is partly genuinely better and partly just lucky on those particular examples. Repeat that over a project and validation slowly turns into a set you have fitted to, one decision at a time. It gets used up. That is not a moral failing, it is arithmetic, and it is why the test set exists and stays sealed.

Tune against the test set even once and it becomes a second validation set. There is no ceremony that restores it.

Split by category, not at random

A random split can betray you, and here is the shape of it.

The word you need first is distribution, which just means the tally of how much of each kind of thing your data contains. The no_robots dataset is a good example. Its 10,000 examples carry a category label, and the tally is lopsided. Generation covers about 4,560 of them. Extract covers about 190.

Now take a random 5 percent test split, so 500 examples. On average about 9 or 10 of those should be Extract examples. On average. A random draw can easily hand you two, or none. If it hands you none, your test score says precisely nothing about extraction, and you will not notice, because the number will look fine.

A stratified split fixes this. Split within each category separately, then combine. Take 5 percent of Generation, 5 percent of Extract, 5 percent of every other category, and glue the pieces together. Every category then appears in every split in the same proportion it appears overall. Do the same by difficulty if you have difficulty labels, and by source if your data came from several places.

Keep one probe that looks nothing like your training data

Your validation and test sets are drawn from the same pile as your training set. So they can only tell you how the model does on data that looks like what you trained on. They are blind to everything else.

Everything else includes the thing most likely to hurt you. Fine-tuning hard on one task can quietly erode abilities the model already had, such as arithmetic or following a formatting instruction. That is called catastrophic forgetting. Your in-distribution test set will report success the whole time it is happening.

The fix is a small out-of-distribution probe: a handful of tasks deliberately unlike your training data, scored before and after the fine-tune. A few maths questions, a few general-knowledge questions, a couple of instruction-following checks. It costs almost nothing and it is the only instrument that sees this failure. Part 15 turns it into a full procedure.

TRAIN model learns · ~90% VALIDATION tune · look often TEST look ONCE decontamination walls — no example (or near-dup) crosses CAPABILITY PROBE (out-of-distribution) a few MMLU / GSM8K / coding items — NOT no_robots-like — to catch catastrophic forgetting (1.5)

The vertical bar is the decontamination wall, and it belongs between your own splits, not only between you and a public benchmark. The separate box at the bottom is the probe, and it deliberately sits outside the whole pile.

The rule to carry away: any number you show someone is worth exactly as much as the wall between the set it was measured on and the set the model trained on. Write the split method down, version it, and never move the wall to make a number look better.

Generate synthetic data without collapsing the model

Human-written examples are slow and expensive. Synthetic data means training examples written by a model instead of a person. It is how many teams reach scale, and it is a genuine capability with a genuine failure mode attached.

Three reasons to generate at all. Scale, because a hundred thousand examples costs API calls rather than annotator hours. Targeted coverage, because if you need 500 examples of one rare edge case you can simply ask for exactly those. And distillation, where a strong model writes high-quality answers and a small model trains to copy them.

Method The idea Landmark work
Self-Instruct seed with a few human-written tasks, then let the model write thousands more instructions and answers from those Wang et al., 2022
Evol-Instruct take an existing prompt and rewrite it to be deeper, harder or more constrained, so you manufacture difficulty on demand Xu et al., 2023
Distillation a stronger teacher model writes the answers and a smaller student trains to match them the prune-and-distil lineage behind Llama-3.2-1B, Part 8

The failure that defines the topic: model collapse

Build this one up from marbles, because the mechanism is simpler than the name suggests.

Put 1,000 marbles in a bag. 990 are red and 10 are blue. Draw 100 at random. On average you expect 1 blue marble. Sometimes you get 2. Often you get 0.

Now throw the original bag away and refill it from what you drew, scaled back up to 1,000. If your 100 draws contained no blue, there is no blue in the new bag. Blue is gone. Not rare, gone. Repeat the whole procedure and the next-rarest colour goes the same way.

A model generating training data behaves like that draw. It produces common patterns often and rare ones rarely. The rare ones live in what statisticians call the tails of the distribution: the edge cases that are individually unlikely and collectively important. Train the next model on that output and the tails thin. Do it again and they vanish.

Shumailov and colleagues published exactly this experiment in Nature in 2024, and named the effect model collapse. They fine-tuned a small language model, OPT-125m, on the wikitext2 dataset. Then they used it to generate a dataset the same size, trained the next generation on that, and repeated for nine generations, five separate times.

Their headline measure was perplexity, which is one score for how surprised a model is by a piece of real text. Lower is better. Before fine-tuning, the model scored about 115 on the real test text. After fine-tuning on the real wikitext2 data it scored about 34, which is the model doing its job well. When each generation trained only on the previous generation’s output, later generations gave back 20 to 28 points of that gain.

The examples are more vivid than the numbers. Given a prompt about fourteenth-century English church architecture, the first model wrote a sensible paragraph about Perpendicular Revival architecture and St. John’s Cathedral in London. By generation nine, the same prompt produced text about the world’s largest populations of black-tailed jackrabbits, white-tailed jackrabbits, blue-tailed jackrabbits, red-tailed jackrabbits and yellow-tailed jackrabbits. The model has not become random. It has fallen into a groove and is repeating variations of one pattern.

gen 0 · real data gen n · tails thinning gen N · collapsed each round of “train on model output” narrows the distribution — rare cases disappear first

Watch the edges of the curve, not the middle. Each generation the low bumps at the sides shrink first, and the remaining mass piles up into one narrow spike.

The paper separates two stages. In early collapse the model starts losing the tails, so rare cases stop appearing. In late collapse the model converges on something with little resemblance to the original data, with much less variety left in it. Early collapse is the dangerous one, because a model in early collapse still looks fine on ordinary examples.

The primary cause they name is plain sampling. Any finite sample of a distribution loses some of the rare events, and information lost at one step can never come back at the next. That is the marble bag. Two secondary causes sit on top: the model cannot represent the original distribution perfectly, and the training procedure itself has its own biases.

One detail worth knowing, because it kills the obvious fix. The collapsed models repeat phrases constantly, so the natural response is to force variety with a repetition penalty. The researchers tried it. The models produced worse continuations, and perplexity doubled compared with the original. Suppressing the symptom made the underlying problem worse.

The defence is mixing, not abstinence

The same paper points at the fix. In a second setting they kept 10 percent of the original real data in every generation’s training set. The degradation dropped to minor. Real data anchors the tails that generated data drops.

So the working recipe for synthetic data is a pipeline, not a prompt.

  1. Keep real data in the mix. Never train a generation purely on the previous generation’s output.
  2. Deduplicate the generated set first. Models repeat themselves far more than people do, so near-duplicate rates in generated data are high.
  3. Verify anything checkable. For maths, check the answer. For code, run the tests. A teacher model’s mistakes get copied faithfully and then amplified.
  4. Then score and filter, and cap per cluster so one dominant pattern cannot swamp the set.

The useful way to hold it: generation sets the ceiling on possible quality, and curation decides how much of that ceiling you actually reach. Producing a million examples is a weekend of API calls. Making them usable is the job.

The licence follows the weights

One compliance point that is not a footnote. Data generated by a commercial model is governed by that model’s terms of service, and several of those explicitly forbid using the outputs to train a competing model. So “we distilled our instruction set from a frontier model” can carry legal exposure that travels with the resulting weights, long after anyone remembers where the data came from.

For a defensible pipeline, use permissively licensed teacher models or your own. Record the provenance of every synthetic batch: which model, which version, which prompt, which date. Keep that record with the dataset. It is precisely what a model-risk review asks for, and it is why a fully documented training-data story has real commercial value.

Poison a copy of your data on purpose and see what each fault breaks

Everything above is easier to believe once you have watched it happen. The protocol is small. Take a clean baseline dataset. Make several copies, and damage each copy in exactly one way, so only one variable moves at a time. Train on each. Compare all of them against the same clean held-out set and the same handful of generation checks.

The damage column below is reasoned from the mechanism rather than measured from one particular run. Treat it as what to expect, and confirm it against your own numbers.

Corruption What you inject Expected damage
Duplication copy about 30 percent of the examples several times over memorisation and verbatim regurgitation, and one style over-weighted
Format corruption break the role markers, or drop the end-of-turn token on some examples the model rambles, will not stop, and its structure drifts
Length bias truncate every answer to a single sentence the model becomes uselessly terse no matter what the task needs
Refusal injection replace about 10 percent of answers with a refusal the model refuses harmless requests, a self-inflicted alignment tax
Contamination copy test examples into the training set the evaluation looks excellent and is fiction

The refusal row introduces one term. An alignment tax is capability lost in exchange for making a model better behaved. Usually it arrives from a deliberate safety decision. This row shows you can install it entirely by accident, through data hygiene alone.

clean-eval quality after training (higher = better): baseline (clean) ✓ good + duplication degraded + format corruption broken output + length bias terse, unhelpful + refusal injection refuses benign + contamination real: bad · reported: FAKE-GOOD

Every line on the left falls, including the ruined runs. The right-hand panel is where the runs separate, and the bar that stays high there is the contaminated one, which is the whole point.

Now the reason this protocol teaches more than any amount of reading. Every training curve in that experiment goes down. All of them. The training loss measures how well the model predicts the text it was given, and it has no opinion about whether that text was any good. Truncate every answer to one sentence and the model learns to produce one sentence, and gets steadily better at it, and the curve looks textbook.

Four of the five corruptions at least fail honestly: run a clean held-out evaluation and the damage shows up as a worse score. Contamination is the exception. It fails dishonestly, because the model memorised the leaked answers, so the reported number looks as good as baseline while the model is worse. That is why decontamination is the control you never skip, and why it has to run before the split.

Audit any dataset in one minute before you train on it

You do not need a purpose-built tool to catch most of this. Three questions cover the bulk of it. How many examples are actually here? How many are exact copies? And did anything in train leak into the evaluation split?

from collections import Counter
from datasets import load_dataset

ds = load_dataset("HuggingFaceH4/no_robots")
train, test = ds["train"], ds["test"]

def norm(example):
    joined = " ".join(m["content"] for m in example["messages"])
    return " ".join(joined.split()).lower()

train_texts = [norm(r) for r in train]
test_texts = [norm(r) for r in test]

print("train examples:  ", len(train_texts))
print("test examples:   ", len(test_texts))
print("exact duplicates:", len(train_texts) - len(set(train_texts)))
print("train/test leak: ", len(set(train_texts) & set(test_texts)))
print("categories:      ", Counter(train["category"]).most_common())

Line by line, assuming you have never used any of these libraries.

  • from collections import Counter pulls in a standard Python helper that tallies how many times each value appears in a list.
  • from datasets import load_dataset pulls in the Hugging Face datasets library, which downloads a named dataset and caches it locally.
  • load_dataset("HuggingFaceH4/no_robots") fetches the dataset. It returns a dictionary-like object with one entry per split.
  • ds["train"], ds["test"] takes the two splits out. Here that is 9,500 and 500 examples.
  • def norm(example) defines a function that flattens one example into a single plain string, so that two examples can be compared.
  • example["messages"] is the list of turns. Each turn is a small dictionary with a role and a content. The join glues every turn’s text together with spaces.
  • " ".join(joined.split()) collapses every run of spaces, tabs and newlines into one space. .lower() then lowercases. Both exist so that two examples differing only in spacing or capitals are treated as the same example.
  • The two list comprehensions run norm over every row and collect the results, giving one string per example.
  • set(train_texts) builds a Python set, which silently discards repeats. So the list length minus the set length is the number of extra copies.
  • set(train_texts) & set(test_texts) is set intersection: the strings that appear in both. Anything above zero is a leak, and it is the one line here you must never ignore.
  • Counter(train["category"]).most_common() reads the whole category column as a list, tallies it, and sorts it by count. This is the lopsided distribution from the splitting section, printed out.

Extend the same shape in three directions when you need more. Add a near-duplicate rate by running min-hash over the normalised text and thresholding on Jaccard similarity, using a library rather than writing it. Add a format check by confirming every example survives the tokenizer’s chat template without an error. Add a length distribution by counting tokens per example with the exact tokenizer you are about to train with, because that is the only count that matches what training will see.

Read the numbers your own run prints rather than trusting any figure quoted here. They depend on the dataset, on how you chose to normalise the text, and on your similarity threshold.

Run it before every training job, not once. It costs a minute, and it catches the class of problem that otherwise appears as an unexplained quality regression three weeks later. This discipline pays off whatever you train next, including the low-rank adapters in the next part on LoRA.

Key takeaways

  • The file format picks your masking, and masking decides what the model learns. Raw text grades the instruction too, which teaches the model to write instructions. Use prompt-and-completion or conversational data for instruction tuning.
  • Packing removes most of the padding waste on short examples. Done naively it lets one example read the one before it, so it needs boundary resets, and it is not a first-run feature.
  • Quality beats quantity because every example gets an equal vote. Adding data helps only while the new data is at least as clean as what you already have.
  • Deduplicate before scoring, and decontaminate before splitting. A duplicate multiplies an example’s pull on the weights. A leak turns your reported score into fiction.
  • Use three sets. Validation is used up by the decisions you make from it, test is opened once, and a separate out-of-distribution probe is the only thing that sees forgetting.
  • Synthetic data scales cheaply and collapses if you train on it recursively, because each generation drops the rare cases. Mix real data back in, deduplicate, verify what is checkable, and record the licence.
  • Corrupted data still produces a falling training curve. Only a clean held-out evaluation shows the damage, and contamination makes even that lie.

You can now

  • Choose between raw text, prompt-and-completion and conversational data by asking which tokens you want graded, from “Pick the file format that grades the tokens you care about”.
  • Decide whether to switch packing on, and name the two boundary resets correct packing has to apply, from “Pack short examples without letting them leak into each other”.
  • Compute a Jaccard similarity by hand and set a near-duplicate threshold from it, from “Run the three cleaning passes, and run them in the right order”.
  • Work out how much a given leak rate inflates a reported benchmark score, from “Run the three cleaning passes, and run them in the right order”.
  • Build a stratified three-way split plus an out-of-distribution probe, and say what each part is for, from “Split the data three ways and keep the wall between the parts”.
  • Design a synthetic-data pipeline that resists model collapse, and justify each step, from “Generate synthetic data without collapsing the model”.
  • Audit an unfamiliar dataset for size, duplicates, leaks and category balance in about a minute, from “Audit any dataset in one minute before you train on it”.

Glossary

Adapter
A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
Alignment tax
The capability a model loses in exchange for being made more helpful or better behaved. Running supervised fine-tuning and then preference tuning is enough on its own to cause it. Lin et al., Mitigating the Alignment Tax of RLHF
Attention
The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
Backward pass
Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
Base model
The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
Batch
A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
Benchmark
A fixed public test set with a published score, such as a maths or general-knowledge exam for models. Useful as a coarse filter and untrustworthy as a decision, because public questions leak into pretraining data. Xu et al., Benchmark Data Contamination survey
Catastrophic forgetting
When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
Chat template
The rule that turns a list of role-and-content messages into the exact text and special tokens the model expects to see. It ships with the tokenizer, and it is where the boundary between prompt and response lives. From messages to tensors
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
Data contamination
Overlap between your training data and your evaluation data, including reworded near-copies. It is the one data problem that fails dishonestly: the score looks excellent because the model memorised the answers. Xu et al., Benchmark Data Contamination survey
Decontamination
Checking that nothing in the training data also appears in the evaluation data, and removing whatever does. Skip it and your reported score measures memorisation rather than skill.
Deduplication
Removing exact and near-duplicate examples from a dataset. Every copy of an example multiplies its pull on the weights, so duplicates cause memorisation and over-weight one style. Lee et al., Deduplicating Training Data Makes Language Models Better
Early stopping
Keeping the checkpoint from the point where held-out loss was lowest, instead of the one from the last step. Past that point the model is memorising rather than learning. Proving the shift
Embedding
The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
End-of-turn token
The special token that marks where an answer stops. If it is left out of the graded region during training, or the wrong one is passed at generation time, the model rambles past the end of its answer. Masking in LLM training
Flash attention
An attention implementation that computes the answer in small tiles inside fast on-chip memory, never building the full score grid. Same result, far less memory, and the term that grew with the square of sequence length becomes linear. Dao et al., FlashAttention
Forward pass
Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
Holdout
Data deliberately kept out of training so you can measure the model on something it has never seen. A number measured on data the model trained on is not evidence of anything. Building a holdout you can trust
Hyperparameter
A setting you choose before training rather than something the model learns, such as the learning rate, the batch size or the number of epochs. The knobs that decide whether it learns
Inference
Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
Instruction tuning
Supervised fine-tuning on instruction-and-answer pairs, so a base model learns to follow instructions and answer in a consistent shape. Zhou et al., LIMA
Knowledge distillation
Training a small model to copy a larger model’s outputs rather than learning from raw data alone. Llama-3.2-1B was built this way, after a larger model had been pruned down. Llama-3.2-1B model card
Label
The correct answer a training example is graded against. For a language model the labels are the token IDs the model was supposed to produce, and a label of -100 means do not grade this position at all. Masking in LLM training
Length bias
The tendency of both human and model annotators to rate longer answers as better, which quietly teaches a preference-tuned model to pad. If a tuned model starts waffling, suspect the data before the algorithm. Meng et al., SimPO
LLM as a judge
Using a strong model to score or compare answers against a written rubric. It agrees with human raters over 80 percent of the time, about as often as humans agree with each other, and it favours the first answer shown, longer answers, and its own style. Zheng et al., Judging LLM-as-a-Judge
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Loss masking
Deciding which positions in an example count toward the loss. For instruction tuning you grade only the assistant’s response and mark everything else with -100, so it contributes no loss and no gradient. Masking in LLM training
Mode collapse
When tuning crushes the variety out of a model’s answers, so it produces the same shapes and phrasings over and over. You catch it by reading many outputs, never a single sample. The six silent failure modes
Model collapse
What happens when models are trained on their own generated output over and over: the rare cases at the edges disappear first, then diversity, then quality. The defence is mixing real data back in, not abstinence. Shumailov et al., AI models collapse on recursively generated data
Out-of-distribution probe
A small evaluation on tasks that look nothing like your training data, kept specifically to catch abilities you have lost. In-distribution evaluation cannot see catastrophic forgetting at all. Evaluating a fine-tune
Overfitting
When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse. The six silent failure modes
Padding
Filler tokens added to short sequences so every sequence in a batch has the same length. Padding is bookkeeping for the hardware, and it has to be masked out of both attention and the loss so it cannot change the answer. Masking in LLM training
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
Pretraining
The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
Reward model
A model trained on human comparisons to give any answer a single score. RLHF optimises against it because you cannot put a human in the loop for millions of steps. Ouyang et al., InstructGPT
Sequence packing
Concatenating several short examples into one full-length sequence so you stop wasting compute on padding. It needs boundary resets, or one example can attend back into another and the model learns nonsense continuations. Hugging Face TRL, SFT Trainer
Stratified split
Splitting a dataset so every category appears in the same proportion in each split, instead of splitting at random. Without it a whole category can land entirely in the training set and your test score says nothing about it.
Structural pruning
Permanently removing whole pieces of a trained model, such as layers or attention heads, to make it smaller. It is normally followed by more training to recover the quality that was lost. Llama-3.2-1B model card
Supervised fine-tuning
Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
Synthetic data
Training examples written by a model rather than by a person. Cheap and endlessly scalable, and it needs curation, a mix of real data alongside it, and a check on the licence of whatever model produced it. Wang et al., Self-Instruct
Test set
The split you look at once, at the very end, for one honest number. Tune against it and it quietly becomes a second validation set and stops being trustworthy.
Token
The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
Tokenizer
The component that splits text into tokens and maps them to integer IDs, and back again. Every model has its own, and it has to match the model you are training. Choosing a base model and building a bench
Validation set
The split you check during training to pick settings and choose the best checkpoint. Because you make decisions from it over and over, it slowly stops being an honest estimate.
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model


Practical exercises

Hand-compute the forensics script’s output on a toy dataset

Using the shape of this part’s forensics script, not the script itself, work out what it would print for a train set of the four normalised strings “a b c”, “a b c”, “d e f”, “g h i”, and a test set of the two strings “d e f”, “j k l”. State the train example count, the test example count, the exact duplicate count in train, and the exact train and eval overlap count.

See the worked solution (opens in a new tab)

Reconcile AlpaGasus’s data reduction against its speedup

AlpaGasus filtered 52,000 examples down to about 9,000, and this part states that cut a 7B fine tune from about 80 minutes to about 14. Compute the dataset size reduction factor and the training time reduction factor separately, compare the two, and explain in one sentence why they land so close to each other.

See the worked solution (opens in a new tab)

Diagnose an eval score that looks too good to be true

After fine tuning, your held-out eval loss looks as good as baseline. Spot checking generations on prompts that closely resemble your test set, several answers are near word-for-word identical to the test set’s own gold answers, phrasing included. Using this part’s corruption table, name the corruption this points to, explain why it is the one row that fails dishonestly rather than showing up as a worse score, and name the pipeline step, and its correct order relative to the split, that should have caught it.

See the worked solution (opens in a new tab)

Quantify the risk an unstratified split poses to a minority category

Assume a rare subcategory makes up about 0.5 percent of a 10,000-example dataset, roughly 50 examples total, and you take a pure random 5 percent test split, 500 examples, with no stratification. Compute the expected number of that subcategory’s examples landing in the test set, then use the Poisson approximation to estimate the probability that the test set ends up with zero examples of it purely by chance. State what that probability tells you about relying on an unstratified split for a minority category this size.

See the worked solution (opens in a new tab)

Design a length-bias corruption experiment and find what actually reveals it

Design the length-bias row from this part’s corruption table as an actual experiment on a copy of no_robots. Describe exactly what you would inject, then reason through what happens to the training loss curve, the held-out eval loss computed the ordinary way, and a before-and-after generation check on a prompt that calls for a genuinely long, multi-step answer. Explain why the first two would not tell you this run went wrong, and what would.

See the worked solution (opens in a new tab)

Frequently asked questions

How much data do I need to fine-tune an LLM?

Far less than most people expect, if it is clean. LIMA reached strong instruction following on exactly 1,000 curated examples, and AlpaGasus filtered a noisy 52,000-example set down to about 9,000 and produced a better model that trained about 5.7 times faster. Start with hundreds to a few thousand high-quality examples rather than tens of thousands of scraped ones.

Is more data always better for fine-tuning?

No. The loss is averaged over examples, so every example gets an equal vote on what the model becomes. Adding data helps only while the new data is at least as clean as what you already have, and past that point each noisy example teaches the model your noise.

What is data contamination and why does it matter so much?

Contamination is any overlap between your training data and the data you evaluate on, including reworded near-copies rather than exact ones. It matters because it is the one data problem that fails dishonestly: the model memorises the leaked answers, so the reported number looks excellent while the model is worse. Leak 10 percent of a test set into training and a model worth 60 percent reports 64.

Can I train on synthetic data generated by another model?

Yes, and many teams blend it into their instruction sets, but two constraints apply. Training each generation only on the last one’s output causes model collapse as the rare cases disappear, so keep real data in the mix and curate hard. And commercial model outputs are governed by that provider’s terms, several of which forbid training a competing model.

Do I need train, validation and test, or is two enough?

Three, if you will ever report a number to anyone. Validation is used up by the repeated decisions you make from it, such as picking settings and choosing a checkpoint, so it stops being an honest estimate. Test is opened once at the end, and a fourth out-of-distribution probe catches forgetting that in-distribution sets cannot see.

Should I use packing for my fine-tuning dataset?

Only once your masks are verified and padding waste is measurably costing you time. Packing concatenates short examples to fill the sequence, which removes most of the padding waste, but it needs boundary resets so that attention and position counting restart at each example. Without them the model learns continuations that run across two unrelated examples.

Sources and further reading

Previous