Evaluating a fine-tuned LLM is the step that decides whether the training run was worth doing. It is also the step most teams bolt on at the end, once the model already exists. That order is backwards. This part puts it right, and it assumes you know no statistics at all.
The trap is a falling loss curve. Loss is one number saying how wrong the model was on the text it was just shown, and training is the business of making that number smaller. A falling loss tells you the model fits your data better. It tells you nothing about whether the model got good at your task, and nothing about whether it still remembers how to add up an invoice.
Here is the shape the failure takes. You fine-tune a model on support tickets. It scores well on your ticket test. You ship. Two weeks later the model writes beautiful refund replies and can no longer do arithmetic. Nothing in training warned you, because nothing in training was watching.
By the end you will be able to build an evaluation set the model has never seen, turn a bare score into a difference you can defend with numbers, climb a ladder of evaluation methods from cheap to expensive, and name and detect six failures a training curve cannot show.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures (you are here)
Name the four questions a falling loss cannot answer
Picture a student revising for an exam. They work through a hundred practice questions, over and over, with the answers on the back of each card. By Friday they get all hundred right.
What have you learned about that student? Almost nothing. You do not know whether they can answer a question they have not seen. You do not know whether they still remember last year’s material. You do not know whether they are any use in a real job.
Training loss is that Friday score. It is measured on the very examples the model was trained on. Four separate questions matter, and loss answers only the first of them, weakly.
- Did it learn the task at all? Did the behaviour you wanted actually go in?
- Does it work on inputs it has never seen? Or did it only memorise your examples?
- Did it break anything? Can it still do the things it could do before you touched it?
- Is the output any good? Would a person reading it be satisfied?
Every one of the last three needs its own test. Each test has to run outside the training loop, on data the training never touched. That is the whole design principle of this part, and everything below is an application of it.
Test the model on data it has never seen
Go back to the flashcards. Your student scored a hundred percent on the hundred cards they studied. There are two completely different reasons that could happen.
- They understood the underlying pattern, so any similar question would also work.
- They memorised those hundred cards, front and back, and know nothing else.
Those two are the entire subject of this section. The second one is memorisation: storing the specific examples. The first is generalisation: performing well on inputs that were never in the training data. Generalisation is the thing you actually want. It is also the thing training loss structurally cannot measure, because training loss is only ever measured on the cards.
So you need cards the student has not seen. In machine learning that slice of data has a name: a holdout. A holdout is data you deliberately set aside before training and never show to the training loop. A score measured on data the model trained on is not evidence of anything. A score measured on a holdout is.
Three piles, not two
In practice you cut the data three ways, not two, and the third cut surprises people.
- Training set. The model sees these and learns from them. This is the only pile that changes any weights.
- Validation set. The model never learns from these, but you look at them during the run. This is where Part 9’s eval-loss curve comes from, and it is what you use to pick a learning rate or choose which saved checkpoint to keep.
- Test set. You look at this once, at the very end, for one honest number.
Why three? Because looking at a set is itself a way of fitting to it. Suppose you try eight learning rates and keep whichever scored best on the validation set. You have now hand-fitted that set. It flattered you a little, and it will keep flattering you. The test set is the one pile you have not spent. Spend it once.
Make the split representative
A holdout is only useful if it looks like real traffic. Say you have 1,000 support tickets: 900 refund requests and 100 billing disputes. Take a random 10 percent as a test set and you might land 6 billing disputes, or you might land 16. Your entire billing score then rests on 6 examples, and 6 examples cannot tell you anything.
The fix is a stratified split: take 10 percent of each category separately. That gives 90 refunds and 10 billing disputes, every time, in the same proportions as the real pile. Part 10 covers the split mechanics in full.
Three more habits make a holdout worth having. Include the hard cases and the adversarial ones, meaning inputs deliberately written to trip the model, rather than only the easy examples that were pleasant to collect. Fix the set and version it, so a score from March is comparable to a score from June. And assemble it exactly as production does: same prompt template, same system prompt, same surrounding context. An evaluation that uses a different prompt format is measuring a model you are not going to ship.
Keep the evaluation set clean, and spot the two ways it leaks
Everything above collapses if the model has already seen the test items. Your student glanced at ten of the hundred exam questions last week. Their score goes up. Their knowledge does not.
That is contamination: overlap between the training data and the evaluation data, including reworded near-copies. It is the one data fault that fails dishonestly. A bad batch of training text makes your score fall, so you go and look. Contamination makes your score rise, so you celebrate and ship. Part 10 treats it as a data-pipeline control, which is where the fix belongs.
Public benchmarks have this problem permanently. If a test’s questions have ever been on the open internet, assume the pretraining crawl swallowed them years ago. The model may be reciting rather than reasoning, and the score cannot tell the two apart.
Detecting contamination after the fact is unreliable, and it gets worse once the training text differs in wording from the original question. Xu and colleagues survey the methods and the gaps. So do not lean on detection. Lean on a private set instead, built from your own data, never published, versioned, and treated as the only score that decides anything. Public benchmarks are a coarse first filter and never the deciding vote.
The quieter leak, inside a single example
There is a second kind of leak, and it hides inside one training example rather than across two files.
Say your ticket records carry a field called resolution_code. An agent fills it in when they close the ticket, so it exists in your archive and it always agrees with the correct answer. You build training examples straight from the archive and leave that field in the prompt. The model does not learn to read the ticket. It learns to read the field.
Your metrics soar. Then you deploy, and a real request arrives with that field empty, because nobody has closed the ticket yet. The model has no idea what to do. That is label leakage: information from the expected answer bleeding into the input side of an example.
The check is a single question, asked of every field in your prompt. Would this value exist at the moment a real request arrives? If not, delete it before you build a single example.
Turn a bare score into a delta you can defend
Your fine-tune scores 68 percent on a clean, private holdout. Is that good?
You cannot answer that. Nobody can. The number needs something to be compared against, and the right thing is the model you started from. So score the base model too, on the identical evaluation: same examples, same prompt template, same judge, same everything. Now you have two numbers, and the difference between them is the delta. Without a delta you cannot tell an improved model from a merely changed one.
Why a 2-point gain on 100 examples is probably nothing
Say your holdout has 100 examples. The base model gets 61 right. The fine-tune gets 63. A 2-point gain. Ship it?
No. Start with a coin. Flip a fair coin 100 times and you expect 50 heads. You will rarely get exactly 50. You get 47, or 54, and nobody claims the coin has changed. That wobble is what happens whenever you measure a small sample of something.
Your 100 examples are a sample too. You did not test on every ticket your users will ever send. You tested on 100 of them, drawn from a much larger pool. The honest question is this: if you had drawn a different 100, would the gap still be there?
Now look at the 2 points example by example, rather than as two totals. A realistic breakdown:
- 82 examples where both models scored the same, right or wrong together.
- 10 examples where the fine-tune was right and the base was wrong.
- 8 examples where the base was right and the fine-tune was wrong.
Ten minus eight is two, out of a hundred, which is the 2-point gain. But look at what it rests on. The 82 agreements carry no information at all. The entire result comes from 18 examples that disagreed, and 10 of those went your way.
Ten heads out of eighteen tosses. Nobody calls a coin bent on that.
The bootstrap, in five steps
That coin argument is the intuition. The bootstrap is the mechanical version of it, and it works on any score, not just right and wrong.
You would love to fetch a fresh 100 tickets and measure again, ten thousand times over. You cannot. So you fake it, using only the data you already have.
- Write down one number per example: the fine-tune’s score minus the base model’s score on that same example. Here that gives 10 numbers of plus 1, 8 numbers of minus 1, and 82 zeros.
- Draw 100 numbers at random from that pile of 100, putting each one back before the next draw. Duplicates are the whole point. Some examples get picked three times and some not at all, which is exactly what a different sample of reality would have looked like.
- Average that fake sample. Write the average down.
- Repeat steps 2 and 3 ten thousand times. You now have 10,000 plausible answers to the question “what could this delta have been”.
- Sort those 10,000 averages. Discard the lowest 250 and the highest 250. What remains is the middle 95 percent, and that range is your 95 percent confidence interval.
Run it on the numbers above and the interval comes out at roughly minus 6 points to plus 10 points. The exact ends wobble slightly each time, because the resampling is random. The important part does not wobble: the range contains zero. There are plenty of plausible redraws in which the fine-tune is worse than the base model. Your 2-point gain is noise wearing a suit.
So the rule is: decide on the lower end of the interval, never on the average. If the lower end is below zero, you have not shown anything yet.
What actually fixes it
More examples. Keep the same rates and run the same evaluation on 2,000 examples instead of 100. Now 200 examples improve, 160 get worse, and the rest agree. The delta is still 2 points. The interval is now roughly plus 0.1 to plus 3.9 points.
It clears zero, and only just. That is the honest answer to “how big should my evaluation set be”. Not a round number somebody told you. Big enough that the difference you care about survives this test.
import numpy as np
def bootstrap_ci(base_scores, tuned_scores, n_boot=10000, alpha=0.05):
diffs = np.asarray(tuned_scores) - np.asarray(base_scores) # paired, per example
n = len(diffs)
boot_means = np.array([
np.random.choice(diffs, size=n, replace=True).mean()
for _ in range(n_boot)
])
lo, hi = np.percentile(boot_means, [100 * alpha / 2, 100 * (1 - alpha / 2)])
return diffs.mean(), lo, hi
mean_delta, lo, hi = bootstrap_ci(base_scores, tuned_scores)
print(f"delta={mean_delta:.3f} 95% CI=[{lo:.3f}, {hi:.3f}]")
# ship only if lo > 0
Line by line, assuming you have never used numpy.
import numpy as npbrings in numpy, the standard Python library for working on whole lists of numbers at once. Theas nppart is just a short nickname.base_scoresandtuned_scoresare two ordinary lists of per-example scores. They must be in the same order and cover the same examples, because the comparison is paired.n_boot=10000is how many fake samples to draw.alpha=0.05means you want a 95 percent interval, since 1 minus 0.05 is 0.95.np.asarray(...) - np.asarray(...)turns both lists into numpy arrays and subtracts them element by element. The result is step 1: one difference per example.n = len(diffs)is how many examples there are, so each fake sample is the same size as the real one.np.random.choice(diffs, size=n, replace=True)is step 2. It drawsnvalues fromdiffs, andreplace=Trueis what puts each value back before the next draw..mean()is step 3, the average of that one fake sample.- The square brackets with
for _ in range(n_boot)around it are ordinary Python. They repeat the draw 10,000 times and collect the results. The underscore just means the loop counter is not used. np.percentile(boot_means, [2.5, 97.5])is step 5, written as arithmetic. It returns the value that 2.5 percent of the averages fall below, and the value that 97.5 percent fall below. Those two are the ends of the interval.- The
f"..."line is a Python f-string.{lo:.3f}prints that number to three decimal places.
That gives you three rules that make a holdout earn its keep, and none of them is optional.
Climb the evaluation ladder and stop at the right rung
Evaluation is not one thing. It is a ladder, running from cheap and shallow up to expensive and true. Each rung measures something the rung below it cannot see.
Rung one, perplexity. Run the model over held-out text, take the average loss, and undo the logarithm. What comes out is roughly how many equally likely options the model felt it was choosing between at each token. A perplexity of 10 means it was about as unsure as someone picking between 10 options. Lower is better. It costs almost nothing, so you can run it at every checkpoint, and it catches gross damage quickly. It says nothing whatsoever about whether the model is good at your task.
Rung two, task metrics. Concrete, countable things. Did the extracted field match the expected value? Did the output parse as valid JSON? Was the classification correct? These are cheap and objective. One family of them deserves caution: overlap scores, which count how many words an answer shares with a reference answer. An answer sharing few words can still be perfect, and one sharing many can still be wrong, so overlap tracks human taste weakly.
Rung three, a model as judge. A strong model reads the answers and scores them against a written guide. The next section is entirely about this rung.
Rung four, human review. The ground truth, and the only rung that is not a proxy for something else. It is slow and expensive, and it needs written guidelines, or two reviewers will score the same answer differently and you will learn nothing.
| Rung | What it measures | Cost | Blind spot |
|---|---|---|---|
| Perplexity | how well the model fits held-out text | trivial | silent on task quality and on safety |
| Task metrics | concrete, countable task success | low | overlap scores track human taste weakly |
| Model as judge | quality against a written rubric | medium | position, length and self-preference bias |
| Human review | real subjective quality | high | slow, and needs clear written guidelines |
The craft is running the cheap rungs constantly and the expensive rungs at milestones. The mistake is mistaking a low rung for the whole picture.
Use a model as a judge without being fooled by it
Here is the mechanism, because the phrase gets used more often than it gets explained. You take one prompt and two answers to it, usually one from the base model and one from your fine-tune. You hand both to a strong model, along with a rubric: a short written scoring guide, such as “score 1 to 5 on whether this resolves the customer’s problem, and 1 to 5 on tone”. The judge returns a score or picks a winner.
Why this caught on is worth stating precisely. Zheng and colleagues measured how often a strong judge picks the same winner a human picks, on the same pairs, and found agreement above 80 percent. Two humans agree with each other about that often as well. So the judge is roughly as consistent as a second reviewer, at a tiny fraction of the cost and time. That is genuinely useful.
It is also an instrument, and instruments have known errors. This one has three, and all three are easy to trip over.
- Position bias. The judge tends to prefer whichever answer it reads first, whatever the answers say. Always putting your fine-tune second quietly costs it points in every comparison you ever run.
- Verbosity bias. The judge tends to prefer the longer answer. A model that learned to pad will collect points it has not earned.
- Self-preference bias. The judge tends to prefer text written in its own style, which is a problem when the judge and one of the contestants come from the same family.
Each has a defence, and each defence is cheap. For position, score every pair in both orders and average the two results, or randomise the order per comparison. For verbosity, compare at matched output lengths, so quality is being read rather than word count. For self-preference and for drift generally, pin the exact judge model and the exact rubric version, then record both alongside the score. Otherwise next quarter’s improvement may just be a judge upgrade in disguise.
One more habit. Before you trust the judge on a thousand pairs, score fifty of them by hand and compare. If you and the judge disagree often, the rubric is usually the problem rather than the model. Treat the judge like any other measuring instrument in an eval harness: calibrate it first, then read it.
Catch the six failures a training curve cannot show
This is the spine of the part. Each failure below can sit happily beside a beautiful loss curve, and several can sit beside a strong headline metric too.
That pattern is not a coincidence. Training cannot surface these, because training is busy optimising the very number that hides them.
1. Overfitting
Symptom. Training loss keeps falling while performance on unseen data stalls or gets worse.
Cause. The model has started memorising your examples instead of learning the pattern behind them. On a set of ten thousand examples this happens quickly, often within two or three passes.
Check. You already have this one for free. Watch the gap between the training curve and the validation curve. While both fall, the model is generalising. The moment validation turns up while training keeps falling, you have crossed into memorisation. Keep the checkpoint from the low point, not the last one. Fewer passes over the data, a lower learning rate, more and more varied data, and regularisation all push the crossing point later.
2. Catastrophic forgetting
Symptom. The model is better at your task and quietly worse at things it used to handle: arithmetic, following a formatting instruction, refusing a request it should refuse.
Cause. This one is worth understanding properly, because it is the failure the whole series has been building toward. Nothing inside the model is labelled. There is no arithmetic module and no politeness module. The same 1.54 billion numbers do every job at once. Training pushes those numbers toward writing good refund replies, and some of the numbers it moved were also doing the carrying in a column of addition. Your loss only ever grades refund replies, so nothing objects.
Two things make it worse. Luo and colleagues found forgetting increasing with model size across the 1B to 7B range they studied. And running supervised fine-tuning followed by preference tuning induces it on its own, a pattern usually called the alignment tax.
Two further results are worth holding while you design the check. Allen-Zhu and Li found that a model can hold a fact and still fail to answer a question about it, depending on how varied the phrasings were when it first met that fact. Their follow-up put a rough ceiling on how much a model can store at all, near 2 bits of knowledge per parameter. Both say the same thing to anyone measuring a fine-tune: what you installed is mostly habit rather than new knowledge, so testing only your own task’s phrasing measures the smallest part of what changed.
Check. Score the same general-capability tests on the base model and on your candidate, with an identical setup, and compare. Three public tests are the standard probes, and it helps to know what each one is:
- MMLU is multiple-choice exam questions across many school and university subjects. It probes general knowledge.
- GSM8K is grade-school maths word problems. The answer is a number, so it is either right or it is not.
- IFEval is instructions with checkable constraints, such as “answer in exactly three bullet points”. A short program can verify compliance, with no judge involved.
The community tool for running them is EleutherAI’s lm-evaluation-harness. One command scores all three:
lm_eval --model hf \
--model_args pretrained=./qwen15b-merged \
--tasks mmlu,gsm8k,ifeval \
--batch_size auto
Every part of that command, for someone meeting it for the first time:
lm_evalis the command the harness installs. It downloads each test, runs your model over it and prints the scores.--model hfsays the model is a local Hugging Face model on disk, rather than something behind an API.--model_args pretrained=./qwen15b-mergedis the folder holding it. The word merged matters. If you trained a LoRA adapter, fold it into the base weights first, so the tool sees one plain model. Part 11 covers merging.--tasks mmlu,gsm8k,ifevalis the list of tests, comma separated.--batch_size autolets the harness pick how many examples to push through at once, based on what fits on your card.- The backslashes at the end of each line are shell syntax for “this command continues below”. Delete them and put it on one line if you prefer.
Run it twice. Once against the base checkpoint, once against the fine-tune. Then compare. Suppose, as an illustration, your ticket score rises 12 points while maths falls 6 and instruction following falls 3. That model has not improved. It has traded, and you now have to decide consciously whether the trade is one you want. Pin your library versions before you build any of this into a pipeline, because harness task names and defaults do move between releases.
Four mitigations, in rough order of how much they help. Prefer LoRA or QLoRA, because a frozen base forgets far less, and that is a large part of why these methods dominate. Lower the learning rate. Use fewer passes over the data, or a smaller adapter. And mix 5 to 10 percent general instruction data into your training set, so the model keeps rehearsing what it already knew.
3. Reward hacking
Symptom. The headline metric climbs steadily. You read twenty outputs and they are worse.
Cause. Whenever you optimise against a stand-in for quality, the model can find ways to score well without being good. The stand-in might be a reward model, which is a separate model trained to score answers, or a model judge, or a keyword-matching metric. Flattery, padding, and exploiting a gap in the rubric all raise the score. Part 13 covers where this starts.
Check. Partly manual, unavoidably. Read the outputs. Metric up plus eyeball bad is the signature, and there is no automated substitute for the eyeball half.
4. Length bias
Symptom. The tuned model waffles. Answers get longer and say no more.
Cause. Human annotators and model judges both tend to rate longer answers as better. If your preference data carries that habit, the model learns that length is quality.
Check. Run a length-controlled comparison. Compare quality only between answers of similar length, rather than on raw preference. If the win rate collapses once lengths are matched, the gain was padding.
5. Mode collapse
Symptom. Every answer starts to sound the same. The same opening, the same three-part structure, the same closing sentence.
Cause. Tuning too hard, for too long, crushes the variety out of the model’s answers. It settles into one safe template.
Check. Read many outputs, never a single sample. One answer always looks fine. Fifty answers reveal the template. Counting how often the same opening phrase appears across a hundred prompts takes ten minutes and finds it immediately.
6. Sycophancy
Symptom. The model agrees with the user. Even when the user is wrong.
Cause. Preference data collected from people rewards agreement, because being agreed with feels good in the moment and gets the thumbs up.
Check. Write prompts that assert a false premise, then see whether the model holds its ground. “This invoice totals 240 pounds, so why did you refund 260?” when the invoice does not total 240. A model that goes along with the premise has learned to please rather than to help. This is the same probing discipline as testing guardrails.
Know when a measure has stopped measuring
Pick any number and push on it hard enough and it stops telling you what it used to tell you. A school judged on pass rates starts teaching the exam. A support team judged on tickets closed starts closing tickets. The number rises. The thing the number stood for does not.
That has a name: Goodhart’s law. When a measure becomes a target, it stops being a good measure.
You have now met three faces of it in one article. Contamination is the benchmark being gamed by the pretraining data. Reward hacking is the model gaming your proxy. And there is a third, quieter one: you, tuning until the number rises. Each of them is the same phenomenon. A proxy under optimisation pressure drifts away from the thing it was meant to stand for. Loss is a proxy for capability. A benchmark is a proxy for skill. A reward model is a proxy for human judgement.
This is not a bug you fix once. It is a permanent condition you manage, and the defences are structural. Hold out fresh evaluations the model has never been optimised against. Rotate them over time. Combine several metrics, so gaming any one of them does not win. Keep a person in the loop on flagged cases. An evaluation you have been optimising against for six months is no longer measuring what it measured on day one.
Measure the thing your users actually need
General benchmarks tell you that you did not lose general competence. They tell you nothing about whether the model is good at your job. The evaluation that decides whether to ship is always built from your own data. Match the metric to the shape of the output.
| Fine-tune type | What to measure | Watch especially |
|---|---|---|
| Classification | accuracy, a confusion matrix, agreement across repeated runs | whether the same input gives the same label twice |
| Open generation | a judge rubric, whether claims are supported by the source, human review on flagged cases | length bias, sycophancy, invented facts |
| Structured output | does it parse, are the field values right, how slow is the worst request | silent drift in the output shape |
| Tool or function use | did the call succeed, were the arguments right, was the order right | call shape breaking before text quality does |
Two of those terms need unpacking. A confusion matrix is a grid: real category down the side, predicted category across the top, counts in the cells. It shows which categories the model mixes up, which is far more useful than one accuracy number. Turning refund tickets into billing tickets is a different problem from turning both into nonsense, and one number cannot tell you which you have.
The second is worth a rule. Classification is the easiest thing in the world to evaluate honestly, because correctness is objective. If your task can be reshaped from open generation into a choice among a fixed set of valid answers, do it. You get a hard metric and a confusion matrix for free.
Two habits then make any task evaluation trustworthy. Build a regression suite, a fixed set of cases run against every checkpoint, so a fix in one area that breaks another is caught the same day. And run the evaluation under production conditions: the same prompt template, the same context assembly, the same routing. An evaluation that does not mirror production is measuring a model you will not deploy, which is exactly the trap described in the playbook for getting an LLM demo into production.
Run evaluation as a loop that starts before training
All of this assembles into one loop, and the order matters more than any individual metric.
Define what success means, and the number that clears the bar, before you train. Do it afterwards and you will rationalise whatever number you happen to get. Everyone does. Writing the threshold down first is the only defence.
Then baseline the base model on the same evaluation, so every result is a difference rather than a bare score. Train. Run the ladder. Compare the delta with its confidence interval. Read the failures. Ship or go round again.
Two pieces of plumbing make the loop stick. Wire the evaluation into your build, so every checkpoint is gated automatically against the thresholds you already committed to. And send a small slice of real traffic to the new model before a full rollout, so production gets a vote before it gets the whole load.
One framing worth carrying past this series. For anything client-facing or regulated, the evaluation is the evidence. A fixed holdout, a versioned rubric, a pinned judge, a base-versus-tuned delta with a confidence interval, and a stability check across runs is what turns a claim about your model into something you can defend. A claim without that artefact is just an assertion.
That closes the arc. You can now size a run before you start it, pick a base model, build the bench, run supervised fine-tuning end to end, fix the data that actually drives it, fit a large model on a small card with LoRA and QLoRA, teach judgement with preference tuning, split a job across cards when one is not enough, and prove whether any of it worked. None of those decisions rests on a black box, because every layer under them was built up in the parts before this one.
Where next. The training side is done, and the serving side is a separate discipline with its own arithmetic: the complete path an LLM takes to answer a question picks the story up at the moment your checkpoint goes into production. Go and fine-tune something, then measure it properly.
Key takeaways
- A falling loss is a score on cards the model has already seen. It says nothing about generalisation, nothing about regressions, and nothing about real quality.
- A holdout is data kept out of training on purpose. Use three piles: train, validation for decisions during the run, and a test set you spend exactly once.
- Contamination raises a score dishonestly, so a private, versioned set built from your own data beats any public leaderboard number.
- Always score the base model on the identical evaluation. A bare number cannot tell an improved model from a merely changed one.
- Decide on the lower end of a bootstrap confidence interval, not the average. A 2-point gain on 100 examples usually rests on about 18 disagreements and is indistinguishable from a coin flip.
- The ladder runs perplexity, task metrics, model judge, human review. A model judge agrees with humans over 80 percent of the time and prefers the first answer, the longer answer and its own style.
- Six failures hide from the training loop: overfitting, catastrophic forgetting, reward hacking, length bias, mode collapse and sycophancy. Forgetting is the dangerous one, so score general tests base versus tuned on every single fine-tune.
You can now
- Split your data into three piles that each answer a different question, from “Test the model on data it has never seen”.
- Audit a training example for label leakage with one question about every field, from “Keep the evaluation set clean, and spot the two ways it leaks”.
- Decide whether a measured gain is real by reading the lower end of a bootstrap confidence interval, from “Turn a bare score into a delta you can defend”.
- Run a model judge with position, verbosity and self-preference bias controlled, from “Use a model as a judge without being fooled by it”.
- Detect catastrophic forgetting with one command run twice, and read the result as a trade rather than an improvement, from “Catch the six failures a training curve cannot show”.
- Set a pass threshold before training and gate every checkpoint against it automatically, from “Run evaluation as a loop that starts before training”.
Glossary
- Activation
- Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Alignment tax
- The capability a model loses in exchange for being made more helpful or better behaved. Running supervised fine-tuning and then preference tuning is enough on its own to cause it. Lin et al., Mitigating the Alignment Tax of RLHF
- Base model
- The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- Benchmark
- A fixed public test set with a published score, such as a maths or general-knowledge exam for models. Useful as a coarse filter and untrustworthy as a decision, because public questions leak into pretraining data. Xu et al., Benchmark Data Contamination survey
- bf16
- A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
- Bootstrap confidence interval
- Resampling your per-example results thousands of times to see how much the average could have moved by luck alone. Ship the change only when the lower end of the interval is still above zero.
- Catastrophic forgetting
- When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
- Checkpoint
- A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
- Data contamination
- Overlap between your training data and your evaluation data, including reworded near-copies. It is the one data problem that fails dishonestly: the score looks excellent because the model memorised the answers. Xu et al., Benchmark Data Contamination survey
- Decontamination
- Checking that nothing in the training data also appears in the evaluation data, and removing whatever does. Skip it and your reported score measures memorisation rather than skill. Data is the actual job
- Epoch
- One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one. Supervised fine-tuning end to end
- Full fine-tuning
- Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
- Generalisation
- How well a model does on inputs it never saw during training. It is the thing you actually want, and training loss does not measure it.
- Goodhart’s law
- When a measure becomes a target, it stops being a good measure. Every metric you optimise hard enough eventually gets gamed, by the model, by you, or by the whole field fitting the same benchmark.
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Holdout
- Data deliberately kept out of training so you can measure the model on something it has never seen. A number measured on data the model trained on is not evidence of anything.
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- Label leakage
- When information from the expected answer sneaks into the input field of a training example, so the model reads the answer off the prompt. Metrics soar and the deployed model is useless, because real inputs carry no such hint.
- Learning rate
- A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
- Length bias
- The tendency of both human and model annotators to rate longer answers as better, which quietly teaches a preference-tuned model to pad. If a tuned model starts waffling, suspect the data before the algorithm. Meng et al., SimPO
- LLM as a judge
- Using a strong model to score or compare answers against a written rubric. It agrees with human raters over 80 percent of the time, about as often as humans agree with each other, and it favours the first answer shown, longer answers, and its own style. Zheng et al., Judging LLM-as-a-Judge
- LoRA
- Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Mode collapse
- When tuning crushes the variety out of a model’s answers, so it produces the same shapes and phrasings over and over. You catch it by reading many outputs, never a single sample.
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Out-of-distribution probe
- A small evaluation on tasks that look nothing like your training data, kept specifically to catch abilities you have lost. In-distribution evaluation cannot see catastrophic forgetting at all.
- Overfitting
- When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse.
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- Perplexity
- A measure of how surprised a model is by a piece of text, worked out from its loss. Lower is better, it is the cheapest evaluation there is, and it says nothing about whether the model is good at your task.
- Preference tuning
- Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- QLoRA
- LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
- Regularisation
- Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
- Reinforcement learning
- Training by trial and reward: the model produces something, a score comes back saying how good it was, and the model is pushed toward whatever scored well. It differs from supervised training, where the right answer is handed over in advance. Schulman et al., Proximal Policy Optimization
- Reward hacking
- The model finding ways to score well without being good: flattery, padding, keyword stuffing, or exploiting a gap in the rubric. The metric climbs and reading the outputs reveals the game. Preference tuning: RLHF, DPO and verifiable rewards
- Reward model
- A model trained on human comparisons to give any answer a single score. RLHF optimises against it because you cannot put a human in the loop for millions of steps. Ouyang et al., InstructGPT
- RLHF
- Reinforcement learning from human feedback. Collect human comparisons, train a reward model on them, then use reinforcement learning to push the model toward higher-scoring answers. It holds four models in memory, which is why it is a big-team tool. Ouyang et al., InstructGPT
- Supervised fine-tuning
- Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
- Sycophancy
- The model telling users what they want to hear rather than what is true. You detect it by asserting a wrong premise in the prompt and checking whether the model holds its ground.
- System prompt
- A message at the start of a conversation that sets the model’s role and rules. During training it is masked out of the loss along with the user’s turns. Masking in LLM training
- Test set
- The split you look at once, at the very end, for one honest number. Tune against it and it quietly becomes a second validation set and stops being trustworthy. Splits, done honestly
- Validation set
- The split you check during training to pick settings and choose the best checkpoint. Because you make decisions from it over and over, it slowly stops being an honest estimate. Splits, done honestly
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
Practical exercises
Extend the bootstrap confidence interval with a probability-of-improvement estimate
Add one line to the article’s bootstrap_ci function that also reports the fraction of the 10,000 bootstrap resamples whose mean delta is positive, as a friendlier “probability of improvement” number for a non-technical audience. Then explain the exact mathematical relationship between that number and the function’s existing ship-or-not rule, lo > 0, and whether adding it changes the decision rule you would actually use to ship.
See the worked solution (opens in a new tab)
Diagnose a judge score that was never actually order-controlled
Your team always feeds the base model’s answer first and the fine-tuned model’s answer second into the judge, for every comparison in every run. Two months later you rerun the identical evaluation, swapping the order so the fine-tuned model goes first, and the fine-tuned model’s win rate drops by 15 points. Diagnose what happened, and describe the protocol you should have used from the start to avoid ever reporting the first, uncorrected number.
See the worked solution (opens in a new tab)
Tell sycophancy apart from length bias in an ambiguous result
After DPO tuning, your model’s judge score improves by 9 points. Reading the outputs, you notice most responses now open with a compliment to the user and rarely push back even when the user’s stated premise is factually wrong. You also run a length-controlled comparison, matching quality at similar output lengths, and the tuned model’s win rate there is almost identical to the base model’s. Name the failure mode that is actually present, the one this rules out, and the specific test you would run next to confirm the one that remains.
See the worked solution (opens in a new tab)
Wire a forgetting check and a task-specific bootstrap gate into one build
The article gives an lm_eval command that scores mmlu, gsm8k and ifeval on a merged checkpoint, and a bootstrap_ci function for your own private holdout. Sketch, in prose plus a short code fragment, how you would combine both into a single automated checkpoint gate that fails the build if either a general-capability benchmark drops by more than 2 points against the base model, or the private holdout’s bootstrapped lower bound on the delta is not above zero.
See the worked solution (opens in a new tab)
Decide between LoRA and QLoRA on Qwen2.5-7B using evidence rather than folklore
On your one 32 GB card, Qwen2.5-7B fine-tuning as plain LoRA needs about 18 GB static state, comfortable but leaving less headroom, while QLoRA needs about 8 GB, leaving far more headroom but running at roughly 40 percent less throughput. Design the single evaluation that would tell you, for your specific task, whether the extra headroom QLoRA buys is worth its throughput cost, using this part’s holdout rules and the six silent failure modes where relevant. State the two outcomes that would each justify a different choice.
Frequently asked questions
Why is training loss not enough for evaluating a fine-tuned LLM?
Because it is measured on the examples the model was trained on, so it is a score on cards the student has already memorised. It cannot tell you whether the model works on unseen inputs, whether it lost abilities it used to have, or whether the output is any good to a reader. Each of those needs a separate test run outside the training loop.
How do I detect catastrophic forgetting?
Score the same general-capability tests on the base model and on your fine-tune with an identical setup, then compare. General knowledge, grade-school maths, instruction following and your own safety set are the standard probes, and one lm-evaluation-harness command runs the public ones. A rise on your task alongside a fall elsewhere is a trade rather than an improvement, and you should decide consciously whether you want it.
How large should my evaluation set be?
Large enough that the difference you care about survives a bootstrap confidence interval. A 2-point gain on 100 examples typically rests on about 18 examples where the two models disagreed, which is no more convincing than 10 heads in 18 coin tosses. Bootstrap the per-example difference and read the lower end: if it sits below zero, the set is too small or the gain is not real.
Can I use an LLM as a judge, and what does it get wrong?
Yes, and it is standard practice, because a strong judge agrees with human raters over 80 percent of the time, about as often as two humans agree with each other. It has three known biases: it prefers the answer shown first, it prefers longer answers, and it prefers text in its own style. Score every pair in both orders and average, compare at matched lengths, and pin the judge model and rubric version so a score change is not just a judge upgrade.
Should I trust public benchmark scores?
Only as a coarse first filter. If a test’s questions have ever been on the open internet, assume the pretraining data swallowed them, which is why popular benchmarks saturate and why scores drop on genuinely private subsets. Detecting that after the fact is unreliable, especially once wording differs. Build a private, versioned evaluation from your own data and let that be the only score that decides anything.
When should I build the evaluation, before or after training?
Before. Define the metrics, the thresholds and the held-out set, and score the base model, all before any training starts. Then the result is a pass or fail you committed to in advance rather than a number you rationalise afterwards. Wire it into your build so every checkpoint is gated, and send a small slice of real traffic to the new model before a full rollout.
Sources and further reading
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, the source of the human-agreement figure and of the three judge biases.
- Liu et al., G-Eval, on judging against an explicit written rubric.
- Luo et al., An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning, the source of the scale finding across the 1B to 7B range.
- Lin et al., Mitigating the Alignment Tax of RLHF, on the forgetting that preference tuning induces on its own.
- Xu et al., Benchmark Data Contamination of Large Language Models: A Survey, on how the leak happens and why detecting it is unreliable.
- Allen-Zhu and Li, Physics of Language Models 3.1, Knowledge Storage and Extraction, on why storing a fact and being able to answer a question about it are two different things.
- Allen-Zhu and Li, Physics of Language Models 3.3, Knowledge Capacity Scaling Laws, for the rough ceiling on how much a model can hold, which bounds what a fine-tune can add.
- Hendrycks et al., Measuring Massive Multitask Language Understanding, the first of the three standard forgetting probes.
- Cobbe et al., Training Verifiers to Solve Math Word Problems, the GSM8K probe, where the answer is a number and grading needs no judge.
- Zhou et al., Instruction-Following Evaluation for Large Language Models, the IFEval probe, whose constraints a short program can verify.
