Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures

Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures

Evaluating a fine-tuned LLM is the step that decides whether the training run was worth doing. It is also the step most teams bolt on at the end, once the model already exists. That order is backwards. This part puts it right, and it assumes you know no statistics at all.

The trap is a falling loss curve. Loss is one number saying how wrong the model was on the text it was just shown, and training is the business of making that number smaller. A falling loss tells you the model fits your data better. It tells you nothing about whether the model got good at your task, and nothing about whether it still remembers how to add up an invoice.

Here is the shape the failure takes. You fine-tune a model on support tickets. It scores well on your ticket test. You ship. Two weeks later the model writes beautiful refund replies and can no longer do arithmetic. Nothing in training warned you, because nothing in training was watching.

By the end you will be able to build an evaluation set the model has never seen, turn a bare score into a difference you can defend with numbers, climb a ladder of evaluation methods from cheap to expensive, and name and detect six failures a training curve cannot show.

Name the four questions a falling loss cannot answer

Picture a student revising for an exam. They work through a hundred practice questions, over and over, with the answers on the back of each card. By Friday they get all hundred right.

What have you learned about that student? Almost nothing. You do not know whether they can answer a question they have not seen. You do not know whether they still remember last year’s material. You do not know whether they are any use in a real job.

Training loss is that Friday score. It is measured on the very examples the model was trained on. Four separate questions matter, and loss answers only the first of them, weakly.

  1. Did it learn the task at all? Did the behaviour you wanted actually go in?
  2. Does it work on inputs it has never seen? Or did it only memorise your examples?
  3. Did it break anything? Can it still do the things it could do before you touched it?
  4. Is the output any good? Would a person reading it be satisfied?
1 · Did it learn the task? the only one loss speaks to — and only on data like the training set 2 · Does it generalize? beyond the exact examples seen — held-out, unseen, realistic cases 3 · Did it break anything? math, reasoning, safety, format — the silent regressions 4 · Is it actually good? by real-world / human standards, not just a proxy number going up

Loss has a weak opinion on the first question and no opinion at all on the other three. The third one is the one that ends launches, because nothing in the training run raises a hand.

Every one of the last three needs its own test. Each test has to run outside the training loop, on data the training never touched. That is the whole design principle of this part, and everything below is an application of it.

Test the model on data it has never seen

Go back to the flashcards. Your student scored a hundred percent on the hundred cards they studied. There are two completely different reasons that could happen.

  • They understood the underlying pattern, so any similar question would also work.
  • They memorised those hundred cards, front and back, and know nothing else.

Those two are the entire subject of this section. The second one is memorisation: storing the specific examples. The first is generalisation: performing well on inputs that were never in the training data. Generalisation is the thing you actually want. It is also the thing training loss structurally cannot measure, because training loss is only ever measured on the cards.

So you need cards the student has not seen. In machine learning that slice of data has a name: a holdout. A holdout is data you deliberately set aside before training and never show to the training loop. A score measured on data the model trained on is not evidence of anything. A score measured on a holdout is.

Three piles, not two

In practice you cut the data three ways, not two, and the third cut surprises people.

  1. Training set. The model sees these and learns from them. This is the only pile that changes any weights.
  2. Validation set. The model never learns from these, but you look at them during the run. This is where Part 9’s eval-loss curve comes from, and it is what you use to pick a learning rate or choose which saved checkpoint to keep.
  3. Test set. You look at this once, at the very end, for one honest number.

Why three? Because looking at a set is itself a way of fitting to it. Suppose you try eight learning rates and keep whichever scored best on the validation set. You have now hand-fitted that set. It flattered you a little, and it will keep flattering you. The test set is the one pile you have not spent. Spend it once.

Make the split representative

A holdout is only useful if it looks like real traffic. Say you have 1,000 support tickets: 900 refund requests and 100 billing disputes. Take a random 10 percent as a test set and you might land 6 billing disputes, or you might land 16. Your entire billing score then rests on 6 examples, and 6 examples cannot tell you anything.

The fix is a stratified split: take 10 percent of each category separately. That gives 90 refunds and 10 billing disputes, every time, in the same proportions as the real pile. Part 10 covers the split mechanics in full.

Three more habits make a holdout worth having. Include the hard cases and the adversarial ones, meaning inputs deliberately written to trip the model, rather than only the easy examples that were pleasant to collect. Fix the set and version it, so a score from March is comparable to a score from June. And assemble it exactly as production does: same prompt template, same system prompt, same surrounding context. An evaluation that uses a different prompt format is measuring a model you are not going to ship.

Keep the evaluation set clean, and spot the two ways it leaks

Everything above collapses if the model has already seen the test items. Your student glanced at ten of the hundred exam questions last week. Their score goes up. Their knowledge does not.

That is contamination: overlap between the training data and the evaluation data, including reworded near-copies. It is the one data fault that fails dishonestly. A bad batch of training text makes your score fall, so you go and look. Contamination makes your score rise, so you celebrate and ship. Part 10 treats it as a data-pipeline control, which is where the fix belongs.

Public benchmarks have this problem permanently. If a test’s questions have ever been on the open internet, assume the pretraining crawl swallowed them years ago. The model may be reciting rather than reasoning, and the score cannot tell the two apart.

Public benchmark questions are on the web → swept into pretraining data → model has seen the answers score is inflated, not skill Private held-out set never published, never in training frontier models routinely drop 5–8 points here vs public this is the honest number the gap = contamination

The leak happens upstream, during pretraining, long before you arrive. That is why a model can look strong on a public test and mediocre on a private one built from the same kind of question.

Detecting contamination after the fact is unreliable, and it gets worse once the training text differs in wording from the original question. Xu and colleagues survey the methods and the gaps. So do not lean on detection. Lean on a private set instead, built from your own data, never published, versioned, and treated as the only score that decides anything. Public benchmarks are a coarse first filter and never the deciding vote.

The quieter leak, inside a single example

There is a second kind of leak, and it hides inside one training example rather than across two files.

Say your ticket records carry a field called resolution_code. An agent fills it in when they close the ticket, so it exists in your archive and it always agrees with the correct answer. You build training examples straight from the archive and leave that field in the prompt. The model does not learn to read the ticket. It learns to read the field.

Your metrics soar. Then you deploy, and a real request arrives with that field empty, because nobody has closed the ticket yet. The model has no idea what to do. That is label leakage: information from the expected answer bleeding into the input side of an example.

The check is a single question, asked of every field in your prompt. Would this value exist at the moment a real request arrives? If not, delete it before you build a single example.

Turn a bare score into a delta you can defend

Your fine-tune scores 68 percent on a clean, private holdout. Is that good?

You cannot answer that. Nobody can. The number needs something to be compared against, and the right thing is the model you started from. So score the base model too, on the identical evaluation: same examples, same prompt template, same judge, same everything. Now you have two numbers, and the difference between them is the delta. Without a delta you cannot tell an improved model from a merely changed one.

Why a 2-point gain on 100 examples is probably nothing

Say your holdout has 100 examples. The base model gets 61 right. The fine-tune gets 63. A 2-point gain. Ship it?

No. Start with a coin. Flip a fair coin 100 times and you expect 50 heads. You will rarely get exactly 50. You get 47, or 54, and nobody claims the coin has changed. That wobble is what happens whenever you measure a small sample of something.

Your 100 examples are a sample too. You did not test on every ticket your users will ever send. You tested on 100 of them, drawn from a much larger pool. The honest question is this: if you had drawn a different 100, would the gap still be there?

Now look at the 2 points example by example, rather than as two totals. A realistic breakdown:

  • 82 examples where both models scored the same, right or wrong together.
  • 10 examples where the fine-tune was right and the base was wrong.
  • 8 examples where the base was right and the fine-tune was wrong.

Ten minus eight is two, out of a hundred, which is the 2-point gain. But look at what it rests on. The 82 agreements carry no information at all. The entire result comes from 18 examples that disagreed, and 10 of those went your way.

Ten heads out of eighteen tosses. Nobody calls a coin bent on that.

The bootstrap, in five steps

That coin argument is the intuition. The bootstrap is the mechanical version of it, and it works on any score, not just right and wrong.

You would love to fetch a fresh 100 tickets and measure again, ten thousand times over. You cannot. So you fake it, using only the data you already have.

  1. Write down one number per example: the fine-tune’s score minus the base model’s score on that same example. Here that gives 10 numbers of plus 1, 8 numbers of minus 1, and 82 zeros.
  2. Draw 100 numbers at random from that pile of 100, putting each one back before the next draw. Duplicates are the whole point. Some examples get picked three times and some not at all, which is exactly what a different sample of reality would have looked like.
  3. Average that fake sample. Write the average down.
  4. Repeat steps 2 and 3 ten thousand times. You now have 10,000 plausible answers to the question “what could this delta have been”.
  5. Sort those 10,000 averages. Discard the lowest 250 and the highest 250. What remains is the middle 95 percent, and that range is your 95 percent confidence interval.

Run it on the numbers above and the interval comes out at roughly minus 6 points to plus 10 points. The exact ends wobble slightly each time, because the resampling is random. The important part does not wobble: the range contains zero. There are plenty of plausible redraws in which the fine-tune is worse than the base model. Your 2-point gain is noise wearing a suit.

So the rule is: decide on the lower end of the interval, never on the average. If the lower end is below zero, you have not shown anything yet.

What actually fixes it

More examples. Keep the same rates and run the same evaluation on 2,000 examples instead of 100. Now 200 examples improve, 160 get worse, and the rest agree. The delta is still 2 points. The interval is now roughly plus 0.1 to plus 3.9 points.

It clears zero, and only just. That is the honest answer to “how big should my evaluation set be”. Not a round number somebody told you. Big enough that the difference you care about survives this test.

import numpy as np

def bootstrap_ci(base_scores, tuned_scores, n_boot=10000, alpha=0.05):
    diffs = np.asarray(tuned_scores) - np.asarray(base_scores)  # paired, per example
    n = len(diffs)
    boot_means = np.array([
        np.random.choice(diffs, size=n, replace=True).mean()
        for _ in range(n_boot)
    ])
    lo, hi = np.percentile(boot_means, [100 * alpha / 2, 100 * (1 - alpha / 2)])
    return diffs.mean(), lo, hi

mean_delta, lo, hi = bootstrap_ci(base_scores, tuned_scores)
print(f"delta={mean_delta:.3f}  95% CI=[{lo:.3f}, {hi:.3f}]")
# ship only if lo > 0

Line by line, assuming you have never used numpy.

  • import numpy as np brings in numpy, the standard Python library for working on whole lists of numbers at once. The as np part is just a short nickname.
  • base_scores and tuned_scores are two ordinary lists of per-example scores. They must be in the same order and cover the same examples, because the comparison is paired.
  • n_boot=10000 is how many fake samples to draw. alpha=0.05 means you want a 95 percent interval, since 1 minus 0.05 is 0.95.
  • np.asarray(...) - np.asarray(...) turns both lists into numpy arrays and subtracts them element by element. The result is step 1: one difference per example.
  • n = len(diffs) is how many examples there are, so each fake sample is the same size as the real one.
  • np.random.choice(diffs, size=n, replace=True) is step 2. It draws n values from diffs, and replace=True is what puts each value back before the next draw.
  • .mean() is step 3, the average of that one fake sample.
  • The square brackets with for _ in range(n_boot) around it are ordinary Python. They repeat the draw 10,000 times and collect the results. The underscore just means the loop counter is not used.
  • np.percentile(boot_means, [2.5, 97.5]) is step 5, written as arithmetic. It returns the value that 2.5 percent of the averages fall below, and the value that 97.5 percent fall below. Those two are the ends of the interval.
  • The f"..." line is a Python f-string. {lo:.3f} prints that number to three decimal places.

That gives you three rules that make a holdout earn its keep, and none of them is optional.

1 · Strict isolation no overlap, no near- duplicates, no leakage through a synthetic generator that saw eval leak → inflated, useless score 2 · Score the base too same examples, same rubric, same judge, same week the fine-tune’s value is the DELTA vs base 3 · Bootstrap the CI a 0.3-pt lift on 50 examples is noise resample the per-example delta 1000× → ship on the lower bound, not the mean

The middle rule is the one people skip. A fine-tune’s score in isolation says nothing, because there is nothing to compare it against.

Climb the evaluation ladder and stop at the right rung

Evaluation is not one thing. It is a ladder, running from cheap and shallow up to expensive and true. Each rung measures something the rung below it cannot see.

Rung one, perplexity. Run the model over held-out text, take the average loss, and undo the logarithm. What comes out is roughly how many equally likely options the model felt it was choosing between at each token. A perplexity of 10 means it was about as unsure as someone picking between 10 options. Lower is better. It costs almost nothing, so you can run it at every checkpoint, and it catches gross damage quickly. It says nothing whatsoever about whether the model is good at your task.

Rung two, task metrics. Concrete, countable things. Did the extracted field match the expected value? Did the output parse as valid JSON? Was the classification correct? These are cheap and objective. One family of them deserves caution: overlap scores, which count how many words an answer shares with a reference answer. An answer sharing few words can still be perfect, and one sharing many can still be wrong, so overlap tracks human taste weakly.

Rung three, a model as judge. A strong model reads the answers and scores them against a written guide. The next section is entirely about this rung.

Rung four, human review. The ground truth, and the only rung that is not a proxy for something else. It is slow and expensive, and it needs written guidelines, or two reviewers will score the same answer differently and you will learn nothing.

▲ higher fidelity, closer to real quality · higher cost ▲ 4 · Human evaluation gold standard for subjective quality · slow, costly · LMSYS Arena = ~5M pairwise votes 3 · LLM-as-judge a strong model scores vs a rubric · 80–90% human agreement at ~500–5000× lower cost · has biases 2 · Task metrics on a held-out set accuracy · F1 · exact-match (classification) · BLEU/ROUGE (generation, weak proxy) 1 · Held-out loss / perplexity cheap, automatic, necessary — but only measures fit on similar data

Cost and fidelity rise together, which is why the ladder is a ladder and not a menu. Run the bottom rungs constantly and the top rung on the cases that matter most.
Rung What it measures Cost Blind spot
Perplexity how well the model fits held-out text trivial silent on task quality and on safety
Task metrics concrete, countable task success low overlap scores track human taste weakly
Model as judge quality against a written rubric medium position, length and self-preference bias
Human review real subjective quality high slow, and needs clear written guidelines

The craft is running the cheap rungs constantly and the expensive rungs at milestones. The mistake is mistaking a low rung for the whole picture.

Use a model as a judge without being fooled by it

Here is the mechanism, because the phrase gets used more often than it gets explained. You take one prompt and two answers to it, usually one from the base model and one from your fine-tune. You hand both to a strong model, along with a rubric: a short written scoring guide, such as “score 1 to 5 on whether this resolves the customer’s problem, and 1 to 5 on tone”. The judge returns a score or picks a winner.

Why this caught on is worth stating precisely. Zheng and colleagues measured how often a strong judge picks the same winner a human picks, on the same pairs, and found agreement above 80 percent. Two humans agree with each other about that often as well. So the judge is roughly as consistent as a second reviewer, at a tiny fraction of the cost and time. That is genuinely useful.

It is also an instrument, and instruments have known errors. This one has three, and all three are easy to trip over.

  • Position bias. The judge tends to prefer whichever answer it reads first, whatever the answers say. Always putting your fine-tune second quietly costs it points in every comparison you ever run.
  • Verbosity bias. The judge tends to prefer the longer answer. A model that learned to pad will collect points it has not earned.
  • Self-preference bias. The judge tends to prefer text written in its own style, which is a problem when the judge and one of the contestants come from the same family.

Each has a defence, and each defence is cheap. For position, score every pair in both orders and average the two results, or randomise the order per comparison. For verbosity, compare at matched output lengths, so quality is being read rather than word count. For self-preference and for drift generally, pin the exact judge model and the exact rubric version, then record both alongside the score. Otherwise next quarter’s improvement may just be a judge upgrade in disguise.

One more habit. Before you trust the judge on a thousand pairs, score fifty of them by hand and compare. If you and the judge disagree often, the rubric is usually the problem rather than the model. Treat the judge like any other measuring instrument in an eval harness: calibrate it first, then read it.

Catch the six failures a training curve cannot show

This is the spine of the part. Each failure below can sit happily beside a beautiful loss curve, and several can sit beside a strong headline metric too.

① Overfitting memorizes train data, fails on new detect: train-loss ≪ val-loss gap ② Catastrophic forgetting gains task, loses old abilities detect: base-vs-tuned on general benchmarks ③ Reward hacking games the metric, not the goal detect: metric up + qualitative read ④ Length / verbosity bias longer scored better regardless detect: length-controlled eval ⑤ Mode / format collapse repetitive, degenerate, low diversity detect: qualitative + diversity check ⑥ Sycophancy agrees with the user over the truth detect: adversarial / leading prompts

Read down the right-hand column. Every single detection runs outside the training loop, on data or in a direction the training never optimised.

That pattern is not a coincidence. Training cannot surface these, because training is busy optimising the very number that hides them.

1. Overfitting

Symptom. Training loss keeps falling while performance on unseen data stalls or gets worse.

Cause. The model has started memorising your examples instead of learning the pattern behind them. On a set of ten thousand examples this happens quickly, often within two or three passes.

Check. You already have this one for free. Watch the gap between the training curve and the validation curve. While both fall, the model is generalising. The moment validation turns up while training keeps falling, you have crossed into memorisation. Keep the checkpoint from the low point, not the last one. Fewer passes over the data, a lower learning rate, more and more varied data, and regularisation all push the crossing point later.

2. Catastrophic forgetting

Symptom. The model is better at your task and quietly worse at things it used to handle: arithmetic, following a formatting instruction, refusing a request it should refuse.

Cause. This one is worth understanding properly, because it is the failure the whole series has been building toward. Nothing inside the model is labelled. There is no arithmetic module and no politeness module. The same 1.54 billion numbers do every job at once. Training pushes those numbers toward writing good refund replies, and some of the numbers it moved were also doing the carrying in a column of addition. Your loss only ever grades refund replies, so nothing objects.

Two things make it worse. Luo and colleagues found forgetting increasing with model size across the 1B to 7B range they studied. And running supervised fine-tuning followed by preference tuning induces it on its own, a pattern usually called the alignment tax.

Two further results are worth holding while you design the check. Allen-Zhu and Li found that a model can hold a fact and still fail to answer a question about it, depending on how varied the phrasings were when it first met that fact. Their follow-up put a rough ceiling on how much a model can store at all, near 2 bits of knowledge per parameter. Both say the same thing to anyone measuring a fine-tune: what you installed is mostly habit rather than new knowledge, so testing only your own task’s phrasing measures the smallest part of what changed.

Check. Score the same general-capability tests on the base model and on your candidate, with an identical setup, and compare. Three public tests are the standard probes, and it helps to know what each one is:

  • MMLU is multiple-choice exam questions across many school and university subjects. It probes general knowledge.
  • GSM8K is grade-school maths word problems. The answer is a number, so it is either right or it is not.
  • IFEval is instructions with checkable constraints, such as “answer in exactly three bullet points”. A short program can verify compliance, with no judge involved.

The community tool for running them is EleutherAI’s lm-evaluation-harness. One command scores all three:

lm_eval --model hf \
        --model_args pretrained=./qwen15b-merged \
        --tasks mmlu,gsm8k,ifeval \
        --batch_size auto

Every part of that command, for someone meeting it for the first time:

  • lm_eval is the command the harness installs. It downloads each test, runs your model over it and prints the scores.
  • --model hf says the model is a local Hugging Face model on disk, rather than something behind an API.
  • --model_args pretrained=./qwen15b-merged is the folder holding it. The word merged matters. If you trained a LoRA adapter, fold it into the base weights first, so the tool sees one plain model. Part 11 covers merging.
  • --tasks mmlu,gsm8k,ifeval is the list of tests, comma separated.
  • --batch_size auto lets the harness pick how many examples to push through at once, based on what fits on your card.
  • The backslashes at the end of each line are shell syntax for “this command continues below”. Delete them and put it on one line if you prefer.

Run it twice. Once against the base checkpoint, once against the fine-tune. Then compare. Suppose, as an illustration, your ticket score rises 12 points while maths falls 6 and instruction following falls 3. That model has not improved. It has traded, and you now have to decide consciously whether the trade is one you want. Pin your library versions before you build any of this into a pipeline, because harness task names and defaults do move between releases.

Four mitigations, in rough order of how much they help. Prefer LoRA or QLoRA, because a frozen base forgets far less, and that is a large part of why these methods dominate. Lower the learning rate. Use fewer passes over the data, or a smaller adapter. And mix 5 to 10 percent general instruction data into your training set, so the model keeps rehearsing what it already knew.

3. Reward hacking

Symptom. The headline metric climbs steadily. You read twenty outputs and they are worse.

Cause. Whenever you optimise against a stand-in for quality, the model can find ways to score well without being good. The stand-in might be a reward model, which is a separate model trained to score answers, or a model judge, or a keyword-matching metric. Flattery, padding, and exploiting a gap in the rubric all raise the score. Part 13 covers where this starts.

Check. Partly manual, unavoidably. Read the outputs. Metric up plus eyeball bad is the signature, and there is no automated substitute for the eyeball half.

4. Length bias

Symptom. The tuned model waffles. Answers get longer and say no more.

Cause. Human annotators and model judges both tend to rate longer answers as better. If your preference data carries that habit, the model learns that length is quality.

Check. Run a length-controlled comparison. Compare quality only between answers of similar length, rather than on raw preference. If the win rate collapses once lengths are matched, the gain was padding.

5. Mode collapse

Symptom. Every answer starts to sound the same. The same opening, the same three-part structure, the same closing sentence.

Cause. Tuning too hard, for too long, crushes the variety out of the model’s answers. It settles into one safe template.

Check. Read many outputs, never a single sample. One answer always looks fine. Fifty answers reveal the template. Counting how often the same opening phrase appears across a hundred prompts takes ten minutes and finds it immediately.

6. Sycophancy

Symptom. The model agrees with the user. Even when the user is wrong.

Cause. Preference data collected from people rewards agreement, because being agreed with feels good in the moment and gets the thumbs up.

Check. Write prompts that assert a false premise, then see whether the model holds its ground. “This invoice totals 240 pounds, so why did you refund 260?” when the invoice does not total 240. A model that goes along with the premise has learned to please rather than to help. This is the same probing discipline as testing guardrails.

Know when a measure has stopped measuring

Pick any number and push on it hard enough and it stops telling you what it used to tell you. A school judged on pass rates starts teaching the exam. A support team judged on tickets closed starts closing tickets. The number rises. The thing the number stood for does not.

That has a name: Goodhart’s law. When a measure becomes a target, it stops being a good measure.

You have now met three faces of it in one article. Contamination is the benchmark being gamed by the pretraining data. Reward hacking is the model gaming your proxy. And there is a third, quieter one: you, tuning until the number rises. Each of them is the same phenomenon. A proxy under optimisation pressure drifts away from the thing it was meant to stand for. Loss is a proxy for capability. A benchmark is a proxy for skill. A reward model is a proxy for human judgement.

This is not a bug you fix once. It is a permanent condition you manage, and the defences are structural. Hold out fresh evaluations the model has never been optimised against. Rotate them over time. Combine several metrics, so gaming any one of them does not win. Keep a person in the loop on flagged cases. An evaluation you have been optimising against for six months is no longer measuring what it measured on day one.

Measure the thing your users actually need

General benchmarks tell you that you did not lose general competence. They tell you nothing about whether the model is good at your job. The evaluation that decides whether to ship is always built from your own data. Match the metric to the shape of the output.

Fine-tune type What to measure Watch especially
Classification accuracy, a confusion matrix, agreement across repeated runs whether the same input gives the same label twice
Open generation a judge rubric, whether claims are supported by the source, human review on flagged cases length bias, sycophancy, invented facts
Structured output does it parse, are the field values right, how slow is the worst request silent drift in the output shape
Tool or function use did the call succeed, were the arguments right, was the order right call shape breaking before text quality does

Two of those terms need unpacking. A confusion matrix is a grid: real category down the side, predicted category across the top, counts in the cells. It shows which categories the model mixes up, which is far more useful than one accuracy number. Turning refund tickets into billing tickets is a different problem from turning both into nonsense, and one number cannot tell you which you have.

The second is worth a rule. Classification is the easiest thing in the world to evaluate honestly, because correctness is objective. If your task can be reshaped from open generation into a choice among a fixed set of valid answers, do it. You get a hard metric and a confusion matrix for free.

Two habits then make any task evaluation trustworthy. Build a regression suite, a fixed set of cases run against every checkpoint, so a fix in one area that breaks another is caught the same day. And run the evaluation under production conditions: the same prompt template, the same context assembly, the same routing. An evaluation that does not mirror production is measuring a model you will not deploy, which is exactly the trap described in the playbook for getting an LLM demo into production.

Run evaluation as a loop that starts before training

All of this assembles into one loop, and the order matters more than any individual metric.

1 · define successmetrics + thresholds FIRST 2 · baselinescore the BASE model 3 · trainthe fine-tune run 4 · eval the laddertask + general + safety 5 · compute the deltatuned − base, with CI 6 · inspect failuresread outputs, red-team 7 · ship or iterategate on lower-bound delta iterate → re-train with fixes

Steps one and two are the ones people skip, and they are the ones that make everything after them honest. Set the thresholds before training, and score the base model so every later result is a delta.

Define what success means, and the number that clears the bar, before you train. Do it afterwards and you will rationalise whatever number you happen to get. Everyone does. Writing the threshold down first is the only defence.

Then baseline the base model on the same evaluation, so every result is a difference rather than a bare score. Train. Run the ladder. Compare the delta with its confidence interval. Read the failures. Ship or go round again.

Two pieces of plumbing make the loop stick. Wire the evaluation into your build, so every checkpoint is gated automatically against the thresholds you already committed to. And send a small slice of real traffic to the new model before a full rollout, so production gets a vote before it gets the whole load.

One framing worth carrying past this series. For anything client-facing or regulated, the evaluation is the evidence. A fixed holdout, a versioned rubric, a pinned judge, a base-versus-tuned delta with a confidence interval, and a stability check across runs is what turns a claim about your model into something you can defend. A claim without that artefact is just an assertion.

0 Base✓ 1 SFT✓ 2 Data✓ 3 LoRA✓ 4 QLoRA✓ 5 Preference✓ 6 Multi-GPU✓ 7 Eval✓ done complete — a base model, adapted, aligned, scaled, and proven

From an untouched base model to a fine-tune backed by evidence. The gap between a model that works and a model you can prove works is the whole subject of this part.

That closes the arc. You can now size a run before you start it, pick a base model, build the bench, run supervised fine-tuning end to end, fix the data that actually drives it, fit a large model on a small card with LoRA and QLoRA, teach judgement with preference tuning, split a job across cards when one is not enough, and prove whether any of it worked. None of those decisions rests on a black box, because every layer under them was built up in the parts before this one.

Where next. The training side is done, and the serving side is a separate discipline with its own arithmetic: the complete path an LLM takes to answer a question picks the story up at the moment your checkpoint goes into production. Go and fine-tune something, then measure it properly.

Key takeaways

  • A falling loss is a score on cards the model has already seen. It says nothing about generalisation, nothing about regressions, and nothing about real quality.
  • A holdout is data kept out of training on purpose. Use three piles: train, validation for decisions during the run, and a test set you spend exactly once.
  • Contamination raises a score dishonestly, so a private, versioned set built from your own data beats any public leaderboard number.
  • Always score the base model on the identical evaluation. A bare number cannot tell an improved model from a merely changed one.
  • Decide on the lower end of a bootstrap confidence interval, not the average. A 2-point gain on 100 examples usually rests on about 18 disagreements and is indistinguishable from a coin flip.
  • The ladder runs perplexity, task metrics, model judge, human review. A model judge agrees with humans over 80 percent of the time and prefers the first answer, the longer answer and its own style.
  • Six failures hide from the training loop: overfitting, catastrophic forgetting, reward hacking, length bias, mode collapse and sycophancy. Forgetting is the dangerous one, so score general tests base versus tuned on every single fine-tune.

You can now

  • Split your data into three piles that each answer a different question, from “Test the model on data it has never seen”.
  • Audit a training example for label leakage with one question about every field, from “Keep the evaluation set clean, and spot the two ways it leaks”.
  • Decide whether a measured gain is real by reading the lower end of a bootstrap confidence interval, from “Turn a bare score into a delta you can defend”.
  • Run a model judge with position, verbosity and self-preference bias controlled, from “Use a model as a judge without being fooled by it”.
  • Detect catastrophic forgetting with one command run twice, and read the result as a trade rather than an improvement, from “Catch the six failures a training curve cannot show”.
  • Set a pass threshold before training and gate every checkpoint against it automatically, from “Run evaluation as a loop that starts before training”.

Glossary

Activation
Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
Adapter
A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
Alignment tax
The capability a model loses in exchange for being made more helpful or better behaved. Running supervised fine-tuning and then preference tuning is enough on its own to cause it. Lin et al., Mitigating the Alignment Tax of RLHF
Base model
The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
Batch
A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
Benchmark
A fixed public test set with a published score, such as a maths or general-knowledge exam for models. Useful as a coarse filter and untrustworthy as a decision, because public questions leak into pretraining data. Xu et al., Benchmark Data Contamination survey
bf16
A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
Bootstrap confidence interval
Resampling your per-example results thousands of times to see how much the average could have moved by luck alone. Ship the change only when the lower end of the interval is still above zero.
Catastrophic forgetting
When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
Data contamination
Overlap between your training data and your evaluation data, including reworded near-copies. It is the one data problem that fails dishonestly: the score looks excellent because the model memorised the answers. Xu et al., Benchmark Data Contamination survey
Decontamination
Checking that nothing in the training data also appears in the evaluation data, and removing whatever does. Skip it and your reported score measures memorisation rather than skill. Data is the actual job
Epoch
One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one. Supervised fine-tuning end to end
Full fine-tuning
Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
Generalisation
How well a model does on inputs it never saw during training. It is the thing you actually want, and training loss does not measure it.
Goodhart’s law
When a measure becomes a target, it stops being a good measure. Every metric you optimise hard enough eventually gets gamed, by the model, by you, or by the whole field fitting the same benchmark.
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
Holdout
Data deliberately kept out of training so you can measure the model on something it has never seen. A number measured on data the model trained on is not evidence of anything.
Inference
Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
Label leakage
When information from the expected answer sneaks into the input field of a training example, so the model reads the answer off the prompt. Metrics soar and the deployed model is useless, because real inputs carry no such hint.
Learning rate
A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
Length bias
The tendency of both human and model annotators to rate longer answers as better, which quietly teaches a preference-tuned model to pad. If a tuned model starts waffling, suspect the data before the algorithm. Meng et al., SimPO
LLM as a judge
Using a strong model to score or compare answers against a written rubric. It agrees with human raters over 80 percent of the time, about as often as humans agree with each other, and it favours the first answer shown, longer answers, and its own style. Zheng et al., Judging LLM-as-a-Judge
LoRA
Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Mode collapse
When tuning crushes the variety out of a model’s answers, so it produces the same shapes and phrasings over and over. You catch it by reading many outputs, never a single sample.
Optimizer
The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
Out-of-distribution probe
A small evaluation on tasks that look nothing like your training data, kept specifically to catch abilities you have lost. In-distribution evaluation cannot see catastrophic forgetting at all.
Overfitting
When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse.
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
Perplexity
A measure of how surprised a model is by a piece of text, worked out from its loss. Lower is better, it is the cheapest evaluation there is, and it says nothing about whether the model is good at your task.
Preference tuning
Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
Pretraining
The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
QLoRA
LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
Regularisation
Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
Reinforcement learning
Training by trial and reward: the model produces something, a score comes back saying how good it was, and the model is pushed toward whatever scored well. It differs from supervised training, where the right answer is handed over in advance. Schulman et al., Proximal Policy Optimization
Reward hacking
The model finding ways to score well without being good: flattery, padding, keyword stuffing, or exploiting a gap in the rubric. The metric climbs and reading the outputs reveals the game. Preference tuning: RLHF, DPO and verifiable rewards
Reward model
A model trained on human comparisons to give any answer a single score. RLHF optimises against it because you cannot put a human in the loop for millions of steps. Ouyang et al., InstructGPT
RLHF
Reinforcement learning from human feedback. Collect human comparisons, train a reward model on them, then use reinforcement learning to push the model toward higher-scoring answers. It holds four models in memory, which is why it is a big-team tool. Ouyang et al., InstructGPT
Supervised fine-tuning
Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
Sycophancy
The model telling users what they want to hear rather than what is true. You detect it by asserting a wrong premise in the prompt and checking whether the model holds its ground.
System prompt
A message at the start of a conversation that sets the model’s role and rules. During training it is masked out of the loss along with the user’s turns. Masking in LLM training
Test set
The split you look at once, at the very end, for one honest number. Tune against it and it quietly becomes a second validation set and stops being trustworthy. Splits, done honestly
Validation set
The split you check during training to pick settings and choose the best checkpoint. Because you make decisions from it over and over, it slowly stops being an honest estimate. Splits, done honestly
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model


Practical exercises

Extend the bootstrap confidence interval with a probability-of-improvement estimate

Add one line to the article’s bootstrap_ci function that also reports the fraction of the 10,000 bootstrap resamples whose mean delta is positive, as a friendlier “probability of improvement” number for a non-technical audience. Then explain the exact mathematical relationship between that number and the function’s existing ship-or-not rule, lo > 0, and whether adding it changes the decision rule you would actually use to ship.

See the worked solution (opens in a new tab)

Diagnose a judge score that was never actually order-controlled

Your team always feeds the base model’s answer first and the fine-tuned model’s answer second into the judge, for every comparison in every run. Two months later you rerun the identical evaluation, swapping the order so the fine-tuned model goes first, and the fine-tuned model’s win rate drops by 15 points. Diagnose what happened, and describe the protocol you should have used from the start to avoid ever reporting the first, uncorrected number.

See the worked solution (opens in a new tab)

Tell sycophancy apart from length bias in an ambiguous result

After DPO tuning, your model’s judge score improves by 9 points. Reading the outputs, you notice most responses now open with a compliment to the user and rarely push back even when the user’s stated premise is factually wrong. You also run a length-controlled comparison, matching quality at similar output lengths, and the tuned model’s win rate there is almost identical to the base model’s. Name the failure mode that is actually present, the one this rules out, and the specific test you would run next to confirm the one that remains.

See the worked solution (opens in a new tab)

Wire a forgetting check and a task-specific bootstrap gate into one build

The article gives an lm_eval command that scores mmlu, gsm8k and ifeval on a merged checkpoint, and a bootstrap_ci function for your own private holdout. Sketch, in prose plus a short code fragment, how you would combine both into a single automated checkpoint gate that fails the build if either a general-capability benchmark drops by more than 2 points against the base model, or the private holdout’s bootstrapped lower bound on the delta is not above zero.

See the worked solution (opens in a new tab)

Decide between LoRA and QLoRA on Qwen2.5-7B using evidence rather than folklore

On your one 32 GB card, Qwen2.5-7B fine-tuning as plain LoRA needs about 18 GB static state, comfortable but leaving less headroom, while QLoRA needs about 8 GB, leaving far more headroom but running at roughly 40 percent less throughput. Design the single evaluation that would tell you, for your specific task, whether the extra headroom QLoRA buys is worth its throughput cost, using this part’s holdout rules and the six silent failure modes where relevant. State the two outcomes that would each justify a different choice.

See the worked solution (opens in a new tab)

Frequently asked questions

Why is training loss not enough for evaluating a fine-tuned LLM?

Because it is measured on the examples the model was trained on, so it is a score on cards the student has already memorised. It cannot tell you whether the model works on unseen inputs, whether it lost abilities it used to have, or whether the output is any good to a reader. Each of those needs a separate test run outside the training loop.

How do I detect catastrophic forgetting?

Score the same general-capability tests on the base model and on your fine-tune with an identical setup, then compare. General knowledge, grade-school maths, instruction following and your own safety set are the standard probes, and one lm-evaluation-harness command runs the public ones. A rise on your task alongside a fall elsewhere is a trade rather than an improvement, and you should decide consciously whether you want it.

How large should my evaluation set be?

Large enough that the difference you care about survives a bootstrap confidence interval. A 2-point gain on 100 examples typically rests on about 18 examples where the two models disagreed, which is no more convincing than 10 heads in 18 coin tosses. Bootstrap the per-example difference and read the lower end: if it sits below zero, the set is too small or the gain is not real.

Can I use an LLM as a judge, and what does it get wrong?

Yes, and it is standard practice, because a strong judge agrees with human raters over 80 percent of the time, about as often as two humans agree with each other. It has three known biases: it prefers the answer shown first, it prefers longer answers, and it prefers text in its own style. Score every pair in both orders and average, compare at matched lengths, and pin the judge model and rubric version so a score change is not just a judge upgrade.

Should I trust public benchmark scores?

Only as a coarse first filter. If a test’s questions have ever been on the open internet, assume the pretraining data swallowed them, which is why popular benchmarks saturate and why scores drop on genuinely private subsets. Detecting that after the fact is unreliable, especially once wording differs. Build a private, versioned evaluation from your own data and let that be the only score that decides anything.

When should I build the evaluation, before or after training?

Before. Define the metrics, the thresholds and the held-out set, and score the base model, all before any training starts. Then the result is a pass or fail you committed to in advance rather than a number you rationalise afterwards. Wire it into your build so every checkpoint is gated, and send a small slice of real traffic to the new model before a full rollout.

Sources and further reading

Previous