Exercise Solutions: Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures

Exercise solutions: Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures

These are the worked solutions for the exercises in Part 15, Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures. Read the exercise first; coming here before you have tried it defeats the point.

Exercise 1

import numpy as np

def bootstrap_ci(base_scores, tuned_scores, n_boot=10000, alpha=0.05):
    diffs = np.asarray(tuned_scores) - np.asarray(base_scores)
    n = len(diffs)
    boot_means = np.array([
        np.random.choice(diffs, size=n, replace=True).mean()
        for _ in range(n_boot)
    ])
    lo, hi = np.percentile(boot_means, [100 * alpha / 2, 100 * (1 - alpha / 2)])
    p_positive = (boot_means > 0).mean()   # fraction of resamples favouring the tuned model
    return diffs.mean(), lo, hi, p_positive

The relationship is exact rather than approximate, because lo is defined as the 2.5th percentile of the bootstrap distribution when alpha is 0.05. If lo is above zero, then by definition at most 2.5 percent of the bootstrap means fall below zero, so p_positive must be at least about 0.975 whenever the existing rule already says ship. The new number restates part of the same distribution in a more intuitive unit, and it adds no information beyond what the existing lower bound already captures.

It does not change the decision rule, and it should not replace it. A high p_positive alone, say 0.9, does not guarantee lo > 0, because the lower tail of the bootstrap distribution can still dip below zero even when 90 percent of resamples are positive, particularly with a skewed or heavy-tailed per-example delta. The lower bound remains the correct thing to gate on. The probability-of-improvement number is a communication aid for reporting the result to people who find a plain percentage more intuitive than a confidence interval, nothing more.

Exercise 2

This is the judge’s position bias, one of the three biases the article names explicitly: a model judge tends to favour whichever answer it sees first, independent of which answer is actually better. Always showing the base model first meant every comparison in the original run gave the fine-tuned model the less-favoured slot, so its true win rate was understated in that first run. In the second run, with the fine-tuned model now in the favoured first slot, a large chunk of the 15-point swing is presentation order rather than genuine quality difference. Neither number in isolation is trustworthy. The 15-point gap between them is roughly the size of the position bias itself for this judge and this rubric.

The correct protocol, stated in the article, is to randomise which model’s answer appears first for every individual comparison, or to score every pair in both orders and average the two scores. Either approach means the reported number already has position bias averaged out before anyone sees a headline win rate, rather than discovering two months later that a fixed presentation order was quietly deciding part of the result. Pin the judge model and rubric version too, so a later score change can be attributed to the model rather than an unnoticed judge upgrade.

Exercise 3

The failure present is sycophancy: telling the user what they want to hear rather than what is true, shown here by the flattering openers and the reluctance to correct a false premise, exactly the pattern the article describes.

Length bias is ruled out, not merely unlikely. Length bias means the model learned that longer responses score better, and the diagnostic for it is a length-controlled comparison: if the win rate holds up once length is matched, the gain is not simply padding dressed up as quality. Since the length-controlled win rate here is almost identical to the base model’s, the 9-point headline gain is not an artefact of verbosity, and whatever is driving it is a real behavioural change rather than length.

To confirm sycophancy specifically, run the test the article prescribes for it: a set of leading and adversarial prompts that assert a wrong premise, checking whether the model holds the true answer or goes along with the user. If the tuned model concedes the false premise noticeably more often than the base model does on the same prompts, that confirms sycophancy directly, rather than inferring it only from the tone of a handful of outputs.

Exercise 4

The gate needs two independent checks run against the same candidate checkpoint, both compared to the same base-model baseline, with the build failing if either one fails, since a checkpoint that wins on your task but quietly loses general capability is exactly the trade the article warns is not an improvement.

import subprocess, json
import numpy as np

def run_lm_eval(model_path):
    subprocess.run(
        ["lm_eval", "--model", "hf", "--model_args", f"pretrained={model_path}",
         "--tasks", "mmlu,gsm8k,ifeval", "--batch_size", "auto",
         "--output_path", f"{model_path}-results.json"],
        check=True,
    )
    return json.load(open(f"{model_path}-results.json"))["results"]

base_results = run_lm_eval("./qwen15b-base")
tuned_results = run_lm_eval("./qwen15b-merged")

forgetting_failed = any(
    tuned_results[task]["acc"] - base_results[task]["acc"] < -0.02
    for task in ("mmlu", "gsm8k", "ifeval")
)

mean_delta, lo, hi = bootstrap_ci(base_task_scores, tuned_task_scores)
task_failed = lo <= 0

if forgetting_failed or task_failed:
    raise SystemExit(f"checkpoint gate failed: forgetting={forgetting_failed} task={task_failed}")

Both checks run against the identical pair of checkpoints, base and candidate, so every number is a delta rather than a bare score, per the article's second holdout rule. The forgetting check is a simple threshold on each benchmark's own accuracy field. The task check reuses the bootstrap function unchanged, since it already returns exactly the lower bound the gate needs. Wiring this into continuous integration means every checkpoint is judged by the same pre-committed thresholds automatically, the workflow the article's final section argues for: define the thresholds before training, then let the gate decide rather than a person eyeballing a number after the fact.

Exercise 5

Fine-tune the identical training data and hyperparameters twice, once as plain LoRA and once as QLoRA, changing only what QLoRA structurally requires: the quantisation config and the k-bit training preparation step. Score both candidates on the same private, versioned, task-specific holdout, under the same production conditions, prompt template and all, per the article's rules for a trustworthy holdout. Baseline the untouched base model on the same holdout too, so every number is a delta rather than a bare score, and bootstrap a confidence interval specifically on the per-example delta between the two candidates themselves, since the real question here is whether they differ from each other rather than whether either one beats the base.

Also run the same general-capability benchmarks, mmlu, gsm8k and ifeval, base versus each candidate, since both are still adapters and catastrophic forgetting is a live risk for either one, not just for a full fine-tune.

Two outcomes, two different answers. If the bootstrapped confidence interval on the LoRA-versus-QLoRA delta straddles zero, meaning neither reliably beats the other on your task, take QLoRA: this matches the published result that QLoRA fine-tuning matches 16-bit LoRA quality, and the extra roughly 10 GB of headroom is then free to spend on a larger batch size or a longer context window, at the cost of the roughly 40 percent lower throughput this part quantifies. If instead the interval's lower bound sits reliably above zero in LoRA's favour, this specific task is one of the exceptions where the four-bit compute tax also costs you something in quality, not just speed, and the right call is to keep plain LoRA and live with its narrower headroom rather than trade quality away for VRAM you may not need.