Exercise solutions: Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
These are the worked solutions for the exercises in Part 13, Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
The reference model’s own weight memory: 1.54 billion parameters times 2 bytes, about 3.08 GB. It needs no gradients and no optimizer state, since it is frozen.
Full-parameter DPO: policy at 24.64 GB, 16 bytes per parameter, plus the reference’s 3.08 GB, is 27.72 GB, an increase of 3.08 divided by 24.64, exactly 12.5 percent. That matches “closer to 12 percent, not half” precisely. It stays close to 12.5 percent regardless of model size, because the ratio is purely a bytes-per-parameter comparison, 2 extra bytes against a 16-byte baseline, independent of how many parameters the model has.
LoRA DPO: the policy’s own static state is about 4 GB, already including the frozen base plus a small adapter and optimizer overhead. Add a separate 3.08 GB reference copy and the total is 7.08 GB, so removing the reference brings you back down to 4 GB, a reduction of 3.08 divided by 7.08, about 43.5 percent. That falls short of a literal half, while still sitting in the same ballpark as the phrase “roughly halves.” The two cases differ this much because in the full-parameter case the reference’s 2 bytes per parameter is a rounding error against 16, while in the LoRA case the reference is a full bf16 copy competing against a static state that is mostly small adapter and optimizer state to begin with, so it makes up nearly half the total.
Exercise 2
The team already has exactly the ingredient the article calls out as rare and valuable: a programmatic verifier, the symbolic solver, that can check correctness directly. That is precisely the case the verifiable-reward branch exists for, and it makes the reward free, objective and far harder to game than either a learned reward model or a set of human pairwise judgments. Paying for thousands of human comparisons to build an offline DPO dataset throws that advantage away. It replaces it with the slowest, most expensive preference source the article lists, for a task where the cheapest and most reliable source, the solver itself, is already sitting there.
The better investment is to harden the verifier and run online reinforcement learning against it, in the spirit of the group relative policy optimisation approach the article describes: sample several responses to the same maths problem, score every one of them with the solver, and treat above-the-group-average as the signal to reinforce. This needs no reward model and no human labelling at all. DPO and its relatives remain the right tool for the parts of the assistant a solver cannot check, such as how it explains a step or how politely it corrects a wrong assumption. For the correctness of the maths itself, the verifiable branch is the one built for exactly this situation, and the pairwise-preference plan should be redirected there instead.
Exercise 3
cfg = DPOConfig(
beta = 0.1,
precompute_ref_log_probs = True, # compute once, then drop the reference
learning_rate = 5e-7,
num_train_epochs = 1,
per_device_train_batch_size = 2,
optim = "paged_adamw_8bit",
)
With ref_model=None and no precompute flag, the reference model sits in memory for the entire run, since every training step needs its log probabilities for the chosen and rejected responses.
With precompute_ref_log_probs=True, the trainer instead makes one full pass over the training set with the reference model before the main training loop begins, computing and caching the log probabilities the loss needs for every example. Once that pass finishes, the reference model is no longer needed for anything and can be released from memory entirely. The rest of training runs against the cached numbers.
The trade is a one-time forward pass over the whole dataset added to the front of the job, in exchange for the reference model’s memory being gone for the entire remainder of training rather than only in principle. This is the right move specifically when even a temporarily doubled memory footprint at the start of the run would not fit, since it converts a standing memory cost into a bounded, one-time compute cost instead.
Exercise 4
The first mistake is the epoch count. The article is explicit that one epoch is usually enough for preference tuning and that training longer produces degenerate, repetitive output, exactly the templated quality being described here. The fix is to cut this back to one epoch and re-evaluate before doing another pass.
The second mistake is a regime mix-up rather than simply too high or too low in the abstract. A learning rate of 1e-5 is roughly the right value for LoRA-based DPO, where only a small adapter is moving and needs a comparatively larger step to register at all. This run is full-parameter DPO, though, which wants roughly 5e-7, two orders of magnitude smaller, because every weight in the model is moving and a large step drifts far past the frozen reference’s leash. Using the LoRA-sized learning rate on a full-parameter run lets the policy wander much further from the reference than beta 0.1 was ever meant to allow, consistent with losing a capability, arithmetic, that the reference model still had. The fix is to drop the learning rate to roughly 5e-7 for this full-parameter run, or, if the faster learning rate is actually wanted, switch the run itself to a peft_config-based LoRA setup where 1e-5 is the appropriate value to begin with.
Exercise 5
Two separate bf16 copies: 2 times 15.24 GB, 7.62 billion parameters at 2 bytes each, is 30.48 GB, plus 0.6 GB adapter and optimizer overhead, 31.08 GB total. That is already over the card’s roughly 29.8 GiB ceiling before any safety margin is even subtracted, so this naive setup does not fit.
One shared frozen base serving as both backbone and reference: 15.24 GB plus 0.6 GB, 15.84 GB. Against the roughly 29.8 GiB ceiling and the standard 2 to 3 GB margin, that leaves on the order of 11 to 13 GB free for activations, comfortably fitting.
The same shared base quantised with NF4 and double quantisation: about 3.93 GB, from the base-memory arithmetic used elsewhere in this series, plus 0.6 GB, is 4.53 GB. That leaves roughly 22 to 24 GB free after the margin, an enormous amount of headroom for larger batches or longer sequences.
Only the naive two-copy setup fails to fit. The article’s own advice, to let one frozen base do both jobs and to shrink that base with QLoRA if you want more room still, is not a marginal optimisation here. It is the difference between a run that does not fit at all and one with either comfortable or very generous headroom.