Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch

Inside LLM Fine-Tuning, part 13 of 15: Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch

Preference tuning is the training stage that teaches a model that one answer is better than another. Not better than a wrong answer. Better than a second answer that is also fluent, also on topic, and also perfectly well spelled.

That gap is where tone lives. It is where helpfulness lives, and knowing when to refuse, and knowing when to stop typing. None of it can be installed by showing the model good examples, and this part explains exactly why.

It assumes you have never met reinforcement learning. So it teaches the handful of words you need, builds the two methods that matter from those words, and then covers the branch where a program hands out the score instead of a person.

By the end you will be able to write the DPO loss down and check its arithmetic by hand, say what raising or lowering its beta dial does, budget the memory for a preference run on one card, and recognise reward hacking when it shows up in your own outputs.

See why supervised fine-tuning cannot say “better”

Supervised fine-tuning, SFT for short, is the stage before this one. You give the model a prompt and one good answer, and the loss makes that exact answer more likely, token by token. Done across enough tasks, that alone teaches a model to follow instructions it never trained on, which Wei and colleagues demonstrated with FLAN. Part 9 runs one end to end. Its structural limit is easiest to see on a real pair of answers.

Take the prompt “My order arrived broken. What should I do?” Here are two replies.

  • Answer A. “Sorry about that. Email [email protected] with your order number and a photo of the damage. A replacement ships within two working days.”
  • Answer B. “I am so sorry to hear this! That must be really frustrating, and nobody wants a broken order. I want you to know we take this seriously. Please do reach out to our support team, who will be delighted to help you with the next steps.”

A is better. It is shorter, it names the address, and it says when the replacement arrives. B is warmer and says nothing you can act on. Now try to teach that preference with SFT. Four steps, and the wall arrives at step three.

  1. An SFT file holds a prompt and one target answer. So you write down the prompt and answer A.
  2. Training raises the probability the model gives to A’s tokens. Good. That is what you wanted.
  3. Nothing in that file mentions B. Nothing in the loss mentions B. There is no slot for it and no term for it. The loss has one direction, and that direction is up.
  4. If you add B to the file as well, you have told the model that B is also a correct answer. It raises B too, which is the opposite of the lesson.

Someone always asks whether raising A pushes B down for free. It nearly does, and the “nearly” is the whole problem. Probabilities have to add up to 1, so softmax does shave a little off everything else when A goes up. It shaves the same little off every other token in the vocabulary. It cannot single out B, because it was never told B exists.

What you need is a signal with two directions. Push the better answer up. Push the worse answer down. That two-sided signal is a preference, and learning from it is what this whole part is about.

SFT — imitation one gold answer “become more like this” — positive only Preference — contrast chosen ↑ rejected ↓ “widen the gap” — push up good, down bad

SFT has one arrow and it points up. Preference tuning has two, and the gap it opens between them is the thing being learned.

One line to hold on to. SFT is handing a student a shelf of good essays and saying “write like these”. Preference tuning is handing them two essays and saying “this one is better, work out why”. Only the second teaches judgement.

Read a preference dataset, and spot the bias baked into it

Every method in this part eats the same shape of data. One prompt, two answers to it, and a label saying which one won. Three fields:

  • prompt, the user’s message. “My order arrived broken. What should I do?”
  • chosen, the answer that won. Answer A above.
  • rejected, the answer that lost. Answer B above.

The important word in that list is “lost”. The rejected answer is not gibberish and it is not wrong. It is plausible. That is what makes the pair worth training on. If the rejected answer were obvious nonsense, the only thing the model would learn is not to write nonsense, which it already knew.

SFT sample prompt: “Summarize this audit finding.” completion: “The control lacked evidence of quarterly review…” ✓ the one answer one prompt · one target Preference sample prompt: “Summarize this audit finding.” chosen: “Control CA-7 lacked evidence of quarterly review (NIST AI RMF)…” ✓ rejected: “There were some issues with the review process maybe.” ✗ vague

An SFT row has two fields and a preference row has three. That third field is the entire difference between imitation and judgement.

Where do the labels come from? Three sources, in rough order of cost.

  • People. An annotator reads the prompt and both answers and picks one. This is the best signal for anything subjective, and it is slow and expensive. Two annotators will also disagree with each other more often than you expect, so the labels are noisy even when everyone is trying.
  • A judge model. A strong model reads the same pair against a written rubric, meaning a short checklist of what good looks like, and picks a winner. It is cheap and it scales. It also has its own tastes, and training on its verdicts installs those tastes in your model.
  • A program. For maths, code and structured output you can check correctness directly, with no opinion involved. This one gets a section of its own further down.

Now the flaw you should look for before anything else. Human raters and judge models both tend to pick the longer answer, a verbosity bias Zheng and colleagues measured directly. Not always, and not deliberately, but often enough to show up in the averages. So the dataset quietly encodes “longer is better”, and the model dutifully learns it.

That single bias explains a lot of downstream design. It is why two of the methods below divide a score by the answer’s length. It is why you should deliberately curate crisp short answers into the chosen column. And it gives you a first diagnosis: if a preference-tuned model suddenly pads every reply, suspect the data before you blame the algorithm. Part 10, on dataset quality, is the fuller treatment.

Learn the four reinforcement-learning words you actually need

The next two sections need vocabulary that this series has not used yet. Start with the difference it all rests on, using something other than a model.

Suppose you are teaching somebody to cook. There are two ways to do it.

  1. Supervised. You hand over a recipe and a photo of the finished dish. They follow it. You compare their result against the photo, step by step. The right answer was in their hands the whole time.
  2. Trial and reward. You hand over nothing. They cook something. You taste it and say “six out of ten”. They cook again. No recipe ever appears, and no photo. All they ever get back is a number.

The second one is reinforcement learning, RL for short. That is the whole idea. The learner produces something, a score comes back saying how good it was, and the learner is pushed toward whatever scored well. Compare it with training that hands over the correct answer in advance and grades against it. SFT is the first kind. Everything in this section is the second.

Four words follow from that picture.

  1. Policy. The model being trained. That is all the word means here. RL borrowed it from settings where the thing being trained has to decide which action to take next, and for a language model the action is which token to write. When a paper says “the policy”, read “the model we are changing”.
  2. Reward. One number scoring a finished answer. Higher is better. Notice how different this is from a loss. A loss is computed at every token against a token you already knew was correct. A reward is computed once, over the whole answer, with no correct answer anywhere in sight.
  3. Sampling. The policy actually writing an answer, one token at a time, the same slow generation you get when you serve a model. RL cannot score an answer that does not exist yet, so every training step has to generate one first.
  4. Reward model. A second model whose job is to produce the reward. The rest of this section explains why you would build such a thing.

Why the score has to come from a model

The obvious plan is to let people hand out the rewards. Do the arithmetic on that plan and it dies immediately.

A modest RL run is 100,000 training steps. Sample four answers per step and that is 400,000 answers needing a score. At 30 seconds of reading per answer, that is 12,000,000 seconds. Divide by 3,600 and you get about 3,333 hours. Divide by 24 and that is roughly 139 days of somebody reading without sleeping, for one run, that you will want to repeat with different settings.

So you cannot put a person in the loop. You put a model there instead. Collect the human comparisons once, as a fixed dataset of maybe tens of thousands of pairs, and train a model to reproduce those verdicts. Now it scores day and night for free.

Building it is less exotic than it sounds. Take a copy of the SFT model. Its top layer normally emits one score per token in the vocabulary, about 151,000 numbers. Replace that layer with one that emits a single number. Then train it on your pairs with one instruction: give the chosen answer a higher number than the rejected one.

Say it scores a pair at 2.1 for chosen and 1.4 for rejected. The margin is 0.7, chosen wins, and the loss is small. Say it scores them 0.3 and 1.9 instead. Chosen loses, and the loss is large, so the gradient pushes the two scores apart. Run that over enough pairs and you have a scorer.

Here is the sentence to remember. Once trained, it will happily score answers no human ever labelled, including strange ones the policy invents mid-training. That generalisation is the entire point of building it. It is also the entire danger, and we come back to it shortly.

Follow RLHF’s three stages and find where each cost comes from

Reinforcement learning from human feedback, RLHF, is the method that made modern assistants possible. Every method after it is a reaction to one of its costs, so it is worth walking properly. It has three stages and you have already met two of them.

  1. SFT. Teach a base model to answer in the right shape at all. Part 9.
  2. Reward model. Train the scorer on human comparisons, exactly as above.
  3. The loop. Improve the policy against that scorer.
Stage 1 · SFT model already built (Parts 9-12) Stage 2 · Reward Model train on (prompt, chosen, rejected) pairs via Bradley-Terry: learn a scalar reward r(prompt, response) → higher for chosen output: a model that SCORES any response Stage 3 · PPO (the RL) policy generates RM scores it KL penalty → stay near theSFT reference update policy to raise reward… …without drifting off the reference (KL leash) 4 models live in memory here

Stage two is the clever one. Human judgement is squeezed into a model once, and stage three then optimises against that model instead of against people.

Stage three runs the same six moves over and over.

  1. Pull a prompt from a pool.
  2. The policy samples an answer to it.
  3. The reward model scores that answer. Say it comes back with 1.8.
  4. A fourth model predicts what score this prompt was likely to earn. Say it predicts 1.2. Subtract: 1.8 minus 1.2 is 0.6. That difference is called the advantage, and it says the answer beat expectations. This fourth model is the value model, and its only job is producing that expectation.
  5. Push the policy’s probabilities up on the tokens of an above-average answer, and down on a below-average one.
  6. Repeat, a hundred thousand times.

PPO is the specific recipe for step 5. Its contribution is caution: it caps how far one update is allowed to move the policy, because RL updates built on a single noisy score can otherwise wreck a model in a handful of steps.

Why the policy has to be kept on a leash

Reward is a number, and an optimizer will chase a number anywhere it leads. Nothing in those six steps says “and remain a working language model”.

Picture a reward model that happens to score enthusiasm slightly well. A policy with nothing holding it back can drift toward exclamation marks, then toward paragraphs of them. The score climbs the whole way. The model stops being usable long before the score stops rising.

The fix is a penalty for drifting away from where you started. You keep a frozen copy of the SFT model, called the reference model, and measure the distance between the two at every step. The measure is called KL divergence. It compares two lists of probabilities over the same tokens and returns 0 when they agree exactly, growing as they pull apart. So the thing being optimised is not the raw reward:

objective = reward(answer) - lambda * KL(policy, reference)

Every symbol in words. reward(answer) is the reward model’s score for what the policy just wrote. KL(policy, reference) is how far the policy’s next-token probabilities have moved from the frozen reference’s. lambda is a number you set, deciding how much that drift costs. Raise it and the leash is short. Lower it and the policy roams.

There is a deeper reason to want the leash, beyond fluency. Preference data is thin. Tens of thousands of comparisons is a rounding error against the trillions of tokens the base model saw during pretraining. Anything your comparisons do not cover is capability the model already has and can only lose. Part 15, on evaluating a fine-tune, is where you catch that happening.

Reward hacking, in one concrete example

Now back to the reward model’s generalisation. Suppose the people who built your preference set very slightly favoured answers that ended by offering more help. Nobody decided this. It just came out that way, the way length bias does.

The reward model learns it. In practice that means it adds roughly 0.4 to any answer ending with a sentence like “I hope this helps. Let me know if you would like me to go deeper on any part of this.”

The policy is now running a hundred thousand rounds of trial and error against that scorer. It will find those 0.4 points. Within a few thousand steps every single answer ends with that sentence. It ends “What is the capital of France?” with that sentence. Average reward is up and to the right. Actual quality is down. Nothing in the training logs mentions it.

That is reward hacking: scoring well without being good. Flattery, padding, keyword stuffing, and anything else that exploits a gap between what the scorer measures and what you meant. Two defences exist, and you want both. The KL leash caps how far the model can chase the gap. Reading real outputs, rather than watching the reward curve, is the one that actually catches it.

Count the models RLHF holds at once

Stage three keeps four models resident. Put the running example’s numbers on them. Qwen2.5-1.5B has 1.54 billion parameters, and this series prices a trained parameter at 16 bytes of static state under standard mixed-precision AdamW, meaning weights plus gradients plus optimizer state. Part 2 derives all sixteen of those bytes. A frozen model pays only its 2 bytes of weights, because it takes no gradients and gets no optimizer state.

Model in the loop What it does Trained? Bytes per parameter Static state at 1.54B
Policy the model you are improving yes 16 about 24.6 GB
Reference anchors the KL penalty no, frozen 2 about 3.1 GB
Reward model scores each sampled answer no, frozen 2 about 3.1 GB
Value model predicts the expected score yes 16 about 24.6 GB
Total about 55.4 GB

Every figure there is static state only. Activations sit on top of all of them, and four models mean four sets. Real setups vary, and a good implementation will share a backbone between the policy and the value model or use a smaller reward model. The order of magnitude does not move: this is a multi-card job for a 1.5B model, on hardware where a plain SFT run of the same model needs about 33 GB in total.

Memory is only the first cost. Step 2 generates fresh text every step, which is far slower than reading text off disk. And PPO carries a pile of settings that interact with each other, so runs that look identical on paper behave differently. For a small team, full RLHF is out of reach. Everything below is an attempt to keep the signal and drop the machinery.

Train on preferences directly with DPO

Direct preference optimization, DPO, collapsed that whole stack into one training step that looks like ordinary supervised training. No reward model. No sampling loop. No value model. The argument runs in four moves.

Move 1. Write down what stage three was actually chasing. Highest average reward, minus lambda times the drift from the reference. That is the objective from the last section, nothing new.

Move 2. For that objective, there is one best possible policy, and it can be written down exactly. It gives every answer the reference’s probability for that answer, multiplied by a factor that grows with the answer’s reward. Better answers get a bigger multiplier. Everything is then rescaled so the probabilities still add to 1.

Move 3. Read move 2 backwards. If the best policy equals the reference times a function of the reward, then the reward equals a function of the ratio between that policy and the reference. In words: an answer’s reward is beta times the logarithm of how much more likely the policy makes it than the reference did, plus a term that depends only on the prompt.

Move 4. Now recall what the reward model was ever asked to do. Only one thing: score chosen above rejected, for the same prompt. That is a difference of two rewards on one prompt, so the prompt-only term appears in both and cancels. Substitute move 3’s expression into the reward model’s own training loss, and the reward model itself is gone. What is left is a loss over the policy, computed straight from preference pairs.

RLHF · 4 models + RL loop policy · reference · reward model · critic online sampling every step unstable · expensive · hard to tune DPO · 2 models, 1 loss policy + frozen reference (the SFT model) no reward model · no sampling loop stable · ~20 lines in TRL · LoRA-friendly

Three of the four models in the RLHF picture disappear. What survives is the policy, one frozen reference, and a single loss over pairs you already had on disk.

The loss, with every symbol named

Two words first, because both are new.

Log-probability. A model gives a whole answer a probability by multiplying together the probability of each of its tokens. For a 60-token answer that product is a fantastically small number, awkward to work with. Take its logarithm and the product becomes a sum, landing somewhere sensible like -35. Closer to zero means the model likes the answer more.

Sigmoid. A function that squashes any number at all into the range 0 to 1. Feed it 0 and you get 0.5. Feed it a large positive number and you get something just under 1. Feed it a large negative number and you get something just above 0.

Now the loss.

L = - log sigmoid( beta * ( (log pi(chosen) - log ref(chosen)) - (log pi(rejected) - log ref(rejected)) ) )

Every piece, in words, from the inside out.

  • pi is the policy, the model being trained. ref is the frozen reference, a copy of where it started.
  • log pi(chosen) is the log-probability the policy gives the chosen answer. log ref(chosen) is the same number from the reference.
  • The first bracket subtracts those two. It reads as “how much more the policy likes the chosen answer than the reference did”. Positive means the training has pushed it up.
  • The second bracket is exactly the same quantity for the rejected answer.
  • Subtracting the second bracket from the first gives the margin. That one number is the whole learning signal. It is large when the policy has lifted the chosen answer and dropped the rejected one, both measured against the reference.
  • beta multiplies the margin before it goes any further. Its own subsection is below.
  • sigmoid turns the scaled margin into a number between 0 and 1. Read it as “how sure this loss is that the chosen answer wins”.
  • The leading - log is the same negative-logarithm penalty Part 1 built for cross-entropy. It is small when the number inside is near 1 and large when it is near 0.

Check it with numbers. Take beta at 0.1, the usual starting value.

Chosen, vs reference Rejected, vs reference Margin beta x margin sigmoid Loss
+2.0 -1.0 3.0 0.30 0.574 0.554
0 0 0 0 0.500 0.693
-1.0 +2.0 -3.0 -0.30 0.426 0.854

Row two is the untrained state, where the policy is still identical to the reference. The margin is 0 and the loss is 0.693, which is the negative logarithm of one half. Row one is the pair learned correctly and row three is the pair learned backwards. Notice the loss never actually reaches 0. There is always a little more gradient pushing the two answers further apart, which is why the number of passes over the data matters so much here.

for each weight update, on a (prompt, chosen, rejected) triple: chosen: how much more likely under policy vs reference? log[ π(chosen) / π_ref(chosen) ] → push UP rejected: same ratio for the bad one log[ π(rejected) / π_ref(rejected) ] → push DOWN Loss = −log σ( β · [ chosen-ratio − rejected-ratio ] ) “make the chosen ratio exceed the rejected ratio by a healthy margin” — a classification loss β controls how hard: high β = stay close to the reference, low β = drift further to satisfy preferences

The whole thing is a two-way classifier. One number goes in, the margin, and the loss is only satisfied when the policy has separated chosen from rejected relative to where it started.

Why the frozen reference appears in every term

Every quantity in that loss is measured against the reference. That is not decoration. It blocks two ways of satisfying the loss without learning anything.

  1. Inflate everything. A model could raise the chosen answer’s probability by becoming more confident about all text at once. Subtracting the reference cancels a general shift, so that move earns nothing.
  2. Collapse. A model could crush the rejected answer’s probability by ceasing to be a coherent model. Drifting from the reference shows up in both brackets, so that move costs as much as it pays.

The reference is doing the same job the KL penalty did in RLHF. RLHF added the leash to the objective as a separate term. DPO folded it into the loss itself.

Set beta, the dial that decides how far the model may move

Beta multiplies the margin, and that is its entire mechanism. Everything else follows from arithmetic you can do here.

Take a margin of 3.0, meaning the policy has already separated the pair fairly well.

  • Beta 0.1. The sigmoid sees 0.30, returns 0.574, and the loss is 0.554. Still far from satisfied. The gradient stays strong, so training keeps pushing the pair further apart, which means the model keeps moving away from the reference.
  • Beta 1.0. The sigmoid sees 3.0, returns 0.953, and the loss is 0.049. Nearly satisfied. The gradient is small, training mostly stops, and the model stays where it was.

So read it this way. High beta means the loss is contented by a small separation, so the model stays near the reference. That is a short leash. Low beta means the loss demands a large separation, so the model travels further to get there. That is a long leash.

Two symptoms tell you which way to turn it.

  • The tuned model is indistinguishable from the SFT checkpoint you started with. Beta is too high, and nothing was learned.
  • The tuned model has picked up the style you wanted but can no longer do arithmetic it used to handle. Beta is too low, or you ran too many passes. That loss of unrelated ability is catastrophic forgetting, and it does not show up anywhere in the training curve.

Start at 0.1. Most published work sits between 0.01 and 0.5.

Three limits worth carrying

  1. It is off-policy. DPO only ever sees answers that were written before training began. It never scores anything the current model produces, so it cannot discover a behaviour that nobody put in the file. It reweights what is already there.
  2. It inherits length bias. Whatever verbosity was in the labels comes straight through into the weights.
  3. Thin data over-optimises. With few pairs, the model can find an extreme setting of the weights that satisfies every one of them and behaves badly everywhere else. One epoch, meaning one pass over the data, is usually correct. More than that tends to give repetitive, templated output.

Pick a DPO variant only when you hit the cost it removes

Once DPO showed that preference learning could be a plain supervised loss, a family followed. Each one deletes a different cost. None of them changes the basic character of the thing.

DPOpairs + reference + 2 stages KTOdrops PAIRING —needs only thumbsup/down per answer ORPOdrops REFERENCE +MERGES stages —SFT+pref in one pass SimPOdrops REFERENCE +length-normalizes —half the memory IPOSOFTENS theBradley-Terryoverfit tendency all four are still OFFLINE — they train on fixed data, no generation loop (that’s covered separately)

Read each box as DPO minus one cost. All four stay offline and supervised, and all four trade a little robustness for cheaper data, less memory, or a shorter pipeline.
  • KTO drops the pairing requirement, taking single answers tagged good or bad instead of matched pairs. That matters because live products collect thumbs up and thumbs down by the thousand and almost never collect clean pairs.
  • ORPO merges SFT and preference tuning into one pass, and needs no reference model, so exactly one model sits in memory.
  • SimPO also drops the reference model, and divides an answer’s score by its length, which attacks the verbosity bias head on.
  • IPO pulls harder toward the reference, which curbs the habit described in limit three above.
Method Data it needs Reference model What it buys Reach for it when
DPO pairs: chosen and rejected yes, frozen the simple, stable default you have clean pairs
KTO single answers, good or bad yes, frozen uses cheap thumbs feedback your feedback is unpaired
ORPO pairs no one stage, one model in memory one card, and you want a single pass
SimPO pairs no less memory, length-fair scoring VRAM is tight or answers are padding
IPO pairs yes, frozen less over-fitting on few pairs DPO is landing on extreme behaviour

The advice the field settled on is blunt. Start with DPO, or KTO if your feedback is unpaired. Move to a variant only once you have hit the specific problem it was built to solve. Do not switch because something scored higher on a leaderboard. Switch because you are out of memory or fighting verbosity. These methods are close enough in practice that data quality decides the outcome, exactly as it did in the part on data.

Swap the human judge for a program when the answer can be checked

Everything above shares one ceiling. Offline preference methods reweight what is already in the data. They cannot invent a skill nobody demonstrated.

The biggest shift in post-training over the last two years came from a different direction. Go back to the RL loop, keep the online sampling, and replace the learned reward model with a verifier: a program that scores an output with no opinion in it.

  • Maths. Compare the final answer to the known one. Correct or not.
  • Code. Run the unit tests. Count how many pass.
  • Structured output. Does it parse, and does it carry the required fields.

Maths and code are special cases, and it is worth being clear about why. Both have a ground truth that is cheap to check, independent of style, and settled before the model ever answers. Most tasks have nothing like that. No program decides which of two summaries reads better, and none ever will.

Where a verifier does exist, three things follow. It is free, so there is no labelling budget. It is objective, so it does not quietly favour longer or friendlier answers. And it is much harder to hack, because you cannot flatter a test runner.

Be honest about that third point. A weak verifier is still gameable. A model rewarded for passing tests can learn to special-case the inputs those tests use. The verifier has to be as strict as your real requirement, or you have simply moved the gap somewhere new.

GRPO, which removes the value model too

The value model existed to answer one question: was this answer better than expected? GRPO answers it by sampling instead of by predicting. Five steps.

  1. Take a prompt. Have the policy write several answers to it, say 8.
  2. Score all 8 with the verifier.
  3. Average those 8 scores. That average is the expectation you needed.
  4. Each answer’s advantage is its own score minus the group average.
  5. Push the policy up on the above-average answers and down on the below-average ones.

Work one group by hand. Of the 8 answers, 3 are correct and score 1, and 5 are wrong and score 0. The average is 3 divided by 8, which is 0.375. So each correct answer carries an advantage of 1 minus 0.375, which is 0.625. Each wrong answer carries 0 minus 0.375, which is -0.375. Nothing predicted anything, and there is no value model anywhere in that list.

1 prompte.g. a math problem sample K responses (8–64) answer 1 → ✓ 1.0 answer 2 → ✗ 0.0 answer 3 → ✓ 1.0 answer 4 → ✗ 0.0 verifier scoresunit test / parser /exact-match check —no reward model advantage =each score minusthe GROUP mean(÷ group std)no critic needed update policy toward above-average answerspush up answers 1 & 3, push down 2 & 4

The group is its own baseline. Sampling several answers to one prompt gives you the expected score for free, which is the whole reason the fourth model can be deleted.

The result that made this famous is DeepSeek’s. SFT plus RL against verifiable rewards, with no learned reward model in the loop, produced a model that writes long chains of working before it answers. Nobody wrote those chains as training targets. Longer, more careful working verifiably got more answers right, so the training pushed toward it on its own.

The scope limit is the same one as before. This branch has nothing to optimise where correctness cannot be checked. Tone, helpfulness and which of two summaries is better all still need the preference methods. The two branches are complementary, and frontier labs run both.

One reframing is worth more than any algorithm choice here. If you can turn an open-ended generation task into a check against a fixed set of valid answers, you have manufactured a verifier, and this branch just became available where it was not before.

Choose a method and budget the memory for one card

Two questions decide it. What preference signal do you have, and how much memory can you spare.

What signal do you have? correctness ischeckable→ GRPO clean chosen/rejected pairs→ DPO only thumbsup / down→ KTO pairs, but VRAM tight /want single stage→ ORPO / SimPO subjective signal, frontier compute,need a learned reward model→ full RLHF / PPO

The default path for most teams runs SFT then DPO. The branches are for unpaired feedback, tight memory, and the case where a program can mark the answer.

Full RLHF stays a frontier-lab tool for broad subjective alignment. Most teams never need it, and now you know exactly which four models they are paying for.

Here is a DPO run on top of an SFT checkpoint.

# DPO on top of a supervised checkpoint.
# Pin your library versions: these trainer arguments were
# current in TRL as of August 2026 and they move quickly.
from trl import DPOConfig, DPOTrainer

cfg = DPOConfig(
    beta                        = 0.1,     # leash length, from the section above
    learning_rate               = 5e-7,    # full-parameter DPO
    num_train_epochs            = 1,       # one pass; more turns repetitive
    per_device_train_batch_size = 2,
    optim                       = "paged_adamw_8bit",
)

trainer = DPOTrainer(
    model         = policy,      # the SFT checkpoint you are training
    ref_model     = None,        # None: TRL freezes its own copy of model
    args          = cfg,
    train_dataset = pref_data,   # columns: prompt, chosen, rejected
)
trainer.train()

Line by line, assuming you have never written training code before.

  • from trl import DPOConfig, DPOTrainer pulls two names out of TRL, the Hugging Face library that implements the preference methods in this part. One holds the settings and the other runs the training.
  • DPOConfig(...) builds the settings object. Every argument inside it is a hyperparameter, meaning a value you pick before training rather than something the model learns.
  • beta = 0.1 is the leash from the section above.
  • learning_rate = 5e-7 sets how large each weight change is. 5e-7 is Python for 5 times 10 to the power minus 7, which is 0.0000005. A supervised fine-tune usually runs near 2e-5, or 0.00002, roughly forty times larger. DPO is nudging a model that already works, so it moves in smaller steps. If you train adapters instead, by passing a peft_config, raise this to about 1e-5: only a small adapter is moving, so its updates need a larger step to register at all. That is the same regime warning the part on LoRA gives, now applied to DPO. Copy the full-parameter number into an adapter run and the loss will barely move.
  • num_train_epochs = 1 is one pass over the dataset, for the reason in limit three above.
  • per_device_train_batch_size = 2 is how many preference triples go through the card at once. Watch this one. Each triple is two full sequences, chosen and rejected, so a batch of 2 puts four sequences of activations on the card, not two.
  • optim = "paged_adamw_8bit" picks the optimizer. The 8-bit part stores AdamW’s two running averages in 1 byte each rather than 4. The paged part lets that state spill out to system RAM during a spike instead of crashing the run.
  • DPOTrainer(...) builds the trainer. model= is your SFT checkpoint. ref_model=None tells TRL to make and freeze its own copy of that model as the reference.
  • train_dataset= takes a dataset with the three columns from earlier: prompt, chosen and rejected.
  • trainer.train() runs the loop.

What the reference model actually costs, in two regimes

The reference takes no gradients and gets no optimizer state. So it costs 2 bytes per parameter of frozen weights, plus its own forward-pass activations. It runs no backward pass, so it stores far fewer of those than the policy does. Where that 2 bytes lands depends entirely on which regime you are in.

Full-parameter DPO on Qwen2.5-1.5B. The policy pays 1.54 billion times 16 bytes, which is about 24.6 GB of static state. The reference adds 1.54 billion times 2 bytes, which is about 3.1 GB. Total about 27.7 GB of static state, so the reference raised the bill by about 12.5 percent. Add roughly 8 GB of activations, an order-of-magnitude figure at batch 4 and sequence 1,024, and the run wants around 36 GB against the 29.8 GiB a 32 GB card actually gives you. It does not fit, and it did not fit before the reference arrived either.

LoRA-based DPO on the same model. LoRA freezes the base and trains a small adapter beside it, which takes the policy’s static state to about 4 GB. Load a second bf16 copy of the base as the reference and you add 3.1 GB, so 7.1 GB. That is a 77 percent increase. Here the reference is most of what you are paying for, which is why removing it very nearly halves the bill.

Three ways to remove it, in order of how little work they cost.

  1. Share one frozen base. With LoRA, the reference is precisely the base model with the adapters switched off. So TRL can disable them for the reference pass rather than hold a second copy. That takes the 7.1 GB back to about 4 GB.
  2. Precompute. Set precompute_ref_log_probs=True. Every log-probability the reference will ever be asked for is computed in one pass before training starts and cached. The reference model is then dropped from memory entirely.
  3. Use a reference-free method. SimPO or ORPO delete the second model by construction.

At the 7B scale the same shape of arithmetic applies with bigger numbers, and the previous part, on QLoRA, is how you shrink that shared frozen base from 2 bytes per parameter to about 0.5. If none of those close the gap, the next question is whether a second card helps, and the next part, on multi-GPU fine-tuning, covers which strategies survive a slow link between cards.

Then prove it worked. For preference tuning specifically that means two measurements, not one: did the model get better on the thing you tuned for, and did it quietly lose something else. The final part is about exactly that.

Key takeaways

  • SFT has one direction and it points up. There is no slot in its file and no term in its loss for a worse answer, which is why relative quality has to be taught in a separate stage.
  • The data unit is a triple: prompt, chosen, rejected. The rejected answer is plausible, not bad, and both human and model labellers quietly favour the longer one.
  • Reinforcement learning means producing something, getting a score back, and being pushed toward what scored well. The policy is the model, the reward is the score, and the reward model is what produces it once you work out that a person cannot.
  • RLHF holds four models at once, about 55 GB of static state for a 1.5B model, and runs a slow sampling loop. Reward hacking is its signature failure: the score climbs while the answers get worse.
  • DPO writes the reward in terms of the policy against a frozen reference, which deletes the reward model, the value model and the loop. Its margin is the whole learning signal.
  • Beta multiplies that margin. High beta satisfies the loss early and keeps the model near its start. Low beta demands a wide separation and lets it drift. Start at 0.1 and run one epoch.
  • Where a program can mark the answer, use a verifier and GRPO instead. The group’s own average replaces the value model, and the reward cannot be flattered.

You can now

  • Explain to a colleague why adding a bad answer to an SFT file teaches the wrong lesson, from “See why supervised fine-tuning cannot say better”.
  • Define policy, reward and reward model without hand-waving, and show the arithmetic for why a human cannot sit in the training loop, from “Learn the four reinforcement-learning words you actually need”.
  • Compute a DPO loss by hand from four log-probabilities and a beta, from “Train on preferences directly with DPO”.
  • Turn a beta value up or down from a symptom you observed in the tuned model, from “Set beta, the dial that decides how far the model may move”.
  • Spot reward hacking in your own outputs and name the two defences against it, from “Reward hacking, in one concrete example”.
  • Budget a DPO run on one 32 GB card and pick the right way to shed the reference model, from “Choose a method and budget the memory for one card”.

Glossary

8-bit optimizer
AdamW with its two running averages stored in 1 byte each instead of 4, using block-wise quantisation. The algorithm and its behaviour are unchanged, and it reclaims most of 8 bytes per parameter. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
AdamW
Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
Adapter
A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
Base model
The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
Batch
A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
Beta (DPO)
The single most important dial in DPO, setting how far the trained model may drift from the frozen reference model. High keeps it close and safe, low allows stronger preference fitting and risks losing capabilities. Around 0.1 is typical. Hugging Face TRL, DPO Trainer
Catastrophic forgetting
When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
DPO
Direct preference optimization. It learns from chosen-versus-rejected pairs using one supervised-style loss, which removes the reward model and the reinforcement learning loop that RLHF needs. Rafailov et al., Direct Preference Optimization
Epoch
One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one. Supervised fine-tuning end to end
Frozen weights
Weights marked as not trainable, so they never receive an update. A frozen weight needs no gradient and no optimizer state, which removes 14 of its 16 bytes, though its activations are still stored. LoRA explained
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
GRPO
Group relative policy optimization. It samples several answers to the same prompt and treats better than the group average as the learning signal, which removes the extra scoring model that reinforcement learning normally needs. Shao et al., DeepSeekMath
Hyperparameter
A setting you choose before training rather than something the model learns, such as the learning rate, the batch size or the number of epochs. The knobs that decide whether it learns
IPO
A preference method that regularises harder toward the reference model, which curbs DPO’s tendency to over-fit a small set of preference pairs and land on an extreme policy. Azar et al., A General Theoretical Paradigm
KL penalty
A term that punishes the model for drifting too far from the version it started as. It is the leash that stops reinforcement learning from wrecking fluency in pursuit of reward. Ouyang et al., InstructGPT
KTO
A preference method that takes single answers tagged good or bad instead of matched pairs, which fits the cheap thumbs-up and thumbs-down feedback real products already collect. Ethayarajh et al., KTO
Learning rate
A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
Length bias
The tendency of both human and model annotators to rate longer answers as better, which quietly teaches a preference-tuned model to pad. If a tuned model starts waffling, suspect the data before the algorithm. Meng et al., SimPO
LoRA
Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Low-rank
A matrix is low-rank when it can be rebuilt exactly from two much smaller matrices multiplied together. LoRA assumes the update a fine-tune needs is low-rank, which is why two skinny matrices can stand in for a full one. Aghajanyan et al., Intrinsic Dimensionality of Fine-Tuning
OOM
Out of memory, the error you get when a run needs more VRAM than the card has. In training it almost always strikes where the forward pass ends and the backward pass begins, which points straight at activations. Activation memory and gradient checkpointing
Optimizer
The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
ORPO
A preference method that merges supervised fine-tuning and preference tuning into one pass and needs no reference model, so only one model sits in memory. Hong et al., ORPO
Overfitting
When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse. The six silent failure modes
Paged optimizer
An optimizer that can move its state out to ordinary system RAM when GPU memory spikes, then bring it back when the pressure passes. It turns a hard crash into a brief slowdown. Dettmers et al., QLoRA
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
Policy
The model being trained, in reinforcement-learning language. Calling it the policy just emphasises that it is the thing whose behaviour is being changed. Ouyang et al., InstructGPT
PPO
Proximal policy optimization, the reinforcement-learning algorithm used in classic RLHF. It is the machinery DPO showed you could do without. Schulman et al., Proximal Policy Optimization
Preference pair
The unit of preference data: one prompt, two answers to it, and a label saying which is better. The rejected answer is plausible rather than bad, which is what makes the comparison informative. Hugging Face TRL, DPO Trainer
Preference tuning
Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour.
QLoRA
LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
Quantisation
Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
Reference model
A frozen copy of the model you started from, kept during preference tuning so the loss can measure how far the trained model has moved. It adds 2 bytes per parameter and takes no gradients or optimizer state. Hugging Face TRL, DPO Trainer
Regularisation
Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
Reinforcement learning
Training by trial and reward: the model produces something, a score comes back saying how good it was, and the model is pushed toward whatever scored well. It differs from supervised training, where the right answer is handed over in advance. Schulman et al., Proximal Policy Optimization
Reward hacking
The model finding ways to score well without being good: flattery, padding, keyword stuffing, or exploiting a gap in the rubric. The metric climbs and reading the outputs reveals the game.
Reward model
A model trained on human comparisons to give any answer a single score. RLHF optimises against it because you cannot put a human in the loop for millions of steps. Ouyang et al., InstructGPT
RLHF
Reinforcement learning from human feedback. Collect human comparisons, train a reward model on them, then use reinforcement learning to push the model toward higher-scoring answers. It holds four models in memory, which is why it is a big-team tool. Ouyang et al., InstructGPT
SimPO
A preference method that needs no reference model and scores an answer by the average probability the model gave each of its tokens, divided by its length. That division attacks the tendency of preference data to reward longer answers. Meng et al., SimPO
Static state
Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
Supervised fine-tuning
Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
Token
The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
Verifiable reward
A reward a program can check, such as unit tests passing or a maths answer matching the known one. It cannot be flattered, which is why it resists the gaming a learned reward model invites. DeepSeek-R1
VRAM
The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model


Practical exercises

Compute exactly how much the DPO reference model costs, for both regimes

Using Qwen2.5-1.5B’s own numbers, about 24.6 GB full fine-tuning static state and about 4 GB for LoRA, treat the DPO reference model as one extra frozen bf16 copy of the base, at 2 bytes per parameter for the model’s 1.54 billion parameters. Compute the exact static-state total with the reference included, for both a full-parameter DPO run and a LoRA-based DPO run, and state each increase as a percentage. Then say which of the article’s two claims, “closer to 12 percent” for full-parameter and “roughly halves” for LoRA, your numbers actually support.

See the worked solution (opens in a new tab)

Recognise when a verifiable reward beats collecting preference pairs

A team is building a maths-tutoring fine-tune where every answer can be checked by a symbolic solver. They plan to collect several thousand human pairwise judgments, chosen versus rejected, and run DPO on the result. Using this part’s own framework for choosing a preference method, explain why this is probably the wrong investment, and what you would build instead.

See the worked solution (opens in a new tab)

Rewrite the DPO call to drop the reference from memory before training starts

The article’s DPO code sets ref_model=None, which makes TRL keep a separate frozen copy of the policy in memory as the reference for the whole run. Rewrite the DPOConfig to use precompute_ref_log_probs=True instead, and describe exactly what changes in the memory profile over the course of the job: what has to happen before training starts, and when the reference model’s memory is actually freed.

See the worked solution (opens in a new tab)

Diagnose a preference-tuned model that both forgot arithmetic and turned repetitive

A colleague runs full-parameter DPO on Qwen2.5-1.5B for three epochs, with beta at 0.1 and learning rate 1e-5. The resulting model scores well on the preference benchmark, but it now fails arithmetic it previously handled, and its answers read as repetitive and templated. Identify both mistakes in this run and give the concrete fix for each.

See the worked solution (opens in a new tab)

Choose a memory strategy for LoRA-based DPO on Qwen2.5-7B

You want to run LoRA-based DPO on Qwen2.5-7B on one 32 GB card. First, compute the static state if you naively load two separate bf16 copies of the 7.62 billion parameter base, one as the policy’s frozen backbone and one as the DPO reference, plus the roughly 0.6 GB of adapter and optimizer overhead from a rank-16 adapter. Then compute the static state if one shared frozen base instead serves as both the backbone and the reference, per the article’s own recommendation. Finally, compute the static state if that shared base is additionally quantised with QLoRA’s NF4 format. State which of the three actually fits under the card’s roughly 29.8 GiB usable ceiling once the standard safety margin is included.

See the worked solution (opens in a new tab)

Frequently asked questions

What is preference tuning and why is SFT not enough?

Preference tuning trains a model on comparisons rather than on single gold answers, so it can raise the probability of a better response and lower a worse one. Supervised fine-tuning stores one target answer per prompt, so its loss has no term at all for a second answer that is fluent but worse. Tone, helpfulness, refusal calibration and length control all need that comparison.

What is the difference between RLHF and DPO?

RLHF trains a separate reward model on human comparisons, then uses reinforcement learning to push the policy toward higher scores, holding four models in memory and generating fresh text every step. DPO proves the reward can be written in terms of the policy measured against a frozen reference, so the reward model, the value model and the sampling loop all disappear. What is left is two models and one supervised-style loss over the same pairs.

What does beta control in DPO?

Beta multiplies the margin between the chosen and rejected answers before the loss is computed, which sets how far the policy may travel from the frozen reference. A high beta satisfies the loss at a small separation, so the model stays close and learns little. A low beta demands a wide separation, so the model drifts further and risks losing capabilities it already had. Around 0.1 is the usual starting point.

Which preference method should I use?

Start with DPO if you have clean pairs, or KTO if your feedback is unpaired thumbs up and thumbs down. Move to ORPO or SimPO when memory is the binding constraint, since both remove the reference model and SimPO also divides scores by length. Use IPO when DPO is over-fitting a small set of pairs, and a verifier with GRPO when correctness can be checked by a program.

Why does my preference-tuned model produce longer answers?

Almost certainly length bias in the preference data. Both human raters and judge models tend to pick the longer answer, so the dataset encodes verbosity as a preference and the model learns it faithfully. The fixes are to curate good short answers into the chosen column and to use a method that normalises for length.

Do I need a GPU cluster for preference tuning?

Not for the offline methods. DPO’s frozen reference adds only 2 bytes per parameter of weights, so a full-parameter run grows by about 12 percent rather than doubling. On a LoRA run the reference is most of what you are paying for, so sharing one frozen base, precomputing its log-probabilities, or picking a reference-free method nearly halves the bill. Full RLHF, with its reward model, value model and sampling loop, is the one that genuinely wants a cluster.

Sources and further reading

Previous