LLM Fine-Tuning Explained: What Actually Changes Inside the Model

Inside LLM Fine-Tuning, part 1 of 15: LLM Fine-Tuning Explained: What Actually Changes Inside the Model

LLM fine-tuning sounds like it should be complicated, and the vocabulary around it certainly is. The operation underneath is small and specific. You take a model that already exists, show it examples of the behaviour you want, and let a training loop nudge the numbers inside it until it produces more of that behaviour. No component is added. No component is replaced. The numbers move.

This is the ground floor of a fifteen-part series on how that works and what it costs. It assumes you have never trained a model, and it does not use a single term it has not defined first.

By the end you will be able to say precisely which numbers change during a fine-tune, score a single prediction the way a training run scores it, work out from one multiplication whether a given model will fit on the graphics card you have, and tell apart a training job that is genuinely out of memory from a serving process that only looks like it.

Read a model’s guess the way a training loop reads it

Start with the smallest complete thing a language model does. You hand it the text “The cat sat on the” and it tells you what it thinks comes next. That single act, repeated, is the whole job, at training time and at serving time alike. Three ideas make it up.

A parameter is one number

Inside the model is a very large pile of numbers. Each one is called a parameter, or equivalently a weight. They are ordinary decimal numbers, things like 0.0271 and -0.4413. Qwen2.5-1.5B, the model this series uses for every calculation, holds 1.54 billion of them, arranged into 28 stacked layers. Those figures come from the Qwen team’s technical report.

Those numbers are the model. Every fact it can recall and every habit it has is stored as their particular values. When somebody says a model has 7 billion parameters, they are telling you how many numbers are in the pile. Nothing more.

A token is a chunk of text

Models do not read characters, and they do not quite read words either. They read tokens, which are chunks of text of roughly word size. Common words are usually one token each. Rarer ones get split, so “unhappiness” might arrive as three separate tokens. Every model ships with a fixed list of all the tokens it knows, called its vocabulary. Qwen2.5’s vocabulary holds about 151,000 entries.

The output is one score for every token in the vocabulary

Here is the part that surprises people. The model does not output a word. It outputs one number for every single token in its vocabulary, so about 151,000 numbers, all at once. Each number is a raw score meaning “how much do I like this token as the next one”.

Those raw scores are called logits. A logit can be negative, it can be large, and the whole set of them does not add up to anything in particular. It is a scoreboard. Nothing on it is a probability yet.

To turn a scoreboard into probabilities you apply softmax. Softmax does two things in order. First it makes every number positive, by raising e (about 2.718) to the power of each score. Then it makes them add up to exactly 1, by dividing each result by the total. Take a toy vocabulary of three tokens with scores of 2.0, 1.0 and 0.1:

  • Raise e to each score: 7.39, 2.72, 1.11.
  • Add those up: 11.22.
  • Divide each one by that total: 0.66, 0.24, 0.10.

Those three numbers sit between 0 and 1 and they sum to 1, so they are probabilities. The model is saying it is 66 percent confident in the first token, 24 percent in the second and 10 percent in the third. A real vocabulary has 151,000 entries instead of 3, and the arithmetic is identical.

That is the complete result of one pass through the model, which is called a forward pass: a probability for every token the model knows. Serving is what you do with those probabilities when you want text out, and the complete path an LLM takes to answer a question covers that side of the story. Training is what you do with those probabilities when you already know which token was correct.

Score that guess: what cross-entropy loss actually measures

Training needs a single number saying how badly the model just did. One number, because the entire training loop is an attempt to make one number smaller. That number is called the loss.

Building it takes three small steps.

Step 1. Look up one probability. You are training on real text, so you know which token actually came next. Say it was “mat”. Ignore the other 150,999 numbers the model produced and look up only one: the probability it gave to “mat”. Call that number p.

Step 2. Turn that probability into a penalty. You want a score that is near zero when p is close to 1, and that grows as p shrinks. The function that does exactly this is the negative natural logarithm, written as minus log p. You never need to compute a logarithm by hand. You need its shape, and this table gives you the whole shape.

Probability the model gave the correct token Loss at that position How to read it
0.99 0.01 near certain and right, almost no penalty
0.90 0.11 confident and right
0.50 0.69 a coin flip
0.10 2.30 the correct token was an outsider
0.01 4.61 confident and wrong, punished hard

Notice the asymmetry. Moving from 0.99 to 0.90 costs a tenth of a point. Moving from 0.10 to 0.01 costs more than two full points. As p heads toward zero the penalty heads toward infinity. That is deliberate. A model that is confidently wrong is worse than one that is merely unsure, and the loss is built to say so.

Step 3. Add up over the positions you are grading. A training example is a whole sequence, not one token. So you repeat steps 1 and 2 at every position you care about and add the results. Divide by the count if you prefer an average. Either way you end up with one number for the whole sequence.

That total has a name: cross-entropy. It is the loss every part of this series refers to, and the loss in every training script you will read. It has a second name as well, negative log-likelihood, which describes the same arithmetic from the other end. The likelihood is the probability the model assigned to the whole of the real text, which means every position’s probability multiplied together. The logarithm turns that long chain of multiplications into a plain sum, and the minus sign flips it so that lower means better. Cross-entropy and negative log-likelihood are the same number. When a paper uses one name and your library uses the other, nothing has changed.

Here is a version you can run right now. It needs nothing but the Python standard library.

import math

# The only input that matters: the probability the model gave
# to the token that actually came next.
for p in (0.99, 0.90, 0.50, 0.10, 0.01):
    print(f"p = {p:.2f}   loss = {-math.log(p):.2f}")

# A whole sequence is just the average over the graded positions.
probs = [0.72, 0.31, 0.95, 0.04]
print("sequence loss:", sum(-math.log(p) for p in probs) / len(probs))

The last line prints about 1.19. Three of those four positions were fine and the fourth, at p equal to 0.04, contributed most of the total on its own. That is the asymmetry doing its job.

The same thing written as a formula, with every symbol named

You will meet this idea written compactly. This is the compact form:

L(theta) = - sum over graded t of log p_theta( y_t | y_<t )

Every symbol in that line, in words:

  • L is the loss, the single number produced by step 3. Lower is better.
  • theta is the Greek letter theta, and it stands for all 1.54 billion parameters at once. So L(theta) reads as “the loss you get when the parameters are set to these particular values”.
  • t is a position in the sequence. Position 1, position 2, and so on.
  • y_t is the token that actually appeared at position t. The right answer.
  • y_<t is everything before position t. The context the model was allowed to see.
  • The vertical bar reads as the word “given”. So p_theta( y_t | y_<t ) is “the probability this model gives to the true token at position t, given the tokens before it”. That is exactly the number p from step 1.
  • log is the natural logarithm, and the minus sign at the front flips its sign so that a high probability produces a low loss.
  • “sum over graded t” is step 3. The word graded is load-bearing. You choose which positions go into that sum.
L(θ) = − Σ log pθ( yt | y<t ) “how surprised was the model by the correct next token, averaged over the tokens we grade” y<t = the context (given) y_t = the true next token (target) t ∈ graded tokens only

Every symbol on the top line is named in the list above. The three coloured notes underneath are the ones to hold on to: the context, the true next token, and the fact that only graded positions enter the sum.

That last point is the first of the two knobs this entire field turns. In a chat fine-tune you normally grade only the assistant’s tokens and skip the user’s, because you want the model to learn how to answer rather than how to ask. Choosing which positions count is called loss masking, and Part 7 is entirely about it. The second knob is which text the sequence came from in the first place, which is Part 10. There is no third knob. Fine-tuning runs the same loss function as the original pretraining run, the next-token objective Brown et al. used to train GPT-3, on different text and with a different set of graded positions.

Say exactly what fine-tuning changes, and what it cannot

Now the question in the title. A pretrained model is a fixed pile of numbers. Three things can change what comes out of it, and only one of them touches the pile.

  1. Prompting changes what you put in front of the numbers.
  2. Retrieval changes what context you put in front of the numbers.
  3. Fine-tuning changes the numbers.

That distinction explains everything downstream. Fine-tuning is the only one of the three that leaves you holding a different model file at the end. Send it the same prompt, with no special context and no clever wording, and you get a different answer, because 1.54 billion numbers are now slightly different.

What it installs reliably

Behaviour. The shape of the answer, the tone, when to stop talking, when to refuse, which of several fluent continuations to prefer. If you have a prompt that says “always reply as JSON with these four keys, never add commentary, keep it under 80 words” and the model obeys four times out of five, a fine-tune on a few thousand examples of that exact behaviour is the right tool. It moves the model’s default rather than arguing with it on every call.

What it does not install

Knowledge. This is the most expensive misconception in the field, so it is worth being blunt about it. The model learned its facts during pretraining, from trillions of tokens, over weeks, on thousands of GPUs. A supervised fine-tune on ten thousand examples is a rounding error against that. It will not install your product catalogue, your internal documentation or last quarter’s numbers in any way you can depend on. It will cheerfully produce text that sounds as though it has, which is worse than failing outright, because the failure is silent.

Zhou et al.’s LIMA result is the clearest published evidence for that split. A strong base model plus a thousand carefully written examples produced a well-behaved assistant. A thousand examples is far too few to teach anything factual, so what the fine-tune supplied was format and style. The knowledge was already sitting in the weights.

Allen-Zhu and Li arrive at the same place from the other direction. In their controlled experiments, whether a model can answer questions about a fact is settled by how that fact appeared during pretraining, and a later fine-tune on question-and-answer pairs does not rescue one that was stored badly. Their follow-up puts the storage capacity at roughly 2 bits of knowledge per parameter.

When the problem is which facts the model has in front of it, the tool is retrieval: fetch the relevant documents at question time and put them in the context window. Our walkthrough of retrieval augmented generation with pgvector builds one end to end. The two combine well in practice. Retrieval supplies the facts, and a fine-tune enforces the shape of the answer wrapped around them.

Place fine-tuning against prompting, retrieval and preference tuning

Fine-tuning is not one thing either. Modern instruction-following models are built in three stages, and each stage does something the previous one structurally cannot.

1 · Pretraining predict next token on trillions of tokens → knows language 2 · SFT imitate good answers → follows instructions, right format 3 · Preference tuning learn better-vs-worse → helpful, safe, well-calibrated, ◄ this lesson

Three stages, three different questions. Can it produce language at all, will it do as asked, and does it pick the better of two acceptable answers. The mark on the third box points at Part 13, which is where that stage is covered.
  1. Pretraining teaches raw competence: grammar, facts, reasoning patterns, the ability to continue text at all. It costs millions of dollars and months of compute. Kaplan et al. and then Hoffmann et al. mapped how that cost trades off against model size and training tokens. You will almost certainly never do this.
  2. Supervised fine-tuning, usually shortened to SFT, teaches the model to imitate good answers in a particular format. Costs run from a few dollars to a few hundred. This is what the word “fine-tuning” means in nearly every sentence you will read, and Parts 9 to 12 do it hands on.
  3. Preference tuning teaches judgement: given two answers that are both fluent and both plausible, which one is better. SFT structurally cannot express that, because SFT only ever shows the model one target answer per prompt. Helpfulness, house tone and refusal calibration get installed here. Part 13 covers it.

The practical rule is that preference tuning runs on top of an SFT checkpoint and never on a raw base model, and that pretraining is somebody else’s job. A checkpoint, in case the word is new, is just a saved copy of all the model’s weights at one point in training.

Approach What it changes Good at Bad at Cost to iterate
Prompting the input fast experiments, task framing, showing a few examples consistency at scale, long instructions eating the context window seconds
Retrieval the context fresh or private facts, citations, data that changes style, format discipline, tone minutes
Supervised fine-tuning the weights format, tone, task-specific behaviour, shorter prompts installing new knowledge, anything that changes daily minutes to hours per run
Preference tuning the weights better-versus-worse judgement, refusal calibration, verbosity anything you cannot express as a preference or verify hours

Three signals that say fine-tune, and two that say do not

The honest answer for most teams is that a better prompt plus retrieval beats a mediocre fine-tune. Three situations genuinely justify moving the weights.

  1. Your prompt has grown into a specification that you paste into every call, and the model still drifts off it.
  2. You need a specific output shape every single time, and a schema check with a retry loop is not getting you there.
  3. Your inference bill is dominated by a long system prompt that a fine-tune could bake into the weights, so every request gets shorter.

Two situations say do not, at least not yet.

  1. The knowledge you need changes weekly. Weights are the wrong place for it. Use retrieval.
  2. You cannot yet measure whether the model got better. A fine-tune will then hand you a change you have no way to evaluate. Build the eval first. Our notes on building evals for LLM features in Python are the right prerequisite, and Part 15 covers evaluating a fine-tune specifically.

Follow one weight through one training step

Everything so far has described a single forward pass. Training is a loop, and the loop has four steps. Walk one full turn of it slowly, because every memory number in the rest of this article falls straight out of what the loop has to hold at once.

Picture the model as a mixing desk with 1.54 billion sliders. Each slider is one parameter. Above the desk is a single readout showing the loss, the “how wrong was that” number from the last section. The job is to get the readout down.

  1. Forward pass. Push a batch of training text through the model and get probabilities out. A batch is simply several examples processed together, because a GPU is far more efficient that way. On the way through, each of the 28 layers computes intermediate results and hands them to the next layer. Those intermediate results are called activations. The forward pass does not discard them, because step 3 is going to need them.
  2. Loss. Compare the probabilities against the tokens that actually came next, exactly as the last section described. One number.
  3. Backward pass. For every slider on the desk, answer one question: if I push this slider up by a hair, does the readout go up or down, and how fast? That answer is one number per slider, and it is called that slider’s gradient. A large gradient means this slider matters a great deal right now. A gradient near zero means it barely moves the readout, so leave it alone. Working this out for all 1.54 billion sliders in one sweep, running backwards from the readout through the layers, is the backward pass. It is also why the activations had to be kept: you cannot work out how a layer’s input affected the loss without knowing what that input was. Part 5 draws the same idea as a ball rolling down into a valley, and the walkthrough of how a neural network learns, from a single neuron up derives it from first principles.
  4. Optimizer step. The gradient says which way to push each slider. It says nothing about how far. That decision belongs to a separate component called the optimizer. The default across the field is AdamW, and AdamW is not satisfied with the current gradient alone. For every single slider it keeps a small running record: an average of that slider’s recent gradients, and an average of their squares. It uses that record to give each slider its own step size, so a slider with a wild and noisy history moves cautiously while a steady one is allowed to move faster.

What AdamW keeps between steps is called the optimizer state, and it is the reason this section exists at all. Two running averages per parameter, plus a high-precision copy of the parameter itself, is three extra numbers for every one of the 1.54 billion sliders. Nobody notices that until they try to fit a run on a card. Part 5 explains what those running averages buy, and Part 6 walks the update rule line by line.

Then the loop starts again with the next batch. A fine-tuning run is a few thousand turns of it.

Size any fine-tuning run with one multiplication

This is the single most useful thing in the series, and the rest of this section exists only to make you believe it.

The memory a job needs is the parameter count multiplied by the bytes each parameter costs. Both halves matter, and people only ever remember the first one. The per-parameter rate swings by a factor of sixteen depending on which job you are running, which is a bigger swing than most of the model-size gaps anyone worries about.

What “a byte per parameter” means

Every parameter is a decimal number, and a computer stores a decimal number using a fixed number of bits. Eight bits make one byte. Two formats matter here.

  • bf16 uses 16 bits, which is 2 bytes. It is the standard format for a model’s weights during training.
  • fp32 uses 32 bits, which is 4 bytes. Twice the size, and it holds far more decimal places.

Part 4 covers why several different 16-bit formats exist and why bf16 won for training. For this article, 2 bytes and 4 bytes is everything you need.

Count what the training loop holds at once

Go back through the four steps and count. Everything the loop needs has to sit in the graphics card’s own memory, called VRAM, which is a separate and much smaller pool than your machine’s system RAM. At the moment the optimizer takes its step, all of the following are in VRAM at once.

  1. The weights themselves, in bf16. One number per parameter, 2 bytes each. That is 2 bytes per parameter.
  2. The gradients from the backward pass. Exactly one per weight, also in bf16. Another 2 bytes per parameter.
  3. The optimizer state. AdamW keeps three fp32 numbers for every parameter: a high-precision master copy of the weight, the running average of its gradient, and the running average of its squared gradient. Three numbers at 4 bytes each is 12 bytes per parameter.

2 plus 2 plus 12 is 16. That is the figure quoted everywhere: standard mixed-precision AdamW costs 16 bytes per parameter. Mixed precision is the name for the arrangement just described, where the fast arithmetic happens in bf16 while a 32-bit copy of each weight is kept alongside it, so that thousands of tiny updates accumulate instead of rounding away to nothing. Part 2 derives all sixteen of those bytes one tenant at a time, and Part 5 explains why the field cheerfully pays the optimizer’s twelve.

FULL FINE-TUNING — 16 bytes / param 2 2 4 4 4 weights bf16 grads bf16 master fp32 Adam m fp32 Adam v fp32 └─ optimizer states = 12 bytes ─────────────────────────────┘ LoRA — base frozen, ~2 bytes / param 2 + tiny adapter (grad+optim on <1% of params) QLoRA — base in 4-bit NF4, ~0.5 bytes / param ½ + adapter · this is what fits a 7–13B onto one card

Read the top bar left to right: two bytes of weights, two of gradients, then twelve of optimizer state split into three four-byte pieces. The two short bars underneath are the second and third rungs of the ladder in the last section of this article.

There is a fourth thing in memory and it does not obey the per-parameter rule at all: the activations from step 1. Their size depends on how many sequences you push through at once and how long those sequences are, so they cannot be written as bytes per parameter. This is why the series always calls the 16-byte figure the static state, meaning weights plus gradients plus optimizer state, the part that stays the same size all run long. Activations are counted separately, every time. Part 3 is about them.

Those four things, weights, gradients, optimizer state and activations, are the four tenants of the card. The word is worth keeping. It is the frame Part 2 builds on, and it is what lets you tell two identical-looking memory problems apart a little further down this page.

Why serving the same model is so much cheaper

Serving is a different job with a different set of tenants. There is no backward pass, so there are no gradients. There is no optimizer, so there is no optimizer state. Only the weights have to be resident.

INFERENCE (one 27B) weights (FP8 = 1 byte) — no gradients — no optimizer + KV cache (grows w/ context) ≈ 1 byte / param FULL FINE-TUNE (the 1B) weights bf16 (2) gradients (2) optimizer states (12) activations ≈ 16 bytes / param

Count the coloured rows on each side. Serving keeps one thing plus a cache. Training keeps four. Everything below the top row on the right exists only to work out how each weight should change.

Serving does add one thing back. While a model is generating an answer it saves some intermediate work for every token it has already produced, so it does not have to redo that work for the next one. That saved pile is called the KV cache. It grows with the length of the conversation rather than with the size of the model, and it comes back in the diagnostic further down.

So serving costs 2 bytes per parameter with the weights in bf16, or about 1 byte if they have been compressed into an 8-bit format. Training the same model costs 16. That sixteen-fold gap produces results which look impossible until you do the multiplication.

27B · FP8 inference 1 byte/param — weights only  27 GB of weights 1B · full fine-tune 16 bytes/param + activations ~22 GB total weights (2B) gradients (2B) optimizer (12B) activations 27 ≈ 22. The 27B is bigger, but its per-param cost is 16× cheaper.

The bar on top is a model twenty-seven times larger than the one below it, and the two land in roughly the same place. Per-parameter rate beats parameter count.

A 27 billion parameter model being served at 1 byte per parameter needs about 27 GB for its weights. A 1 billion parameter model being fully fine-tuned needs 16 GB of static state plus its activations, which lands around 22 GB in total. The larger model sits in the same memory bracket as the smaller one, because its per-parameter rate is sixteen times lower.

The running example, worked in full

Qwen2.5-1.5B has 1.54 billion parameters. Full fine-tuning it under standard mixed-precision AdamW:

1.54e9 parameters x 16 bytes = 24.64e9 bytes = about 24.6 GB of static state

Then add roughly 8 GB of activations, at a batch of 4 sequences of 1,024 tokens each. Treat that 8 as an order of magnitude rather than a measurement, because the exact bytes depend on which intermediate results the framework decides to keep, and different frameworks and settings keep different sets. So the run is about 33 GB in total: 24.6 GB of static state plus roughly 8 GB of activations.

The target hardware for every calculation in this series is one 32 GB consumer card. The run wants 33 and the card offers 32, so it does not fit. The gap is also worse than that single gigabyte suggests, once the units are handled honestly.

GB and GiB, because the difference decides this case

Two different units share the same spoken name and differ by about 7 percent, which happens to be exactly the size of the margin you are working in.

  • A gigabyte, written GB, is 1,000,000,000 bytes in the decimal sense, which is 10 to the power 9. This is what you get when you multiply parameters by bytes, as above.
  • A gibibyte, written GiB, is 1,073,741,824 bytes, which is 2 to the power 30. It is about 7 percent larger than a GB.

GPU vendors, drivers and nvidia-smi all work in the binary unit, even when the box says GB. So a card sold as 32 GB is really 32 GiB of chips, and once the driver and the CUDA context have taken their share you get roughly 29.8 GiB to actually use.

Now compare like with like. Divide each decimal figure by 1.074 to get gibibytes.

  • 24.6 GB of static state is 22.9 GiB.
  • Roughly 8 GB of activations is roughly 7.5 GiB.
  • Together that is about 30.4 GiB, against about 29.8 GiB of usable card.

It does not fit, and that is before you have left any safety margin at all. Leave 2 to 3 GB of headroom in every estimate you make. Memory fragmentation and one unusually long batch both live in that last gigabyte, and both of them produce the same out-of-memory crash at 3 a.m.

The five-second screen

All of this collapses into one multiplication you can do in your head. Parameters times 16 gives decimal GB. Divide by 1.074 for gibibytes. Then add activations separately, because nothing in that multiplication accounts for them.

Model Parameters Full fine-tune static state at 16 bytes each The same in GiB Fits on one 32 GB card?
Qwen2.5-0.5B 0.49B 7.8 GB 7.3 GiB yes, with room left for activations
Qwen2.5-1.5B 1.54B 24.6 GB 22.9 GiB no, once roughly 8 GB of activations is added
Qwen2.5-7B 7.62B 122 GB 114 GiB no, and this is the reason LoRA exists
14B class 14B 224 GB 209 GiB no
70B class 70B 1,120 GB 1,043 GiB no, this is a multi-GPU job and Part 14 covers it

Every entry in the third column is one multiplication. Every entry in it is also static state only. None of them includes activations, and forgetting that is the most common sizing mistake in the field.

Diagnose a card that is using far more memory than the model’s size

Here is a situation you will meet within a week of touching a GPU. Someone loads a 1B-class model onto a 32 GB card, and nvidia-smi reports 22 GB in use. The model’s weights are a couple of gigabytes. Something looks badly wrong.

There are two common causes. They produce an identical symptom and they call for completely different fixes.

A · you’re FINE-TUNING it 16 bytes/param → ~16 GB + activations = the ~22 GB you saw fix: LoRA → ~4 GB QLoRA → ~2 GB (drops grad + optimizer off the base) B · you’re SERVING it (vLLM) 1B weights ≈ 2 GB — tiny. engine pre-grabs ~90% of the card for an empty KV-cache pool fix: gpu_memory_utilization ↓ max_model_len ↓ the 22 GB is reservation, not the model

The same symptom on the left and the right, from two unrelated causes. The left-hand column is memory genuinely in use. The right-hand column is a reservation the serving engine made at startup.

Cause A: the model is being trained. Then the 22 GB is the four tenants doing precisely what the last section said they would do, at 16 bytes per parameter plus activations. Nothing is broken and no memory is being wasted. If you need that number smaller, the move is to remove tenants entirely, which is what the ladder in the next section does.

Cause B: the model is being served through an engine such as vLLM, which grabs most of the card at startup to hold its KV cache pool, and keeps hold of it whether or not any request has arrived. The 22 GB is a reservation, not consumption. The fix is a configuration number: lower gpu_memory_utilization, or lower max_model_len so that each sequence needs less cache. Our part on continuous batching and paged attention explains why serving engines reserve in a block rather than allocating on demand.

Telling the two apart takes one question: is anything on this card computing gradients? If yes, it is A. If the process is a serving engine, it is B. The symptom is identical and the diagnosis is trivial, once you know that the two jobs keep different lists of tenants.

Choose between full fine-tuning, LoRA and QLoRA

Almost every practical fine-tuning decision is a choice between three rungs of one ladder. Each rung removes a cost that the rung below it paid.

Rung 1: full fine-tuning

Every parameter is trainable, so every parameter carries a gradient and an optimizer state. That is where all sixteen bytes come from. You get maximum flexibility and three costs with it. It is the most expensive option in memory. It produces a complete multi-gigabyte model copy for every task you tune. And because it moves every weight, it can quietly degrade abilities the base model already had.

That degradation has a name, catastrophic forgetting. The thing to know now is that your training loss will look excellent the whole time it is happening, because the loss only measures performance on your data. Part 15 covers how to catch it before your users do.

Rung 2: LoRA

LoRA starts from a question. If the model already knows the language and you are only teaching it a behaviour, does the change to the weights really need 1.54 billion degrees of freedom? The empirical answer turned out to be no.

So LoRA freezes the base model. A frozen weight is one the training loop has been told never to change. It still takes part in the forward pass, so it still costs its 2 bytes of storage. It receives no gradient and gets no optimizer state, so 14 of its 16 bytes simply vanish.

Then, beside each weight matrix it targets, LoRA adds a small trainable patch. The arithmetic is worth doing once. Qwen2.5-1.5B moves rows of 1,536 numbers between its layers, so many of its weight matrices, which are simply rectangular grids of numbers, come out at 1,536 by 1,536. That is 2,359,296 numbers in one matrix. Rather than learn a change to all of them, LoRA learns two skinny matrices instead: one that is 1,536 by 16, and one that is 16 by 1,536. Count them: 24,576 plus 24,576, so 49,152 numbers, which is about 2 percent of the original. That 16 is called the rank, written r, and it is the main knob you turn. Multiply the two skinny matrices together and the result has exactly the same shape as the original, so the learned patch can be laid straight on top of it.

Two things change as a result. The static state for a Qwen2.5-1.5B fine-tune drops from about 24.6 GB to about 4 GB. And the file you ship at the end is just those skinny matrices, a few tens of megabytes, which is called an adapter. Twenty different task adapters can sit beside one copy of the base model instead of twenty multi-gigabyte checkpoints. Part 11 is the full treatment.

Rung 3: QLoRA

LoRA removed the gradients and the optimizer state for the frozen base. It could not remove the base’s own 2 bytes per parameter, because the forward pass still has to run through those weights. QLoRA attacks that last cost by storing them in 4 bits instead of 16.

Squeezing a number into fewer bits is called quantisation. You give up decimal places to buy space. QLoRA uses a specific 4-bit code called NF4, designed around the fact that a trained model’s weights cluster near zero in a bell-curve shape, so it spends its sixteen available codes where the weights actually are rather than spreading them evenly. That takes the base from about 2 bytes per parameter down to about 0.5.

The catch sits on the critical path. A 4-bit number cannot be multiplied directly, so before each matrix multiply the values have to be expanded back into bf16. That expansion is called dequantisation, and it costs real time on every forward pass and every backward pass. QLoRA is a fitting tool. Reach for it when a model would not otherwise fit at all. Running it on a model that already fits buys you nothing and costs you speed, and Part 12 puts a number on that slowdown.

Method Trainable parameters Bytes per parameter on the base Qwen2.5-1.5B static state Qwen2.5-7B static state
Full fine-tuning 100 percent 16 about 24.6 GB about 122 GB, impossible on one card
LoRA, r=16, all linear layers about 1 percent 2 about 4 GB about 18 GB
QLoRA, NF4 base about 1 percent about 0.5 works, but slower for no gain about 8 GB, with headroom

Every figure in that table is static state only, meaning weights, gradients and optimizer state. Activations are in none of them, and activations do not shrink under LoRA or QLoRA. Both methods run the same forward and backward pass as a full fine-tune, over the same batch and the same sequence length, so the same roughly 8 GB is still owed on top of every row in the 1.5B column. Add it back before comparing any row against a 32 GB card. At the 7B scale the activation figure moves with batch size and with which memory-saving options are on, so there is no single number worth quoting. Part 3 covers those options.

One honest limit on the ladder before you leave it. For format, tone and instruction following, LoRA usually matches full fine-tuning at a fraction of the cost. For teaching a genuinely new capability such as heavy mathematics or code, Biderman et al. measured LoRA underperforming full fine-tuning, while also forgetting less of what the base model already knew. Both halves of that finding are real, and Part 11 covers the trade in full.

Key takeaways

  • A model is a pile of ordinary numbers called parameters. Fine-tuning changes those numbers. Prompting and retrieval change what you put in front of them, and leave the model file untouched.
  • Cross-entropy loss is the probability the model gave to the token that actually came next, run through a negative logarithm. High probability gives a low score, and being confidently wrong is punished far harder than being unsure. Negative log-likelihood is the same number under another name.
  • Fine-tuning uses exactly the same loss as pretraining. The only two decisions are which positions get graded (loss masking, Part 7) and which text the sequence came from (data, Part 10).
  • It installs behaviour reliably and knowledge poorly. Use retrieval for facts and a fine-tune for format, tone and task discipline.
  • Memory is bytes per parameter multiplied by parameters. The rate runs from about 1 for 8-bit serving to 16 for standard mixed-precision AdamW full fine-tuning, and 12 of those 16 belong to the optimizer.
  • Multiply parameters by 16 to screen a full fine-tune in five seconds. Then add activations separately, convert to GiB by dividing by 1.074, and keep 2 to 3 GB of margin.
  • The three rungs are full fine-tuning, LoRA and QLoRA. Each removes a cost the one below it paid, every static-state figure quoted for them excludes activations, and QLoRA is a fitting tool rather than a speed tool.

You can now

  • Explain what a logit is, what softmax does to it, and read a model’s output as probabilities, from “Read a model’s guess the way a training loop reads it”.
  • Compute the loss for a single prediction by hand and explain why a confident wrong answer costs so much more than an unsure one, from “Score that guess: what cross-entropy loss actually measures”.
  • Decide whether a given problem wants a better prompt, retrieval, a fine-tune or preference tuning, from “Place fine-tuning against prompting, retrieval and preference tuning”.
  • Name the four things a training loop holds in GPU memory at once and say which one is not measured per parameter, from “Follow one weight through one training step”.
  • Estimate the static state of any full fine-tune from its parameter count, convert it to GiB and compare it honestly against a real card, from “Size any fine-tuning run with one multiplication”.
  • Separate a training job that genuinely needs the memory from a serving engine that has merely reserved it, using one question, from “Diagnose a card that is using far more memory than the model’s size”.

Glossary

Activation
Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
Activation memory
The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
AdamW
Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
Adapter
A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
Attention
The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
Backpropagation
The procedure that computes a gradient for every weight in one sweep backwards through the model, from the loss at the end to the first layer. It works by applying the chain rule one step at a time. The calculus behind backpropagation
Backward pass
Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
Base model
The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
Batch
A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
bf16
A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
Catastrophic forgetting
When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
Causal language model
A model that only ever looks left, predicting each token from the ones before it. Every model in this series is one, as opposed to a model such as BERT that reads in both directions. The complete inference path
Causal mask
The rule inside attention that stops each position seeing anything later in the sequence. It is built into the architecture, you never set it, and it is what lets one pass train next-token prediction at every position without cheating. Masking in LLM training
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
Collator
The small piece of code that takes several examples and assembles one batch out of them: padding them to a common length, building the attention mask, and setting padded labels to -100. Masking in LLM training
Cross-entropy
A score for how wrong a prediction was. It is small when the model gave high probability to the token that actually came next, and large when it did not. How a neural network learns
De-quantisation
Expanding quantised numbers back to a wider format so arithmetic can be done on them. QLoRA de-quantises its 4-bit weights to bf16 for every matrix multiply, and that extra work is where its speed cost comes from. QLoRA explained
DPO
Direct preference optimization. It learns from chosen-versus-rejected pairs using one supervised-style loss, which removes the reward model and the reinforcement learning loop that RLHF needs. Rafailov et al., Direct Preference Optimization
Forward pass
Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
Frozen weights
Weights marked as not trainable, so they never receive an update. A frozen weight needs no gradient and no optimizer state, which removes 14 of its 16 bytes, though its activations are still stored. LoRA explained
Full fine-tuning
Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
GiB
A gibibyte, 2 to the power 30 bytes, which is how drivers and vendors report GPU memory. Byte arithmetic such as 16 times the parameter count lands in decimal GB instead, so 24.6 GB is about 22.9 GiB and a 32 GB card gives roughly 29.8 GiB. Training memory and the 16 bytes per parameter
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
Hidden size
The width of the list of numbers that flows between layers, written d_model. The example model’s is 1,536, and it multiplies straight into activation memory. Inside one transformer block
Holdout
Data deliberately kept out of training so you can measure the model on something it has never seen. A number measured on data the model trained on is not evidence of anything. Building a holdout you can trust
Inference
Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
KV cache
The keys and values of past tokens, kept during generation so they are not recomputed for every new token. It exists only at inference; training has no generation loop and therefore no KV cache. Continuous batching and paged attention
Layer
One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
Logits
The raw scores a model produces for every possible next token, before they are turned into probabilities. One number per token in the vocabulary, and higher means the model favours that token. The complete inference path
LoRA
Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Loss masking
Deciding which positions in an example count toward the loss. For instruction tuning you grade only the assistant’s response and mark everything else with -100, so it contributes no loss and no gradient. Masking in LLM training
Low-rank
A matrix is low-rank when it can be rebuilt exactly from two much smaller matrices multiplied together. LoRA assumes the update a fine-tune needs is low-rank, which is why two skinny matrices can stand in for a full one. Aghajanyan et al., Intrinsic Dimensionality of Fine-Tuning
Matrix multiply
The operation that dominates all the arithmetic in a transformer: multiply a block of inputs by a block of weights to get a block of outputs. Usually shortened to matmul. Inside one transformer block
Memory tenant
One of the four things sharing the card during training: weights, gradients, optimizer state and activations. All four are live at the same moment, so peak memory is their sum rather than the largest of them. The four tenants and the 16 bytes per parameter
Mixed precision
Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
Negative log-likelihood
Another name for the same score as cross-entropy in this setting. Take the probability the model gave the correct token, take its logarithm, and flip the sign, so a low probability becomes a large penalty. PyTorch, CrossEntropyLoss
Neural network
A stack of layers that multiply their input by learned weights and pass the result on. Training means adjusting those weights until the output is closer to what you wanted. How a neural network learns
Next-token prediction
The one thing a language model is trained to do: given the text so far, put a probability on every possible next token. Fine-tuning uses exactly the same objective as pretraining.
NF4
The 4-bit storage format QLoRA uses. Its 16 values sit at the quantiles of a bell curve rather than at even spacings, because model weights cluster near zero. It has no exponent or mantissa fields and nothing computes in it directly. Dettmers et al., QLoRA
OOM
Out of memory, the error you get when a run needs more VRAM than the card has. In training it almost always strikes where the forward pass ends and the backward pass begins, which points straight at activations. Activation memory and gradient checkpointing
Optimizer
The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
Optimizer state
The numbers an optimizer keeps between steps, such as running averages of past gradients. Under standard mixed-precision AdamW it is 12 of the 16 bytes per parameter, which makes it the largest memory tenant. Rajbhandari et al., ZeRO
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
Perplexity
A measure of how surprised a model is by a piece of text, worked out from its loss. Lower is better, it is the cheapest evaluation there is, and it says nothing about whether the model is good at your task. The evaluation ladder
Preference tuning
Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
Pretraining
The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result.
QLoRA
LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
Quantisation
Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
RLHF
Reinforcement learning from human feedback. Collect human comparisons, train a reward model on them, then use reinforcement learning to push the model toward higher-scoring answers. It holds four models in memory, which is why it is a big-team tool. Ouyang et al., InstructGPT
Softmax
The step that turns a list of raw scores into probabilities that add up to 1. Larger scores get larger probabilities, and the gaps between the scores decide how confident the result looks. The complete inference path
Static state
Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
Supervised fine-tuning
Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
System prompt
A message at the start of a conversation that sets the model’s role and rules. During training it is masked out of the loss along with the user’s turns. Masking in LLM training
Tensor
A grid of numbers with any number of dimensions. One number is a scalar, a row of them is a vector, a table is a matrix, and anything past that is still a tensor with more dimensions. How a neural network learns
Token
The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
VRAM
The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes.


Practical exercises

Screen a 3B model against a 32 GB card

A colleague wants to full fine-tune a 3 billion parameter model (not a model this series has used, just a round number) under standard mixed-precision AdamW. Using the one-line cost model from this part, 16 bytes per parameter, compute the static state in decimal GB, then convert it to GiB and compare it against the roughly 29.8 GiB a 32 GB card actually gives you. Does it fit as static state alone, before a single activation tensor is counted?

See the worked solution (opens in a new tab)

Find what the ladder table leaves out

Take the same 3 billion parameter model. Using the bytes-per-parameter figures from the three-rung ladder, about 2 for LoRA and about 0.5 for QLoRA, estimate the static state under each method. Then say plainly what both numbers are missing, and why that omission is a bigger deal at 3B than it was for the 1.5B example this part actually measured.

See the worked solution (opens in a new tab)

Diagnose a serving-memory alarm

A teammate loads your fine-tuned 1.5B model onto a spare 32 GB card purely to serve it, and is alarmed that nvidia-smi reports about 22 GB in use for a model whose bf16 weights are only a couple of gigabytes. They suspect the checkpoint got saved in some corrupted training state. This part gives two genuine causes for that symptom. Say which one actually applies here, given that the card is serving rather than training, and describe one check that would confirm it inside a minute.

See the worked solution (opens in a new tab)

Choose a path for a 7B maths fine-tune on one card

You need to teach a 7B model heavy multi-step maths reasoning, and you have exactly one 32 GB card. Full fine-tuning needs about 122 GB of static state, ruled out immediately. QLoRA gets static state down to about 8 GB, which fits comfortably. But this part’s FAQ notes that LoRA-family methods measurably underperform full fine-tuning specifically on maths and code, while forgetting less of what the base model already knew. Given only the one card, name the two realistic paths forward and what each one actually costs you. Do not invent a third option that lets you have full fine-tuning’s quality on this hardware.

See the worked solution (opens in a new tab)

Frequently asked questions

What actually changes inside a model when you fine-tune it?

The parameters, which are the billions of ordinary decimal numbers the model is made of. Nothing is added and no component is replaced. The training loop measures how wrong each prediction was, works out which direction every number should move to make that measurement smaller, and moves it a little. Send the same prompt afterwards and you get a different answer, because the numbers are different.

What is cross-entropy loss, in plain language?

It is a score for how wrong a prediction was. Look up the probability the model gave to the token that actually came next, take the natural logarithm and flip the sign. A probability of 0.90 scores 0.11 and a probability of 0.01 scores 4.61, so confident and wrong is punished far harder than unsure. Negative log-likelihood is another name for the same number, and the two are interchangeable.

Does fine-tuning teach a model new facts?

Not reliably. Knowledge comes from pretraining over trillions of tokens, and a supervised fine-tune on a few thousand examples mostly teaches format, style and task behaviour. When you need the model to know something specific and current, retrieval is the correct tool. A fine-tune can then teach it to use those retrieved facts in the shape you want.

Should I fine-tune or use RAG?

Use retrieval when the problem is which facts the model sees, and fine-tuning when the problem is how the model behaves. They are complementary, and many production systems run both: retrieval supplies the context and a fine-tune enforces the output shape and tone. If your knowledge changes weekly, keep it out of the weights.

How much VRAM do I need to fine-tune a 7B model?

Full fine-tuning a 7B model needs roughly 122 GB of static state at 16 bytes per parameter, so it will not fit on one consumer card. LoRA brings the static state to about 18 GB because the frozen base only pays for its own weights, and QLoRA brings it to about 8 GB by storing that base in 4 bits. Activations sit on top of all three figures, and on a 32 GB card plain LoRA is the simpler and faster choice.

Is LoRA as good as full fine-tuning?

For format, style and instruction adaptation it usually matches full fine-tuning at a fraction of the cost. For teaching a genuinely new capability such as heavy maths or code, Biderman et al. found that LoRA underperforms full fine-tuning while also forgetting less of what the base model already knew, so full fine-tuning still earns its cost on those tasks. Part 11 covers the trade in detail.

Sources and further reading

Previous