LoRA Explained: Freeze the Model, Learn a Low-Rank Patch

Inside LLM Fine-Tuning, part 11 of 15: LoRA Explained: Freeze the Model, Learn a Low-Rank Patch

LoRA fine-tuning is the reason a 7 billion parameter model can be trained on one consumer graphics card. The idea underneath is small. Leave the model’s own numbers frozen. Learn a tiny patch beside them. Add the patch in when the model runs.

Part 1 sketched that in a paragraph. This part builds it properly. It starts from what a grid of numbers is, and what it means for one to be low-rank. No linear algebra is assumed, and none is needed.

By the end you will be able to rebuild a grid of numbers from two smaller grids by hand. You will be able to count exactly how many numbers a LoRA adapter trains. You will know why one of its two grids starts at zero, how to set rank and alpha with a reason instead of a copied config, and which parts of the memory bill LoRA cuts and which it leaves exactly where they were.

See the three bills full fine-tuning keeps paying

Full fine-tuning means training every number in the model. It works, and it is expensive in three separate ways. Seeing all three stacked up is what makes the rest of this article feel inevitable rather than clever.

1 · MEMORY 16 bytes / param to train. Qwen-1.5B 24.6 GB optimizer alone = 18.5 GB 2 · STORAGE / TASK every fine-tune = a full model copy. 10 tasks 10 × 3 GB = 30 GB of near-identical weights 3 · FORGETTING moving every weight erodes prior skills. math · code quietly degrade the forgetting hazard

Three separate costs with three separate victims. LoRA goes after all three at once, which is why it became the default rather than one option among several.

Bill one is training memory. Under standard mixed-precision AdamW, every parameter costs 16 bytes: 2 for the weight, 2 for its gradient, and 12 for the optimizer’s records. A gradient is one number per weight saying which way to nudge it. The optimizer is the component that decides how far. For Qwen2.5-1.5B that comes to about 24.6 GB. That figure is static state only, meaning weights, gradients and optimizer records. Activations are extra, and they are the subject of Part 3. Part 2 derives all sixteen of those bytes one at a time.

Bill two is storage. Every task you tune produces a complete new copy of the model. Qwen2.5-1.5B in bf16, the compact 16-bit number format training uses, is about 3 GB on disk. Ten tasks means ten copies, so about 30 GB of files that are almost identical to each other.

Bill three is forgetting. Moving every weight to teach one format can quietly erode abilities the model already had. Arithmetic gets worse. Code gets worse. That is called catastrophic forgetting, and the cruel part is that your training loss looks excellent the whole time, because the loss only measures performance on your own data.

Full fine-tuning still earns its cost in two cases. The first is teaching a genuinely new capability, such as a new language or serious mathematics. The second is when you ship one flagship model and storage never comes up. For style, format, tone and instruction following, which is most of what supervised fine-tuning actually does, there is a much cheaper route.

Rebuild a big grid from two skinny ones, and read off its rank

This is the one idea the rest of the article rests on. It takes five minutes and some arithmetic you can check on paper.

A matrix is a grid of numbers in rows and columns. That is the whole definition. A weight matrix inside a model is a rectangle of the ordinary decimal numbers Part 1 called parameters. Nothing mysterious lives in the word.

A grid that is secretly a multiplication table

Here is a 4 by 4 grid. Sixteen numbers.

2 5 1 3
4 10 2 6
6 15 3 9
8 20 4 12

Look at it before reading on. There is a pattern. Row 2 is row 1 doubled. Row 3 is row 1 tripled. Row 4 is row 1 times four.

So the grid is a multiplication table. Write one short list down the side and one along the top.

  • Down the side: 1, 2, 3, 4.
  • Along the top: 2, 5, 1, 3.

Every cell is its row’s number times its column’s number. Row 3, column 2 is 3 times 5, which is 15. Row 4, column 4 is 4 times 3, which is 12. Check any cell you like. They all work.

Now count what you stored. The grid holds 16 numbers. The two lists hold 4 plus 4, so 8. You stored half as much and lost nothing. The grid can be rebuilt exactly, cell by cell.

Grids that need more than one table

Most grids are not that obliging. Take a second multiplication table, built the same way from two more lists.

  • Down the side: 1, 0, 0, 1.
  • Along the top: 10, 0, 0, 10.

That table has 10 in its four corners and 0 everywhere else. Now add the two tables together, cell by cell.

12 5 1 13
4 10 2 6
6 15 3 9
18 20 4 22

Row 1, column 1 is 2 plus 10, which is 12. Row 4, column 4 is 12 plus 10, which is 22. The middle of the grid is untouched, because the second table was zero there.

This new grid is no longer a multiplication table. No single pair of lists can produce it. Two pairs can, and you just watched them do it.

That count has a name. The rank of a grid is how many multiplication tables you have to add together to rebuild it exactly. The first grid has rank 1. The second has rank 2. A 4 by 4 grid can need as many as 4. In general, a grid never needs more tables than the length of its shorter side.

Where the saving comes from, and where it stops

Now count storage again, because the whole method lives here.

Rebuilding a square grid that is d by d from r tables needs r pairs of lists. Each pair is d numbers down the side and d along the top, so 2d numbers. All r tables together cost r times 2d.

Set d to 4 and try it.

  • One table: 8 numbers against the grid’s 16. You save half.
  • Two tables: 16 numbers. Exactly break-even.
  • Three tables: 24 numbers. Worse than writing the grid out.

So on a tiny grid the trick barely pays. Break-even sits at rank d divided by 2. On a large grid it pays enormously, because d is large and the rank you use is small. Qwen2.5-1.5B has weight grids that are 1,536 by 1,536, a shape the Qwen2.5 technical report sets out, so break-even sits at rank 768. Real LoRA runs use rank 16.

A grid is called low-rank when a small number of tables rebuilds it. Low compared to the length of its shorter side. That is the operational definition, and it is the only one you need.

The bet LoRA actually makes

Here is the move that decides whether any of this is useful, and the old shorthand often gets it wrong.

Nobody claims a trained model’s weight grids are low-rank. They are not. The claim is about the change.

Fine-tuning takes a weight grid and turns it into a slightly different grid. Subtract one from the other and you get a third grid: the difference, written delta W, which is the patch you would have to add to the original to get the fine-tuned version. LoRA’s bet is that delta W is low-rank, or close enough to it that the gap does not matter.

That bet has published evidence behind it. Aghajanyan and colleagues asked how many numbers you genuinely have to be free to choose in order to fine-tune a model well. They tuned a small set of numbers and spread them back across the full weight set using a fixed random recipe. A model reached 90 percent of its full fine-tuning score while only a few hundred numbers were actually free. If adaptation really needs a few hundred free choices, forcing all 1.54 billion weights to move is enormous overkill.

Attach a low-rank patch to a frozen weight

Two steps turn the last section into a training method.

Step 1. Freeze the base. A frozen weight is one the training loop has been told never to update. It still runs in the forward pass, because the frozen weights are what computes the answer. It just never changes. Freezing removes its gradient and its optimizer records, so 14 of its 16 bytes vanish. The 2 bytes of the weight itself stay, because the forward pass still has to multiply by it.

Step 2. Add a patch beside it. Instead of storing delta W as a full grid, store it as r pairs of lists. Stack the r down-the-side lists into one skinny grid and call it B. Stack the r along-the-top lists into another and call it A. Multiply B by A and you get every one of those multiplication tables, added together, in a grid the same shape as W.

So the layer now does two things with the same input, and adds the results.

  1. Push the input x, a list of 1,536 numbers, through the frozen grid W, exactly as before. Call the result Wx.
  2. Push the same x through A. A has r rows, so it turns 1,536 numbers into 16. That squeeze is the bottleneck, and r is how wide it is.
  3. Push those 16 numbers through B, which expands them back out to 1,536.
  4. Multiply that result by a fixed number, alpha divided by r. The next section but one explains what that does.
  5. Add it onto Wx.

Written compactly, that is:

h = Wx + (alpha/r) * B * A * x

Every symbol in words:

  • x is the layer’s input, a list of numbers.
  • W is the original weight grid. Frozen. It never changes.
  • A is the skinny grid that squeezes x down to r numbers. Trainable.
  • B is the skinny grid that expands those r numbers back to full width. Trainable.
  • alpha and r are two numbers you choose. Their ratio scales the patch.
  • h is what the layer hands to the next layer. It is the sum of the old behaviour and the new patch.
x W (frozen) d × k · no gradient · ❄ A r × k · random B d × r · zeros the “bottleneck” — squeeze through rank r, then back out + × (α/r) h h = Wx + (α/r)·B·A·x — only B and A are trained

Both branches read the same input. The top one is the frozen base model, untouched. The bottom one squeezes the input down to r numbers, expands it back out, scales it, and adds it on.

Notice the last word of that formula’s structure. The patch is added. It is not spliced into the middle of the layer and it is not glued onto the end. It is summed into the same place, in the same shape. Two useful things fall out of that later: you can fold the patch permanently into W, and you can attach and detach it at will.

The pair of skinny grids, saved as a file, is called an adapter. For this model it is tens of megabytes against the base model’s 3 GB.

Say why one of the two grids starts at all zeros

Every trainable number has to start somewhere. Its starting value is its initialisation. In LoRA, A gets small random numbers and B gets all zeros. That asymmetry is doing real work.

Follow the arithmetic. B is all zeros. Anything multiplied by zero is zero. So B times A times x is a list of 1,536 zeros, whatever x happened to be. The scale does not rescue it either, because alpha over r times zero is still zero.

So at step zero the patch contributes nothing at all. The formula collapses to h = Wx. The model behaves exactly like the base model, to the last decimal place. The adapter starts life as a no-op.

Three consequences follow, and all three are practical.

  1. Training begins at a known-good model. You never start from a randomly disturbed copy of something that worked. You start from the working model and move away from it deliberately.
  2. Your first loss value is a free sanity check. Step zero should report the base model’s own loss on your data. If it reports something wildly different, the adapter is not the problem. Look at your data format, your chat template or your masking instead.
  3. It is why LoRA wants a much larger learning rate. A full fine-tune nudges numbers that are already good, so big steps break things. LoRA grows a patch out of literally nothing, so small steps leave it sitting near nothing.

Why not set both grids to zero, or both to random

Both questions have clean answers, and answering them is the fastest way to see what the choice buys.

Set both to zero and nothing ever learns. Part 3 showed that a weight’s gradient is the message arriving from above multiplied by that weight’s stored input. Work it through here. B‘s input is whatever A produced. If A is random, its output is not zero, so B gets a real gradient on the very first step and starts to move. A‘s message from above has to travel back through B. While B is exactly zero, that message arrives as zero, so A waits. As soon as B has moved off zero, A starts moving too. Zero both grids and neither ever gets a non-zero gradient. The pair sits at zero for the entire run.

Set both to random and you have vandalised a working model before training has begun. The patch would inject noise into every adapted layer on step one. That is precisely the instability the zero start was designed to avoid.

Count the parameters you stopped training

Now put numbers on the saving. Take one square weight grid, 2,048 rows by 2,048 columns, and work it out in full.

  • Full grid. 2,048 times 2,048 is 4,194,304 numbers.
  • LoRA at rank 16. A is 16 by 2,048, so 32,768 numbers. B is 2,048 by 16, so another 32,768. Together that is 65,536.
  • The ratio. 4,194,304 divided by 65,536 is exactly 64.
full W (2048 × 2048): 4,194,304 params LoRA r=16 (16 × 4096): 65,536 params 64× fewer trainable params — per matrix, across every targeted layer

Read the two bars as counts of numbers, not as sizes on the card. The green sliver is everything that actually trains, and the gap between the bars is the d over 2r shortcut drawn out.

There is a shortcut hiding in that arithmetic. In general, a grid with d rows and k columns holds d times k numbers. Its LoRA patch holds r times d plus k. Divide the first by the second and you have the saving factor.

When the grid is square, d and k are the same, and the whole thing simplifies to d / (2r). Check it against the numbers above: 2,048 divided by 2 times 16, so 2,048 divided by 32, which is 64. It agrees. Two routes, one answer, and you can now size any square grid in your head.

Feed-forward grids are not square, so they use the general form. Here is one of Qwen2.5-1.5B’s, worked the same way.

Weight grid Shape Numbers in the full grid Numbers in a rank 16 patch Times fewer
Illustrative square grid 2,048 by 2,048 4,194,304 65,536 64
One feed-forward grid in Qwen2.5-1.5B 1,536 by 8,960 13,762,560 167,936 about 82

The second row is 16 x (1,536 + 8,960), which is 16 x 10,496, which is 167,936. Then 13,762,560 divided by 167,936 comes to about 82. Wide grids save more, because the full count grows with the two sides multiplied while the patch grows with them added.

Sum that across every adapted grid in all 28 layers and a rank 16 adapter for this model trains about 18.5 million numbers. Against 1.54 billion parameters, that is about 1.2 percent of the model. Across models and settings the figure usually lands somewhere between 0.5 and 2 percent.

Set rank, alpha and target modules on purpose

LoRA has far fewer knobs than a full fine-tune. Four of them matter, and they interact, so copying a config without understanding them goes wrong in ways that look like LoRA being weak.

Rank sets how much the patch can express

Rank, written r, is the width of the bottleneck. It is how many multiplication tables the patch is allowed to add together. More tables means more the adapter can express, and more numbers to train.

  • 4 to 8: cheap, limited, fine for light style and format work.
  • 16: the standard starting point, and where this article’s arithmetic sits.
  • 32 to 64: more room for harder tasks.

Bigger is not reliably better. Past a point extra rank adds numbers without adding quality, and with the standard scaling it can even hurt. The next subsection explains exactly why. Start at 16.

Alpha sets how loudly the patch speaks

Alpha is the other half of that scale factor. The patch’s output is multiplied by alpha divided by r before it is added in. Only the ratio matters, never alpha on its own.

Work an example. At rank 16 with alpha 32, the factor is 32 divided by 16, which is 2. Every number the patch produces is doubled on its way into the layer. The common convention is alpha equal to twice the rank, which always gives a factor of 2.

Now watch the trap. Raise the rank to 64 and leave alpha at 32. The factor becomes 32 divided by 64, which is 0.5. You have just quartered the patch’s strength at the exact moment you tried to give it more capacity. The patch got wider and quieter at the same time, so the run barely changes and you conclude that rank does not help.

The fix is to move alpha with the rank. At rank 64, set alpha to 128 and the factor stays at 2. Why does the dial exist in this shape at all? Because it separates two questions. Rank asks how wide the bottleneck is. The ratio asks how strongly the result is mixed in. Keeping them apart means you can change one without retuning the other.

Target modules decides which grids get a patch

A linear layer is one weight grid with a job: take a list of numbers in, multiply by the grid, hand a list of numbers out. That is the operation every section above has been describing.

One transformer block in Qwen2.5-1.5B holds seven of them. Four belong to attention, the step where each position in the text looks at the other positions and mixes in what it finds useful. They are named q_proj, k_proj, v_proj and o_proj. Three belong to the feed-forward part, where each position is processed on its own: gate_proj, up_proj and down_proj.

one transformer block’s linear layers: attention: q_proj k_proj v_proj o_proj MLP: gate_proj up_proj down_proj attention-only (PEFT default: q,v) + MLP = “all-linear” (QLoRA-style) PEFT default adapts only q_proj, v_proj — cheapest, fine for light adaptation. target_modules=”all-linear” adapts every linear layer — QLoRA showed this can match full fine-tuning quality. More parameters, more capacity, still tiny vs full FT.

The four blue boxes are attention. The three grey ones are the feed-forward part, and they hold roughly seven eighths of a block’s weights. Adapting only the blue ones leaves the grey ones with no patch at all.

The original LoRA paper, by Hu and colleagues, adapted only two of the four attention grids. Later work found that adapting every linear layer is what closes the gap to full fine-tuning, and the QLoRA paper by Dettmers and colleagues is where that became the standard advice.

The arithmetic says why. In one block of this model, the four attention grids hold about 5.5 million numbers between them. The three feed-forward grids hold about 41.3 million. So the feed-forward part is roughly 88 percent of the block. Adapt attention only and you have left seven eighths of each block with no patch on it.

Use all linear layers unless you are deliberately economising. Libraries let you write that as a single keyword rather than listing module names, which is safer, because names differ between model families. It targets the linear layers inside the transformer blocks and leaves the token lookup table and the final output layer alone.

Dropout is a minor dial you set once

Dropout means randomly zeroing a fraction of the values flowing through the patch during training, so the adapter cannot lean too hard on any single path. It is a mild form of regularisation, which is the general name for any deliberate constraint that stops a model fitting its training data too closely. Fitting too closely is called overfitting: the model memorises your examples instead of learning the pattern in them.

A value of 0.05 is a fine default. Nudge it up for very small datasets and down to zero for large ones. Set it once and move on.

Two variants worth knowing about

Variant What it changes When to reach for it
Rank-stabilised LoRA Scales by alpha divided by the square root of the rank, instead of by the rank itself When you want to push rank to 64 or beyond and actually feel the difference
DoRA, weight-decomposed low-rank adaptation Splits the update into a learned size per column plus a low-rank direction, which behaves more like a full fine-tune When a low-rank run underperforms and you intend to merge for serving

The rank-stabilised variant follows straight from the alpha trap above. Dividing by r means every step up in rank is scaled down in proportion. Dividing by the square root of r shrinks the patch far more gently as rank grows, so the extra capacity survives. If raising the rank once did nothing for you, the standard scaling is the first suspect. Try the stabilised version before deciding the task did not need the capacity.

DoRA needs one more idea, and an arrow is the easiest way to hold it. An arrow has two independent properties: how long it is, and which way it points. A column of numbers in a weight grid is the same. It has an overall size, and it has a pattern of proportions between its entries, which acts like a direction. A full fine-tune moves both freely. Plain LoRA moves them together in one lump. DoRA splits them: it keeps a separate trainable size for each column and applies the low-rank patch only to the direction. Liu and colleagues measured that this tracks full fine-tuning more closely, especially at low rank. It costs extra bookkeeping on every step, so it is a considered choice rather than a default.

A sensible starting config, then: rank 16, alpha 32, all linear layers, dropout 0.05, no adaptation of bias terms. Change one dial at a time and measure. For format adaptation on a clean instruction dataset, that default lands close to a full fine-tuning baseline.

Run a LoRA fine-tune and read its memory bill honestly

Run the same job as the full fine-tune, LoRA style, and watch the static state collapse.

32 GB card full FT 24.6 GB static LoRA ~4 GB base weights (bf16) gradients optimizer states grad+optim now on ~1% of params

The blue segment is the frozen base weights, and it is the same length in both bars. What LoRA removes is the gradient and optimizer segments. Both bars are static state, with activations still to be added on top.

Do the arithmetic yourself, in three lines.

  1. The frozen base. 1.54 billion parameters at 2 bytes each in bf16 is about 3.08 GB. Every one of those weights is still needed for the forward pass, so none of it goes away.
  2. The adapter. About 18.5 million trainable numbers, each paying the full 16 bytes, because these are the numbers that do get gradients and optimizer records. That is 18.5e6 x 16, so about 0.3 GB.
  3. Add them. About 3.4 GB, which the series rounds to about 4 GB once you keep the 2 to 3 GB of headroom every estimate here carries.

Against 24.6 GB for the full fine-tune, that is roughly six times less. Every number in that list is static state. Not one of them includes activations.

The same collapse is what makes a 7B model trainable on one card. Qwen2.5-7B has 7.62 billion parameters, so full fine-tuning wants about 122 GB of static state, which is impossible on any single consumer card. Under LoRA the frozen base is 7.62 billion at 2 bytes, so about 15.2 GB, plus roughly 0.6 GB of adapter and optimizer state at rank 16. That is about 15.8 GB before the safety margin, which this series quotes as about 18 GB of static state once the margin is in.

The config, line by line

The code is the supervised fine-tuning script from Part 9 plus one config object and one changed number.

from peft import LoraConfig
from trl import SFTConfig, SFTTrainer

lora = LoraConfig(
    task_type="CAUSAL_LM",
    r=16,
    lora_alpha=32,
    target_modules="all-linear",
    lora_dropout=0.05,
    bias="none",
    # use_rslora=True,      # if you push the rank high
    # use_dora=True,        # if low rank underperforms
)

cfg = SFTConfig(
    output_dir="qwen15b-lora",
    assistant_only_loss=True,
    max_length=1024,
    learning_rate=2e-4,
    lr_scheduler_type="cosine",
    warmup_ratio=0.03,
    num_train_epochs=1,
    per_device_train_batch_size=8,
    gradient_accumulation_steps=2,
    bf16=True,
    gradient_checkpointing=True,
    logging_steps=10,
    eval_strategy="steps",
    eval_steps=100,
    seed=42,
)

trainer = SFTTrainer(
    model=model,
    args=cfg,
    train_dataset=ds["train"],
    eval_dataset=ds["test"],
    processing_class=tok,
    peft_config=lora,
)
trainer.train()
trainer.model.print_trainable_parameters()

Every line, assuming you have never written PyTorch.

  • from peft import LoraConfig pulls in PEFT, the library that implements LoRA. PEFT stands for parameter-efficient fine-tuning, which is the umbrella name for any method that trains a few new numbers and freezes the rest.
  • from trl import SFTConfig, SFTTrainer pulls in TRL, the library that runs the training loop for you.
  • task_type="CAUSAL_LM" tells PEFT this is a next-token language model, so it knows which layers exist and how to wrap them.
  • r=16 is the bottleneck width from the section above.
  • lora_alpha=32 gives the scale factor of 32 over 16, which is 2.
  • target_modules="all-linear" is the single keyword that adapts every linear layer instead of a hand-written list.
  • lora_dropout=0.05 zeroes 5 percent of the patch’s values at random during training.
  • bias="none" leaves the model’s bias terms, which are small per-layer offsets, out of training. They are tiny and rarely worth adapting.
  • The two commented lines switch on the rank-stabilised scaling and DoRA. Leave them off until you have a reason.
  • output_dir is the folder the adapter and its checkpoints are written to. A checkpoint is a saved copy of the trainable weights at one point in the run.
  • assistant_only_loss=True grades only the assistant’s tokens and skips the user’s. Part 7 covers why.
  • max_length=1024 caps each training example at 1,024 tokens, a token being a chunk of text roughly the size of a word.
  • learning_rate=2e-4 is the single number that scales every weight change. 2e-4 means 0.0002. This is the line that matters most, and the next subsection is about it.
  • lr_scheduler_type="cosine" makes the learning rate glide smoothly down toward zero as the run goes on, along a cosine curve. It is the reliable default.
  • warmup_ratio=0.03 ramps the learning rate up from zero over the first 3 percent of steps, so the earliest updates cannot do damage.
  • num_train_epochs=1 means one full pass over the training data. One to three is normal for instruction tuning.
  • per_device_train_batch_size=8 processes 8 examples at once on the card. LoRA frees enough memory that you can usually raise this.
  • gradient_accumulation_steps=2 runs two of those batches, adds their gradients together and updates once. The result behaves like a batch of 16, which is called the effective batch, while only 8 examples are ever in memory at a time.
  • bf16=True does the arithmetic in the 16-bit format.
  • gradient_checkpointing=True throws away most stored activations and recomputes them during the backward pass. It costs roughly 20 to 30 percent more time and saves a great deal of memory. Part 3 covers it in full.
  • logging_steps, eval_strategy and eval_steps control how often numbers are printed and how often the model is scored on held-out data.
  • seed=42 fixes the random numbers so two runs of the same script match.
  • SFTTrainer(...) assembles everything: the model, the settings, the training and evaluation data, the tokenizer that turns text into tokens, and the LoRA config. Passing peft_config is the whole difference from a full fine-tune. PEFT wraps the model, freezes the base and inserts the adapters.
  • trainer.train() runs the loop.
  • print_trainable_parameters() prints how many numbers are actually training. Expect roughly 1 percent. If it prints 100 percent, the adapters were not wired in and you are quietly paying for a full fine-tune.

Library flags drift. Those are current as of August 2026, so pin your versions and check the current documentation before a real run.

The one mistake that ruins most LoRA runs

Copying the learning rate from a full fine-tuning script. Full fine-tuning wants roughly 1e-5 to 2e-5, meaning 0.00001 to 0.00002. LoRA wants roughly 1e-4 to 3e-4, meaning 0.0001 to 0.0003, about ten times more.

The reason is the zero start from earlier. A full fine-tune is nudging weights that are already good. LoRA is growing a patch from nothing. Use the small value and the loss barely moves, the run looks like a failure, and the wrong conclusion is that LoRA is weak.

The activation bill does not move at all

FULL FT all layers stored LoRA SAME — backward stillflows through frozen base QLoRA same + de-quant (transient) INFERENCE one layer, freed each step + KV cache LoRA/QLoRA cut optimizer & gradient memory on the base — not activations. Checkpointing still matters with LoRA. activations de-quant KV cache

Three of these four bars are the same height. Only inference escapes the activation bill, because only inference has no backward pass.

This is the most common misconception about LoRA, so it is worth being blunt. An activation is any intermediate value the forward pass computes and holds on to. LoRA cuts gradient and optimizer memory. It does nothing whatsoever to activations.

Two reasons, both structural. The forward pass still runs through every frozen layer, because the frozen layers are what produces the answer. And the backward pass still travels all the way back through them, because the adapter in layer 1 has to be reached somehow. Frozen means those layers do not update. It never means the work skips them. Part 3 works through the whole argument.

So state the bill honestly. For Qwen2.5-1.5B at a batch of 4 sequences of 1,024 tokens, a LoRA run is about 4 GB of static state plus roughly 8 GB of activations, so about 12 GB in total. Treat the 8 as an order of magnitude rather than a measurement. That total fits a 32 GB card with real room to spare, which the full fine-tune’s roughly 33 GB does not. Both halves of the comparison have to include activations, or the comparison is not a comparison.

Decide between merging the adapter and swapping it at serve time

The payoff of “the patch is just added” is operational. Once training is done, an adapter is a small standalone file, and you have two ways to serve it.

A · MERGED (merge_and_unload) W ← W + (α/r)·B·A one fused weight matrix ✓ zero inference overhead ✓ deploy like any model (Bedrock CMI!) ✗ one task per served copy B · ADAPTER-SWAP one shared frozen base (loaded once) adapter: legal adapter: SQL adapter: support ✓ N tasks, ~one base’s worth of VRAM ✓ hot-swap / route per request ~ small per-adapter switch overhead

Merging gives one fused grid and one task. Swapping keeps one base in memory and picks which small patch to add for each request.

Merge when you serve one specialised model. Merging computes W + (alpha/r) * B * A once, ahead of time, and stores the result as the layer’s new weight grid. The patch has the same shape as W, so the sum is just a grid again. You get a plain model file with no adapter machinery and no extra work at request time. It loads anywhere a normal model loads.

Swap when you serve many specialisations from one base. Load the base model once and keep the small adapters beside it, picking one per request. Ten tasks then cost roughly one model’s worth of VRAM instead of ten, because the multi-gigabyte part is shared. This interacts directly with the batching and cache economics in the part on continuous batching and paged attention.

# fold the adapter into the base, giving a plain model that needs no PEFT to serve
from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "qwen15b-lora")
merged = model.merge_and_unload()
merged.save_pretrained("qwen15b-merged")

Line by line.

  • AutoModelForCausalLM.from_pretrained(...) downloads or loads the base model by name and puts its weights in memory in bf16.
  • PeftModel.from_pretrained(base, "qwen15b-lora") reads the adapter folder your training run wrote and attaches the two skinny grids to the right layers. Nothing is folded in yet, and the base is untouched.
  • merge_and_unload() does the folding. It adds the scaled patch into each adapted weight grid and then removes the adapter machinery, leaving an ordinary model object.
  • save_pretrained(...) writes that model to a folder as a normal checkpoint. Anything that can load the base can load this.

Because the patch is added rather than spliced in, the operation is reversible. Subtracting the same term recovers the original grid exactly, and libraries track this so an adapter can be switched on and off. You can also attach several adapters and blend them with weights, which is tempting for combining a domain adapter with a tone adapter. Treat that as an experiment. Adapters trained separately can interfere with each other, and nothing guarantees the blend behaves like either one.

Say what LoRA gives up as well as what it saves

Put both methods on the same evaluation for a format-adaptation task and they usually land close together. Do not over-generalise from that result. The careful finding, from Biderman and colleagues, measured on Llama 2 models across code and mathematics, is that LoRA learns less and forgets less, and both halves are real.

Dimension Full fine-tuning LoRA
Static state, Qwen2.5-1.5B about 24.6 GB about 4 GB, roughly six times less
Activation memory roughly 8 GB at batch 4, sequence 1,024 the same roughly 8 GB, unchanged
Storage per task a full copy, about 3 GB an adapter, tens of megabytes
Format, style and tone strong matches full fine-tuning
Hard new skills such as code or maths stronger, it learns more weaker, it learns less
Catastrophic forgetting more less
Sensitivity to the learning rate forgiving needs the higher LoRA range

One mental model explains both halves at once. LoRA regularises. It restricts how much the model is allowed to change, so it changes less. Changing less is exactly why it forgets less, and it is exactly why it learns less. For adaptation, that trade is close to free. For teaching a genuinely new capability, full fine-tuning still earns its cost.

Neither half shows up in a loss curve. Part 15 gives the procedure: score general tasks on the base model and on your adapter, and compare. The drop you do not see is the forgetting LoRA saved you from.

One last placement note. LoRA is also the foundation the next two parts build on. Part 12 stores the frozen base in 4 bits so a larger model fits, and preference tuning in Part 13 is almost always run as a LoRA on top of a supervised fine-tune.

Key takeaways

  • A matrix is a grid of numbers. Its rank is how many multiplication tables you must add together to rebuild it exactly, and low-rank means a small number of tables is enough.
  • LoRA does not claim a model’s weights are low-rank. It claims the change a fine-tune needs is, and it only ever learns two skinny grids instead of the full change.
  • One of those grids starts at all zeros, so the patch adds exactly nothing at step zero and training begins at the base model. Zero both and neither ever moves.
  • Rank sets how much the patch can express. Alpha over rank sets how loudly it speaks. Raise rank without raising alpha and you quietly turn the patch down.
  • Adapting every linear layer, rather than only attention, is what closes the gap to full fine-tuning, because the feed-forward grids hold about 88 percent of a block’s weights.
  • Static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Roughly 8 GB of activations still sits on top of both figures, so the LoRA run totals about 12 GB.
  • LoRA needs roughly ten times a full fine-tune’s learning rate. Copying the smaller value is the single most common way to ruin a run.
  • Merge the adapter to serve one task with no overhead. Keep it detached to serve many tasks from one shared base.

You can now

  • Rebuild a grid of numbers from two shorter lists and read off its rank, from “Rebuild a big grid from two skinny ones, and read off its rank”.
  • Explain in plain words what LoRA adds to a frozen layer and in what order, from “Attach a low-rank patch to a frozen weight”.
  • Say what would happen if both LoRA grids started at zero, and why one of them does, from “Say why one of the two grids starts at all zeros”.
  • Compute the trainable-parameter saving for any weight grid, square or not, from “Count the parameters you stopped training”.
  • Pick a rank, move alpha with it, and justify targeting all linear layers, from “Set rank, alpha and target modules on purpose”.
  • State a LoRA run’s static state and its total separately, without letting one stand in for the other, from “Run a LoRA fine-tune and read its memory bill honestly”.
  • Choose between merging an adapter and swapping adapters at serve time, from “Decide between merging the adapter and swapping it at serve time”.

Glossary

Activation
Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
Activation memory
The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
Adapter
A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy.
Alpha
The number that scales a LoRA adapter’s output as it is added into the layer, through the ratio of alpha to rank. Only that ratio matters, and alpha of twice the rank is the common convention. Hugging Face PEFT, LoRA guide
Attention
The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
Backward pass
Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
Base model
The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
Batch
A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
bf16
A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
Catastrophic forgetting
When fine-tuning on one task quietly erodes abilities the model already had, such as arithmetic or instruction following. Nothing in the loss curve shows it, so you have to score general tasks before and after. Luo et al., Catastrophic Forgetting During Continual Fine-tuning
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
Cosine schedule
A learning-rate schedule that decays smoothly from the peak down toward zero along a cosine curve. It is the reliable default for fine-tuning. Hugging Face TRL, SFT Trainer
DoRA
A LoRA variant that splits the learned update into a size and a low-rank direction, which behaves closer to full fine-tuning especially at low rank. Liu et al., DoRA
Dropout
Randomly zeroing a fraction of values during training so the model cannot lean on any single path. A mild form of regularisation, and LoRA applies 0.05 of it on the adapter branch by default. Hugging Face PEFT, LoRA guide
Effective batch
The number of examples that actually go into one weight update: the per-device batch times the accumulation steps times the number of GPUs. This is the number that matters for training behaviour, rather than the batch that happens to fit. Batch size, accumulation and the effective batch
Epoch
One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one. Supervised fine-tuning end to end
Forward pass
Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
Frozen weights
Weights marked as not trainable, so they never receive an update. A frozen weight needs no gradient and no optimizer state, which removes 14 of its 16 bytes, though its activations are still stored.
Full fine-tuning
Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
Gradient accumulation
Running several small batches, adding their gradients together, and only updating the weights once at the end. You get the steadier signal of a big batch while holding just one small batch in memory. Batch size, accumulation and the effective batch
Gradient checkpointing
Throwing away most stored activations and recomputing them during the backward pass. Peak activation memory drops a long way in exchange for roughly 20 to 30 percent more time. Chen et al., Training Deep Nets with Sublinear Memory Cost
Inference
Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
Layer
One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
Learning rate
A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
Learning-rate schedule
A rule that changes the learning rate over the course of a run, typically warming up and then decaying. The schedule is a separate thing from the optimizer, and both act on every step. Warmup and schedule
LoRA
Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Low-rank
A matrix is low-rank when it can be rebuilt exactly from two much smaller matrices multiplied together. LoRA assumes the update a fine-tune needs is low-rank, which is why two skinny matrices can stand in for a full one. Aghajanyan et al., Intrinsic Dimensionality of Fine-Tuning
Merging an adapter
Folding a trained adapter’s update permanently into the model’s weights, giving one plain checkpoint with no runtime cost. It works because the adapter’s output is added, and with QLoRA you must load the base in bf16 first.
Optimizer
The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
Optimizer state
The numbers an optimizer keeps between steps, such as running averages of past gradients. Under standard mixed-precision AdamW it is 12 of the 16 bytes per parameter, which makes it the largest memory tenant. Rajbhandari et al., ZeRO
Overfitting
When the model starts memorising the training examples instead of learning the pattern. Training loss keeps falling while performance on unseen data stalls or gets worse. The six silent failure modes
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
Parameter-efficient fine-tuning
Any method that trains a small number of new parameters and leaves the pretrained ones frozen. LoRA and QLoRA are the ones in practical use, and PEFT is the library that implements them. Hugging Face PEFT, LoRA guide
Preference tuning
Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
Pretraining
The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
QLoRA
LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
Rank
The width of LoRA’s bottleneck, written r. It sets how much the adapter can change: 4 to 8 is light, 16 is the usual starting point, and bigger is not reliably better. Hugging Face PEFT, LoRA guide
Regularisation
Any deliberate constraint that stops a model fitting its training data too closely, so it generalises better. Weight decay and dropout are the two you meet in this series. Loshchilov and Hutter, Decoupled Weight Decay Regularization
Static state
Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
Supervised fine-tuning
Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
Transformer
The architecture behind every model in this series: a stack of blocks that alternate attention with a feed-forward network. Vaswani et al., Attention Is All You Need
VRAM
The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
Warmup
Ramping the learning rate up from zero over the first few percent of steps, so the earliest updates cannot damage the model while the optimizer still has no history. Set it as a ratio rather than a step count. Warmup and schedule
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model


Practical exercises

Redo the parameter-reduction arithmetic for Qwen2.5-1.5B’s own attention weights

The article’s parameter-count figure uses an illustrative 2048 by 2048 matrix to show a 64 times reduction at rank 16. Redo the same arithmetic for two of Qwen2.5-1.5B’s own attention projections instead. First, q_proj and o_proj, which are square at 1536 by 1536, since 12 query heads times head_dim 128 equals the 1536 hidden size. Then k_proj and v_proj, which are 1536 by 256, since GQA gives only 2 KV heads times head_dim 128. Use the article’s general form, the product of the dimensions divided by rank times their sum, for both, at rank 16, and state both ratios. Then explain in a sentence or two why the two ratios come out so differently even though both matrices sit in the same attention block.

See the worked solution (opens in a new tab)

Predict what changes if you target only q_proj and v_proj instead of all linear layers

Take the LoRA config from the article and change target_modules="all-linear" to target_modules=["q_proj", "v_proj"], the original LoRA paper’s choice shown in the figure. A Qwen2.5-1.5B decoder block has seven linear projections: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj. The MLP projections are 1536 by 8,960 or 8,960 by 1536. Using the same method as the previous exercise, work out the total trainable-parameter count per block for all-linear versus q_proj-and-v_proj only, at rank 16. Then say plainly what actually changes and what does not: does the VRAM bill move, and does the model’s quality relative to full fine-tuning move.

See the worked solution (opens in a new tab)

Diagnose a LoRA run that trains at the wrong rate and does not respond to more rank

A colleague runs LoRA on Qwen2.5-1.5B with rank 16, alpha 32, learning rate 2e-5, copied from a full fine-tuning script. After three epochs the loss has barely moved. They double the rank to 64, keeping alpha at 32 and the same learning rate, hoping more capacity will help. It still barely moves. Name both mistakes in this run and give the concrete fix for each.

See the worked solution (opens in a new tab)

Work out whether 12 client adapters fit on one card without merging

You need to serve 12 different client-specific LoRA adapters for Qwen2.5-1.5B from one 32 GB card, swapping the small adapter per request rather than merging. Using the trainable-parameter count for an all-linear, rank 16 adapter from the second exercise, work out the bf16 storage size of one adapter, then the total for the frozen base plus all 12 adapters resident at once. Compare that against merging each adapter into its own standalone checkpoint and trying to keep all 12 of those loaded simultaneously instead.

See the worked solution (opens in a new tab)

Extend the 7B LoRA memory math to a much higher rank

The article gives Qwen2.5-7B’s LoRA static state as about 18 GB: 7.62 billion frozen parameters at 2 bytes each is about 15.2 GB, plus roughly 0.6 GB of adapter and optimizer state at rank 16. Adapter and optimizer state scales roughly linearly with rank, since the trainable parameter count itself is linear in rank for a fixed set of target modules. Estimate the static state at rank 64 and at rank 256, and check both against the card’s roughly 29.8 GiB usable capacity once you include the standard 2 to 3 GB safety margin. State plainly, using the static-versus-total distinction the series insists on, whether either estimate actually tells you the run will fit.

See the worked solution (opens in a new tab)

Frequently asked questions

What is LoRA in fine-tuning?

LoRA, short for low-rank adaptation, freezes the pretrained weights and learns a small patch beside each targeted weight grid. The patch is stored as two skinny grids whose product has the same shape as the original, so it can simply be added on. About 1 percent of the model’s numbers train, and the frozen ones need no gradient and no optimizer state.

What does low-rank actually mean, without the linear algebra?

Think of a multiplication table: write one list down the side and one along the top, and every cell is its row’s number times its column’s number. A grid’s rank is how many such tables you have to add together to rebuild it exactly. Low-rank means a small number of tables is enough, so the grid can be stored as a few pairs of short lists instead of every cell.

What rank and alpha should I use for LoRA?

Start at rank 16 with alpha 32, which gives a scale factor of 2, targeting all linear layers with dropout 0.05. Bigger rank is not reliably better, and if you raise rank you must raise alpha with it or you shrink the patch’s strength at the same time. To push rank to 64 or beyond, either set alpha to twice the new rank or switch on the rank-stabilised variant.

Why does LoRA need a higher learning rate than full fine-tuning?

Because one of its two grids starts at all zeros, so the adapter starts from nothing rather than from converged pretrained values. Full fine-tuning wants roughly 1e-5 to 2e-5, since large moves damage what pretraining built, while LoRA wants roughly 1e-4 to 3e-4. Use the full fine-tuning value in a LoRA run and the loss barely moves.

Should I merge my LoRA adapter or keep it separate?

Merge when you serve one specialised model, because folding the patch into the weights gives a plain checkpoint with no runtime overhead that any serving stack can load. Keep it separate when you serve many specialisations from one base, because a few-megabyte adapter per task means ten tasks cost roughly one model’s VRAM rather than ten. Merging is reversible, since the patch is added rather than spliced in.

Does LoRA reduce training time as well as memory?

Somewhat, and less than people expect. You skip computing gradients and optimizer updates for the frozen weights, which is real work saved, but the forward pass and the backward pass still traverse every layer and the activation cost is unchanged. The dramatic saving is in static state and storage rather than in wall-clock time.

Sources and further reading

Previous