QLoRA is the trick that puts a 7 billion parameter model on a single consumer graphics card. It is LoRA with one change. The frozen copy of the model is stored at 4 bits per number instead of 16, so the cost of simply holding it drops by about four times.
That change is not free, and the price tag is the part most write-ups leave out. Nothing on a GPU does arithmetic on a 4-bit weight directly. Every multiplication has to expand those weights back to 16 bits first. The expansion is real work, on every layer, on every step.
By the end of this part you will be able to quantise a handful of numbers by hand and say exactly how much accuracy you lost, explain where NF4 puts its sixteen values and why, work out a 4-bit model’s memory to the nearest tenth of a gigabyte, and decide from a model’s size alone whether QLoRA is the right tool or a slowdown you inflicted on yourself.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs (you are here)
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
Name the one memory cost LoRA could not remove
Start with what LoRA already fixed, because QLoRA only attacks what was left over.
A full fine-tune holds three things on the card for every number in the model. The number itself. One gradient saying which way to move it. And the optimizer’s running notes about it. That comes to 16 bytes per parameter, and Part 1 derives all sixteen. Those three together are called the static state, because their size does not change from the first training step to the last.
LoRA freezes the model. A frozen number is one the training loop has been told never to change. It needs no gradient and no optimizer notes, so fourteen of its sixteen bytes vanish.
Two bytes survive, and they survive for a solid reason. The frozen weights are still what computes the answer. Every one of them has to be sitting on the card, in bf16, a 16-bit number format that costs 2 bytes per number.
Put real numbers on that residue.
- Qwen2.5-1.5B, this series’ small running example: 1.54 billion parameters times 2 bytes is about 3.1 GB. Nobody cares.
- Qwen2.5-7B: 7.62 billion times 2 bytes is about 15.2 GB. Add roughly 0.6 GB for the adapter and its optimizer state and you are near 16 GB, which this series rounds to about 18 GB of static state once the standard safety margin goes in. On a 32 GB card that is workable and no longer comfortable.
- A 70 billion parameter model: about 140 GB. You cannot load it, let alone train beside it.
That residue is the wall. It is not gradients and it is not the optimizer. It is the plain cost of keeping the weights where the maths can reach them.
One warning before the arithmetic starts. Every number in this part is static state. Activations, the intermediate values the forward pass has to keep for the backward pass, sit on top of all of it and they do not shrink under LoRA or QLoRA. Part 3 works out why.
Store a weight in 4 bits, and see exactly what you give up
Four bits is four binary digits. Each is 0 or 1, so there are 2 x 2 x 2 x 2 patterns, from 0000 to 1111. Sixteen patterns, and therefore sixteen possible values. That is the entire budget.
Here is the part that catches people out. The bit pattern is not the number. It is a row number into a table of sixteen decimal values. Storing a weight means finding the closest row and writing down its index. Reading it back means looking that row up.
Squeezing numbers onto a small set of allowed values like this is called quantisation. Expanding them again is called dequantisation. Both words appear constantly from here on, and neither means anything more than that.
The table, and the number that stretches it
Try the simplest possible table first: sixteen values spread evenly from -1 to 1. The gaps between them are all the same size, which is 2 divided by 15, or about 0.133. So the rows read -1.0, -0.867, -0.733, and on up to 1.0.
Real weights are nowhere near that range. A trained model is full of numbers like 0.0271 and -0.0413. Drop those onto this table and they all land on the same two or three rows. Fourteen of your sixteen values do nothing.
So quantisation always has two halves. A table of allowed values, and a scale factor that stretches the table onto your actual numbers. That pairing is not new to 4 bits. Jacob and colleagues set out the same scale-and-round arrangement for 8-bit integer inference in 2017, and the shape of it has not changed since.
Follow one weight all the way through. Take a group of 64 weights from a real layer. The largest of them, ignoring sign, is 0.08.
- Find the scale. The scale factor is that largest value, 0.08. Divide every weight in the group by it and the group now spans -1 to 1, which is exactly what the table covers.
- Scale one weight. Take
w = 0.0271. Divide by 0.08 and you get 0.339. - Round to a row. The two nearest rows are 0.2 and 0.333. Closest is 0.333. Store that row’s index. Four bits, done.
- Read it back. Look up the row: 0.333. Multiply by the scale: 0.333 x 0.08 = 0.0267.
The weight went in as 0.0271 and came out as 0.0267. You lost about 0.0004. For a frozen weight that is never going to be updated again, that is a fine trade for cutting its storage by four.
You do have to store the scale factor too, or the numbers can never be recovered. Hold that thought. It comes back with a bill attached.
Why the groups are small
Why 64 weights per group instead of the whole matrix at once? Because the scale is the largest value in the group. A single freak weight of 3.0 hiding in a matrix of 2.4 million would set the scale for all of them. Every ordinary weight would then collapse onto the first row or two.
Keeping groups small keeps an outlier’s damage inside its own group. QLoRA quantises in blocks of 64, each with its own scale factor. This is called block-wise quantisation, and the same idea is what makes 8-bit optimizer state work.
Outliers cause the same trouble one level up, among the values flowing through the network rather than the weights. Dettmers and colleagues found that 8-bit inference on large models only holds up if a few outsized feature dimensions are kept in 16 bits and computed separately. That paper is the direct predecessor of the 4-bit work in this part, by the same lead author.
The two failures that motivate the next section
Run a small weight through the same four steps and the picture changes. Take w = 0.004. Divide by 0.08 and you get 0.05. The nearest rows are -0.0667 and 0.0667, so it rounds to 0.0667. Multiply back and 0.004 comes out as 0.0053. That is 33 percent off.
Now try a weight of exactly 0. Divide by the scale and it is still 0. The nearest rows are the same pair. So zero comes back as 0.0053 as well. Sixteen evenly spaced values that are symmetric around zero cannot include zero, because there is no middle row when the row count is even.
Both failures happen right where the weights are densest. That is the problem NF4 was designed to solve.
One thing to be clear about first. That table is all there is. NF4 has no exponent field and no mantissa field, the two parts that make up an ordinary floating-point number. Part 4 covers both fields and the difference between storing and computing. A 4-bit NF4 weight is a row index, and no piece of hardware multiplies row indexes.
See why even spacing wastes most of a 4-bit budget
The table above put its sixteen values at even spacings. That single choice is what QLoRA changes. Build the reason before the name arrives.
What a trained model’s weights actually look like
Read every number out of one weight matrix. Count how many fall into each narrow band of values. You do not get a flat picture. You get a tall pile near zero that thins out fast in both directions.
Adult height behaves the same way. Most people sit within a few inches of the average. A few are noticeably taller or shorter. Almost nobody is under four feet or over seven. Plot the counts and you get a hump in the middle with thin tails on each side.
That shape has a name. It is called a normal distribution, or a bell curve. Trained network weights come out with roughly that shape. Nobody arranged it. It is just what training does.
Count the wasted rows
Now lay the even table over that pile and count. Take 1,000 weights from one block, already divided by their scale, so they run from -1 to 1. The sixteen even rows sit 0.133 apart.
Suppose 900 of the 1,000 land between -0.2 and 0.2. For a trained layer that is a realistic split. Which rows are inside that band? Four of them: -0.2, -0.0667, 0.0667 and 0.2.
So 900 different weights get flattened onto four values. The other twelve rows share the remaining 100 weights between them. Three quarters of your budget is spent on territory that is nearly empty.
Put the rows where the weights are
The fix falls straight out of that count.
- Sort the 1,000 weights from smallest to largest.
- Cut the sorted list into 16 piles of about 62 each.
- Read off a value from the middle of each pile.
- Use those sixteen values as your table.
Now every row stands for roughly the same number of weights. About 62 each, instead of 225 in some rows and 3 in others.
Those cut points have a name. A quantile is a cut point in a sorted list, with a fixed fraction of the data below it. The median is the quantile with half the data below it. Cut a list into sixteen equal piles and the fifteen boundaries between the piles are quantiles.
That is NF4
NF4, short for NormalFloat-4, is exactly this. Its sixteen values sit at the quantiles of a bell curve rather than at even spacings.
Two details make it work in practice.
First, the values are not computed from your model. They are computed once, from a standard bell curve, and baked into the library as a fixed sixteen-entry table. Your block’s scale factor is what stretches that fixed table onto your actual weights, exactly as in the last section. So there is no per-model calibration step and no extra data to store.
Second, the table is arranged so that one of the sixteen entries is exactly 0. A weight of zero comes back as zero. The steps are noticeably finer near zero than the even table’s, and noticeably coarser out in the tails, which is the whole point.
| Table of sixteen values | Bits stored per weight | Where the levels sit | Exact zero available? | What it is for |
|---|---|---|---|---|
| Even spacing | 4 | equal gaps across the whole range | no | the obvious first try |
| NF4 | 4 | at the quantiles of a bell curve: fine near zero, coarse in the tails | yes | QLoRA’s frozen base |
| bf16 | 16 | set by exponent and mantissa fields | yes | the adapter, and everything that computes |
Same four bits. Same memory. The only change is where the sixteen values sit, and it buys back a large part of the accuracy that even spacing threw away.
There is a second reason quality holds up, and it matters more than the table does. Nothing computes in NF4. The frozen base takes a one-time rounding hit when it is loaded, and every matrix multiply afterwards runs on values expanded back to bf16. The trainable adapter is bf16 from start to finish. The part of the model that is actually learning never gets rounded to 4 bits at all. Dettmers and colleagues report that fine-tuning on an NF4 base matches fine-tuning on a 16-bit base.
Shrink the scale factors too, and total up the real bill
Every block of 64 weights carries its own scale factor, and a scale factor is a number you have to store. Charge it to the budget honestly.
Store it as an ordinary 32-bit float and one scale per 64 weights costs 32 divided by 64, so 0.5 bits per weight. Against a 4-bit budget that is a 12.5 percent surcharge. On a 70 billion parameter model it is over 4 GB of scale factors on their own.
Once you see it written that way, the fix is obvious. The scale factors are just a list of numbers. So quantise them.
That is double quantisation. Each scale factor is stored in 8 bits instead of 32, which is 8 divided by 64, so 0.125 bits per weight. The 8-bit scales are themselves grouped, 256 at a time, with one 32-bit constant per group. That second layer adds 32 divided by 16,384, about 0.002 bits per weight, and rounds away to nothing.
Add up what one NF4 weight really costs.
- 4 bits of weight.
- 0.125 bits of quantised scale factor.
- About 0.002 bits for the scale factor’s own scale factor.
Call it 4.125 bits. Divide by 8 and you get 0.515625 bytes per parameter, which this series quotes as about 0.51 bytes.
Now run the 7B. 7.62 billion parameters times 0.515625 bytes is about 3.9 GB of frozen base weights, against 15.2 GB in bf16. That is the fourfold cut, and it is the cut that decides whether the model loads at all.
Add the adapter and its optimizer state, roughly 0.6 GB at rank 16, and the 2 to 3 GB safety margin this series keeps in every estimate. You land at about 8 GB of static state for a 7B QLoRA run. Say static state out loud every time you quote that figure. Activations are not inside it.
Double quantisation saves just under 0.4 bits per weight, and that saving per weight is fixed. So it is a rounding error on a small model and several gigabytes on a very large one. It costs you nothing to leave on, and at 70B it is the difference between fitting and not.
Turn a memory spike into a slowdown with a paged optimizer
Fine-tuning runs do not usually die at a steady, predictable memory level. They die in a spike.
Picture the failure concretely. Your run has been fine for four thousand steps. Then a batch comes along holding the four longest examples in the dataset. Activation memory for that one step is far above the average. The allocator asks the card for a block it cannot give, and the process dies with an out-of-memory error. Nothing since the last checkpoint is saved.
A paged optimizer is the safety valve for exactly that moment. The name is borrowed from operating systems, so start there.
What paging means
Your laptop pretends to have more memory than it does. It splits memory into fixed-size chunks called pages. Pages that have not been touched in a while get written out to disk. When a program reads one of those addresses again, the operating system pauses it, fetches the page back, and lets it continue. The program never notices except that the read was slow.
NVIDIA offers the same arrangement between the card and the machine’s system RAM. You allocate a buffer with one address that both the GPU and the CPU can use, and the driver decides where the bytes physically live from moment to moment.
What it does when memory spikes
Put the optimizer state in that kind of buffer and this happens.
- The forward and backward pass hits its spike and asks for more memory than the card has free.
- The driver looks for pages it can move out. The optimizer state is the ideal candidate, because nothing is reading it during the forward and backward pass.
- Those pages are copied over PCIe, the wire between the card and the motherboard, into system RAM. The card now has room and the step completes.
- The optimizer step then reads the state, which faults the pages straight back onto the card before the read completes.
The run survives. It survives slowly, because PCIe is far slower than the card’s own memory. That is the trade, and it is a good one for a rare event.
Two honest limits. Paging is a valve, never a memory plan. If every step pages, your run is crawling and the real fix is a smaller batch or a shorter sequence. And under QLoRA the optimizer state belongs only to the adapter, which is about one percent of the parameters, so there is not much to page in the first place. Its job here is to absorb spikes, not to make room.
You also need the system RAM to page into. A fine-tuning machine wants generous host memory for this reason alone.
Predict the throughput tax before you pay it
Now the cost. Storing weights in 4 bits means every matrix multiply has to expand them back to 16 bits first. That is extra work, inserted before the maths, on every layer of every step.
Why the expansion cannot be skipped
Almost all of a GPU’s speed comes from its tensor cores. A tensor core is a dedicated block on the chip that multiplies small grids of numbers in one shot. It accepts a short, fixed list of input formats, and it decodes them in hardware.
NF4 is not on that list and never will be. A tensor core can decode an exponent-and-mantissa layout because the meaning of the bits is fixed. It cannot decode a row index into a lookup table.
So for every matrix multiply, a kernel, meaning the small program the GPU runs for one operation, has to do this:
- Read the 4-bit codes for a block of weights.
- Look each one up in the sixteen-entry table.
- Multiply by that block’s scale factor.
- Write the resulting
bf16values into fast on-chip memory. - Only then hand them to the tensor core.
The number, and how to read it
Measured on an RTX 5090, 4-bit training throughput sits at about 58 percent of 16-bit throughput. That holds roughly steady across model sizes, so it behaves like a fixed tax rather than a scaling effect.
Read that figure carefully, because it is easy to misquote. Losing 42 percent of your tokens per second does not mean a run takes 42 percent longer. Time is the reciprocal of throughput, so the wall-clock increase is always the larger number. The exercises at the end of this part ask you to work out how much larger.
Why it is ever worth paying
The saving that offsets the tax is memory bandwidth, meaning how many bytes per second the card can pull out of its own memory. Every training step has to read every weight at least once. For a 7B model that is 15.2 GB of reading per pass in bf16, against 3.9 GB in NF4.
So the question is what the card was waiting on.
- A large model spends much of its step waiting for weights to arrive. Reading a quarter as many bytes is a genuine win, and it can outweigh the lookup work.
- A small model was never waiting on reads. You have added expansion work with nothing to offset it. The tax is pure loss.
Measured end to end on that card, the crossover sits between the two. Below about 3B, 4-bit training costs more energy per token than it saves. At 7B and above it comes out ahead. The same bandwidth argument drives serving speed, and the part on why decode is memory-bound develops it in full.
Decide when QLoRA is worth it, and when it is a mistake
Put the memory and the tax together and the rule is short. Reach for QLoRA when a model would not otherwise fit. Never by default.
Here is what that means on the target hardware for this series, one 32 GB consumer card, which gives about 29.8 GiB you can actually use.
| Model | Full fine-tune, static state | LoRA, bf16 base, static state | QLoRA, 4-bit base, static state | What to do |
|---|---|---|---|---|
| Qwen2.5-1.5B | about 24.6 GB, tight | about 4 GB, easy | works, and is slower for nothing | plain LoRA |
| Qwen2.5-7B | about 122 GB, impossible | about 18 GB | about 8 GB | either, and QLoRA if activations are tight |
| 13B class | impossible | about 28 GB, very tight | about 12 GB | QLoRA |
| 30B class, dense | impossible | about 62 GB, impossible | about 18 to 20 GB | QLoRA |
| 70B class | impossible | impossible | about 38 GB, still impossible | multiple cards |
Every cell in that table is static state. Weights, gradients and optimizer state, and nothing else. Activations sit on top of all of them.
For Qwen2.5-1.5B at a batch of 4 sequences of 1,024 tokens, activations run to roughly 8 GB, order of magnitude. At 7B and above there is no single number worth quoting. The figure moves with your batch size, your sequence length, whether gradient checkpointing is on, and whether attention runs on a flash or SDPA backend rather than building the full score grid. Those four settings can swing it by a factor of several. Work out your own from Part 3 rather than trusting a headline total.
Reading the table row by row:
- 1.5B. Plain LoRA, every time. QLoRA would give you headroom you did not need in exchange for a large slice of your throughput.
- 7B. Plain LoRA fits at about 18 GB of static state, and activations decide it. If a long sequence or a bigger batch pushes you over, QLoRA’s 8 GB buys the room. This is also roughly where the 4-bit tax starts paying for itself.
- 13B and 30B. QLoRA territory. These are the models plain LoRA cannot hold and QLoRA can, and they are large enough that the bandwidth saving is real.
- 70B. Even at 0.51 bytes per parameter the base alone is about 36 GB. One card is not enough, and Part 14 covers splitting a run across several.
Turn a LoRA script into a QLoRA script with one config object
The code is the LoRA script from the previous part plus a settings object telling the loader to hold the base in 4 bits. The adapter, the training loop and the loss masking are all unchanged. QLoRA is a loading detail rather than a new way of training.
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig, prepare_model_for_kbit_training
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # NormalFloat-4
bnb_4bit_use_double_quant=True, # quantise the constants too
bnb_4bit_compute_dtype=torch.bfloat16, # de-quantisation target for the matmuls
)
model = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-7B", # a model full fine-tuning could never touch here
quantization_config=bnb, dtype=torch.bfloat16,
)
model = prepare_model_for_kbit_training(model) # norms to fp32, grad-checkpoint hooks
lora = LoraConfig(task_type="CAUSAL_LM", r=16, lora_alpha=32,
target_modules="all-linear", lora_dropout=0.05, bias="none")
# then SFTConfig and SFTTrainer exactly as in the LoRA part, but with:
# optim="paged_adamw_8bit" the paged optimizer, survives spikes
# learning_rate=2e-4 still the LoRA regime
Line by line, assuming you have never written PyTorch.
import torchpulls in PyTorch, the library that owns the numbers and the GPU memory they sit in.BitsAndBytesConfigis a small settings object. It carries your quantisation choices, and the model loader reads them.load_in_4bit=Truesays quantise the weights while loading, layer by layer. You never need room for the full 16-bit version, which matters when the 16-bit version would not fit.bnb_4bit_quant_type="nf4"picks the quantile table. Set it explicitly. The library’s other 4-bit option is a small floating-point layout that is not shaped around how weights are distributed, and it costs exactly the same memory.bnb_4bit_use_double_quant=Trueswitches on the second round of quantisation on the scale factors.bnb_4bit_compute_dtype=torch.bfloat16names the format the 4-bit weights get expanded into before each multiply. This one argument is the storage-versus-compute split written down.AutoModelForCausalLM.from_pretraineddownloads and builds the model. Causal LM means a next-token-prediction model, which is the kind this series trains.quantization_config=bnbhands the settings object to the loader. Without it the model loads in 16 bits as normal.dtype=torch.bfloat16sets the format for everything that is not quantised.prepare_model_for_kbit_training(model)does the housekeeping a quantised base needs. Skip it and you get instability or missing-gradient errors that look like adapter bugs.LoraConfig(...)is the adapter, identical to the previous part.r=16is its size,lora_alpha=32scales its output, andtarget_modules="all-linear"puts one beside every linear layer.optim="paged_adamw_8bit"is the paged optimizer from two sections ago.learning_rate=2e-4is the LoRA regime, roughly ten times a full fine-tune’s rate. Copying a full fine-tune’s rate into a LoRA run is the classic way to waste a day.
The preparation helper earns a sentence of its own, because it is easy to drop. It casts the normalisation layers and the output layer to fp32 for numerical stability. It turns off the model’s generation-time key and value cache, which has no role during training. And it makes sure gradients can flow back through the frozen 4-bit layers to reach the adapters.
Library APIs and quantisation support drift faster here than anywhere else in this series. Pin your versions, and test the quantisation library on its own before you build a run on top of it. The flags above are current as of August 2026.
One live limitation worth knowing. The NF4 kernel expects two-dimensional weight matrices. Some mixture-of-experts models pack their experts into three-dimensional tensors instead, and those cannot be 4-bit quantised this way yet on consumer cards. Dense models are fine. If you meet a quantisation error on such a base, that is the cause rather than your code, so check before you commit to one.
Ship the result, and plan for the merge before you need it
Two things bite at the end of an otherwise smooth QLoRA project. Both are easier to handle in the plan than in the last hour.
The 4-bit you train in is not the 4-bit the hardware multiplies
Recent GPUs advertise native 4-bit floating point as a headline feature, under names like NVFP4 and MXFP4. It is tempting to assume that is what QLoRA uses. It is not, and the two serve opposite purposes.
NF4 is a storage format. A software kernel expands it before any arithmetic happens, and its job is to make a frozen model small during training. NVFP4 and MXFP4 are compute formats. They are number layouts the tensor cores multiply directly with no expansion step, and their job is to make a served model fast.
So treat the format you train in and the format you serve in as two separate decisions. Fine-tune with NF4. If you want fast 4-bit serving afterwards, convert the finished model into a hardware compute format inside your serving engine, as its own documented step.
Why a 16-bit adapter will not fold into a 4-bit base
Merging means folding the adapter’s learned change permanently into the model’s weights, so you ship one ordinary model file with no adapter attached. With a plain LoRA run that is a clean addition, and Part 11 covers when to do it.
With a 4-bit base it is not clean, and the reason is the table from earlier. Merging means computing the frozen weight plus the adapter’s change, then storing the result. But a 4-bit weight can only be stored as one of sixteen rows. The sum has to be rounded back onto the nearest row, and the adapter’s change is usually far smaller than the gap between two neighbouring rows. So it rounds straight back to where it started. You would run the merge and get your original model.
The standard path avoids the problem entirely. Load the base in bf16, merge into that, and save a 16-bit checkpoint. Quantise afterwards if you want to. The usual route for that second step is post-training quantisation, where a finished model is compressed in one pass without any further training. Frantar and colleagues set out the best known version of it, GPTQ, and most serving engines will take a model in that form.
# serving a QLoRA result: load the base in bf16, merge, save full precision
from peft import PeftModel
from transformers import AutoModelForCausalLM
import torch
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-7B", dtype=torch.bfloat16) # bf16, NOT 4-bit
merged = PeftModel.from_pretrained(base, "qwen7b-qlora").merge_and_unload()
merged.save_pretrained("qwen7b-merged") # then quantise for fast serving
Three lines matter there. from_pretrained loads the base in bf16 with no quantisation config, which is the whole trick. PeftModel.from_pretrained attaches your trained adapter to it. merge_and_unload folds the adapter in and hands back a plain model with no adapter machinery left.
Notice what that first line costs. It needs enough memory to hold the base in bf16, which is precisely the memory QLoRA was avoiding. For a 7B that is about 15.2 GB, and fine on a 32 GB card. For a 30B it is about 60 GB, and it is not. Plan for a larger machine, offload to system RAM, or skip the merge and keep the adapter as a separate file at serving time.
With shipping handled, the next question is whether the model picks the better of two acceptable answers, which is where Part 13 on preference tuning starts.
Key takeaways
- QLoRA is LoRA with the frozen base stored in 4 bits. It attacks the one cost LoRA left standing, which is the plain cost of holding the weights where the maths can reach them.
- Quantisation is a sixteen-row lookup table plus a scale factor per block of 64 weights. Dequantisation is looking the row up and multiplying by that scale.
- Trained weights cluster near zero in a bell curve, so evenly spaced levels waste most of their budget on empty tails. NF4 puts its sixteen values at the quantiles of that curve instead, including an exact zero.
- NF4 is storage only. It has no exponent or mantissa fields, nothing computes in it, and the trainable adapter stays in bf16 throughout.
- Double quantisation takes the scale factors from 0.5 bits per weight to about 0.13, giving 4.125 bits total, about 0.51 bytes per parameter. A 7B base lands at about 3.9 GB.
- QLoRA on the 7B is about 8 GB of static state. Activations sit on top and depend on batch size, sequence length, gradient checkpointing and the attention backend, so no single total is worth quoting.
- Four-bit throughput is about 58 percent of 16-bit on an RTX 5090, and time is the reciprocal of throughput. Use QLoRA to make a model fit, never by default.
- Merging needs the base back in bf16 first, because a 16-bit change rounds away when folded into a 4-bit weight. Budget the memory for that step in advance.
You can now
- Quantise a number by hand through the scale-and-round steps and state the error you introduced, from “Store a weight in 4 bits, and see exactly what you give up”.
- Explain to a colleague why sixteen evenly spaced values waste a 4-bit budget on trained weights, using the count of rows in the crowded band, from “See why even spacing wastes most of a 4-bit budget”.
- Compute any model’s NF4 memory from its parameter count at 0.515625 bytes each, and say what that figure excludes, from “Shrink the scale factors too, and total up the real bill”.
- Predict whether a given run will gain or lose from 4-bit storage before launching it, from “Predict the throughput tax before you pay it”.
- Pick between full fine-tuning, LoRA and QLoRA for a given model on a given card, from “Decide when QLoRA is worth it, and when it is a mistake”.
- Plan the memory and the sequence of steps for shipping a merged QLoRA model, from “Ship the result, and plan for the merge before you need it”.
Glossary
- 8-bit optimizer
- AdamW with its two running averages stored in 1 byte each instead of 4, using block-wise quantisation. The algorithm and its behaviour are unchanged, and it reclaims most of 8 bytes per parameter. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
- Activation
- Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
- Activation memory
- The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
- AdamW
- Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Alpha
- The number that scales a LoRA adapter’s output as it is added into the layer, through the ratio of alpha to rank. Only that ratio matters, and alpha of twice the rank is the common convention. Hugging Face PEFT, LoRA guide
- Base model
- The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions. Choosing a base model and building a bench
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- bf16
- A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
- Block-wise quantisation
- Quantising numbers in small blocks, each with its own scale factor, instead of using one scale for a whole tensor. It handles local variation, and it is how both 8-bit optimizer state and NF4 weights work. Dettmers et al., 8-bit Optimizers via Block-wise Quantization
- Checkpoint
- A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
- De-quantisation
- Expanding quantised numbers back to a wider format so arithmetic can be done on them. QLoRA de-quantises its 4-bit weights to bf16 for every matrix multiply, and that extra work is where its speed cost comes from.
- DoRA
- A LoRA variant that splits the learned update into a size and a low-rank direction, which behaves closer to full fine-tuning especially at low rank. Liu et al., DoRA
- Double quantisation
- Quantising the scale factors that quantisation itself produces. In QLoRA it takes their overhead from about half a bit per parameter down to about 0.13 bits, which is several gigabytes on a large model. Dettmers et al., QLoRA
- fp32
- 32-bit floating point, with 8 exponent bits and 23 mantissa bits. It is the precise reference format, used for the master copy of the weights and for the optimizer’s running averages. Number formats for training
- FP8
- 8-bit floating point, which ships in two versions because the range-against-precision trade is too tight to settle once. E4M3 for weights and activations, E5M2 for gradients. Micikevicius et al., FP8 Formats for Deep Learning
- Frozen weights
- Weights marked as not trainable, so they never receive an update. A frozen weight needs no gradient and no optimizer state, which removes 14 of its 16 bytes, though its activations are still stored. LoRA explained
- Full fine-tuning
- Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
- GiB
- A gibibyte, 2 to the power 30 bytes, which is how drivers and vendors report GPU memory. Byte arithmetic such as 16 times the parameter count lands in decimal GB instead, so 24.6 GB is about 22.9 GiB and a 32 GB card gives roughly 29.8 GiB. Training memory and the 16 bytes per parameter
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- KV cache
- The keys and values of past tokens, kept during generation so they are not recomputed for every new token. It exists only at inference; training has no generation loop and therefore no KV cache. Continuous batching and paged attention
- Layer
- One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
- Layer norm
- A step that rescales the numbers flowing through a layer so they stay in a sensible range, which keeps training stable. Modern LLMs use a cheaper version of it called RMSNorm. Inside one transformer block
- Learning rate
- A single number that scales every weight change. Too high and the loss spikes or blows up, too low and the model barely moves off the base. Learning rate, the master dial
- LoRA
- Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Low-rank
- A matrix is low-rank when it can be rebuilt exactly from two much smaller matrices multiplied together. LoRA assumes the update a fine-tune needs is low-rank, which is why two skinny matrices can stand in for a full one. Aghajanyan et al., Intrinsic Dimensionality of Fine-Tuning
- Matrix multiply
- The operation that dominates all the arithmetic in a transformer: multiply a block of inputs by a block of weights to get a block of outputs. Usually shortened to matmul. Inside one transformer block
- Memory tenant
- One of the four things sharing the card during training: weights, gradients, optimizer state and activations. All four are live at the same moment, so peak memory is their sum rather than the largest of them. The four tenants and the 16 bytes per parameter
- Merging an adapter
- Folding a trained adapter’s update permanently into the model’s weights, giving one plain checkpoint with no runtime cost. It works because the adapter’s output is added, and with QLoRA you must load the base in bf16 first. Merge, swap, and serving many tasks
- Mixture of experts
- A model where each token is routed through a few of many parallel feed-forward blocks instead of all of them. Some pack their experts into three-dimensional tensors, which the current 4-bit quantisation code cannot handle.
- Neural network
- A stack of layers that multiply their input by learned weights and pass the result on. Training means adjusting those weights until the output is closer to what you wanted. How a neural network learns
- NF4
- The 4-bit storage format QLoRA uses. Its 16 values sit at the quantiles of a bell curve rather than at even spacings, because model weights cluster near zero. It has no exponent or mantissa fields and nothing computes in it directly. Dettmers et al., QLoRA
- OOM
- Out of memory, the error you get when a run needs more VRAM than the card has. In training it almost always strikes where the forward pass ends and the backward pass begins, which points straight at activations. Activation memory and gradient checkpointing
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Optimizer state
- The numbers an optimizer keeps between steps, such as running averages of past gradients. Under standard mixed-precision AdamW it is 12 of the 16 bytes per parameter, which makes it the largest memory tenant. Rajbhandari et al., ZeRO
- Paged optimizer
- An optimizer that can move its state out to ordinary system RAM when GPU memory spikes, then bring it back when the pressure passes. It turns a hard crash into a brief slowdown. Dettmers et al., QLoRA
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- Parameter-efficient fine-tuning
- Any method that trains a small number of new parameters and leaves the pretrained ones frozen. LoRA and QLoRA are the ones in practical use, and PEFT is the library that implements them. Hugging Face PEFT, LoRA guide
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- QLoRA
- LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
- Quantisation
- Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
- Rank
- The width of LoRA’s bottleneck, written r. It sets how much the adapter can change: 4 to 8 is light, 16 is the usual starting point, and bigger is not reliably better. Hugging Face PEFT, LoRA guide
- Static state
- Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
- Supervised fine-tuning
- Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
- Tensor
- A grid of numbers with any number of dimensions. One number is a scalar, a row of them is a vector, a table is a matrix, and anything past that is still a tensor with more dimensions. How a neural network learns
- Tensor core
- The part of an NVIDIA GPU built to do matrix multiplies in low precision very fast. It is why bf16 training beats fp32, and why a format tensor cores cannot multiply directly, such as NF4, costs throughput. NVIDIA, Accelerating AI training with TF32 tensor cores
- Token
- The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
- VRAM
- The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
Practical exercises
Verify the 0.51 bytes per parameter figure, and see why double quantisation matters more at scale
NF4 stores 4 bits per parameter, and double quantisation adds roughly 0.125 bits per parameter for the block scale constants: 8 bits per constant across blocks of 64 weights is 8 divided by 64. Compute the resulting bytes per parameter, then use it to compute Qwen2.5-7B’s base-weight memory in this format, at 7.62 billion parameters. Recompute both numbers as if double quantisation were switched off, so block constants cost the naive 32 bits each instead, 32 divided by 64 bits per parameter of overhead, and compare the two base-weight totals. Then do the same comparison for a 70 billion parameter model and say what changes about how much double quantisation actually buys you.
See the worked solution (opens in a new tab)
Predict what happens if you swap NF4 for plain 4-bit floats
In the BitsAndBytesConfig from the article, change bnb_4bit_quant_type="nf4" to bnb_4bit_quant_type="fp4", a plain 4-bit float with uniformly spaced levels, leaving everything else the same, including double quantisation. Predict what happens to the base model’s memory footprint and to its quality, and explain your answer using the article’s own reasoning for why NF4 exists.
See the worked solution (opens in a new tab)
Quantify what QLoRA actually costs you on a model that already fits
Your plain bf16 LoRA run on Qwen2.5-1.5B processes 4,000 tokens per second on your 32 GB card. Out of curiosity you quantise the same model to NF4 and run QLoRA instead, expecting the extra headroom to help. Using the article’s figure that 4-bit throughput sits at about 58 percent of 16-bit on this class of card, predict the QLoRA throughput, then compute how much longer a fixed 500,000-token epoch takes under each, in seconds. State the result as a percentage increase in wall-clock time, and say whether that number matches “40 percent” the way you might expect.
See the worked solution (opens in a new tab)
Work out whether merging a QLoRA-tuned 13B model fits on the same card you trained it on
You QLoRA fine-tune a 13 billion parameter model on your 32 GB card comfortably, at around 12 GB per the article’s table. Now you want to ship a merged, standalone checkpoint, which requires de-quantising the base back to bf16 first. Work out the bf16 base’s memory footprint, compare it against the card’s roughly 29.8 GiB usable ceiling minus the standard safety margin, and say how much room is actually left for the adapter and the merge operation itself. Contrast that with the equivalent calculation for the 7B case the article describes as comfortable.
See the worked solution (opens in a new tab)
Build the QLoRA 7B total from the ground up and see where 8 GB actually comes from
The article’s table gives Qwen2.5-7B QLoRA as comfortable, about 8 GB with headroom. Build that number yourself rather than taking it on faith. Start from the NF4 double-quantised base-weight figure from the first exercise, add the roughly 0.6 GB of adapter and optimizer overhead the article gives for a rank-16 LoRA adapter on this same 7B model, then add the series’ standard 2 to 3 GB safety margin. Does your bottom-up total match the article’s headline figure, and what does that tell you about what the headline figures in this series actually include?
Frequently asked questions
What is QLoRA and how is it different from LoRA?
QLoRA is LoRA with the frozen base stored in 4 bits instead of 16. The adapter, the training loop and the loss masking are all identical. The single change is quantisation at load time, which cuts the cost of holding the base weights by about four times and is what lets a 13B or 30B model train on one consumer card.
What is NF4, and why not use evenly spaced 4-bit values?
NF4 is a 4-bit storage format whose sixteen values sit at the quantiles of a bell curve, meaning each value covers roughly the same number of weights. Trained weights cluster near zero, so evenly spaced values put most of their resolution out in the near-empty tails and leave the crowded middle sharing three or four levels. NF4 also includes an exact zero, which an even sixteen-value table cannot.
Does QLoRA reduce model quality?
Very little, because no arithmetic happens in 4 bits. The frozen weights are stored in NF4 and expanded back to bf16 for every matrix multiply, and the trainable adapter stays in bf16 the whole time. Dettmers and colleagues report that fine-tuning on a 4-bit base matches fine-tuning on a 16-bit base, so the price you pay is throughput rather than accuracy.
Is QLoRA slower than LoRA?
Yes. Four-bit training throughput measures about 58 percent of 16-bit on an RTX 5090, because tensor cores cannot multiply NF4 and every matrix multiply has to expand the weights first. Note that time is the reciprocal of throughput, so the increase in wall-clock time is larger than the drop in tokens per second.
Should I use QLoRA on a small model?
No. If a model already fits with plain LoRA, quantising it gives you headroom you do not need in exchange for a large slice of your throughput, and below about 3B it costs more energy per token than it saves. QLoRA is a fitting tool rather than a speed tool.
How do I deploy a model I fine-tuned with QLoRA?
Load the base in bf16 rather than 4 bits, attach the adapter, merge, and save a 16-bit checkpoint, then quantise that into your serving engine’s own format. A 16-bit adapter change cannot be folded into a 4-bit weight, because the sum rounds straight back onto the level it started from. Plan for the merge needing enough memory to hold the whole base in bf16.
Sources and further reading
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, the source of NF4, double quantisation and paged optimizers, and of the result that a 4-bit base matches a 16-bit one.
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, the adapter method QLoRA builds on.
- Dettmers et al., 8-bit Optimizers via Block-wise Quantization, where the per-block scale factor idea used by NF4 comes from.
- Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference, the fundamentals of the scale-and-round scheme every section above uses.
- Dettmers et al., LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale, the 8-bit predecessor to this work and the source of the outlier-dimension problem.
- Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, the main alternative route, applied to a finished model rather than during training.
- bitsandbytes documentation, for the 4-bit configuration options used in the code above.
- Hugging Face PEFT, the LoRA guide, for the adapter configuration the QLoRA script reuses unchanged.
- Micikevicius et al., FP8 Formats for Deep Learning, for the contrast between a hardware compute format and storage-only quantisation.
