Exercise solutions: QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
These are the worked solutions for the exercises in Part 12, QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
With double quantisation: 4 bits of data plus 0.125 bits of quantised constants is 4.125 bits, or 4.125 divided by 8, 0.515625 bytes per parameter. For Qwen2.5-7B’s 7.62 billion parameters that is 7.62 billion times 0.515625 bytes, about 3.93 GB, matching the article’s “about 4 GB” figure for the base weights.
Without double quantisation: 4 bits of data plus 0.5 bits of full-precision constants is 4.5 bits, or 0.5625 bytes per parameter. For the same model that is about 4.29 GB.
The difference at 7B is about 0.36 GB, a few hundred megabytes, not something you would plan a run around.
For a 70 billion parameter model the same difference is 70 billion times the 0.046875 byte-per-parameter gap, about 3.28 GB. That is the “several gigabytes” the article promises on a very large model. The saving per parameter is fixed, so it only becomes a large absolute number once the parameter count itself is large. Double quantisation is a rounding error at 7B and a real design decision at 70B.
Exercise 2
Memory does not change at all. Both fp4 and nf4 store 4 bits per parameter, so the byte accounting from the first exercise is identical either way: the Qwen2.5-7B base still lands at about 3.93 GB with double quantisation on. The quantisation type is a question of where the 16 available levels sit, not how many bits they cost.
Quality is where the difference shows up, and only in one direction, for the worse. Plain 4-bit floats space their levels uniformly across the representable range. Network weights are not uniform. They cluster near zero in a roughly normal distribution. Uniform spacing wastes most of its resolution out in the sparse tails and under-resolves the dense region near zero where almost all the weights actually live. NF4 exists specifically to fix this, by placing its 16 levels at the quantiles of a normal distribution so each level covers roughly equal probability mass. Switching to fp4 throws that design away for no memory benefit, so it is a change that can only hurt and never help, which is why bitsandbytes defaults to nf4 for training rather than fp4.
Exercise 3
QLoRA throughput: 4,000 tokens per second times 0.58, about 2,320 tokens per second.
Epoch time in bf16: 500,000 divided by 4,000 is 125 seconds.
Epoch time in QLoRA: 500,000 divided by 2,320 is about 215.5 seconds.
That is roughly 90.5 seconds longer, about 72 percent more wall-clock time for the identical epoch, not 40 percent. The 40-percent figure describes the drop in throughput, a 42-percentage-point loss from 100 to 58 that the article rounds to “about 40 percent less throughput.” The slowdown in time is a different, larger number, because time is the reciprocal of throughput. Losing 42 percent of your tokens per second costs you 1 divided by 0.58 minus 1, about 72 percent more time, not 42 percent. For a 1.5B model that already fits comfortably in plain bf16 LoRA at about 4 GB, this entire 72 percent is pure loss. No model that would not otherwise fit got unlocked, so nothing was bought with it.
Exercise 4
13 billion parameters in bf16 at 2 bytes each is 26 GB. Against a ceiling of about 29.8 GiB, that leaves roughly 3.8 GB before the adapter or the margin are even accounted for. Subtract the standard 2 to 3 GB safety margin and there is well under a gigabyte to a couple of gigabytes left for the adapter’s own footprint and whatever transient buffers the merge operation needs. That is technically possible, but it is genuinely tight, the kind of step that can fail with an out-of-memory error on a bad day.
Compare the 7B case: 7.62 billion parameters in bf16 is about 15.24 GB, leaving roughly 14.5 GB before the margin and around 11.5 to 12.5 GB after it, comfortable enough that the article calls it a non-issue. The article’s own warning, that this step is a common surprise for large models, starts biting well before the 30B case it explicitly calls out as needing a bigger machine. By the 13B tier the merge is already the tightest part of the whole workflow, not the fine-tuning run that produced the adapter in the first place.
Exercise 5
Base weights, NF4 with double quantisation: about 3.93 GB, from the first exercise.
Adapter and optimizer state for the trainable LoRA parameters: about 0.6 GB, the same figure the article gives for LoRA on this model, since the adapter itself trains in full precision regardless of what format the frozen base is stored in.
Static state before any margin: 3.93 plus 0.6, about 4.53 GB.
Add the standard 2 to 3 GB safety margin, say 3 GB for a round upper estimate: 4.53 plus 3, about 7.53 GB, which rounds cleanly to “about 8 GB.”
The article’s headline figure is exactly this same sum: frozen weights, plus adapter and optimizer state, plus the series’ standard margin, already folded in and rounded, rather than a separately measured number. Any “about N GB” total quoted elsewhere in this series can be reconstructed the same way from the underlying static-state pieces plus the standard 2 to 3 GB cushion, which means you can sanity-check any of them yourself rather than treating them as a black box. It is also why the table calls this case “with headroom”: the 8 GB already has the safety cushion built in, on top of the roughly 29.8 GiB the card actually has to give.