Exercise Solutions: LoRA Explained: Freeze the Model, Learn a Low-Rank Patch

Exercise solutions: LoRA Explained: Freeze the Model, Learn a Low-Rank Patch

These are the worked solutions for the exercises in Part 11, LoRA Explained: Freeze the Model, Learn a Low-Rank Patch. Read the exercise first; coming here before you have tried it defeats the point.

Exercise 1

Square case, q_proj or o_proj, 1536 by 1536: full parameters are 1536 times 1536, which is 2,359,296. LoRA parameters at rank 16 are 16 times (1536 plus 1536), which is 49,152. The ratio is 2,359,296 divided by 49,152, exactly 48. That matches the shortcut d over 2r used elsewhere in the article: 1536 divided by 32 is 48. This is the real number for this model’s square attention projections, and it is smaller than the figure’s illustrative 64, because the figure’s hypothetical matrix used dimension 2048 rather than this model’s actual hidden size of 1536. The ratio scales directly with that dimension.

Non-square case, k_proj or v_proj, 1536 by 256: full parameters are 1536 times 256, which is 393,216. LoRA parameters at rank 16 are 16 times (1536 plus 256), which is 28,672. The ratio is about 13.7, far below the square case’s 48.

The shortcut d over 2r only applies when both matrix dimensions are equal. When one dimension shrinks to 256, because grouped-query attention gives only 2 KV heads against 12 query heads, the LoRA replacement’s size is dominated by the sum of the two dimensions, which barely changes, while the full matrix’s size is dominated by their product, which shrinks a great deal. The saving LoRA buys is largest exactly where the original matrix was large in both directions, and grouped-query attention already made the K and V projections small in one direction before LoRA ever touched them.

Exercise 2

Per block, at rank 16: q_proj and o_proj each cost 49,152 trainable parameters, from the first exercise. k_proj and v_proj each cost 28,672. Each of gate_proj, up_proj and down_proj costs 16 times (1536 plus 8,960), which is 167,936.

All linear layers together, per block: 2 times 49,152, plus 2 times 28,672, plus 3 times 167,936, is 659,456 trainable parameters. Across the model’s 28 layers that is about 18.46 million parameters, about 1.2 percent of the 1.54 billion total, comfortably inside the 0.5 to 2 percent range the article quotes, a useful check that this arithmetic is right.

q_proj and v_proj only: 49,152 plus 28,672 is 77,824 per block, about 2.18 million across 28 layers, about 0.14 percent of the total.

That is roughly 8.5 times fewer trainable parameters, 18.46 million against 2.18 million. The VRAM bill barely moves. The frozen base already dominates the LoRA memory budget at about 3 GB out of the roughly 4 GB total for this model, so shrinking the already-small adapter by a further 8.5 times saves well under a gigabyte, nothing you would notice on a 32 GB card. Quality is what actually moves: the article is explicit that adapting only the original two attention projections leaves a real gap to full fine-tuning, while adapting all linear layers is what closes it. This change trades a slightly simpler adapter footprint for that gap. On a card that already fits the model comfortably, it is usually not worth making.

Exercise 3

The first mistake is the learning rate. A full fine-tune wants roughly 1e-5 to 2e-5, because it nudges weights that are already converged. LoRA wants roughly ten times more, 1e-4 to 3e-4, because it trains a small adapter from a zero start and needs a much bigger step to move at all. Copying the full fine-tune value into a LoRA run is exactly the mistake the article calls the single most common way to ruin a LoRA run, and it alone explains why the loss barely moved even before the rank change.

The second mistake is subtler and shows up in how the rank was raised. Doubling rank from 16 to 64 while leaving alpha at 32 changes the alpha over rank ratio from 2 down to 0.5. Under the standard LoRA scaling, that ratio sets the adapter’s effective strength, so quartering it actively suppresses most of the extra capacity the higher rank was supposed to buy, which is why raising rank did not help either. The fix is either to raise alpha to 128 to keep the ratio at 2, or to switch on the rank-stabilised variant, use_rslora=True, which divides by the square root of rank instead of rank itself and keeps the effective magnitude roughly stable as rank grows. Combine either fix with the corrected learning rate of roughly 1e-4 to 3e-4 and the run should behave as expected.

Exercise 4

One adapter’s trainable parameters, stored in bf16 for inference at 2 bytes each: 18.46 million times 2 bytes is about 36.9 megabytes, matching the article’s own description of an adapter as tens of megabytes. Twelve adapters together: about 443 megabytes, well under half a gigabyte.

The frozen base in bf16 is 1.54 billion parameters times 2 bytes, about 3.08 GB, matching the article’s own “about 3 GB” figure for a full copy.

Base plus all 12 unmerged adapters resident at once: about 3.08 GB plus 0.44 GB, roughly 3.5 GB total, trivial against a 32 GB card and leaving enormous room for activations and KV cache during serving.

Merging each adapter into its own full checkpoint instead means 12 separate models at about 3.08 GB each, roughly 37 GB in total. That already exceeds a single 32 GB card’s usable capacity before a single request’s activations are loaded, so keeping 12 merged models resident at once does not fit. This is the concrete version of the article’s claim that ten tasks cost roughly one model’s VRAM instead of ten, expressed as an exact number for this case: the difference between about 3.5 GB and about 37 GB, not just a ratio.

Exercise 5

Rank 64 is four times rank 16, so adapter and optimizer overhead scales to about 0.6 times 4, roughly 2.4 GB. Static state: 15.2 plus 2.4 is 17.6 GB. With the standard 2 to 3 GB margin, that lands around 19.6 to 20.6 GB, comfortably under the roughly 29.8 GiB ceiling.

Rank 256 is sixteen times rank 16, so overhead scales to about 0.6 times 16, roughly 9.6 GB. Static state: 15.2 plus 9.6 is 24.8 GB. With the margin, that lands around 26.8 to 27.8 GB, technically still under 29.8 GiB, but only just.

Here is what the exercise is really testing. Both of those numbers are static state only: frozen weights plus adapter plus optimizer. Activations still sit on top, unchanged from a full fine-tune at the same batch size and sequence length, exactly as the article insists a static figure must never be read as a total. At rank 64 there is roughly 9 to 10 GiB of headroom left for activations after the margin, which is workable. At rank 256 there is at most 2 to 3 GiB left, unlikely to cover activation memory for any usable batch size. Despite the static figure technically fitting under the card’s ceiling, the run as a whole almost certainly would not. Pushing rank this high on a 7B model is a bad idea on one card. The adapter itself never gets expensive relative to the frozen base. The real problem is that it quietly eats the exact margin the activations need.