Exercise solutions: LLM Fine-Tuning Explained: What Actually Changes Inside the Model
These are the worked solutions for the exercises in Part 1, LLM Fine-Tuning Explained: What Actually Changes Inside the Model. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
Multiply parameters by bytes per parameter: 3,000,000,000 times 16 is 48,000,000,000 bytes, which is 48 GB in the decimal sense this series uses for byte math. Converting to the binary units your driver actually reports, 48e9 divided by 2 to the 30 is about 44.7 GiB.
A 32 GB card gives you roughly 29.8 GiB once the driver and CUDA context take their cut. 44.7 GiB is nowhere near 29.8 GiB, so this does not fit, and it does not fit by a wide enough margin that no amount of the usual 2 to 3 GB safety cushion changes the answer. Note too that this is the static-state figure alone, weights, gradients and optimizer. Activations still have not been added. A model that fails this screen fails before you even get to the harder question of batch size and sequence length.
Exercise 2
LoRA: 3,000,000,000 times 2 is 6,000,000,000 bytes, about 6 GB. QLoRA: 3,000,000,000 times 0.5 is 1,500,000,000 bytes, about 1.5 GB. Both comfortably clear a 32 GB card on paper.
What both numbers miss is activation memory, and that is the whole point of the exercise. This part’s own running example states plainly that LoRA and QLoRA freeze the base model but do not touch the forward pass, so the same activation bill the full fine-tune paid is still owed. For the 1.5B model in this series that bill runs about 8 GB at batch 4 and sequence 1024, an order-of-magnitude figure that Part 3 derives properly. The trap at 3B is exactly the one this part warns against: quoting a static-state number next to a fits-or-not verdict without saying whether activations are included. A 3B model very likely has a larger hidden size and possibly more layers than the 1.5B example, so its activation footprint at the same batch and sequence would be larger than 8 GB too, not smaller and not the same. You cannot answer whether 6 GB of LoRA static state actually fits until you have that number, which this part deliberately does not hand you. That is Part 3’s job.
Exercise 3
The two causes this part gives are: the card is genuinely training, so the extra memory is the other three tenants doing real work, or the card is serving through an engine that pre-allocates a KV cache pool at startup, which looks alarming in nvidia-smi but is mostly a configuration number rather than memory actually in use for the current request.
Since the scenario is explicitly serving, not training, the second explanation is the one that applies. A training process would need to be running for the first explanation to hold, and nothing in the scenario says training is happening. The one-minute check: watch whether the reported usage moves when a request is sent versus when the card is idle. A pre-allocated KV cache pool sits at roughly the same size whether or not a request is in flight, because it was reserved once at startup for the engine’s maximum batch and context settings. Real training memory, by contrast, tracks the training loop, rising through the forward pass and draining through the backward pass. A quicker version of the same check is to open the serving engine’s own startup logs or config for a memory or cache size setting; that number is usually printed at boot and will match the alarming figure almost exactly.
Exercise 4
Path one: accept the quality gap and ship QLoRA or LoRA anyway. You get a model that fits on the card you have, trains in a reasonable time, and, per Biderman et al., forgets less of what the base model already knew than a full fine-tune would. If maths is one capability among several the model needs rather than the only thing that matters, that forgetting-less property can be a genuine upside, not just a consolation prize.
Path two: get an actual full-rank fine-tune of the 7B model by moving to more than one GPU, so the 122 GB of static state is sharded across cards instead of squeezed onto one. This is the multi-GPU territory Part 14 covers, DDP, ZeRO and FSDP. It costs more hardware and more operational complexity, not more cleverness. There is no single-GPU trick, no quantisation scheme and no clever batching arrangement that gives you full fine-tuning’s maths and code quality on a 7B model from one 32 GB card. The paper this part cites is explicit that the gap is real for exactly this kind of task, and the honest answer is to pick which cost you are willing to pay: quality, or hardware.