Exercise solutions: Choosing a Base Model and Building a Fine-Tuning Bench
These are the worked solutions for the exercises in Part 8, Choosing a Base Model and Building a Fine-Tuning Bench. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
Standard mixed-precision AdamW costs 16 bytes per parameter of static state. For 3 billion parameters: 3e9 * 16 bytes = 4.8e10 bytes = 48 GB. A 32 GB card has about 29.8 GiB usable. Forty-eight gigabytes is well beyond that, more than 18 GB over even before counting any activation memory on top of it, so a 3B model cannot be fully fine tuned on this card at all. It would need LoRA or QLoRA instead, exactly the point at which this part’s fourth criterion, small enough to fully fine tune on one card, stops being satisfied and the choice moves to a different training method rather than a different hyperparameter.
Exercise 2
Qwen2.5-1.5B’s tied embedding matrix has 151,000 * 1,536 = 231,936,000 parameters, about 232 million, which is 232 / 1,540 = 0.151, about 15.1 percent of its 1.54 billion total. SmolLM2-1.7B’s stated 100 million parameter embedding matrix is 100 / 1,710 = 0.058, about 5.8 percent of its 1.71 billion total. Qwen spends roughly two and a half times the fraction of its parameter budget on the embedding table that SmolLM2 does, about 15 percent against about 6 percent, even though SmolLM2 is the larger model overall, because Qwen’s vocabulary is about three times larger. This is the double-edged detail this part names directly: a larger vocabulary buys multilingual and code coverage at the direct cost of parameters that could otherwise sit in the transformer body doing the actual reasoning work.
Exercise 3
The most likely missing piece is nvrtc, CUDA’s runtime compiler, absent from some installers on very new architectures, which breaks anything that tries to compile a kernel on the fly rather than loading one already compiled ahead of time. This part’s two defenses: prefer prebuilt wheels compiled ahead of time for your exact architecture over anything that compiles from source at import or run time, and for a first run use PyTorch’s built-in scaled dot product attention rather than an external flash-attention package, since a small supervised run does not need the extra throughput, and this removes the exact code path that hit the missing compiler in the first place.
Exercise 4
Full fine tune static state for Qwen2.5-1.5B is about 24.6 GB, plus an order-of-magnitude activation figure of about 8 GB at batch 4 and sequence 1024, for a total near 32.6 GB before any safety margin. The g6.xlarge’s 24 GB cannot hold even the static state alone, 24.6 GB, let alone the activations on top, so it is out for a full fine tune; it is exactly the right size for LoRA or QLoRA instead. The g6e.xlarge’s 48 GB comfortably covers the full 32.6 GB with about 15 GB of headroom to spare, enough for the two to three gigabyte safety margin this series recommends and then some. The instance to provision for a full fine tune is the g6e.xlarge.
Exercise 5
SmolLM2-1.7B is the only one of the three with a fully disclosed training data mix, which is what a model-risk review actually needs, but it is English first and its ecosystem is the smallest of the three. Qwen2.5-1.5B has the strongest multilingual reach and the deepest tooling, but its data transparency stops at open weights, not an open training corpus. Llama-3.2-1B has very deep tooling too, but its data is not disclosed either, it carries a community licence with its own review overhead, and it covers only eight languages. No candidate clears all three bars because they are genuinely different bets: openness of the weights, openness of the process that produced them, and how mature the surrounding tooling is, do not move together. Given the review’s requirement is specifically about provenance rather than performance, the requirement not to compromise on is data transparency, because a smaller ecosystem is a cost that more engineering time can absorb, while a training corpus that was never disclosed cannot be made auditable after the fact no matter how much time is spent on it.