Exercise solutions: AdamW Explained, Line by Line
These are the worked solutions for the exercises in Part 6, AdamW Explained, Line by Line. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
Each step folds the new gradient in at weight one minus beta and carries the old average forward at weight beta, so the influence of a gradient from n steps back shrinks in proportion to beta multiplied by itself n times. That influence falls to about a third, one over e, once n is close to 1/(1-beta), because the natural log of beta sits close to -(1-beta) whenever beta is close to 1. Checking the formula against the two numbers this part already gives confirms it: beta one at 0.9 gives 1/(1-0.9) = 10 steps, and beta two at 0.999 gives 1/(1-0.999) = 1000 steps, both matching. Applying the same formula to the two new values: beta two at 0.95, the pretraining setting, gives 1/(1-0.95) = 20 steps, a much shorter variance memory than the default 1000, which is exactly why it can track a fast-changing pretraining loss landscape without the brake going stale. Beta one at 0.99 gives 1/(1-0.99) = 100 steps, ten times longer than the usual 0.9, which would make the smoothed direction react far more slowly to a genuine change in where the loss wants to go.
Exercise 2
With the running average starting at zero, one step of the update is the old average times beta one plus the new gradient times one minus beta one. For a constant gradient g this has a closed form: the raw average at step t equals g times one minus beta one raised to the power t, which you can check by substitution. At step one: m1 = 0.9*0 + 0.1*g = 0.1g, matching g*(1-0.9) = 0.1g. At step two: m2 = 0.9*0.1g + 0.1g = 0.19g, matching g*(1-0.81) = 0.19g. At step three: m3 = 0.9*0.19g + 0.1g = 0.271g, matching g*(1-0.729) = 0.271g. Bias correction divides by exactly that same factor, one minus beta one raised to the power t, so the corrected estimate is g at every single one of the three steps, not just approaching g over time. The correction is not something that merely improves with more steps here, it is exact from step one, because bias correction was built to cancel precisely this startup-at-zero bias, and a constant gradient is the case where that cancellation is perfectly clean rather than only asymptotic.
Exercise 3
Two states at four bytes each is eight bytes per parameter for m and v alone, ignoring the fp32 master copy. Qwen2.5-1.5B has 1.54 billion parameters, so 1.54e9 * 8 bytes = 12.32 GB for adamw_torch’s m and v. Paged_adamw_8bit stores each of those two states in about a quarter the space, so 12.32 GB / 4 = 3.08 GB. The saving is 12.32 - 3.08 = 9.24 GB, on the two optimizer states alone. Put next to the full sixteen-bytes-per-parameter accounting this series uses elsewhere: the standard optimizer share is twelve bytes, four for m, four for v, four for the fp32 master weight copy mixed precision keeps. Quantising only m and v to one byte each takes that twelve down to four plus one plus one, six bytes, the same figure the part on running a supervised fine tune gives for this exact switch. The two numbers describe the same saving from two angles: this exercise’s 9.24 GB is the states alone, the six-byte figure includes the master copy that never shrinks.
Exercise 4
The formula from the first exercise gives beta two at 0.5 a memory window of 1/(1-0.5) = 2 steps, against 1/(1-0.95) = 20 steps for the pretraining value that was intended, ten times shorter than what the run needed and five hundred times shorter than the fine tuning default of 0.999. The per-weight brake is the square root of v sitting in the denominator of the update. With a memory of only two steps, v tracks almost nothing but the last gradient or two, so a single unusually large gradient inflates v immediately, shrinking that weight’s next step sharply, then v decays back down within another step or two once the large gradient has passed, letting the next ordinary gradient through against a denominator that is now too small for it. The brake swings between too tight and too loose from one step to the next instead of reflecting a stable estimate of that weight’s typical gradient size, which is exactly the noisy, over-reactive denominator that turns an ordinary run of gradients into an oversized, destabilising step.
Exercise 5
LoRA at rank 16 on Qwen2.5-1.5B brings static state to about 4 GB, per this series’ running numbers. Activations at batch 4 and sequence 1024 are an order-of-magnitude figure of about 8 GB, and LoRA does not change that, since it freezes the base model but leaves the forward pass, and therefore the activations, untouched. Total: 4 + 8 = 12 GB, against a 32 GB card’s roughly 29.8 GiB usable. Even keeping the two to three gigabyte safety margin this series recommends, headroom is on the order of 29.8 - 12 = 17.8 GiB, nowhere near the wall. Paged_adamw_8bit exists to solve VRAM pressure, and there is none here: the memory this run actually needs is well under half of what the card offers. Reaching for it anyway buys nothing but a small amount of quantisation noise in the optimizer state for no memory benefit that matters. The better pick for this run is the plain or fused adamw_torch, and this part’s own guidance backs that: reach for the 8-bit variant only once you have actually hit a wall the first two tricks could not clear, which this run never does.