Exercise solutions: Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
These are the worked solutions for the exercises in Part 4, Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4. Read the exercise first; coming here before you have tried it defeats the point.
Exercise 1
For a binary floating-point format, the gap between representable values near a magnitude v is about 2 to the power of, floor of log base 2 of v, minus the mantissa bit count. Near 0.5, which is 2 to the power of minus 1, and with bf16’s 7 mantissa bits, that gap is 2 to the power of minus 1 minus 7, minus 8, which is about 0.0039, matching this part.
Near 2.0, which is 2 to the power of 1, the same formula gives 2 to the power of 1 minus 7, minus 6, which is about 0.0156, four times coarser than the gap near 0.5, exactly in proportion to 2.0 being four times larger than 0.5.
An update of 0.0003 near a weight of 2.0 is roughly 52 times smaller than that 0.0156 gap, so it vanishes exactly the same way the 0.0001 update near 0.5 did: 2.0000 plus 0.0003, computed and stored in bf16, rounds straight back to 2.0000.
The generalisable point is that bf16’s blind spot is not a fixed absolute number. It is a fixed relative resolution, roughly 2 to the minus 7, about 0.4 percent of the weight’s own magnitude, that scales up in absolute terms as the weight itself gets larger. 0.0039 near 0.5 and 0.0156 near 2.0 are the same relative blind spot expressed at two different magnitudes. This is exactly why an fp32 master copy is not a special-case fix for weights near 0.5; it is needed at every magnitude a weight can take, because the relative gap is always there.
Exercise 2
FP16’s exponent is only 5 bits, which caps its range far tighter than bf16’s 8-bit exponent, and makes values below roughly 6e-8 underflow silently to zero. Gradients late in training, once the model is close to converged, are routinely that small for at least some parameters. If loss scaling is not being used correctly, and the scenario states nobody touched it, then either it defaults to off or to a static constant that no longer suits gradients that have shrunk this far into training. Small gradients underflow to zero, which alone would stall learning quietly rather than produce NaN. NaN specifically points at the opposite edge of the same problem: a fixed scaling constant chosen to rescue small gradients earlier in training can push other, still reasonably sized values in the multiply-then-divide chain past fp16’s maximum of 65,504, producing infinity, which then poisons the loss with NaN on the next arithmetic step. bf16 never needs this rescue in the first place because it shares fp32’s 8-bit exponent and therefore fp32’s range, so it experiences neither the underflow nor the overflow side of the same fragile mechanism.
The concrete fix, short of abandoning fp16 for bf16, is dynamic loss scaling rather than a single static constant: the scale factor is adjusted automatically step by step, reduced whenever an overflow is detected in that step’s gradients and increased gradually when steps run clean for a while. This removes the need to hand-pick one constant that has to stay correct across the entire, shrinking range of gradient magnitudes a full training run passes through.
Exercise 3
E4M3 has 3 mantissa bits, giving a relative resolution of about 2 to the minus 3, about 1 part in 8, roughly 12.5 percent between representable steps. E5M2 has 2 mantissa bits, giving about 2 to the minus 2, 1 part in 4, roughly 25 percent between steps. Switching this tensor from E4M3 to E5M2 roughly doubles the relative rounding error on every value it stores.
E4M3’s documented maximum magnitude is about 448. This tensor’s actual range, about minus 2.1 to 3.4, sits nowhere near that ceiling; E4M3 already has far more headroom than this tensor will ever use. E5M2’s maximum of about 57,344 is enormously larger still, but since E4M3 was never the constraint to begin with, none of that additional range buys anything real for this specific tensor. The result of the swap is a clean loss: about twice the rounding error per value, in exchange for dynamic range that was already unused before the swap and remains unused after it. This is the general shape of the E4M3 versus E5M2 decision this part describes: give exponent bits to whichever tensor actually needs the range, gradients under FP8 training being the usual example, and give mantissa bits to whichever tensor’s values already sit comfortably within a bounded range, which weight tensors like this one normally do.
Exercise 4
Weights move from bf16’s 2 bytes per parameter to FP8 E4M3’s 1 byte per parameter. Gradients move from bf16’s 2 bytes to FP8 E5M2’s 1 byte. The optimizer tenant is untouched: it is still a 4-byte fp32 master copy plus two 4-byte fp32 running moments, 12 bytes per parameter, because nothing about the optimizer’s own state depends on what dtype the forward and backward compute happened to run in.
New total: 1 plus 1 plus 12 is 14 bytes per parameter, down from 16. For Qwen2.5-1.5B at 1.54 billion parameters: 14 times 1.54 billion is 21.56 billion bytes, about 21.6 GB, against the previous about 24.6 GB, a saving of about 3 GB, roughly a 12 percent reduction.
The tenant this change can never touch is the optimizer. It is 12 of the 16 original bytes, three quarters of the total, and it is fixed by the optimizer’s own numeric requirements, not by whatever precision the compute happens to run in. Since only the weights and gradients tenants are compute-precision-dependent, and together they are only 4 of the 16 bytes, the maximum possible saving from any forward-and-backward compute-precision trick, bf16 to FP8 or anything smaller still, is capped at 4 bytes per parameter, 25 percent of the total. To meaningfully shrink static-state memory beyond that ceiling, you have to attack the optimizer directly, by quantising its state as 8-bit Adam does, or by freezing the weight entirely so it never needs optimizer state at all, which is what LoRA does.