Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4

Inside LLM Fine-Tuning, part 4 of 15: Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4

The bf16 vs fp16 question looks like a detail about storage. It is the difference between a training run that finishes and one that dies at step 4,000 with a screen full of NaN. Both formats are 16 bits wide. Both halve your memory against the 32-bit default. One of them needs a workaround bolted on the side, and the other does not, which is why current recipes set bf16=True and move on.

This part explains why, starting from what a bit is. It assumes you have never thought about how a computer stores a decimal number. Everything is built from one toy format and one worked example you can check on paper.

By the end you will be able to take any format name apart on sight, work out the smallest change it can record, and tell a format that shrinks your memory from one that only changes how fast the arithmetic runs.

The previous part counted activation memory in bytes. This part is about what one of those bytes holds.

Take a number apart into sign, exponent and mantissa

A computer has to store 0.0271 and 3.4 billion and 0.00000006 in the same fixed-size box. The scheme it uses is called floating point, and every format in this article is one version of it.

A bit is one yes or no

A bit is a single slot holding either 0 or 1. Eight bits make one byte. The useful thing about bits is that each extra one doubles how many different patterns you can write down. 8 bits give 256 patterns. 16 bits give 65,536. 32 bits give about 4.3 billion.

That is the honest way to read a 16-bit number format. It does not mean “16 bits of accuracy”. It means the format has exactly 65,536 named values, and its whole design is the decision about which 65,536 numbers those patterns stand for.

Every float copies scientific notation

You already know the scheme. Write 6,250 as 6.25 times 10 to the power 3. Two separate facts are sitting there: the digits 6.25, and the scale 3. Write 0.00625 as 6.25 times 10 to the power minus 3, and the digits are identical while the scale has changed. A computer does the same in base 2, and stores the two facts in separate fields.

The three fields

  1. The sign bit. Always exactly 1 bit. A 0 means positive, a 1 means negative. No format spends more than one bit here.
  2. The exponent. This is the scale. It records roughly how big the number is, in powers of two.
  3. The mantissa. This is the detail. It records where the number sits inside the size the exponent named. You will also see it called the significand. The two words mean the same field.

The sign bit is fixed at 1 forever, so the only real decision a format makes is how to split the rest between exponent and mantissa.

± EXPONENT → range MANTISSA → precision value = (−1)^sign × 2^exponent × 1.mantissa more exponent bits → wider reach · more mantissa bits → finer resolution

The whole article lives in the last two boxes. Exponent bits say how far the format reaches. Mantissa bits say how sharply it sees.

Shelves and slots, the picture to hold on to

Picture a ladder of shelves. One shelf covers everything from 1 up to 2. The next covers 2 to 4. Then 4 to 8, then 8 to 16, and so on upward. Going down: 0.5 to 1, then 0.25 to 0.5, then 0.125 to 0.25. Each shelf covers twice as much ground as the one below.

Now the two fields have jobs you can see.

  • The exponent picks the shelf. More exponent bits means more shelves, so the ladder reaches further up and further down.
  • The mantissa picks the slot on that shelf. Every shelf is cut into the same number of evenly spaced slots, and a mantissa of M bits gives 2 to the power M slots. More mantissa bits means finer detail everywhere.

Notice what falls out of that. Higher shelves are wider but hold the same number of slots, so their slots are further apart. A format’s precision is never one fixed number. It is always a fraction of whatever value you are near.

One worked example, in a toy format

Invent an 8-bit format: 1 sign bit, 3 exponent bits, 4 mantissa bits. Store 6.25 in it.

  1. 6.25 is positive, so the sign bit is 0.
  2. 6.25 sits between 4 and 8, so it belongs on the shelf running from 4 to 8. That is what the 3 exponent bits record.
  3. 4 mantissa bits give 16 slots on that shelf. The shelf is 4 units wide, so the slots are 4 divided by 16, which is 0.25 apart: 4.00, 4.25, 4.50, and so on up to 7.75.
  4. Count nine slots up from 4.00 and you land on exactly 6.25. Check it: 4.00 plus 9 times 0.25 is 6.25. Stored exactly, nothing lost.

Now try 6.30. The two nearest slots are 6.25 and 6.50, and 6.30 is closer to the first, so it is stored as 6.25 and the remaining 0.05 disappears. That is the entire mechanism of precision loss. A value that does not land on a slot moves to the nearest one. It is rounding, and here it costs 0.8 percent.

Measure a format’s range and its precision in numbers you can check

Two words get used loosely in every discussion of number formats. Both have definitions you can compute.

Range is how far the ladder reaches. It takes two numbers to state: the largest magnitude the format holds, and the smallest non-zero magnitude. The toy format runs from 0.25 at the bottom up to just under 16 at the top.

Precision is how far apart neighbouring slots are. Because slots widen as you climb, you have to say where you are measuring. Near 6.25 the toy’s slots are 0.25 apart. Near 0.5 they are 0.03125 apart, because that shelf is only 0.25 wide and still holds 16 slots.

The handy version is relative. With M mantissa bits, neighbouring slots differ by roughly one part in 2 to the power M of the number itself. The toy has 4 mantissa bits, so about one part in 16, which is why the gap near 6.25 came out at 4 percent of 6.25. Precision is a percentage, not a count of decimal places.

Move one bit and watch both numbers change

Keep the toy at 8 bits, but move one bit from the mantissa to the exponent. That gives 1 sign bit, 4 exponent bits, 3 mantissa bits.

What precision does. Slots per shelf halve, from 16 to 8. On the shelf from 4 to 8 they are now 0.5 apart: 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5. The number 6.25 is no longer representable at all. It sits exactly halfway between two slots.

What range does. The number of shelves roughly doubles. The top of the ladder moves from just under 16 to just under 256. The smallest non-zero value drops from 0.25 to about 0.0156. Sixteen times further in each direction, bought with a single bit.

One bit, two large and opposite effects. That is the whole design space, and every real format is a point in it. Hold that arithmetic, because it is why a 5-bit exponent tops out at 65,504 while an 8-bit exponent reaches 3.4 followed by 38 digits.

Tell rounding apart from deletion: underflow and overflow

Rounding is not the only thing that can happen to a number. Two other things can, and they are worse.

Underflow is what happens when a value is smaller than the smallest non-zero magnitude the format holds. It does not round to something small. It becomes exactly zero.

One wrinkle matters, because a famous number depends on it. Below the bottom shelf, formats keep a last stretch of evenly spaced values running down to zero. These are called subnormal numbers. They are spaced at the bottom shelf’s smallest step, so they push the floor down by a factor of 2 to the power M. In the toy format the bottom shelf starts at 0.25 with slots 0.015625 apart, so subnormals reach 0.015625 and stop. Anything smaller becomes zero.

Overflow is the same failure at the other end. A value larger than the format’s biggest number becomes infinity. Infinity is contagious: infinity minus infinity is undefined, and undefined operations produce NaN, which stands for “not a number”. Every later operation touching a NaN produces another one. Once one reaches the weights, every output is NaN and the run is dead with no useful error message.

Why one of these is survivable and the other is not

  • Rounding moved 6.30 to 6.25. The value is still approximately what it was.
  • Underflow moved a gradient of 1e-8 to 0. That is not approximately 1e-8. A gradient of zero is a different instruction: it says leave this weight exactly where it is.
  • Overflow moved a large value to infinity, and from there to NaN, which ends the run.

Networks are remarkably tolerant of the first and completely intolerant of the other two. Three reasons explain the tolerance.

  1. A rounded gradient still points the same way. A gradient is one number per weight saying which direction to nudge it, an idea Part 1 introduced. Round 0.0012374 to 0.0012 and it still says “reduce this weight, by roughly this much”. Only the exact size shifted.
  2. The errors cancel. A batch averages many examples, so it averages many roundings too. Some go up, some go down, and the average lands close to where it would have anyway.
  3. The signal was never exact. Each step already uses a random sample of the data, so the gradient is already an estimate with noise on it. Rounding adds a little more noise to something approximate. The walkthrough of how a network learns from one neuron up covers why a slightly wrong gradient still trains.

Now the reason range cannot be given up. One training step holds numbers spanning an enormous span of magnitudes. Weights sit around 0.01 to 1. Activations sit around 1. Gradients run from about 1e-4 down to 1e-8 and below as a run converges. Square one of those gradients, which is exactly what the optimizer does for its own bookkeeping, and 1e-7 becomes 1e-14. That is fourteen orders of magnitude alive at once.

1 · NEEDS RANGE gradients ~1e−8, activations ~1e+3 span many orders of magnitude lose range → value gone 2 · STILL DOWNHILL a slightly-rounded gradient still points roughly the right way 3 · AVERAGES OUT rounding noise is roughly zero-mean; over millions of steps it cancels

The first panel is the range argument and the other two are the precision argument. Together they say to spend bits on the exponent whenever the choice is close.

Which gives the sentence that decides every format choice here. Rounded is not lost. Out of range is lost. A rounded gradient is a slightly wrong instruction, and the training loop absorbs slightly wrong instructions all day. A gradient that underflowed has been deleted, and no amount of averaging brings back a number that is no longer there.

Read FP32, FP16 and BF16 straight off their bit counts

Every tool is on the table, so the three formats that cover almost all training can be read directly.

  • FP32: 1 sign bit, 8 exponent bits, 23 mantissa bits. 32 bits, so 4 bytes per number. The reference format.
  • FP16: 1 sign bit, 5 exponent bits, 10 mantissa bits. 16 bits, so 2 bytes.
  • BF16: 1 sign bit, 8 exponent bits, 7 mantissa bits. Also 16 bits, so also 2 bytes. The name is short for brain float, after the Google Brain team that designed it.
FP32 8 exp 23 mantissa range 3.4e38 · ~7 dig FP16 5 exp 10 mantissa range only 65,504 · ~3–4 dig BF16 8 exp 7 mant range 3.4e38 · ~2–3 dig Dashed line: FP32 & BF16 exponents end at the same place → identical range. FP16’s green bar is visibly shorter.

BF16 is FP32 with 16 mantissa bits deleted and nothing else touched. FP16 spent the same 16-bit budget the other way, buying slots at the cost of shelves.

Work each one out with the shelf rule. None of it needs memorising, because all of it is derivable.

FP32. 8 exponent bits give 256 patterns, of which 254 are usable shelves. The ladder tops out just under 2 to the power 128, about 3.4e38, and its lowest full shelf sits at about 1.2e-38. 23 mantissa bits give 8,388,608 slots per shelf, so near 0.5 the gap between neighbouring values is 0.5 divided by 8,388,608, about 6e-8.

FP16. 5 exponent bits give 32 patterns, of which 30 are usable shelves. That is a short ladder. Its top shelf runs from 32,768 to 65,536, and 10 mantissa bits give 1,024 slots on it, so up there the slots are 32 apart. The highest value FP16 can hold is therefore 65,536 minus 32, which is 65,504. That is where the famous number comes from, and you just derived it. At the bottom its lowest full shelf is about 6.1e-5, and subnormals stretch that down by a factor of 1,024 to about 6e-8, the absolute floor.

BF16. 8 exponent bits, exactly FP32’s, so exactly FP32’s shelves: about 3.4e38 at the top, about 1.2e-38 at the bottom. 7 mantissa bits give only 128 slots per shelf, so near 0.5 the gap is 0.5 divided by 128, which is 0.0039. A thousand times coarser than FP32 at the same magnitude.

In one line: BF16 is FP32 with 16 mantissa bits thrown away and the exponent left alone.

Format Sign / exponent / mantissa Bytes Biggest value Gap near 0.5
FP32 1 / 8 / 23 4 about 3.4e38 about 0.00000006
FP16 1 / 5 / 10 2 65,504 about 0.0005
BF16 1 / 8 / 7 2 about 3.4e38 about 0.0039

Consequence one: FP16’s cliff, and the workaround bolted on to survive it

FP16’s short ladder has a hard edge at each end, and the bottom edge is the dangerous one.

1e−81e−41 65,5043.4e38 FP16 usable CLIFF → ∞ ↓ underflow to 0 below ~1e−8 BF16 & FP32 usable — all the way to 3.4e38 gradients (~1e−8) & activations (~1e+3) live across this whole span — FP16 can’t hold the ends

Every tick is a factor of ten. Where FP16 stops is not a rounding boundary, it is a cliff, and values past it turn into zero or infinity rather than into approximations.

Gradients late in a run are routinely down at 1e-7 and below, because the model is nearly settled and the remaining corrections are tiny. In FP16 those land in the subnormal stretch, where precision is already degrading, and plenty fall clean off the bottom at 6e-8 and become exactly zero. Nothing errors. The loss stops improving and the run burns GPU hours learning nothing.

The fix is called loss scaling. It works in four steps.

  1. Before the backward pass, multiply the loss by a large constant. Call it S. A common starting value is 65,536.
  2. Every gradient in the model then comes out exactly S times larger than it would have been. A gradient that would have been 3e-8, safely below the floor, is now 3e-8 times 65,536, about 0.002. Comfortably inside FP16’s range.
  3. Before the optimizer applies anything, divide every gradient by S again. The maths is unchanged and the small values made the trip in one piece.
  4. Everything now depends on S being the right size, which is the problem.
without scaling (FP16): grad = 1e−8 → underflows → 0 (deleted) with loss scaling (S = 65536): loss × S → grad × S → 6.5e−4 (in range ✓) → ÷ S before optimizer step → update lands BF16 skips this whole dance — its FP32-sized range means the 1e−8 gradient was representable all along.

Loss scaling exists only because of the cliff. Too small an S and gradients still underflow. Too large and they overflow to infinity, which is where the silent NaN comes from.

Pick S too small and the smallest gradients still underflow. Pick S too large and some gradients overflow past 65,504, become infinity, and turn into NaN one operation later. Gradients shrink as a run progresses, so a constant that was right at step 500 can be wrong by step 4,000. The standard defence is dynamic loss scaling: start S high, halve it whenever a step overflows, creep it back up when steps run clean. It works, and it is still a moving part that can be misconfigured.

BF16’s 8 exponent bits make the whole mechanism unnecessary. Its floor is 1.2e-38, so a gradient of 1e-8 is nowhere near trouble. That is the practical reason BF16 became the training default the moment hardware supported it: same 2 bytes, one fewer thing to tune, one fewer way for a run to die at step 4,000 for no visible reason.

Consequence two: BF16’s coarse mantissa, and the FP32 copy that rescues it

BF16 pays for its range with only 7 mantissa bits, and that bill comes due in one specific place.

Take a weight sitting at 0.5 and an update of 0.0001. Near 0.5 the gap between BF16 values is 0.0039, so the next value BF16 can represent above 0.5 is 0.50390625. Adding the update gives 0.5001, nowhere near halfway to that next slot, so it rounds straight back down. The weight stays at exactly 0.5 and the update is gone.

FP32’s gap near 0.5 is about 6e-8, so the same update is about 1,700 times larger than the smallest change it can record. FP32 keeps it easily.

in BF16 step near 0.5 ≈ 0.0039 0.5000 + 0.0001 → 0.5000 update below the step size — vanishes in FP32 (master copy) step near 0.5 ≈ 6e−8 0.5000 + 0.0001 → 0.5001 update preserved ✓

The same update in two formats. This is the whole reason mixed-precision training keeps a separate 32-bit copy of every weight.

One lost update does not matter. Thousands do. A fine-tune is a few thousand steps, and if each step’s correction rounds away, the weight never moves. The fix has a name, mixed precision, and one rule: do the fast arithmetic in BF16 for the speed and the halved memory, but keep each weight’s running total in FP32 so small updates accumulate instead of evaporating. That FP32 master copy is 4 of the 12 bytes of optimizer state the part on the four memory tenants itemises. From the memory side it looks like overhead. From the numerics side it is the only thing between you and a model that quietly stops learning.

None of this has to be taken on trust. The following prints every number in this section.

import torch

# The three formats, straight from the library.
for dt in (torch.float32, torch.bfloat16, torch.float16):
    fi = torch.finfo(dt)
    print(dt, "biggest", fi.max,
          "smallest full-shelf value", fi.smallest_normal,
          "step just above 1.0", fi.eps)

# The update that disappears.
w   = torch.tensor([0.5],    dtype=torch.bfloat16)
upd = torch.tensor([0.0001], dtype=torch.bfloat16)
print("bf16:", (w + upd).item())        # 0.5, the update is gone

w32 = torch.tensor([0.5], dtype=torch.float32)
print("fp32:", (w32 + 0.0001).item())   # 0.5001, the update survives

The eps figures are the step just above 1.0, one shelf up from the 0.5 numbers above, so each is exactly double: 0.0078 for BF16 and about 1.19e-7 for FP32.

What the choice costs on the running example

Qwen2.5-1.5B, the model this series sizes everything against, has 1.54 billion parameters. Multiply that by the bytes each format spends per number.

Weights held in Bytes per parameter Qwen2.5-1.5B weights alone
FP32 4 about 6.2 GB
BF16 or FP16 2 about 3.1 GB
FP8 1 about 1.5 GB
NF4 about 0.5 about 0.8 GB

That column is the weights only. A full fine-tune under standard mixed-precision AdamW costs 16 bytes per parameter once gradients and optimizer state are counted, which is about 24.6 GB of static state here, with roughly 8 GB of activations owed on top as an order of magnitude. Part 2 does that accounting in full. The narrower point: the format decision moves the first row of the bill by a factor of eight.

Explain why there is no BF32

If BF16 is simply “keep FP32’s range and cut the mantissa”, the obvious question is where BF32 is. Walk the bit budget and the answer arrives on its own.

The brain-float recipe has exactly two rules. Keep FP32’s 8 exponent bits. Cut mantissa bits until the number fits the box you are aiming at.

Aim it at a 16-bit box.

  1. Budget: 16 bits.
  2. The sign bit takes 1, leaving 15.
  3. The exponent takes its 8, leaving 7.
  4. The mantissa gets what is left, so 7 bits.

That is BF16, and it is genuinely a new format. FP32’s mantissa was 23 bits and you cut 16 of them away.

Now aim it at a 32-bit box.

  1. Budget: 32 bits.
  2. The sign bit takes 1, leaving 31.
  3. The exponent takes its 8, leaving 23.
  4. The mantissa gets what is left, so 23 bits.

Read that result: 1 sign bit, 8 exponent bits, 23 mantissa bits. That is FP32, exactly. You cut nothing, because there was nothing to cut.

recipe for a brain-float: keep FP32’s 8-bit exponent, then fit into the target width. target 16 bits → 8 exp 7 mant = BF16 ✓ (had to cut mantissa to 7 — a real, new format) target 32 bits → 8 exp 23 mant (nothing to cut) 8-bit exponent + 32 bits total = 1/8/23 = FP32. There’s no mantissa to remove, so no smaller-than-FP32 thing is created. “BF32” is just FP32 with a new label.

Follow the construction to its end and it lands where it started. An 8-bit exponent inside a 32-bit float is the definition of FP32.

So BF32 is a second name for a format that already exists. The recipe only produces something new when the target box is smaller than 32 bits, because only then are you forced to cut. The one way to make it non-trivial would be to start from FP64’s 11 exponent bits and shrink to 32, and nobody wants that, because machine learning has never needed FP64’s reach.

Tell a storage format from a compute mode, and switch TF32 on

There is a real idea in the neighbourhood of BF32, and it did not ship as a format. The idea is FP32’s range with less precision, so 32-bit matrix multiplies run faster. NVIDIA built exactly that as TF32: FP32’s 8 exponent bits with FP16’s 10 mantissa bits, which is 19 bits carrying information once the sign bit is counted.

The part everyone misses is that TF32 is not something you store. This distinction is the most valuable idea in the part, so here it is in full.

  • A storage format answers one question: how many bytes does one number occupy in memory? BF16 answers 2. FP32 answers 4. Change it and your memory footprint changes by a known factor.
  • A compute mode answers a different question: how carefully does the chip do the multiply? Change it and only speed changes. Not one byte moves.

Every format so far has been a storage format. You write dtype=torch.bfloat16 and each number in that tensor takes 2 bytes, and you can watch the memory drop.

TF32 is not one of those. There is no torch.tf32. You never allocate a TF32 tensor. Here is what actually happens on a matrix multiply once TF32 is on.

  1. Your tensor sits in memory as FP32, at 4 bytes per number, exactly as before.
  2. The tensor core reads those FP32 values and throws away 13 of the 23 mantissa bits, keeping 10.
  3. It does the multiplications at that reduced precision.
  4. It adds the products up in full FP32 and writes an FP32 result back to memory.
MEMORY FP32 (1/8/23) TENSOR CORE truncate inputs → TF32 (1/8/10) multiply in TF32 accumulate → FP32 MEMORY FP32 result You never see TF32 in memory — inputs and outputs are FP32; the truncation happens on the fly inside the core.

FP32 going in, FP32 coming out. The truncation happens inside the multiply and never touches memory, so the footprint is unchanged.

Why is it faster? A multiplier circuit that only handles 10 mantissa bits is far smaller in silicon than one handling 23, so the chip fits many more of them and gets through many more multiplies per cycle. A tensor core is the part of an NVIDIA GPU built for exactly this, and TF32 is one of the precisions it accepts.

Why is it safe? TF32 keeps all 8 of FP32’s exponent bits. Nothing underflows or overflows that would not have done so anyway. The only change is rounding, which is the cheap kind of loss.

Switching it on is the whole cost. Stock PyTorch ships TF32 matrix multiplies off. The flag torch.backends.cuda.matmul.allow_tf32 defaults to False, because PyTorch’s default is to give you precisely what FP32 promises and make reduced precision something you ask for.

# Stock PyTorch ships TF32 matmuls OFF. This is the whole cost of turning them on.
torch.backends.cuda.matmul.allow_tf32 = True

# The newer spelling of the same idea, covering more operations:
torch.set_float32_matmul_precision("high")

Pin your versions and check the docs for the one you pinned, since PyTorch has added newer spellings of this setting over time. One thing to know before reaching for it: TF32 only helps tensors that are genuinely FP32. If you already train in BF16, your matrix multiplies are already running at low precision and there is nothing left to speed up. It is fair to call TF32 the useful BF32 that shipped as a maths mode rather than a dtype.

Pick the right FP8 flavour, E4M3 or E5M2

Below 16 bits the range-against-precision trade sharpens, which is why FP8 shipped as two formats rather than one.

The names look like part numbers. They are not. They spell out the layout. E is the number of exponent bits and M is the number of mantissa bits. So E4M3 means 4 exponent bits and 3 mantissa bits, and E5M2 means 5 and 2. Add the sign bit every float carries and both come to 8 bits, which is 1 byte. In that notation the toy format earlier was E3M4, and the variant you built by moving one bit was E4M3, which is a real format in real silicon.

  • E4M3: 3 mantissa bits, so 8 slots per shelf, so about one part in 8 between neighbouring values, roughly 12 percent. Ceiling 448. It favours precision.
  • E5M2: 2 mantissa bits, so 4 slots per shelf, so about one part in 4, roughly 25 percent. Ceiling 57,344, about 128 times higher. It favours range.

One footnote on that 448, since the shelf arithmetic alone predicts 240. E4M3 declines to reserve a whole exponent code for infinity the way the other formats do, which buys it one more shelf. At 8 bits, conventions are too expensive.

E4M3 4 exp 3 mant more precision → weights / activations E5M2 5 exp 2 m more range → gradients range vs FP32: FP8 (max ~448 / ~57k) FP32 range (3.4e38) — neither FP8 comes close; distractor D was false

The same decision made twice at 8 bits. Neither one comes near FP32’s reach, which is why FP8 training still keeps higher-precision copies alongside.

Now the assignment, which you can predict without being told.

  • Weights and activations go in E4M3. Their values sit in a fairly narrow band, so range was never the binding constraint. Spend the bits on mantissa, where they buy something.
  • Gradients go in E5M2. Gradients are the tensors that span orders of magnitude and the ones that underflow. Spend the bits on exponent, because the alternative is deletion.

That is the article’s rule again: give range to the tensor that would otherwise lose values, precision to the tensor that would only get noisier. Neither FP8 format reaches anywhere near FP32, so FP8 training keeps higher-precision master weights and a scale factor per tensor, which is loss scaling generalised.

Understand NF4, a 4-bit format with no exponent at all

Below 8 bits nobody splits the budget into fields any more. They map the values onto a small set of codes instead, which is what quantisation means. Jacob et al. set out the integer version most later work builds on, where a tensor carries a scale and an offset and its values become plain integers. Two results took that into large language models. Dettmers et al. showed with LLM.int8() that a large transformer can be served at 8 bits once a small number of outlier features are kept in higher precision, and Frantar et al. showed with GPTQ that weights can be pushed to 3 or 4 bits after training, one layer at a time, with no retraining at all. Both are inference-side moves on a finished model.

NF4 is the training-side one. It is the format QLoRA stores its frozen base model in, and it is a different animal. It has no sign field, no exponent field and no mantissa field. Put the shelf picture down.

Start from the budget. 4 bits give 16 patterns, so NF4 has 16 values available, full stop. It is a lookup table: each 4-bit code is an index into a fixed list of 16 numbers. The only design question is which 16 numbers go in the list. There are two answers.

Option A, evenly spaced. Take the smallest and largest value in the tensor and cut the distance between them into 16 equal steps.

Option B, quantile spaced. Place the 16 values so an equal share of the tensor’s numbers falls near each one.

Two words there need defining first.

A quantile is a cut point in sorted data. Sort every value from smallest to largest, then chop the sorted list into equally sized groups. The values you chopped at are the quantiles. Cutting into four equal groups gives quartiles, the same idea under a more familiar name.

A normal distribution is the bell curve: the shape data takes when most values cluster near the middle and fewer and fewer appear as you move out to either side. Adult heights follow it. So do the weights of a trained network, and that is the claim the format rests on. Weights start life as small random values drawn from around zero, training moves most of them only a little, and weight decay pulls them back toward zero as the run goes on. Plot a histogram of any trained layer and you get a narrow bell centred on zero with thin tails.

The choice, worked on sixteen numbers

Take sixteen weights from one row, already sorted, and suppose you have only 2 bits, so 4 codes rather than 16.

minus 0.90, minus 0.15, minus 0.10, minus 0.08, minus 0.05, minus 0.03, minus 0.01, 0.00, 0.01, 0.02, 0.04, 0.06, 0.09, 0.13, 0.20, 0.85

Evenly spaced. The values run from minus 0.90 to 0.85, a span of 1.75, so four equal buckets are 0.4375 wide. Count what lands in each: 1 weight, then 5, then 9, then 1. Fourteen of the sixteen crowd into two buckets, and two of your four codes each describe a single outlier.

Quantile spaced. Four equal groups means four weights per code, by construction. The two middle groups come out narrow, because that is where the weights are packed, so values near zero get resolved finely. The outer groups come out wide, so the extremes are stored coarsely. That is deliberate: there are two extremes and fourteen of everything else.

Scale that from 4 codes to 16 and from 16 weights to a block of 64 and you have NF4. In practice it runs block by block, each block carrying its own scale factor, and the part on QLoRA does that accounting properly.

One property matters more than any of it. NF4 is storage only. No tensor core multiplies NF4 values, because there is no arithmetic circuit for a lookup table. Before every matrix multiply the codes have to be expanded back into BF16, which is called de-quantisation, and that is real work on the critical path of every forward and backward pass. NF4 buys memory and spends time.

Read any new format in five seconds

The rule that gets you through a spec sheet is three lines long. Exponent bits give range. Mantissa bits give precision. Total bits give memory.

exponent = range → mantissa = precision → format exp (range) mant (prec) total FP64double115264 bit FP32single82332 bit TF32compute-only81019 bit* FP16half51016 bit BF16brain-float8716 bit FP8-E4M3fp8 precise438 bit FP8-E5M2fp8 wide528 bit

Read down the exponent column and you are reading reach. Read down the mantissa column and you are reading resolution. FP32, TF32 and BF16 share an exponent and differ only in what they surrendered for it.

Notice the families. FP32, TF32 and BF16 all carry the same 8-bit exponent and differ only in precision, which is why moving a value between them is safe for anything range-sensitive. FP16 and both FP8 formats took the smaller-range bet, which is why each needs either a workaround or a scale factor to survive a run.

Then add one question, and it predicts behaviour rather than layout. Do I store it, compute in it, or both? BF16 is both, so memory and speed change. TF32 is compute only, so speed changes and memory does not. NF4 is storage only, so memory drops and speed gets worse, because every use pays for de-quantisation.

A real pipeline runs several formats at once, each placed where its trade fits.

Job in the pipeline Format Why that trade fits
Master weight copy and optimizer accumulation FP32 needs precision so tiny updates survive thousands of steps
Forward and backward compute during training BF16 FP32’s range at half the memory, and no loss scaling
Weights and activations under FP8 training FP8 E4M3 bounded range, so favour precision
Gradients under FP8 training FP8 E5M2 gradients span orders of magnitude, so favour range
FP32 matrix multiplies on tensor cores TF32 one-line opt-in, FP32 range, no memory change
Fast inference serving on recent hardware FP8 native tensor-core compute format
Storage of a frozen base model NF4 4-bit storage, de-quantised to BF16 to do any maths

The store-or-compute question also explains a result that surprises people. Quantising a small model that already fits on your card makes it slower for no benefit, because NF4 is storage only and every multiply expands it back to BF16 first. The next part puts these formats to work on the thing that consumes them: gradients and the optimizer state that tracks them.

Key takeaways

  • A float is a sign bit, an exponent that picks a shelf, and a mantissa that picks a slot on it. Exponent bits buy range, mantissa bits buy precision, and the bit budget is fixed, so more of one is less of the other.
  • Precision is relative. With M mantissa bits, neighbouring values differ by about one part in 2 to the power M of the number itself, so a format’s blind spot grows with the weight it is holding.
  • Rounding moves a number to the nearest representable value and networks absorb it. Underflow moves it to zero and overflow moves it to infinity, and neither is recoverable. Rounded is not lost. Out of range is lost.
  • BF16 and FP16 are both 16 bits. BF16 keeps FP32’s 8-bit exponent, so it needs no loss scaling. FP16’s 5-bit exponent caps it at 65,504 and floors it at about 6e-8, so loss scaling is mandatory and mis-tuning it produces silent NaN.
  • BF16’s gap near 0.5 is 0.0039, so an update of 0.0001 vanishes. That is why mixed precision keeps an FP32 master copy, 4 of the optimizer’s 12 bytes per parameter.
  • There is no BF32 because a brain-float aimed at 32 bits has no mantissa left to cut. The useful version shipped as TF32, a compute mode with 19 significant bits, off by default in stock PyTorch.
  • Ask of any format: do I store it, compute in it, or both. BF16 is both, TF32 is compute only, NF4 is storage only, and that answer predicts what each does to memory and throughput.

You can now

  • Take any number apart into its sign, exponent and mantissa fields and say what each contributes, from “Take a number apart into sign, exponent and mantissa”.
  • Compute a format’s biggest value and its gap between neighbours near any magnitude, from its bit counts alone, from “Measure a format’s range and its precision in numbers you can check”.
  • Explain why a rounded gradient is survivable and an underflowed one is not, and use that to predict which tensor needs which format, from “Tell rounding apart from deletion: underflow and overflow”.
  • Diagnose an FP16 run that stalls or produces NaN, and say what loss scaling was supposed to do about it, from “Read FP32, FP16 and BF16 straight off their bit counts”.
  • Decide whether a format changes your memory, your speed, or both, before reading a line of its documentation, from “Tell a storage format from a compute mode, and switch TF32 on”.
  • Explain why NF4’s sixteen values sit at the quantiles of a bell curve rather than at even spacings, from “Understand NF4, a 4-bit format with no exponent at all”.

Glossary

Activation
Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
Backpropagation
The procedure that computes a gradient for every weight in one sweep backwards through the model, from the loss at the end to the first layer. It works by applying the chain rule one step at a time. The calculus behind backpropagation
Backward pass
Running backwards from the loss through the model to produce a gradient for every weight. It is backpropagation in practice, and it needs the activations the forward pass stored. How a neural network learns
bf16
A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling.
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
De-quantisation
Expanding quantised numbers back to a wider format so arithmetic can be done on them. QLoRA de-quantises its 4-bit weights to bf16 for every matrix multiply, and that extra work is where its speed cost comes from. QLoRA explained
E4M3
One of the two 8-bit floating-point formats: 4 exponent bits and 3 mantissa bits. It favours precision over range, so it is used for weights and activations, whose values stay in a fairly narrow band. Micikevicius et al., FP8 Formats for Deep Learning
E5M2
The other 8-bit floating-point format: 5 exponent bits and 2 mantissa bits. It favours range over precision, so it is used for gradients, which span many orders of magnitude and would otherwise underflow. Micikevicius et al., FP8 Formats for Deep Learning
Exponent
The part of a floating-point number that says how big or small it is, in powers of two. More exponent bits means the format reaches further before values overflow or underflow.
Floating point
The way computers store numbers that span a huge range, as a sign, an exponent and a mantissa. Every format in this series is a different way of splitting a fixed bit budget between those last two.
Forward pass
Running data through the model from input to output to get a prediction and a loss. Along the way it produces the activations that the backward pass will need. The complete inference path
fp16
A 16-bit number format with 5 exponent bits and 10 mantissa bits. Precise for its size but short of reach, so small gradients underflow to zero and loss scaling becomes mandatory.
fp32
32-bit floating point, with 8 exponent bits and 23 mantissa bits. It is the precise reference format, used for the master copy of the weights and for the optimizer’s running averages.
fp32 master weights
The full-precision copy of the weights that the optimizer actually updates in a mixed-precision run. It exists because a 16-bit weight cannot record an update far smaller than itself, so without it the updates round away and training stalls. Micikevicius et al., Mixed Precision Training
FP8
8-bit floating point, which ships in two versions because the range-against-precision trade is too tight to settle once. E4M3 for weights and activations, E5M2 for gradients. Micikevicius et al., FP8 Formats for Deep Learning
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
Inference
Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
LoRA
Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Loss scaling
Multiplying the loss by a large constant before the backward pass so small gradients do not underflow fp16, then dividing them back before the optimizer step. bf16 does not need it at all. Micikevicius et al., Mixed Precision Training
Mantissa
The part of a floating-point number that holds its significant digits. More mantissa bits means finer steps between the values the format can represent.
Matrix multiply
The operation that dominates all the arithmetic in a transformer: multiply a block of inputs by a block of weights to get a block of outputs. Usually shortened to matmul. Inside one transformer block
Mixed precision
Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
NaN
Not a number, the value floating-point arithmetic produces from an undefined operation such as infinity minus infinity. Once one reaches the weights the run is dead, and it is the usual end state of a badly scaled fp16 run.
Neural network
A stack of layers that multiply their input by learned weights and pass the result on. Training means adjusting those weights until the output is closer to what you wanted. How a neural network learns
NF4
The 4-bit storage format QLoRA uses. Its 16 values sit at the quantiles of a bell curve rather than at even spacings, because model weights cluster near zero. It has no exponent or mantissa fields and nothing computes in it directly. Dettmers et al., QLoRA
Optimizer
The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
Overflow
What happens when a number is too large for the format to hold, so it becomes infinity. Once infinity enters the arithmetic the usual result is NaN and a dead run.
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
QLoRA
LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
Quantisation
Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
SGD
Stochastic gradient descent, the simplest optimizer: multiply the gradient by the learning rate and subtract. It stores nothing extra, and it is unreliable on transformers because one global learning rate has to suit every weight. Gradients and optimizers explained
Tensor
A grid of numbers with any number of dimensions. One number is a scalar, a row of them is a vector, a table is a matrix, and anything past that is still a tensor with more dimensions. How a neural network learns
Tensor core
The part of an NVIDIA GPU built to do matrix multiplies in low precision very fast. It is why bf16 training beats fp32, and why a format tensor cores cannot multiply directly, such as NF4, costs throughput. NVIDIA, Accelerating AI training with TF32 tensor cores
TF32
A tensor-core compute mode with fp32’s exponent and a 10-bit mantissa, 19 significant bits in total. Your data stays fp32 in memory, so it changes speed and not memory, and stock PyTorch ships it switched off. NVIDIA, Accelerating AI training with TF32 tensor cores
Underflow
What happens when a number is too small for the format to represent, so it becomes zero. A gradient that underflows has not been made noisy, it has been deleted, and averaging cannot bring it back.
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model


Practical exercises

Find bf16’s blind spot at a different weight value

This part shows that near 0.5, bf16’s gap between representable values is about 0.0039, so an update of 0.0001 vanishes. Work out the equivalent gap near a weight value of 2.0, then say whether an update of 0.0003 to a weight near 2.0 survives or vanishes, and what that tells you about whether bf16’s blind spot is a fixed number or something that depends on the weight’s own magnitude.

See the worked solution (opens in a new tab)

Diagnose why fp16 died and bf16 did not

Two otherwise identical training runs of the same model diverge only in numeric format. The bf16 run trains cleanly to completion. The fp16 run reports loss as NaN around step 4,000, and nobody on the team touched loss scaling in either run. Using this part’s account of what each format’s exponent buys and costs, give the single most likely mechanism that produces exactly this split, bf16 fine and fp16 NaN, and name one concrete fix short of simply switching to bf16.

See the worked solution (opens in a new tab)

Pick the wrong FP8 flavour and count the cost

A weight tensor for one layer has values ranging from about minus 2.1 to 3.4. The recommended choice for weights under FP8 training is E4M3. Suppose you stored this tensor in E5M2 instead. Using the mantissa bit counts and the documented maximum magnitudes for each format, about 448 for E4M3 and about 57,344 for E5M2, say how much relative precision you lose per value, and whether the extra range you gained is actually being used by this tensor.

See the worked solution (opens in a new tab)

Work out what FP8 compute would actually save

Suppose a team full fine-tunes Qwen2.5-1.5B and switches the forward and backward compute from bf16 to FP8, storing weights in E4M3 and gradients in E5M2, while keeping the same fp32 master-copy-plus-two-moments optimizer scheme from Part 2. Work out the new bytes-per-parameter figure and the new static state in GB for the 1.54 billion parameter model, and say which of the four tenants this change can never touch, no matter how aggressive the compute format gets.

See the worked solution (opens in a new tab)

Frequently asked questions

What is the difference between bf16 and fp16?

Both are 16 bits, and they split those bits differently. BF16 uses 8 exponent bits and 7 mantissa bits, so it reaches as far as FP32 with coarse steps between values. FP16 uses 5 exponent bits and 10 mantissa bits, so it resolves finely inside a much shorter reach that stops at 65,504. BF16 is preferred for training because its range removes the need for loss scaling entirely.

Why does fp16 training need loss scaling?

Because FP16’s 5-bit exponent puts its absolute floor at roughly 6e-8, and gradients late in training are routinely that small or smaller. Loss scaling multiplies the loss by a large constant before the backward pass, which lifts every gradient by that same constant into FP16’s representable band, then divides them back before the optimizer step. BF16 never needs it, because it shares FP32’s 8-bit exponent and therefore FP32’s reach.

Why does mixed-precision training keep an fp32 copy of the weights?

Because BF16 cannot record a small update. Near 0.5 its gap between neighbouring values is 0.0039, so adding 0.0001 rounds straight back to the original value and the update is gone. FP32’s gap near 0.5 is about 6e-8, so it accumulates thousands of tiny updates correctly while BF16 handles the fast arithmetic.

Is there a bfloat32 format?

No, and the reason is structural rather than historical. A brain-float means keeping FP32’s 8 exponent bits and cutting mantissa bits to fit a smaller box, and at a 32-bit target the leftover mantissa is 23 bits, which is FP32 itself. Nothing was cut. The genuinely useful idea, FP32’s range with less precision for faster maths, shipped as NVIDIA’s TF32.

What is TF32 and do I need to enable it?

TF32 is a tensor-core compute mode with FP32’s 8 exponent bits and FP16’s 10 mantissa bits, 19 significant bits in total. Your data stays FP32 in memory and the truncation happens inside the multiply, so it costs no extra memory. You do need to enable it: stock PyTorch ships TF32 matmuls off by default, so set torch.backends.cuda.matmul.allow_tf32 to True yourself, or call torch.set_float32_matmul_precision with “high”.

Why does FP8 come in two versions?

Because at 8 bits the range-against-precision trade is too tight to resolve once. In both names the E is the exponent bit count and the M is the mantissa bit count. E4M3 spends more on mantissa and is used for weights and activations, whose values stay in a fairly narrow band. E5M2 spends more on exponent and is used for gradients, which span orders of magnitude and would otherwise underflow to zero.

Sources and further reading

Previous