Tag: mixed precision
-

Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
bf16 vs fp16, why loss scaling exists, why mixed precision keeps an fp32 master copy, and the difference between a storage format…
-

Training Memory: The Four Tenants and the 16 Bytes Per Parameter
Why a 3 GB model needs 25 GB to train. The four memory tenants derived from first principles, where the 16 bytes…