Category: LLM Fine-Tuning
How LLM fine-tuning actually works: training memory, numerics, optimizers, masking, SFT, data, LoRA, QLoRA and preference tuning.
-

Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
What a gradient is, what an optimizer decides, and why Adam’s two running statistics per weight are worth 8 bytes each on…
-

Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
bf16 vs fp16, why loss scaling exists, why mixed precision keeps an fp32 master copy, and the difference between a storage format…
-

Activation Memory: Why the Forward Pass Costs More Than the Weights
Activation memory scales with batch, sequence, hidden size and depth, not parameters. The lifecycle, the quadratic attention trap, and every lever that…
-

Training Memory: The Four Tenants and the 16 Bytes Per Parameter
Why a 3 GB model needs 25 GB to train. The four memory tenants derived from first principles, where the 16 bytes…
-

LLM Fine-Tuning Explained: What Actually Changes Inside the Model
LLM fine-tuning runs the pretraining objective on your data with a loss mask. What actually changes, when to do it, and how…