Tag: LLM fine-tuning
-

Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
Preference tuning teaches better versus worse, which SFT cannot. RLHF, DPO and its beta, the KTO and ORPO and SimPO family, and…
-

QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
QLoRA puts a 7B to 30B fine-tune on one card. NF4, double quantisation, paged optimizers, the throughput tax nobody mentions, and the…
-

LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
LoRA fine-tuning from the mechanic up: the low-rank update, the zero initialisation, rank and alpha, the honest comparison, and merging versus adapter…
-

Choosing a Base Model and Building a Fine-Tuning Bench
How to pick a base model for fine-tuning, three candidates read as datasheets, where to run the bench, and the toolchain stack…
-

AdamW Explained, Line by Line
AdamW dissected: where it runs, the lineage behind the name, the four update lines, why decoupled weight decay was a real fix,…
-

Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
What a gradient is, what an optimizer decides, and why Adam’s two running statistics per weight are worth 8 bytes each on…
-

Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
bf16 vs fp16, why loss scaling exists, why mixed precision keeps an fp32 master copy, and the difference between a storage format…
-

Activation Memory: Why the Forward Pass Costs More Than the Weights
Activation memory scales with batch, sequence, hidden size and depth, not parameters. The lifecycle, the quadratic attention trap, and every lever that…
-

Training Memory: The Four Tenants and the 16 Bytes Per Parameter
Why a 3 GB model needs 25 GB to train. The four memory tenants derived from first principles, where the 16 bytes…
-

LLM Fine-Tuning Explained: What Actually Changes Inside the Model
LLM fine-tuning runs the pretraining objective on your data with a loss mask. What actually changes, when to do it, and how…