Picking a base model for fine-tuning looks like a one-line download. It is not. That one line fixes how much memory your run needs, which tokenizer you are stuck with, what you are allowed to do with the result, and how much of the software stack will simply work. Choose badly and your first week goes on fighting the environment instead of learning the mechanics.
This part covers the whole setup stage. It is deliberately several subjects on one page: what a spec sheet means, what a licence lets you do, what a rented GPU costs, how to get the drivers to cooperate, and which dataset to start on. Nobody needs all of that at once.
So the first section is a map. Read it, take the two or three sections you need today, and skip the rest without guilt. By the end you will have a base model you can defend, somewhere to run it, and a first dataset loaded.
- Part 1. LLM Fine-Tuning Explained: What Actually Changes Inside the Model
- Part 2. Training Memory: The Four Tenants and the 16 Bytes Per Parameter
- Part 3. Activation Memory: Why the Forward Pass Costs More Than the Weights
- Part 4. Number Formats for Training: FP32, FP16, BF16, TF32, FP8 and NF4
- Part 5. Gradients and Optimizers: From SGD to Adam, and Why m and v Cost 8 Bytes
- Part 6. AdamW Explained, Line by Line
- Part 7. Masking in LLM Training: Loss, Causal and Padding
- Part 8. Choosing a Base Model and Building a Fine-Tuning Bench (you are here)
- Part 9. Supervised Fine-Tuning End to End: The First Real Run
- Part 10. Data Is the Actual Job: Formats, Quality, Splits and Synthetic Data
- Part 11. LoRA Explained: Freeze the Model, Learn a Low-Rank Patch
- Part 12. QLoRA Explained: A 4-Bit Base, NF4, and What It Really Costs
- Part 13. Preference Tuning: RLHF, DPO, and the Verifiable-Reward Branch
- Part 14. Multi-GPU Fine-Tuning: DDP, ZeRO, FSDP and the Wire Between the Cards
- Part 15. Evaluating a Fine-Tune: The Ladder, the Delta and Six Silent Failures
Find the sections you need and skip the rest
The sections below cover subjects that have almost nothing to do with each other. They share a page because they all land in the same week. Here is what each one is, and who it is for.
- The spec sheet, and the three candidates. What the numbers on a model card mean, then three real models compared. Everyone needs these two.
- The licence. What each model’s terms let you do. Skip it if the work never leaves your own machine. Read it before a fine-tune of yours reaches a customer, a public repository or a review board.
- Where the bench runs. Your own card, a rented card, or a managed service. Read it if you do not already have a GPU in front of you.
- The toolchain. Drivers, CUDA and PyTorch builds. Skip it until something breaks. Then come straight back, because the answer is very likely here.
- The first dataset. What to train on first, and why. Short, and everyone needs it.
One framing decision runs underneath all of it. A production base model is chosen to get the most task quality per dollar at serving time. A learning base model is chosen to let you see the most machinery per run. Those two goals pull in different directions. This part picks the second.
Say what a base model choice locks in
A base model is the raw pretrained checkpoint. A checkpoint is a saved copy of every weight in the model at one moment. Somebody spent months and a large budget teaching this one to continue text. It has never been taught to hold a conversation.
An Instruct model is that same base model after somebody fine-tuned it to follow instructions. That is the job Part 9 walks through. So the two are the same model at two different stages, and the choice between them is a choice about how much you want to watch.
Four properties matter more than any leaderboard score when the goal is to learn.
- Small enough to fully fine-tune on one card. Full fine-tuning updates every weight, and it is the expensive case. If you can run it, you get to feel the memory wall. Then you get to feel LoRA and QLoRA take that wall away. If your model only ever fits with QLoRA, you never learn what QLoRA saved you from.
- Fast enough to run twenty times a day. A model of 1 to 2 billion parameters finishes a small supervised run in minutes on one modern card. You want to break the run on purpose and fix it, not babysit an overnight job.
- Boringly conventional in shape. Every library, every chat template and every memory trick already supports the standard design. Unusual designs are where unsupported-operation errors live. For learning, boring is a feature.
- Base rather than Instruct. Take the raw checkpoint. Then you impose the chat format yourself, and you watch instruction following appear from nothing. An Instruct checkpoint has had that done to it already, which muddies cause and effect.
Past those four, every base-model choice is a point in the same space. The axes are the same for a production model. Only the weight on each one changes.
| What you are choosing | What it controls | How much it matters for learning |
|---|---|---|
| Parameter count | memory, speed, headroom on the card | highest, it has to allow a full fine-tune |
| Model shape | how much tooling supports it out of the box | high, stay with a standard decoder |
| Licence | what you may do with the weights and the output | low while private, high once anyone else sees it |
| Vocabulary size | how much of the budget goes to the lookup table, and how many languages it reaches | medium, it moves the memory split |
| Context length | the longest sequence allowed, and so the activation cost | low, 8K is plenty for a first run |
| Base or Instruct | whether a chat format is already baked in | high, take Base |
| Published training data | whether you can say what the model was trained on | low for mechanics, high for a review board |
| Ecosystem depth | tutorials, ready-made quantised copies, forum answers | medium to high, fewer dead ends |
Screen a candidate model in one multiplication
Whether a model fits on your card is arithmetic, not judgement. Part 2 derives the number. Standard mixed-precision AdamW costs 16 bytes for every parameter. Mixed precision means the arithmetic runs in a 16-bit format while a 32-bit copy of each weight is kept for the update. AdamW is the optimizer, the component that decides how far each weight moves.
Multiply the parameter count by 16 and you get gigabytes of static state. Static state means the weights, the gradients and the optimizer’s own stored numbers. It stays exactly the same size all run long. Here is the whole screen, worked:
- 1.23 billion parameters times 16 is about 19.7 GB.
- 1.54 billion times 16 is about 24.6 GB.
- 1.71 billion times 16 is about 27.4 GB.
- 7.62 billion times 16 is about 122 GB, which no single card holds.
Every one of those numbers is static state only. None of them includes activations. Activations are the intermediate results the forward pass produces on the way through, and the backward pass then needs them. They are counted separately, every single time. For Qwen2.5-1.5B at a batch of 4 sequences of 1,024 tokens, they land somewhere near 8 GB. Treat that as an order of magnitude rather than a measurement. Part 3 covers what moves it.
Now apply the screen to one 32 GB card. About 29.8 GiB of that is actually usable once the driver has taken its share, and Part 1 explains why those two units differ. A model near 1.5B parameters fits, and fits tightly enough that you feel the wall. A 1B model leaves comfortable headroom. A 1.7B model fits only with care.
Two settings do most of that caring, and both come back later.
- Gradient checkpointing throws away most of the stored activations and recomputes them during the backward pass. Peak activation memory drops a long way, and the run takes roughly 20 to 30 percent longer.
- Gradient accumulation runs several small batches, adds their gradients together, and updates the weights once at the end. You get the steadier signal of a large batch while holding only a small one in memory.
Leave 2 to 3 GB of margin in every estimate you make. Fragmentation and one unusually long batch both live in that last gigabyte.
Read a model spec sheet without an architecture background
This section is about the numbers printed on a model card. It assumes you know nothing about how a transformer is built. It teaches only the parts that change a decision.
Start with the one architectural fact that matters. All three models below are decoder-only transformers of the standard shape. Decoder-only means one stack of identical blocks, where each position in the text may look only at what came before it. Standard means everyone builds them this way. The practical consequence is the whole point: every tool in the ecosystem already loads them, so nothing in your toolchain will refuse.
If you want the names, here they are once. All three use rotary position embeddings, SwiGLU feed-forward blocks, and RMS normalisation placed before each sub-layer rather than after. A feed-forward block is the part of each layer that processes every position on its own. Normalisation is a step that rescales the numbers flowing through, so they stay in a sensible range. You do not need to know any more than that to choose a base model, and this series never asks you to. Those names matter here only as a sign that a model is conventional. Each of them has one paper behind it: Su et al. for the rotary embeddings, Shazeer for SwiGLU, and Zhang and Sennrich for RMS normalisation, a cheaper variant of the layer normalisation of Ba et al. The placement is a published result too: Xiong et al. found that normalising before each sub-layer makes training much better behaved than normalising after it. The companion part on one transformer block takes them apart properly.
Layers, hidden size and context length
Three numbers on every card describe the shape.
Layers is how many blocks are stacked. Qwen2.5-1.5B has 28. More layers means more work per token and a slower loop.
Hidden size is how wide the list of numbers flowing between layers is. Qwen2.5-1.5B’s is 1,536. Papers write it d_model.
Context length is the longest sequence the model was built to handle. It is a ceiling, not a target. A 32K model is perfectly happy training on 1,024-token examples. Training at 32K would cost far more activation memory than you have.
The vocabulary, and why its size is a trade
A tokenizer splits text into tokens and gives each token an integer ID. The vocabulary is the full list of tokens that tokenizer knows. Qwen2.5’s holds about 151,000 entries.
A model cannot do arithmetic on an integer ID. It needs actual numbers. So the first thing it does is look the ID up in a table. That table has one row per vocabulary entry, and each row is as wide as the hidden size. It is called the embedding table. Every number in it is a parameter, exactly like the ones inside the layers.
Do the multiplication on a toy model. Take a vocabulary of 1,000 and a hidden size of 10. The table is 1,000 rows of 10 numbers, so 10,000 parameters. Now double the vocabulary to 2,000. The table doubles to 20,000. The layers have not changed at all.
That is the trade, and it bites hardest on small models. A model has a fixed parameter budget. Every number in the lookup table is a number not in the layers that do the work. A big vocabulary spends a lot of a small model’s budget on the table.
There is a saving available. The model also has to turn its final numbers back into one score per token, and that needs a table of exactly the same shape. Many small models reuse one matrix for both jobs. That is called tied embeddings, and it halves the cost.
SmolLM2-1.7B makes the contrast concrete. Its vocabulary is about 49,000 and it ties its embeddings, so that one matrix holds roughly 100 million parameters. Against a total of 1.71 billion, that is about six percent. Qwen2.5-1.5B carries a vocabulary three times larger inside a narrower hidden size, so its share is much bigger. The exercises at the end ask you to work out how much bigger.
The other half of the trade is what the tokenizer does to your text. A large vocabulary has room for more whole words, and for words in more languages. The same sentence then turns into fewer tokens. Fewer tokens means shorter sequences, and shorter sequences cost less activation memory. A small vocabulary saves parameters and spends sequence length. Neither answer is right in general.
Attention heads, and the one place these models differ
Attention is the step where each position in the text looks at the other positions and pulls in whatever is useful. It runs several times in parallel, over different slices of the numbers. Each parallel run is called a head.
Every head needs three things per token: a query, a key and a value. Loosely, the query is what this position is looking for. The key and the value are what an earlier position offers. In the original design each head had its own keys and values. Grouped-query attention changes that. Several query heads share one set of keys and values.
Qwen2.5-1.5B has 12 query heads over 2 key-value heads. Divide, and six query heads share each key-value pair. That is the entire idea.
Here is why anyone bothers. When a model is generating text it keeps the keys and values of every token it has already produced, so it does not have to work them out again. That saved pile is called the KV cache, and it grows as the conversation gets longer. Storing 2 sets instead of 12 makes it roughly six times smaller.
For your training run this changes almost nothing. Training has no generation loop, so it has no KV cache. It matters later, when you serve what you trained, and it is a real reason this model is cheap to run. The companion part on multi-head attention has the mechanism.
Choose between three concrete candidates
Three real models, read as spec sheets rather than as a ranking. Every one of them can be fully fine-tuned on one 32 GB card. So the decision comes down to the second-order differences.
Qwen2.5-1.5B, the capable generalist
1.54 billion parameters. 28 layers, hidden size 1,536, a vocabulary of about 151,000, tied embeddings, 12 query heads over 2 key-value heads, and a native context of 32K. Released under Apache-2.0 at this size.
Its full fine-tune static state is about 24.6 GB, so a 32 GB card holds it with the activations squeezed. It is the strongest of the three on languages other than English, and on maths and code for its size, which is the case Qwen’s own technical report makes with its evaluation tables. Its ecosystem is the deepest of the three. This series uses it for every calculation.
Llama-3.2-1B, the fast sanity model
1.23 billion parameters. Just 16 layers, a 128K context window, a vocabulary of about 128,000, and roughly 9 trillion training tokens behind it. Its static state is about 19.7 GB, the lightest of the three. Sixteen layers makes it the quickest loop for confirming that a script runs at all.
Its history is the interesting part, and it is worth a minute because it is a clean example of how small models get made. This one was not trained from scratch. It was built in two steps.
Step 1: cut pieces out of a bigger model. Start with a trained 8 billion parameter model, one of the family Grattafiori et al. document in the Llama 3 paper. Remove whole pieces of it, such as layers and attention heads, until what remains is the size you wanted. That is called structural pruning. The result runs, and it is worse than what you started with, because you have thrown away parts that were doing something.
Step 2: repair it by copying a bigger model. Show a large model some text and record the full set of scores it produces for the next token. Part 1 calls those scores logits: one raw number for every token in the vocabulary. Then train the pruned model to produce a similar set of scores at the same positions. That is knowledge distillation. The big model is the teacher and the small one is the student.
Copying the whole score list teaches more than copying only the correct token would. It also shows the student which other tokens the teacher found plausible, and by how much. That is the argument Hinton, Vinyals and Dean made when they introduced distillation: the teacher’s full set of scores carries more information per example than the single correct answer does. None of this changes how you use the model. It is simply an unusually visible example of a model-shaping pipeline, and you are about to run a small one yourself.
The licence is the flag on this one. It gets its own section below.
SmolLM2-1.7B, the auditable one
1.71 billion parameters. 24 layers, an 8K context, a small vocabulary of about 49,000, tied embeddings, Apache-2.0. Trained on 11 trillion tokens from a data mix that is published in full. Its static state is about 27.4 GB, the tightest of the three, so a 32 GB card holds it only if you keep activations down.
Its small vocabulary is the trade from the last section made real. More of its budget sits in the layers. Its tokenizer is less efficient on code and on text that is not English, which costs you sequence length. It is English-first, its ecosystem is the smallest here, and it is the only one of the three whose entire training history you can read.
Side by side
| Criterion | Qwen2.5-1.5B | Llama-3.2-1B | SmolLM2-1.7B |
|---|---|---|---|
| Full fine-tune static state | about 24.6 GB | about 19.7 GB, lightest | about 27.4 GB, tightest |
| Fits one 32 GB card | yes, with care | yes, easiest | yes, manage activations |
| Licence | Apache-2.0 | community licence | Apache-2.0 |
| Training data published | no, weights only | no, weights only | yes, in full |
| Ecosystem depth | deepest | very deep | growing |
| Languages | strongest | eight languages | English first |
| What it teaches you | why shared key-value heads are cheap to serve | how prune-and-distil builds a small model | what a fully documented lineage looks like |
The recommendation for a learning bench is Qwen2.5-1.5B base. It is Apache-2.0, it full-fine-tunes on one 32 GB card at about 24.6 GB of static state, and its shape is supported everywhere.
Keep Llama-3.2-1B as a sanity model. At 19.7 GB and 16 layers it is the fastest way to confirm that a script works before you point it at anything bigger. Reach for SmolLM2-1.7B when the story is provenance, and you need to be able to say that every token the model saw is documented.
Decide whether a model’s licence will cause you trouble
Skip this section if the work stays on your own machine. Nothing here changes what you may do privately. Read it before a fine-tune of yours reaches a customer, a public repository, an app store or a review board. That happens sooner than people expect, because a model picked for learning quietly becomes the production default.
Two kinds of licence turn up on model weights. They are not the same kind of thing.
Apache-2.0, which is what people mean by open source
Qwen2.5-1.5B and SmolLM2-1.7B both ship under it. In plain terms:
- You may use the weights for anything, including a paid product.
- You may change them, fine-tune them, and pass the result on.
- You may keep your own changes closed.
- You must keep the licence text and the copyright notice with any copy you hand on, and note which files you changed.
- You get a patent licence from whoever released the model. It ends if you sue them over a patent covering it.
- There is nobody to ask, nothing to report, and no rule about what you name the result.
That is a short procurement conversation. In many companies it is not a conversation at all.
A community licence, which is free but conditional
Llama-3.2-1B ships under a community licence instead. It is free. It allows commercial use. It allows you to fine-tune and to redistribute. It is also not an open-source licence in the formal sense, because it carries three conditions that Apache-2.0 does not.
- An acceptable-use policy. A written list of uses you agree not to put the model to. It travels with the weights. So anyone you hand your fine-tune to is bound by it as well.
- Naming and attribution. If you publish a model built on it, you have to say so, display the required notice, and carry the Llama name at the front of your own model’s name.
- A large-deployer clause. Above a monthly-active-user threshold written into the licence, you have to go and request a separate licence, which may or may not be granted. The threshold is high enough that it will not touch you unless your product is already one of the largest in the world.
For research and for most commercial work, this is fine. It is not equivalent to Apache-2.0, and a procurement lawyer will notice the difference. Read the licence file itself rather than a summary of it, including this one.
Two things people forget
The first is outputs. Some community licences also say something about what you may do with text the model generates. That matters if you plan to use one model to build training data for another. Read that clause before you build a data pipeline on it. Part 10 covers generated training data and when it earns its keep.
The second is the dataset. A dataset carries its own licence, entirely separate from the model’s. A permissive model plus a dataset restricted to non-commercial use still leaves you with a non-commercial result. Check both, every time.
Run a training bench beside a model that is already serving
This section is for people with two cards in one machine. Skip it if you have one.
You do not have to stop serving in order to start training. A full fine-tune of a 1.5B model needs about 24.6 GB of static state, plus a few gigabytes of activations once gradient checkpointing is on. That fits inside one 32 GB card. So one card can keep answering requests while the other one trains.
Set CUDA_VISIBLE_DEVICES on the training process so it sees only one card, and leave the serving process on the other. That variable controls which GPUs a process is allowed to see. It is the whole isolation mechanism.
Be clear about what this is not. Two cards used this way are two separate benches. They are not one memory pool of double the size. A model too big for one card has to be deliberately split across both, which is a different job with different machinery. Part 14 covers when a second card genuinely helps, and why the interconnect, the link the two cards use to talk to each other, decides the answer.
One caution that has nothing to do with memory. The two cards still share a power supply, a bus and the host’s RAM. A full fine-tune will pull as much power as its card is allowed to draw. If the serving side is sensitive to latency, put a power cap on the training card with nvidia-smi. The memory is already isolated. The electricity is not.
Pick where the bench runs, and know what each option hides
Read this if you do not already have a GPU in front of you. There are three places to put a fine-tuning bench. Only one of them is a classroom. The other two are places to deploy to, and they are worth understanding because clients ask about them. The rule is simple: the more the platform manages for you, the more of the mechanism it hides.
Your own machine
Everything is visible, and everything is your problem. You see the driver, the memory, the speed and every error message. That is precisely the point. The cost is the toolchain section below. It is longer on your own hardware than anywhere else, because a card you bought last month is newer than most of the software that has to support it.
A rented virtual machine with a GPU attached
You still write the training loop yourself, so almost nothing is hidden. What disappears is the hardware problem. Reach for it to run a job without touching a production box, to try a bigger card than you own, to reproduce a run on standard datacentre hardware, or to show a client a clean cloud pipeline.
Cloud GPUs are usually a generation or two behind the newest consumer cards. That is an advantage here. The stable software stack simply works on them, with none of the version drama. Costs get their own section next.
A managed platform
A managed platform offers two different doors, and it matters a great deal which one you walk through. AWS Bedrock is the example used here. Other clouds have equivalents.
The first door is importing your own weights. You fine-tune wherever you like. Then you merge the adapter into the base model. An adapter is the small file that LoRA produces, and merging folds its learned patch permanently into the weights, so you are left with one ordinary model file. You upload that file, and the platform serves it with no machine for you to manage. Support is normally limited to a named list of model families, which is a concrete payoff for having picked a conventional one.
The second door is managed customisation, where the platform runs the training for you. Do not learn on it. There are two reasons. The supported model list tends to be large models only. And it hides the training loop, the masking and the optimizer, which are exactly the things you came to see. Know it exists, because clients will ask.
Rent a cloud GPU without a surprise bill
This section is about money and paperwork rather than machine learning. Skip it if you are training locally.
AWS EC2 is the concrete example. Other clouds have equivalent instance types under different names. Four instances cover the useful range.
| Instance | GPU | VRAM | Good for | Indicative on-demand rate |
|---|---|---|---|---|
| g6.xlarge | one L4 | 24 GB | LoRA and QLoRA, cheapest modern option | under a dollar an hour |
| g5.xlarge | one A10G | 24 GB | LoRA and QLoRA, very well trodden | about a dollar an hour |
| g6e.xlarge | one L40S | 48 GB | full fine-tune of 1B to 1.7B with room | a low single-digit dollar rate |
| p4d.24xlarge | eight A100 | 320 GB | overkill here, multi-GPU work later | tens of dollars an hour |
Treat every price in that table as indicative. Rates differ by region and they change often. Check the console before you provision anything.
Notice what the memory arithmetic does to that table. A 24 GB card cannot hold a 1.5B model’s 24.6 GB of full fine-tune static state, let alone its activations. So on those two instances you use LoRA or QLoRA, which freeze the base model and train a small patch beside it. For a real full fine-tune you step up to the 48 GB instance. The multiplication decides the invoice.
Four habits keep the bill small.
- Ask for quota first. New accounts frequently have a limit of zero on GPU instance families. The request can take a day to be granted. Make it before you need it.
- Stop when idle, terminate when finished. Billing counts seconds of instance life, not seconds of GPU work. A forgotten idle instance costs the same as a busy one.
- Copy artefacts out before you terminate. Push checkpoints and logs to object storage. The instance’s disk goes when the instance does, and it charges you for every hour it exists.
- Use spot capacity for jobs that checkpoint. Spot is commonly 60 to 70 percent cheaper than the on-demand rate, and that figure is indicative too. The catch is that the instance can be taken back with little warning. If your run saves a checkpoint regularly, that is an inconvenience. If it does not, it is a lost day.
One shortcut is worth knowing. A deep-learning machine image ships with the driver, CUDA and PyTorch already matched for that instance’s GPU. Choosing one skips the entire next section.
Fix the toolchain errors a new GPU throws at you
Skip this section until something breaks. Then come straight back. If your card is a year or two old, or you are on a cloud image, the install below simply works and you will never need the rest. If your card is very new, this is the section that saves your week.
A fine-tuning environment is a tower of four floors. Each floor has to match the one below it. Get the bottom two wrong and nothing above them runs, with an error message that points nowhere useful.
- The driver. The software your operating system uses to talk to the card. It comes from the GPU vendor. Running
nvidia-smiprints its version and proves the card is visible at all. - CUDA. The vendor’s toolkit for running your own code on the card. It has its own version number, separate from the driver’s. A newer driver can serve an older CUDA. It does not work the other way round.
- A PyTorch build compiled for your card. This is the load-bearing floor. PyTorch ships as several different builds, one per CUDA version, and you choose which one you install.
- The libraries. Transformers, datasets, accelerate, TRL, PEFT and bitsandbytes. These are ordinary Python packages. They install normally once the floors below them agree.
What compute capability is, and what the error actually means
Every NVIDIA chip carries a version number saying what it can do. It is called the compute capability. It is written two ways, as 12.0 and as sm_120, and they mean the same thing.
Code that runs on a GPU is compiled ahead of time, for a specific list of compute capabilities, and shipped inside the package you install. If your chip’s number is not on that list, the package contains no code your chip can run. Nothing happens.
That is exactly what this message is telling you:
no kernel image is available for execution on the device
A kernel is one small program that runs on the GPU. A kernel image is the compiled form of it. So the message says: I looked in the package for code compiled for your chip, and there is none. It is not a bug in your script. Editing your script cannot fix it. The fix is always to install a build that includes your card’s number.
Concretely: a current Blackwell consumer card reports compute capability 12.0, which is sm_120, and that needs a CUDA 12.8 build of PyTorch.
# 1. is the card visible, and what driver is running?
nvidia-smi
# 2. install a PyTorch build made for your card's CUDA version
# (swap the index URL for the build your card needs)
pip install torch --index-url https://download.pytorch.org/whl/cu128
# 3. confirm the compiled code for your chip is really there,
# rather than only that CUDA loaded
python -c "import torch; print(torch.__version__, torch.cuda.get_device_capability())"
# 4. the fine-tuning libraries, once the floors below them agree
pip install transformers datasets accelerate trl peft bitsandbytes
Line by line:
nvidia-smiprints the driver version and lists the cards it can see. If it prints nothing useful, stop here. Nothing above this floor can work.- The install line uses
--index-url. By default pip downloads from its own package site, which carries one general PyTorch build. PyTorch publishes its CUDA-specific builds on its own site, and that flag points pip there instead. Thecu128in the address is the CUDA version. Change it to match what your card needs. - The one-line Python check does the thing people skip. It does not only prove that CUDA loaded.
get_device_capability()prints your card’s real compute capability as a pair of numbers, so you can compare it against the build you just installed. - The last line installs the fine-tuning libraries on top of a floor you have now verified.
Pin your versions. Write the exact version of every one of those packages into a requirements file the first time the run works. These libraries move quickly and their argument names change between releases. The code in this series is current as of August 2026, and an unpinned environment will drift out from under you.
Two traps on very new hardware
The first trap is compiling on the fly. Some libraries do not ship precompiled code for every chip. Instead they build a kernel at the moment it is first needed. That needs CUDA’s runtime compiler, a component called nvrtc, and some installers leave it out. When it is missing, the on-the-fly build fails and you get the same “no kernel image” message as before, from a completely different cause.
Two defences, both cheap. First, prefer prebuilt wheels over building from source. A wheel is a Python package whose compiled parts were built by somebody else before you downloaded it, so nothing has to compile on your machine. Second, for a first run, use PyTorch’s own built-in attention implementation. It is called scaled dot-product attention, usually shortened to SDPA. The alternative is installing a separate flash-attention package. Flash attention is a faster way of computing attention that never builds the full grid of scores. It is worth having eventually. A small supervised run does not need the extra speed, so add it once the basic loop works.
The second trap is bitsandbytes. It holds the 4-bit code that QLoRA depends on, and its support for a brand new chip usually lands after PyTorch’s. Test it on its own before you build a training script on top of it. Otherwise a broken library looks exactly like a bug in your own code.
import bitsandbytes as bnb
print(bnb.__version__) # import test: fails loudly if the build is broken
import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
qconf = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
m = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B",
quantization_config=qconf, device_map="auto")
print(next(m.parameters()).dtype) # confirms the 4-bit load actually ran, not just imported
What each line does:
- The import and the version print are the blunt test. A broken build fails right there, loudly, before anything else has run.
BitsAndBytesConfig(load_in_4bit=True, ...)asks for the model’s weights to be stored in 4 bits instead of 16. Storing numbers in fewer bits is called quantisation, and Part 12 covers what it costs.bnb_4bit_compute_dtype=torch.bfloat16says the arithmetic still happens in a 16-bit format. Nothing multiplies 4-bit numbers directly, so they are expanded back before every multiply.from_pretrained(...)downloads the model and loads it under that configuration.device_map="auto"lets the library decide where each piece goes.- The last line prints the data type of the first parameter it finds. That is your proof that the 4-bit load actually ran, rather than merely importing without complaint.
If aligning versions on the host is eating your days, run the whole bench inside an official PyTorch container. A container is a packaged filesystem carrying its own copy of the lower floors, so the host’s versions stop mattering. You install the libraries on top. It also makes the bench reproducible, which is worth having the first time somebody asks to see the exact environment a result came out of.
Load a first dataset that shows you what is happening
Short section, and everyone needs it. The dataset for a first supervised run should be small, written by people, and already shaped like a conversation. HuggingFaceH4/no_robots is all three. It holds about 10,000 instruction and response pairs, written by hand, across categories such as generation, open question answering, brainstorming, rewriting, summarising and coding.
Three reasons it is the right thing to start on.
- It is small. One epoch, meaning one full pass over the training data, takes minutes. You can run it before lunch and again afterwards.
- It is written by people. So it gives you a clean baseline for what good supervised behaviour looks like. That baseline matters, because Part 10 asks you to train on bad data on purpose and watch what happens.
- It is already structured as messages. Each example is a list of turns. Each turn carries a role, either user or assistant, and its content. So the boundary between the question and the answer is already marked, which frees you to concentrate on the one mechanic that trips everybody: deciding which tokens count toward the loss.
from datasets import load_dataset
from transformers import AutoTokenizer
ds = load_dataset("HuggingFaceH4/no_robots")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B") # the BASE, not -Instruct
ex = ds["train"][0]
print(ex["messages"]) # already a list of role and content dicts
# render with the chat template so you can SEE the special tokens and the boundary
print(tok.apply_chat_template(ex["messages"], tokenize=False))
Line by line:
load_datasetfetches the dataset and caches it locally. It comes back already split into train and test.AutoTokenizer.from_pretrainedloads the tokenizer belonging to this exact model. The tokenizer has to match the model. A mismatched one maps text to the wrong IDs and the run trains on nonsense. Note that the model name has no-Instructsuffix. That is deliberate.- Printing
ex["messages"]shows you the raw structure: a list of dictionaries, each with a role and a piece of content. apply_chat_templateis the interesting one. A chat template is the rule that turns that list of messages into one flat string, with special marker tokens around each turn, in exactly the layout the model expects. Passingtokenize=Falsehands you the string rather than the IDs, so you can read those markers with your own eyes.
One thing to notice in that output, because it confuses people. A base model’s tokenizer usually ships those chat markers and a template. The base model’s weights have never been trained to respond to them. That is fine, and it is the point. The template is a convention that lives with the tokenizer. Teaching the weights to honour it is the actual job, and Part 9 does exactly that. Do not read “the tokenizer knows the template” as “the model knows how to chat”.
Key takeaways
- Choose a learning base model for how much machinery you can see per run. A production base model is chosen for task quality per dollar. Those are different jobs.
- The four properties that decide it: small enough to full-fine-tune on one card, fast to iterate, conventional in shape, and Base rather than Instruct.
- Multiply parameters by 16 to screen a full fine-tune. On one 32 GB card that makes 1.5B the sweet spot, 1B comfortable, and 7B impossible. Every one of those figures is static state, with activations still owed on top.
- A large vocabulary spends more of a small model’s fixed budget on the embedding lookup table and less on the layers. A small vocabulary reverses that and spends sequence length instead.
- Qwen2.5-1.5B base is the strongest general choice, on Apache-2.0. Llama-3.2-1B is the fastest sanity model and carries a conditional community licence. SmolLM2-1.7B is the one with a fully published training history.
- Apache-2.0 asks for a notice and nothing else. A community licence adds an acceptable-use policy, a naming rule and a large-deployer clause, and the dataset carries its own separate licence.
- The toolchain is a tower: driver, then CUDA, then a PyTorch build compiled for your card’s compute capability, then the libraries. No kernel image available means that build contains no code for your chip.
You can now
- Screen any candidate model against your card with one multiplication, and say what that number leaves out, from “Screen a candidate model in one multiplication”.
- Read the layer count, hidden size, vocabulary size and head configuration off a model card and say what each one costs you, from “Read a model spec sheet without an architecture background”.
- Defend a choice between three real base models on memory, licence and ecosystem, from “Choose between three concrete candidates”.
- Say plainly what Apache-2.0 and a community licence each let you do, and spot the two clauses people forget, from “Decide whether a model’s licence will cause you trouble”.
- Provision a rented GPU without a surprise invoice, using quota, teardown and spot rules, from “Rent a cloud GPU without a surprise bill”.
- Read a no kernel image available error, name which floor of the tower is wrong, and fix it, from “Fix the toolchain errors a new GPU throws at you”.
Glossary
- Activation
- Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
- Activation memory
- The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
- AdamW
- Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
- Adapter
- A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
- Attention
- The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
- Attention head
- One of several parallel copies of the attention computation, each free to look for a different kind of relationship. Their outputs are joined back together at the end. Multi-head attention explained
- Base model
- The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions.
- Batch
- A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
- bf16
- A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
- Chat template
- The rule that turns a list of role-and-content messages into the exact text and special tokens the model expects to see. It ships with the tokenizer, and it is where the boundary between prompt and response lives. From messages to tensors
- Checkpoint
- A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
- Compute capability
- NVIDIA’s version number for what a GPU chip can do, for example 12.0, written sm_120. Your CUDA and PyTorch build has to contain code compiled for it, or nothing runs and you get a no kernel image available error.
- Context length
- The longest sequence a model was built to handle, for example 32K tokens. It caps how long your training examples may be; it is not a target to train at. Qwen2.5-1.5B model card
- CUDA
- NVIDIA’s platform for running code on their GPUs. The driver, the CUDA toolkit and your PyTorch build all have to agree with each other and with the card before anything runs.
- Decoder-only
- The standard LLM architecture: one stack of identical transformer blocks, with a causal mask so each position sees only what came before it. Llama, Qwen and most open models are all this shape. Inside one transformer block
- Embedding
- The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
- Epoch
- One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one. Supervised fine-tuning end to end
- Feed-forward network
- The part of a transformer block that processes each position on its own, widening it to a larger size and squeezing it back. It holds most of a transformer’s weights and produces its largest activation. The transformer feed-forward network
- Flash attention
- An attention implementation that computes the answer in small tiles inside fast on-chip memory, never building the full score grid. Same result, far less memory, and the term that grew with the square of sequence length becomes linear. Dao et al., FlashAttention
- Full fine-tuning
- Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
- Gradient
- One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
- Gradient accumulation
- Running several small batches, adding their gradients together, and only updating the weights once at the end. You get the steadier signal of a big batch while holding just one small batch in memory. Batch size, accumulation and the effective batch
- Gradient checkpointing
- Throwing away most stored activations and recomputing them during the backward pass. Peak activation memory drops a long way in exchange for roughly 20 to 30 percent more time. Chen et al., Training Deep Nets with Sublinear Memory Cost
- Grouped-query attention
- An attention design where several query heads share one set of keys and values, which shrinks the KV cache with little quality loss. The example model has 12 query heads over 2 key-value heads, cutting the cache about sixfold. Ainslie et al., GQA
- Hidden size
- The width of the list of numbers that flows between layers, written d_model. The example model’s is 1,536, and it multiplies straight into activation memory. Inside one transformer block
- Inference
- Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
- Interconnect
- The link the GPUs use to talk to each other. It is the hidden limit on every multi-GPU strategy, and it is decided by the machine rather than by the cards. The wire between the cards
- Knowledge distillation
- Training a small model to copy a larger model’s outputs rather than learning from raw data alone. Llama-3.2-1B was built this way, after a larger model had been pruned down. Llama-3.2-1B model card
- KV cache
- The keys and values of past tokens, kept during generation so they are not recomputed for every new token. It exists only at inference; training has no generation loop and therefore no KV cache. Continuous batching and paged attention
- Layer
- One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
- Layer norm
- A step that rescales the numbers flowing through a layer so they stay in a sensible range, which keeps training stable. Modern LLMs use a cheaper version of it called RMSNorm. Inside one transformer block
- Logits
- The raw scores a model produces for every possible next token, before they are turned into probabilities. One number per token in the vocabulary, and higher means the model favours that token. The complete inference path
- LoRA
- Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
- Loss
- One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
- Mixed precision
- Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
- Multi-head attention
- Attention run as several heads in parallel over different slices of each position’s numbers, then recombined. Grouped-query attention is the memory-saving variant current models use. Multi-head attention explained
- OOM
- Out of memory, the error you get when a run needs more VRAM than the card has. In training it almost always strikes where the forward pass ends and the backward pass begins, which points straight at activations. Activation memory and gradient checkpointing
- Optimizer
- The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
- Parameter
- One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
- PCIe
- The general-purpose bus that connects cards to the rest of the machine. Without NVLink it is also how two GPUs talk to each other, at roughly a tenth of the speed and routed through the CPU. The interconnect is the hinge
- Pre-norm
- Normalising the input to each sub-layer rather than its output. It makes deep transformers much easier to train, and every model in this series uses it. Inside one transformer block
- Preference tuning
- Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
- Pretraining
- The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
- QLoRA
- LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
- Quantisation
- Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
- Queries, keys and values
- The three sets of numbers attention works from. Each position makes a query saying what it is looking for, a key advertising what it offers, and a value carrying what it passes on when its key is matched. Multi-head attention explained
- RMSNorm
- A cheaper version of layer norm that rescales values by their root mean square without first subtracting the average. Standard in current LLMs. Inside one transformer block
- RoPE
- Rotary position embeddings, the standard way modern LLMs encode where each token sits in the sequence, by rotating parts of the query and key numbers by an angle that depends on position. Multi-head attention explained
- SDPA
- PyTorch’s built-in scaled dot-product attention, which picks an efficient backend for you including a flash-attention style one. Using it saves installing a separate attention package.
- Sequence length
- How many tokens are in one training example after tokenisation. Activation memory grows in step with it, and the attention part grows with its square. Activation memory and gradient checkpointing
- Sharding
- Splitting the training state so each GPU holds only a slice of it and fetches the rest when it needs it. It is the fix for a model that will not fit, and it costs traffic between the cards. Rajbhandari et al., ZeRO
- Special token
- A token that stands for structure rather than ordinary text, such as the marker that opens or closes an assistant turn. They come from the tokenizer and are placed by the chat template. From messages to tensors
- Static state
- Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
- Structural pruning
- Permanently removing whole pieces of a trained model, such as layers or attention heads, to make it smaller. It is normally followed by more training to recover the quality that was lost. Llama-3.2-1B model card
- Supervised fine-tuning
- Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
- SwiGLU
- The gated feed-forward design used in most current LLMs, where one branch of the widened layer acts as a gate on the other. It is the reason a model’s feed-forward width is quoted as a single number such as 8,960. The transformer feed-forward network
- Synthetic data
- Training examples written by a model rather than by a person. Cheap and endlessly scalable, and it needs curation, a mix of real data alongside it, and a check on the licence of whatever model produced it. Wang et al., Self-Instruct
- Tied embeddings
- Reusing one weight matrix both to turn tokens into numbers at the input and to score tokens at the output. It saves a large slice of parameters on a small model. Qwen2.5-1.5B model card
- Token
- The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
- Tokenizer
- The component that splits text into tokens and maps them to integer IDs, and back again. Every model has its own, and it has to match the model you are training.
- Transformer
- The architecture behind every model in this series: a stack of blocks that alternate attention with a feed-forward network. Vaswani et al., Attention Is All You Need
- Transformer block
- One repeated unit of the model: attention, then a feed-forward network, with normalisation and residual connections around them. A 28-layer model is 28 of these stacked up. Inside one transformer block
- Vocabulary
- The complete set of tokens a model knows, typically somewhere between 32,000 and 151,000 of them. A larger vocabulary makes text shorter in tokens and spends more of a small model’s parameters on the embedding.
- VRAM
- The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
- Weight
- A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model
Practical exercises
Screen a 3B model against the 32 GB card
A 3 billion parameter model is proposed as the next step up from this part’s recommended base. Using the sixteen-bytes-per-parameter rule this part uses to screen candidates, compute its full fine tune static state in gigabytes and decide whether it fits on the one 32 GB card this series targets, showing the arithmetic rather than a rule of thumb.
See the worked solution (opens in a new tab)
Compute the embedding table’s share of the parameter budget
Qwen2.5-1.5B has a vocabulary of roughly 151,000 and a hidden size of 1,536, with tied embeddings so the input and output projections share one matrix. Compute that matrix’s parameter count and its share of the model’s 1.54 billion total, then compare it against the roughly 100 million parameters, about six percent share, this part gives for SmolLM2-1.7B’s smaller vocabulary. State which model spends the larger fraction of its budget on the embedding table, and by roughly what multiple.
See the worked solution (opens in a new tab)
Diagnose a no kernel image error on new hardware
A fresh install on a very new consumer card throws no kernel image is available for execution on the device the first time a flash-attention-style kernel needs to compile. Using this part’s toolchain section, name the specific missing piece most likely responsible, and the two concrete defenses this part recommends for a first run.
See the worked solution (opens in a new tab)
Pick a cloud instance for a full fine tune
You need to full fine tune Qwen2.5-1.5B on a rented instance rather than local hardware. Using this part’s own static state and activation figures for that model, and the VRAM column of its cloud GPU table, compute whether a g6.xlarge, one L4, 24 GB, can hold the run, then do the same for a g6e.xlarge, one L40S, 48 GB, and state which instance you would actually provision.
See the worked solution (opens in a new tab)
Resolve a three-way tradeoff for a model-risk review
A model-risk review needs a base model with fully documented training data provenance. The same project also needs strong multilingual coverage and the deepest possible tooling ecosystem, since the team is small. Using this part’s three-way comparison, explain why no single candidate satisfies all three requirements at once, then justify which requirement you would refuse to compromise on and why.
Frequently asked questions
Which base model should I use to learn fine-tuning?
A Base checkpoint of 1 to 2 billion parameters, with a conventional decoder shape and a permissive licence. Qwen2.5-1.5B base is a strong default. It fully fine-tunes on one 32 GB card at about 24.6 GB of static state, it is Apache-2.0 at that size, and every tool in the ecosystem already supports it.
Should I fine-tune a Base or an Instruct model?
For learning, take Base. You impose the chat format yourself and watch instruction following appear, which makes cause and effect visible. For shipping, an Instruct checkpoint is often the better start, since it already follows instructions and you are only adjusting behaviour on top of that.
How much VRAM do I need to fully fine-tune a 1.5B model?
About 24.6 GB of static state under standard mixed-precision AdamW, plus activations that land near 8 GB at a batch of 4 and a sequence of 1,024. So a 32 GB card fits it only with the memory tricks on: bf16 compute, gradient checkpointing, and a small per-device batch made up with gradient accumulation.
What does “no kernel image is available for execution on the device” mean?
It means the library you installed contains no GPU code compiled for your chip. Every NVIDIA chip has a compute capability number, and compiled GPU code is built ahead of time for a specific list of them. Your card’s number is not on the list in that build. Editing your script cannot fix it. Install a build that covers your card.
Does the licence of the base model matter?
For private learning, barely. For anything a client or a review board will see, a lot. Apache-2.0 asks you to keep a notice and nothing more. A community licence adds an acceptable-use policy, a naming requirement and a large-deployer clause, and your dataset carries its own separate licence on top.
Why does vocabulary size affect my fine-tune?
Because the lookup table that turns token IDs into numbers has one row per vocabulary entry, and every number in it is a parameter. A large vocabulary therefore spends more of a small model’s fixed budget on that table and less on the layers. A small vocabulary reverses the split, at the cost of splitting non-English text and code into more tokens, which eats sequence length.
Sources and further reading
- Qwen2.5-1.5B model card, for the layer count, head configuration, vocabulary size and licence.
- Qwen2.5 Technical Report, for the family’s training scale and the evaluation numbers behind the “strongest of the three” claim.
- Llama-3.2-1B model card, for its parameter count, context window, licence terms and prune-and-distil history.
- Grattafiori et al., The Llama 3 Herd of Models, for the family the 1B model was cut down from.
- SmolLM2-1.7B model card, for its parameter count, vocabulary size and published training data.
- HuggingFaceH4/no_robots dataset card, the human-written instruction set used throughout this series.
- Ainslie et al., GQA: Training Generalized Multi-Query Transformer Models, for why 12 query heads over 2 key-value heads shrinks the KV cache.
- Su et al., RoFormer, the source for the rotary position embeddings all three candidates use.
- Shazeer, GLU Variants Improve Transformer, the source for SwiGLU feed-forward blocks.
- Ba et al., Layer Normalization and Zhang and Sennrich, Root Mean Square Layer Normalization, for what normalisation does and what the RMS variant drops, with Xiong et al. for why it sits before each sub-layer.
- Hinton et al., Distilling the Knowledge in a Neural Network, for the teacher-and-student step in the Llama-3.2 provenance section.
- bitsandbytes documentation, for the 4-bit library the toolchain section tells you to verify on its own.
