Choosing a Base Model and Building a Fine-Tuning Bench

Inside LLM Fine-Tuning, part 8 of 15: Choosing a Base Model and Building a Fine-Tuning Bench

Picking a base model for fine-tuning looks like a one-line download. It is not. That one line fixes how much memory your run needs, which tokenizer you are stuck with, what you are allowed to do with the result, and how much of the software stack will simply work. Choose badly and your first week goes on fighting the environment instead of learning the mechanics.

This part covers the whole setup stage. It is deliberately several subjects on one page: what a spec sheet means, what a licence lets you do, what a rented GPU costs, how to get the drivers to cooperate, and which dataset to start on. Nobody needs all of that at once.

So the first section is a map. Read it, take the two or three sections you need today, and skip the rest without guilt. By the end you will have a base model you can defend, somewhere to run it, and a first dataset loaded.

Find the sections you need and skip the rest

The sections below cover subjects that have almost nothing to do with each other. They share a page because they all land in the same week. Here is what each one is, and who it is for.

  • The spec sheet, and the three candidates. What the numbers on a model card mean, then three real models compared. Everyone needs these two.
  • The licence. What each model’s terms let you do. Skip it if the work never leaves your own machine. Read it before a fine-tune of yours reaches a customer, a public repository or a review board.
  • Where the bench runs. Your own card, a rented card, or a managed service. Read it if you do not already have a GPU in front of you.
  • The toolchain. Drivers, CUDA and PyTorch builds. Skip it until something breaks. Then come straight back, because the answer is very likely here.
  • The first dataset. What to train on first, and why. Short, and everyone needs it.

One framing decision runs underneath all of it. A production base model is chosen to get the most task quality per dollar at serving time. A learning base model is chosen to let you see the most machinery per run. Those two goals pull in different directions. This part picks the second.

Say what a base model choice locks in

A base model is the raw pretrained checkpoint. A checkpoint is a saved copy of every weight in the model at one moment. Somebody spent months and a large budget teaching this one to continue text. It has never been taught to hold a conversation.

An Instruct model is that same base model after somebody fine-tuned it to follow instructions. That is the job Part 9 walks through. So the two are the same model at two different stages, and the choice between them is a choice about how much you want to watch.

Four properties matter more than any leaderboard score when the goal is to learn.

  1. Small enough to fully fine-tune on one card. Full fine-tuning updates every weight, and it is the expensive case. If you can run it, you get to feel the memory wall. Then you get to feel LoRA and QLoRA take that wall away. If your model only ever fits with QLoRA, you never learn what QLoRA saved you from.
  2. Fast enough to run twenty times a day. A model of 1 to 2 billion parameters finishes a small supervised run in minutes on one modern card. You want to break the run on purpose and fix it, not babysit an overnight job.
  3. Boringly conventional in shape. Every library, every chat template and every memory trick already supports the standard design. Unusual designs are where unsupported-operation errors live. For learning, boring is a feature.
  4. Base rather than Instruct. Take the raw checkpoint. Then you impose the chat format yourself, and you watch instruction following appear from nothing. An Instruct checkpoint has had that done to it already, which muddies cause and effect.

Past those four, every base-model choice is a point in the same space. The axes are the same for a production model. Only the weight on each one changes.

What you are choosing What it controls How much it matters for learning
Parameter count memory, speed, headroom on the card highest, it has to allow a full fine-tune
Model shape how much tooling supports it out of the box high, stay with a standard decoder
Licence what you may do with the weights and the output low while private, high once anyone else sees it
Vocabulary size how much of the budget goes to the lookup table, and how many languages it reaches medium, it moves the memory split
Context length the longest sequence allowed, and so the activation cost low, 8K is plenty for a first run
Base or Instruct whether a chat format is already baked in high, take Base
Published training data whether you can say what the model was trained on low for mechanics, high for a review board
Ecosystem depth tutorials, ready-made quantised copies, forum answers medium to high, fewer dead ends

Screen a candidate model in one multiplication

Whether a model fits on your card is arithmetic, not judgement. Part 2 derives the number. Standard mixed-precision AdamW costs 16 bytes for every parameter. Mixed precision means the arithmetic runs in a 16-bit format while a 32-bit copy of each weight is kept for the update. AdamW is the optimizer, the component that decides how far each weight moves.

Multiply the parameter count by 16 and you get gigabytes of static state. Static state means the weights, the gradients and the optimizer’s own stored numbers. It stays exactly the same size all run long. Here is the whole screen, worked:

  • 1.23 billion parameters times 16 is about 19.7 GB.
  • 1.54 billion times 16 is about 24.6 GB.
  • 1.71 billion times 16 is about 27.4 GB.
  • 7.62 billion times 16 is about 122 GB, which no single card holds.

Every one of those numbers is static state only. None of them includes activations. Activations are the intermediate results the forward pass produces on the way through, and the backward pass then needs them. They are counted separately, every single time. For Qwen2.5-1.5B at a batch of 4 sequences of 1,024 tokens, they land somewhere near 8 GB. Treat that as an order of magnitude rather than a measurement. Part 3 covers what moves it.

Now apply the screen to one 32 GB card. About 29.8 GiB of that is actually usable once the driver has taken its share, and Part 1 explains why those two units differ. A model near 1.5B parameters fits, and fits tightly enough that you feel the wall. A 1B model leaves comfortable headroom. A 1.7B model fits only with care.

Two settings do most of that caring, and both come back later.

  • Gradient checkpointing throws away most of the stored activations and recomputes them during the backward pass. Peak activation memory drops a long way, and the run takes roughly 20 to 30 percent longer.
  • Gradient accumulation runs several small batches, adds their gradients together, and updates the weights once at the end. You get the steadier signal of a large batch while holding only a small one in memory.

Leave 2 to 3 GB of margin in every estimate you make. Fragmentation and one unusually long batch both live in that last gigabyte.

Read a model spec sheet without an architecture background

This section is about the numbers printed on a model card. It assumes you know nothing about how a transformer is built. It teaches only the parts that change a decision.

Start with the one architectural fact that matters. All three models below are decoder-only transformers of the standard shape. Decoder-only means one stack of identical blocks, where each position in the text may look only at what came before it. Standard means everyone builds them this way. The practical consequence is the whole point: every tool in the ecosystem already loads them, so nothing in your toolchain will refuse.

If you want the names, here they are once. All three use rotary position embeddings, SwiGLU feed-forward blocks, and RMS normalisation placed before each sub-layer rather than after. A feed-forward block is the part of each layer that processes every position on its own. Normalisation is a step that rescales the numbers flowing through, so they stay in a sensible range. You do not need to know any more than that to choose a base model, and this series never asks you to. Those names matter here only as a sign that a model is conventional. Each of them has one paper behind it: Su et al. for the rotary embeddings, Shazeer for SwiGLU, and Zhang and Sennrich for RMS normalisation, a cheaper variant of the layer normalisation of Ba et al. The placement is a published result too: Xiong et al. found that normalising before each sub-layer makes training much better behaved than normalising after it. The companion part on one transformer block takes them apart properly.

Layers, hidden size and context length

Three numbers on every card describe the shape.

Layers is how many blocks are stacked. Qwen2.5-1.5B has 28. More layers means more work per token and a slower loop.

Hidden size is how wide the list of numbers flowing between layers is. Qwen2.5-1.5B’s is 1,536. Papers write it d_model.

Context length is the longest sequence the model was built to handle. It is a ceiling, not a target. A 32K model is perfectly happy training on 1,024-token examples. Training at 32K would cost far more activation memory than you have.

The vocabulary, and why its size is a trade

A tokenizer splits text into tokens and gives each token an integer ID. The vocabulary is the full list of tokens that tokenizer knows. Qwen2.5’s holds about 151,000 entries.

A model cannot do arithmetic on an integer ID. It needs actual numbers. So the first thing it does is look the ID up in a table. That table has one row per vocabulary entry, and each row is as wide as the hidden size. It is called the embedding table. Every number in it is a parameter, exactly like the ones inside the layers.

Do the multiplication on a toy model. Take a vocabulary of 1,000 and a hidden size of 10. The table is 1,000 rows of 10 numbers, so 10,000 parameters. Now double the vocabulary to 2,000. The table doubles to 20,000. The layers have not changed at all.

That is the trade, and it bites hardest on small models. A model has a fixed parameter budget. Every number in the lookup table is a number not in the layers that do the work. A big vocabulary spends a lot of a small model’s budget on the table.

There is a saving available. The model also has to turn its final numbers back into one score per token, and that needs a table of exactly the same shape. Many small models reuse one matrix for both jobs. That is called tied embeddings, and it halves the cost.

SmolLM2-1.7B makes the contrast concrete. Its vocabulary is about 49,000 and it ties its embeddings, so that one matrix holds roughly 100 million parameters. Against a total of 1.71 billion, that is about six percent. Qwen2.5-1.5B carries a vocabulary three times larger inside a narrower hidden size, so its share is much bigger. The exercises at the end ask you to work out how much bigger.

The other half of the trade is what the tokenizer does to your text. A large vocabulary has room for more whole words, and for words in more languages. The same sentence then turns into fewer tokens. Fewer tokens means shorter sequences, and shorter sequences cost less activation memory. A small vocabulary saves parameters and spends sequence length. Neither answer is right in general.

Attention heads, and the one place these models differ

Attention is the step where each position in the text looks at the other positions and pulls in whatever is useful. It runs several times in parallel, over different slices of the numbers. Each parallel run is called a head.

Every head needs three things per token: a query, a key and a value. Loosely, the query is what this position is looking for. The key and the value are what an earlier position offers. In the original design each head had its own keys and values. Grouped-query attention changes that. Several query heads share one set of keys and values.

Qwen2.5-1.5B has 12 query heads over 2 key-value heads. Divide, and six query heads share each key-value pair. That is the entire idea.

Here is why anyone bothers. When a model is generating text it keeps the keys and values of every token it has already produced, so it does not have to work them out again. That saved pile is called the KV cache, and it grows as the conversation gets longer. Storing 2 sets instead of 12 makes it roughly six times smaller.

For your training run this changes almost nothing. Training has no generation loop, so it has no KV cache. It matters later, when you serve what you trained, and it is a real reason this model is cheap to run. The companion part on multi-head attention has the mechanism.

Choose between three concrete candidates

Three real models, read as spec sheets rather than as a ranking. Every one of them can be fully fine-tuned on one 32 GB card. So the decision comes down to the second-order differences.

Qwen2.5-1.5B params 1.54 B layers 28 heads Q/KV 12 / 2 GQA ctx 32K vocab ~151K embed tie yes train tok 18 T license Apache-2.0 made by Alibaba / Qwen full-FT static 24.6 GB multilingual · strong Llama-3.2-1B params 1.23 B layers 16 heads GQA ctx 128K vocab 128K embed tie yes train tok ~9 T license Llama 3.2 Comm. made by Meta (pruned+ distilled 8B) full-FT static 19.7 GB SmolLM2-1.7B params 1.71 B layers 24 heads MHA (32/32)* ctx 8K vocab 49K embed tie yes train tok 11 T license Apache-2.0 made by Hugging Face full-FT static 27.4 GB fully-open data recipe

Three specification cards, not a leaderboard. The line worth checking on any candidate before you commit is the attention configuration, because it is the only place these three genuinely differ in shape.

Qwen2.5-1.5B, the capable generalist

1.54 billion parameters. 28 layers, hidden size 1,536, a vocabulary of about 151,000, tied embeddings, 12 query heads over 2 key-value heads, and a native context of 32K. Released under Apache-2.0 at this size.

Its full fine-tune static state is about 24.6 GB, so a 32 GB card holds it with the activations squeezed. It is the strongest of the three on languages other than English, and on maths and code for its size, which is the case Qwen’s own technical report makes with its evaluation tables. Its ecosystem is the deepest of the three. This series uses it for every calculation.

Llama-3.2-1B, the fast sanity model

1.23 billion parameters. Just 16 layers, a 128K context window, a vocabulary of about 128,000, and roughly 9 trillion training tokens behind it. Its static state is about 19.7 GB, the lightest of the three. Sixteen layers makes it the quickest loop for confirming that a script runs at all.

Its history is the interesting part, and it is worth a minute because it is a clean example of how small models get made. This one was not trained from scratch. It was built in two steps.

Step 1: cut pieces out of a bigger model. Start with a trained 8 billion parameter model, one of the family Grattafiori et al. document in the Llama 3 paper. Remove whole pieces of it, such as layers and attention heads, until what remains is the size you wanted. That is called structural pruning. The result runs, and it is worse than what you started with, because you have thrown away parts that were doing something.

Step 2: repair it by copying a bigger model. Show a large model some text and record the full set of scores it produces for the next token. Part 1 calls those scores logits: one raw number for every token in the vocabulary. Then train the pruned model to produce a similar set of scores at the same positions. That is knowledge distillation. The big model is the teacher and the small one is the student.

Copying the whole score list teaches more than copying only the correct token would. It also shows the student which other tokens the teacher found plausible, and by how much. That is the argument Hinton, Vinyals and Dean made when they introduced distillation: the teacher’s full set of scores carries more information per example than the single correct answer does. None of this changes how you use the model. It is simply an unusually visible example of a model-shaping pipeline, and you are about to run a small one yourself.

The licence is the flag on this one. It gets its own section below.

SmolLM2-1.7B, the auditable one

1.71 billion parameters. 24 layers, an 8K context, a small vocabulary of about 49,000, tied embeddings, Apache-2.0. Trained on 11 trillion tokens from a data mix that is published in full. Its static state is about 27.4 GB, the tightest of the three, so a 32 GB card holds it only if you keep activations down.

Its small vocabulary is the trade from the last section made real. More of its budget sits in the layers. Its tokenizer is less efficient on code and on text that is not English, which costs you sequence length. It is English-first, its ecosystem is the smallest here, and it is the only one of the three whose entire training history you can read.

Side by side

Criterion Qwen2.5-1.5B Llama-3.2-1B SmolLM2-1.7B
Full fine-tune static state about 24.6 GB about 19.7 GB, lightest about 27.4 GB, tightest
Fits one 32 GB card yes, with care yes, easiest yes, manage activations
Licence Apache-2.0 community licence Apache-2.0
Training data published no, weights only no, weights only yes, in full
Ecosystem depth deepest very deep growing
Languages strongest eight languages English first
What it teaches you why shared key-value heads are cheap to serve how prune-and-distil builds a small model what a fully documented lineage looks like

The recommendation for a learning bench is Qwen2.5-1.5B base. It is Apache-2.0, it full-fine-tunes on one 32 GB card at about 24.6 GB of static state, and its shape is supported everywhere.

Keep Llama-3.2-1B as a sanity model. At 19.7 GB and 16 layers it is the fastest way to confirm that a script works before you point it at anything bigger. Reach for SmolLM2-1.7B when the story is provenance, and you need to be able to say that every token the model saw is documented.

Decide whether a model’s licence will cause you trouble

Skip this section if the work stays on your own machine. Nothing here changes what you may do privately. Read it before a fine-tune of yours reaches a customer, a public repository, an app store or a review board. That happens sooner than people expect, because a model picked for learning quietly becomes the production default.

Two kinds of licence turn up on model weights. They are not the same kind of thing.

Apache-2.0, which is what people mean by open source

Qwen2.5-1.5B and SmolLM2-1.7B both ship under it. In plain terms:

  • You may use the weights for anything, including a paid product.
  • You may change them, fine-tune them, and pass the result on.
  • You may keep your own changes closed.
  • You must keep the licence text and the copyright notice with any copy you hand on, and note which files you changed.
  • You get a patent licence from whoever released the model. It ends if you sue them over a patent covering it.
  • There is nobody to ask, nothing to report, and no rule about what you name the result.

That is a short procurement conversation. In many companies it is not a conversation at all.

A community licence, which is free but conditional

Llama-3.2-1B ships under a community licence instead. It is free. It allows commercial use. It allows you to fine-tune and to redistribute. It is also not an open-source licence in the formal sense, because it carries three conditions that Apache-2.0 does not.

  1. An acceptable-use policy. A written list of uses you agree not to put the model to. It travels with the weights. So anyone you hand your fine-tune to is bound by it as well.
  2. Naming and attribution. If you publish a model built on it, you have to say so, display the required notice, and carry the Llama name at the front of your own model’s name.
  3. A large-deployer clause. Above a monthly-active-user threshold written into the licence, you have to go and request a separate licence, which may or may not be granted. The threshold is high enough that it will not touch you unless your product is already one of the largest in the world.

For research and for most commercial work, this is fine. It is not equivalent to Apache-2.0, and a procurement lawyer will notice the difference. Read the licence file itself rather than a summary of it, including this one.

Two things people forget

The first is outputs. Some community licences also say something about what you may do with text the model generates. That matters if you plan to use one model to build training data for another. Read that clause before you build a data pipeline on it. Part 10 covers generated training data and when it earns its keep.

The second is the dataset. A dataset carries its own licence, entirely separate from the model’s. A permissive model plus a dataset restricted to non-commercial use still leaves you with a non-commercial result. Check both, every time.

Run a training bench beside a model that is already serving

This section is for people with two cards in one machine. Skip it if you have one.

You do not have to stop serving in order to start training. A full fine-tune of a 1.5B model needs about 24.6 GB of static state, plus a few gigabytes of activations once gradient checkpointing is on. That fits inside one 32 GB card. So one card can keep answering requests while the other one trains.

GPU 0 — PRODUCTION CUDA_VISIBLE_DEVICES=0 inference model + KV cache (APIM) served via your router power-limited via nvidia-smi -pl untouched by training GPU 1 — BENCH CUDA_VISIBLE_DEVICES=1 weights bf16 (~3 GB) grads + optimizer (~21 GB) activations ≈ 28–30 GB peak · full FT Qwen-1.5B keep 2–3 GB headroom no NVLink needed — single-card job, zero cross-GPU traffic

One card serves, one card trains, and neither process can see the other’s memory. Because a 1.5B full fine-tune fits on a single card, none of the multi-GPU machinery from Part 14 is needed here.

Set CUDA_VISIBLE_DEVICES on the training process so it sees only one card, and leave the serving process on the other. That variable controls which GPUs a process is allowed to see. It is the whole isolation mechanism.

Be clear about what this is not. Two cards used this way are two separate benches. They are not one memory pool of double the size. A model too big for one card has to be deliberately split across both, which is a different job with different machinery. Part 14 covers when a second card genuinely helps, and why the interconnect, the link the two cards use to talk to each other, decides the answer.

One caution that has nothing to do with memory. The two cards still share a power supply, a bus and the host’s RAM. A full fine-tune will pull as much power as its card is allowed to draw. If the serving side is sensitive to latency, put a power cap on the training card with nvidia-smi. The memory is already isolated. The electricity is not.

Pick where the bench runs, and know what each option hides

Read this if you do not already have a GPU in front of you. There are three places to put a fine-tuning bench. Only one of them is a classroom. The other two are places to deploy to, and they are worth understanding because clients ask about them. The rule is simple: the more the platform manages for you, the more of the mechanism it hides.

LOCAL GPU AWS EC2 VM BEDROCK you see the driver✓ you see CUDA/PyTorch✓ you write the train loop✓ you own checkpoints✓ you manage the GPU box✓ nothing hidden driver/CUDA on the AMI✓ stable stack (no Blackwell)✓ you write the train loop✓ you own checkpoints✓ AWS manages hardware~ hides the box, not the code no driver/CUDA to see✗ no train loop (managed)✗ import fine-tuned weights✓ serverless invoke✓ AWS manages everything✗ hides the mechanics you want

Learn on the left, deploy on the right. Your own machine and a rented virtual machine both let you write the training loop yourself. A fully managed customisation service deliberately takes it away.

Your own machine

Everything is visible, and everything is your problem. You see the driver, the memory, the speed and every error message. That is precisely the point. The cost is the toolchain section below. It is longer on your own hardware than anywhere else, because a card you bought last month is newer than most of the software that has to support it.

A rented virtual machine with a GPU attached

You still write the training loop yourself, so almost nothing is hidden. What disappears is the hardware problem. Reach for it to run a job without touching a production box, to try a bigger card than you own, to reproduce a run on standard datacentre hardware, or to show a client a clean cloud pipeline.

Cloud GPUs are usually a generation or two behind the newest consumer cards. That is an advantage here. The stable software stack simply works on them, with none of the version drama. Costs get their own section next.

A managed platform

A managed platform offers two different doors, and it matters a great deal which one you walk through. AWS Bedrock is the example used here. Other clouds have equivalents.

The first door is importing your own weights. You fine-tune wherever you like. Then you merge the adapter into the base model. An adapter is the small file that LoRA produces, and merging folds its learned patch permanently into the weights, so you are left with one ordinary model file. You upload that file, and the platform serves it with no machine for you to manage. Support is normally limited to a named list of model families, which is a concrete payoff for having picked a conventional one.

fine-tune on GPU / EC2 merge adapter → full weights upload to S3 bucket import CMI (Qwen/Llama) invoke serverless

Everything that teaches you fine-tuning happens in the first box. Everything after it is packaging and serving, which is genuinely useful as a deliverable and teaches you nothing about training.

The second door is managed customisation, where the platform runs the training for you. Do not learn on it. There are two reasons. The supported model list tends to be large models only. And it hides the training loop, the masking and the optimizer, which are exactly the things you came to see. Know it exists, because clients will ask.

Rent a cloud GPU without a surprise bill

This section is about money and paperwork rather than machine learning. Skip it if you are training locally.

AWS EC2 is the concrete example. Other clouds have equivalent instance types under different names. Four instances cover the useful range.

Instance GPU VRAM Good for Indicative on-demand rate
g6.xlarge one L4 24 GB LoRA and QLoRA, cheapest modern option under a dollar an hour
g5.xlarge one A10G 24 GB LoRA and QLoRA, very well trodden about a dollar an hour
g6e.xlarge one L40S 48 GB full fine-tune of 1B to 1.7B with room a low single-digit dollar rate
p4d.24xlarge eight A100 320 GB overkill here, multi-GPU work later tens of dollars an hour

Treat every price in that table as indicative. Rates differ by region and they change often. Check the console before you provision anything.

Notice what the memory arithmetic does to that table. A 24 GB card cannot hold a 1.5B model’s 24.6 GB of full fine-tune static state, let alone its activations. So on those two instances you use LoRA or QLoRA, which freeze the base model and train a small patch beside it. For a real full fine-tune you step up to the 48 GB instance. The multiplication decides the invoice.

1 · QUOTA request G/P vCPU limit increase 2 · AMI Deep Learning AMI (CUDA+PyTorch ready) 3 · INSTANCE g6e.xlarge + key pair 4 · NETWORK VPC + SG SSH from your IP 5 · STORE EBS gp3 for ckpts 6 · S3 + IAM dataset in S3, role grants read/write 7 · TRAIN your SFT script, checkpoint to EBS→S3 8 · TEARDOWN stop/terminate — GPU bills by the second

Two of these steps cost people a day each. A new account often has a quota of zero on GPU instance families, so ask for the increase before you plan anything. And a GPU instance bills by the second whether the GPU is busy or idle.

Four habits keep the bill small.

  1. Ask for quota first. New accounts frequently have a limit of zero on GPU instance families. The request can take a day to be granted. Make it before you need it.
  2. Stop when idle, terminate when finished. Billing counts seconds of instance life, not seconds of GPU work. A forgotten idle instance costs the same as a busy one.
  3. Copy artefacts out before you terminate. Push checkpoints and logs to object storage. The instance’s disk goes when the instance does, and it charges you for every hour it exists.
  4. Use spot capacity for jobs that checkpoint. Spot is commonly 60 to 70 percent cheaper than the on-demand rate, and that figure is indicative too. The catch is that the instance can be taken back with little warning. If your run saves a checkpoint regularly, that is an inconvenience. If it does not, it is a lost day.

One shortcut is worth knowing. A deep-learning machine image ships with the driver, CUDA and PyTorch already matched for that instance’s GPU. Choosing one skips the entire next section.

Fix the toolchain errors a new GPU throws at you

Skip this section until something breaks. Then come straight back. If your card is a year or two old, or you are on a cloud image, the install below simply works and you will never need the rest. If your card is very new, this is the section that saves your week.

A fine-tuning environment is a tower of four floors. Each floor has to match the one below it. Get the bottom two wrong and nothing above them runs, with an error message that points nowhere useful.

RTX 5090 · Blackwell · sm_120 hardware NVIDIA driver ≥ 570.x exposes the GPU CUDA 12.8 (min) / 12.9 sm_120 kernels PyTorch ≥ 2.7 · cu128 wheels (2.11 current) the floor that matters most transformers · datasets · accelerate · tokenizers HF core TRL · PEFT · bitsandbytes SFT/DPO · LoRA · 4-bit FlashAttention-2 / torch SDPA · W&B · lm-eval attention · tracking · eval your SFT script (Part 9) what you write

Align the bottom three floors and everything above them installs normally. Get them wrong and you spend hours on an error message that never mentions the real problem.
  1. The driver. The software your operating system uses to talk to the card. It comes from the GPU vendor. Running nvidia-smi prints its version and proves the card is visible at all.
  2. CUDA. The vendor’s toolkit for running your own code on the card. It has its own version number, separate from the driver’s. A newer driver can serve an older CUDA. It does not work the other way round.
  3. A PyTorch build compiled for your card. This is the load-bearing floor. PyTorch ships as several different builds, one per CUDA version, and you choose which one you install.
  4. The libraries. Transformers, datasets, accelerate, TRL, PEFT and bitsandbytes. These are ordinary Python packages. They install normally once the floors below them agree.

What compute capability is, and what the error actually means

Every NVIDIA chip carries a version number saying what it can do. It is called the compute capability. It is written two ways, as 12.0 and as sm_120, and they mean the same thing.

Code that runs on a GPU is compiled ahead of time, for a specific list of compute capabilities, and shipped inside the package you install. If your chip’s number is not on that list, the package contains no code your chip can run. Nothing happens.

That is exactly what this message is telling you:

no kernel image is available for execution on the device

A kernel is one small program that runs on the GPU. A kernel image is the compiled form of it. So the message says: I looked in the package for code compiled for your chip, and there is none. It is not a bug in your script. Editing your script cannot fix it. The fix is always to install a build that includes your card’s number.

Concretely: a current Blackwell consumer card reports compute capability 12.0, which is sm_120, and that needs a CUDA 12.8 build of PyTorch.

# 1. is the card visible, and what driver is running?
nvidia-smi

# 2. install a PyTorch build made for your card's CUDA version
#    (swap the index URL for the build your card needs)
pip install torch --index-url https://download.pytorch.org/whl/cu128

# 3. confirm the compiled code for your chip is really there,
#    rather than only that CUDA loaded
python -c "import torch; print(torch.__version__, torch.cuda.get_device_capability())"

# 4. the fine-tuning libraries, once the floors below them agree
pip install transformers datasets accelerate trl peft bitsandbytes

Line by line:

  • nvidia-smi prints the driver version and lists the cards it can see. If it prints nothing useful, stop here. Nothing above this floor can work.
  • The install line uses --index-url. By default pip downloads from its own package site, which carries one general PyTorch build. PyTorch publishes its CUDA-specific builds on its own site, and that flag points pip there instead. The cu128 in the address is the CUDA version. Change it to match what your card needs.
  • The one-line Python check does the thing people skip. It does not only prove that CUDA loaded. get_device_capability() prints your card’s real compute capability as a pair of numbers, so you can compare it against the build you just installed.
  • The last line installs the fine-tuning libraries on top of a floor you have now verified.

Pin your versions. Write the exact version of every one of those packages into a requirements file the first time the run works. These libraries move quickly and their argument names change between releases. The code in this series is current as of August 2026, and an unpinned environment will drift out from under you.

Two traps on very new hardware

The first trap is compiling on the fly. Some libraries do not ship precompiled code for every chip. Instead they build a kernel at the moment it is first needed. That needs CUDA’s runtime compiler, a component called nvrtc, and some installers leave it out. When it is missing, the on-the-fly build fails and you get the same “no kernel image” message as before, from a completely different cause.

Two defences, both cheap. First, prefer prebuilt wheels over building from source. A wheel is a Python package whose compiled parts were built by somebody else before you downloaded it, so nothing has to compile on your machine. Second, for a first run, use PyTorch’s own built-in attention implementation. It is called scaled dot-product attention, usually shortened to SDPA. The alternative is installing a separate flash-attention package. Flash attention is a faster way of computing attention that never builds the full grid of scores. It is worth having eventually. A small supervised run does not need the extra speed, so add it once the basic loop works.

The second trap is bitsandbytes. It holds the 4-bit code that QLoRA depends on, and its support for a brand new chip usually lands after PyTorch’s. Test it on its own before you build a training script on top of it. Otherwise a broken library looks exactly like a bug in your own code.

import bitsandbytes as bnb
print(bnb.__version__)                      # import test: fails loudly if the build is broken

import torch
from transformers import AutoModelForCausalLM, BitsAndBytesConfig

qconf = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
m = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B",
                                          quantization_config=qconf, device_map="auto")
print(next(m.parameters()).dtype)           # confirms the 4-bit load actually ran, not just imported

What each line does:

  • The import and the version print are the blunt test. A broken build fails right there, loudly, before anything else has run.
  • BitsAndBytesConfig(load_in_4bit=True, ...) asks for the model’s weights to be stored in 4 bits instead of 16. Storing numbers in fewer bits is called quantisation, and Part 12 covers what it costs.
  • bnb_4bit_compute_dtype=torch.bfloat16 says the arithmetic still happens in a 16-bit format. Nothing multiplies 4-bit numbers directly, so they are expanded back before every multiply.
  • from_pretrained(...) downloads the model and loads it under that configuration. device_map="auto" lets the library decide where each piece goes.
  • The last line prints the data type of the first parameter it finds. That is your proof that the 4-bit load actually ran, rather than merely importing without complaint.

If aligning versions on the host is eating your days, run the whole bench inside an official PyTorch container. A container is a packaged filesystem carrying its own copy of the lower floors, so the host’s versions stop mattering. You install the libraries on top. It also makes the bench reproducible, which is worth having the first time somebody asks to see the exact environment a result came out of.

Load a first dataset that shows you what is happening

Short section, and everyone needs it. The dataset for a first supervised run should be small, written by people, and already shaped like a conversation. HuggingFaceH4/no_robots is all three. It holds about 10,000 instruction and response pairs, written by hand, across categories such as generation, open question answering, brainstorming, rewriting, summarising and coding.

Three reasons it is the right thing to start on.

  1. It is small. One epoch, meaning one full pass over the training data, takes minutes. You can run it before lunch and again afterwards.
  2. It is written by people. So it gives you a clean baseline for what good supervised behaviour looks like. That baseline matters, because Part 10 asks you to train on bad data on purpose and watch what happens.
  3. It is already structured as messages. Each example is a list of turns. Each turn carries a role, either user or assistant, and its content. So the boundary between the question and the answer is already marked, which frees you to concentrate on the one mechanic that trips everybody: deciding which tokens count toward the loss.
from datasets import load_dataset
from transformers import AutoTokenizer

ds = load_dataset("HuggingFaceH4/no_robots")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B")   # the BASE, not -Instruct

ex = ds["train"][0]
print(ex["messages"])            # already a list of role and content dicts

# render with the chat template so you can SEE the special tokens and the boundary
print(tok.apply_chat_template(ex["messages"], tokenize=False))

Line by line:

  • load_dataset fetches the dataset and caches it locally. It comes back already split into train and test.
  • AutoTokenizer.from_pretrained loads the tokenizer belonging to this exact model. The tokenizer has to match the model. A mismatched one maps text to the wrong IDs and the run trains on nonsense. Note that the model name has no -Instruct suffix. That is deliberate.
  • Printing ex["messages"] shows you the raw structure: a list of dictionaries, each with a role and a piece of content.
  • apply_chat_template is the interesting one. A chat template is the rule that turns that list of messages into one flat string, with special marker tokens around each turn, in exactly the layout the model expects. Passing tokenize=False hands you the string rather than the IDs, so you can read those markers with your own eyes.

One thing to notice in that output, because it confuses people. A base model’s tokenizer usually ships those chat markers and a template. The base model’s weights have never been trained to respond to them. That is fine, and it is the point. The template is a convention that lives with the tokenizer. Teaching the weights to honour it is the actual job, and Part 9 does exactly that. Do not read “the tokenizer knows the template” as “the model knows how to chat”.

Key takeaways

  • Choose a learning base model for how much machinery you can see per run. A production base model is chosen for task quality per dollar. Those are different jobs.
  • The four properties that decide it: small enough to full-fine-tune on one card, fast to iterate, conventional in shape, and Base rather than Instruct.
  • Multiply parameters by 16 to screen a full fine-tune. On one 32 GB card that makes 1.5B the sweet spot, 1B comfortable, and 7B impossible. Every one of those figures is static state, with activations still owed on top.
  • A large vocabulary spends more of a small model’s fixed budget on the embedding lookup table and less on the layers. A small vocabulary reverses that and spends sequence length instead.
  • Qwen2.5-1.5B base is the strongest general choice, on Apache-2.0. Llama-3.2-1B is the fastest sanity model and carries a conditional community licence. SmolLM2-1.7B is the one with a fully published training history.
  • Apache-2.0 asks for a notice and nothing else. A community licence adds an acceptable-use policy, a naming rule and a large-deployer clause, and the dataset carries its own separate licence.
  • The toolchain is a tower: driver, then CUDA, then a PyTorch build compiled for your card’s compute capability, then the libraries. No kernel image available means that build contains no code for your chip.

You can now

  • Screen any candidate model against your card with one multiplication, and say what that number leaves out, from “Screen a candidate model in one multiplication”.
  • Read the layer count, hidden size, vocabulary size and head configuration off a model card and say what each one costs you, from “Read a model spec sheet without an architecture background”.
  • Defend a choice between three real base models on memory, licence and ecosystem, from “Choose between three concrete candidates”.
  • Say plainly what Apache-2.0 and a community licence each let you do, and spot the two clauses people forget, from “Decide whether a model’s licence will cause you trouble”.
  • Provision a rented GPU without a surprise invoice, using quota, teardown and spot rules, from “Rent a cloud GPU without a surprise bill”.
  • Read a no kernel image available error, name which floor of the tower is wrong, and fix it, from “Fix the toolchain errors a new GPU throws at you”.

Glossary

Activation
Any intermediate value the forward pass produces on its way from input to output. Activations are held in memory because the backward pass needs them to work out the gradients. Activation memory, from birth to death
Activation memory
The memory holding the forward pass’s intermediate values until the backward pass consumes them. Its size comes from batch size, sequence length, hidden size and layer count, never from parameter count. Why the forward pass costs more than the weights
AdamW
Adam with the weight decay applied straight to the weight instead of folded into the gradient. It is the default optimizer for essentially every LLM fine-tune, and its stored state is 12 of the 16 bytes per parameter. AdamW explained, line by line
Adapter
A small set of extra trainable weights added beside a frozen model, so you train the adapter and leave the model alone. A LoRA adapter is tens of megabytes against a multi-gigabyte model copy. LoRA explained
Attention
The step where each position in the sequence looks at other positions and mixes in whatever it finds useful. It is what lets a model use context instead of reading each token in isolation. Multi-head attention explained
Attention head
One of several parallel copies of the attention computation, each free to look for a different kind of relationship. Their outputs are joined back together at the end. Multi-head attention explained
Base model
The raw pretrained checkpoint, before anyone has taught it to hold a conversation. An Instruct model is a base model that has already been fine-tuned to follow instructions.
Batch
A group of examples processed together in one step, so the GPU stays busy and the gradient is averaged over several examples instead of one. Bigger batches give a steadier signal and cost more activation memory. Supervised fine-tuning end to end
bf16
A 16-bit number format with 8 exponent bits and 7 mantissa bits, so it reaches as far as fp32 with much coarser steps. It is the training default because it needs no loss scaling. Number formats for training
Chat template
The rule that turns a list of role-and-content messages into the exact text and special tokens the model expects to see. It ships with the tokenizer, and it is where the boundary between prompt and response lives. From messages to tensors
Checkpoint
A saved copy of the model at some point in the run, so you can resume from it, compare it, or ship it. Distinct from gradient checkpointing, which is a memory trick with an unfortunately similar name. Supervised fine-tuning end to end
Compute capability
NVIDIA’s version number for what a GPU chip can do, for example 12.0, written sm_120. Your CUDA and PyTorch build has to contain code compiled for it, or nothing runs and you get a no kernel image available error.
Context length
The longest sequence a model was built to handle, for example 32K tokens. It caps how long your training examples may be; it is not a target to train at. Qwen2.5-1.5B model card
CUDA
NVIDIA’s platform for running code on their GPUs. The driver, the CUDA toolkit and your PyTorch build all have to agree with each other and with the card before anything runs.
Decoder-only
The standard LLM architecture: one stack of identical transformer blocks, with a causal mask so each position sees only what came before it. Llama, Qwen and most open models are all this shape. Inside one transformer block
Embedding
The lookup table that turns each token ID into a list of numbers the model can do arithmetic on. Its size is vocabulary times hidden size, so a large vocabulary spends a lot of a small model’s parameters here. The complete inference path
Epoch
One full pass over the training data. Most instruction fine-tunes need only one to three, and preference tuning usually needs exactly one. Supervised fine-tuning end to end
Feed-forward network
The part of a transformer block that processes each position on its own, widening it to a larger size and squeezing it back. It holds most of a transformer’s weights and produces its largest activation. The transformer feed-forward network
Flash attention
An attention implementation that computes the answer in small tiles inside fast on-chip memory, never building the full score grid. Same result, far less memory, and the term that grew with the square of sequence length becomes linear. Dao et al., FlashAttention
Full fine-tuning
Updating every weight in the model. It costs 16 bytes per parameter of static state under standard mixed-precision AdamW, produces a complete model copy per task, and forgets the most. Training memory and the 16 bytes per parameter
Gradient
One number per weight saying which way to nudge that weight to make the loss smaller, and how steeply the loss responds. Picture the slope under a ball rolling into a valley. How a neural network learns
Gradient accumulation
Running several small batches, adding their gradients together, and only updating the weights once at the end. You get the steadier signal of a big batch while holding just one small batch in memory. Batch size, accumulation and the effective batch
Gradient checkpointing
Throwing away most stored activations and recomputing them during the backward pass. Peak activation memory drops a long way in exchange for roughly 20 to 30 percent more time. Chen et al., Training Deep Nets with Sublinear Memory Cost
Grouped-query attention
An attention design where several query heads share one set of keys and values, which shrinks the KV cache with little quality loss. The example model has 12 query heads over 2 key-value heads, cutting the cache about sixfold. Ainslie et al., GQA
Hidden size
The width of the list of numbers that flows between layers, written d_model. The example model’s is 1,536, and it multiplies straight into activation memory. Inside one transformer block
Inference
Using a trained model to produce output. There is no backward pass and no optimizer, which is why serving a model costs a fraction of what training it does. The complete inference path
Interconnect
The link the GPUs use to talk to each other. It is the hidden limit on every multi-GPU strategy, and it is decided by the machine rather than by the cards. The wire between the cards
Knowledge distillation
Training a small model to copy a larger model’s outputs rather than learning from raw data alone. Llama-3.2-1B was built this way, after a larger model had been pruned down. Llama-3.2-1B model card
KV cache
The keys and values of past tokens, kept during generation so they are not recomputed for every new token. It exists only at inference; training has no generation loop and therefore no KV cache. Continuous batching and paged attention
Layer
One processing stage inside the model, taking a list of numbers in and handing a transformed list out. The example model is 28 transformer layers deep. Inside one transformer block
Layer norm
A step that rescales the numbers flowing through a layer so they stay in a sensible range, which keeps training stable. Modern LLMs use a cheaper version of it called RMSNorm. Inside one transformer block
Logits
The raw scores a model produces for every possible next token, before they are turned into probabilities. One number per token in the vocabulary, and higher means the model favours that token. The complete inference path
LoRA
Low-rank adaptation. Freeze the model and learn a small pair of skinny matrices beside each targeted weight matrix, so about 1 percent of parameters train and the static state for a 1.5B run drops from about 24.6 GB to about 4 GB. Hu et al., LoRA
Loss
One number saying how wrong the model was on this batch. Training is the whole business of making it smaller, and a falling loss on its own proves very little. How a neural network learns
Mixed precision
Doing the arithmetic in a 16-bit format for speed and memory while keeping a 32-bit copy of the weights so small updates are not lost to rounding. Essentially every modern training run works this way. Micikevicius et al., Mixed Precision Training
Multi-head attention
Attention run as several heads in parallel over different slices of each position’s numbers, then recombined. Grouped-query attention is the memory-saving variant current models use. Multi-head attention explained
OOM
Out of memory, the error you get when a run needs more VRAM than the card has. In training it almost always strikes where the forward pass ends and the backward pass begins, which points straight at activations. Activation memory and gradient checkpointing
Optimizer
The part of training that turns gradients into actual weight changes. The gradient says which way to move, and the optimizer decides how far. Gradients and optimizers explained
Parameter
One of the numbers inside the model that training can change. Parameter and weight mean the same thing here, and a 1.5B model has about 1.5 billion of them. Training memory and the 16 bytes per parameter
PCIe
The general-purpose bus that connects cards to the rest of the machine. Without NVLink it is also how two GPUs talk to each other, at roughly a tenth of the speed and routed through the CPU. The interconnect is the hinge
Pre-norm
Normalising the input to each sub-layer rather than its output. It makes deep transformers much easier to train, and every model in this series uses it. Inside one transformer block
Preference tuning
Training on comparisons rather than single right answers, so a model can be taught that one fluent answer is better than another. It is the stage that installs tone, helpfulness, verbosity control and refusal behaviour. Preference tuning: RLHF, DPO and verifiable rewards
Pretraining
The first and by far most expensive training stage, where a model learns language and general capability by predicting the next token over trillions of tokens of text. Fine-tuning starts from its result. Where fine-tuning sits in the pipeline
QLoRA
LoRA with the frozen model stored in 4 bits instead of 16. It cuts the last remaining cost of simply holding the base weights by about four times, at the price of roughly 40 percent less throughput. Dettmers et al., QLoRA
Quantisation
Storing numbers with fewer bits by mapping them onto a small set of allowed values. It saves memory and gives up some accuracy in return. bitsandbytes documentation
Queries, keys and values
The three sets of numbers attention works from. Each position makes a query saying what it is looking for, a key advertising what it offers, and a value carrying what it passes on when its key is matched. Multi-head attention explained
RMSNorm
A cheaper version of layer norm that rescales values by their root mean square without first subtracting the average. Standard in current LLMs. Inside one transformer block
RoPE
Rotary position embeddings, the standard way modern LLMs encode where each token sits in the sequence, by rotating parts of the query and key numbers by an angle that depends on position. Multi-head attention explained
SDPA
PyTorch’s built-in scaled dot-product attention, which picks an efficient backend for you including a flash-attention style one. Using it saves installing a separate attention package.
Sequence length
How many tokens are in one training example after tokenisation. Activation memory grows in step with it, and the attention part grows with its square. Activation memory and gradient checkpointing
Sharding
Splitting the training state so each GPU holds only a slice of it and fetches the rest when it needs it. It is the fix for a model that will not fit, and it costs traffic between the cards. Rajbhandari et al., ZeRO
Special token
A token that stands for structure rather than ordinary text, such as the marker that opens or closes an assistant turn. They come from the tokenizer and are placed by the chat template. From messages to tensors
Static state
Weights, gradients and optimizer state together: the memory a run holds from start to finish, 16 bytes per parameter under standard mixed-precision AdamW. Activations sit on top, so a static-state figure is never a total. Training memory and the 16 bytes per parameter
Structural pruning
Permanently removing whole pieces of a trained model, such as layers or attention heads, to make it smaller. It is normally followed by more training to recover the quality that was lost. Llama-3.2-1B model card
Supervised fine-tuning
Training a pretrained model on curated prompt-and-response examples, using the same next-token objective as pretraining, with a mask so only the response is graded. Usually shortened to SFT. Supervised fine-tuning end to end
SwiGLU
The gated feed-forward design used in most current LLMs, where one branch of the widened layer acts as a gate on the other. It is the reason a model’s feed-forward width is quoted as a single number such as 8,960. The transformer feed-forward network
Synthetic data
Training examples written by a model rather than by a person. Cheap and endlessly scalable, and it needs curation, a mix of real data alongside it, and a check on the licence of whatever model produced it. Wang et al., Self-Instruct
Tied embeddings
Reusing one weight matrix both to turn tokens into numbers at the input and to score tokens at the output. It saves a large slice of parameters on a small model. Qwen2.5-1.5B model card
Token
The unit a language model actually reads and writes: a short piece of text, often a word or part of a word. Every length and cost in training is counted in tokens. The complete inference path
Tokenizer
The component that splits text into tokens and maps them to integer IDs, and back again. Every model has its own, and it has to match the model you are training.
Transformer
The architecture behind every model in this series: a stack of blocks that alternate attention with a feed-forward network. Vaswani et al., Attention Is All You Need
Transformer block
One repeated unit of the model: attention, then a feed-forward network, with normalisation and residual connections around them. A 28-layer model is 28 of these stacked up. Inside one transformer block
Vocabulary
The complete set of tokens a model knows, typically somewhere between 32,000 and 151,000 of them. A larger vocabulary makes text shorter in tokens and spends more of a small model’s parameters on the embedding.
VRAM
The memory on the GPU itself. Everything a training step touches has to fit inside it, which is what most of the arithmetic in this series is about. Training memory and the 16 bytes per parameter
Weight
A single learned number inside the model, used to multiply an input on its way through a layer. Weights are what fine-tuning changes, and the only thing it changes. What actually changes inside the model


Practical exercises

Screen a 3B model against the 32 GB card

A 3 billion parameter model is proposed as the next step up from this part’s recommended base. Using the sixteen-bytes-per-parameter rule this part uses to screen candidates, compute its full fine tune static state in gigabytes and decide whether it fits on the one 32 GB card this series targets, showing the arithmetic rather than a rule of thumb.

See the worked solution (opens in a new tab)

Compute the embedding table’s share of the parameter budget

Qwen2.5-1.5B has a vocabulary of roughly 151,000 and a hidden size of 1,536, with tied embeddings so the input and output projections share one matrix. Compute that matrix’s parameter count and its share of the model’s 1.54 billion total, then compare it against the roughly 100 million parameters, about six percent share, this part gives for SmolLM2-1.7B’s smaller vocabulary. State which model spends the larger fraction of its budget on the embedding table, and by roughly what multiple.

See the worked solution (opens in a new tab)

Diagnose a no kernel image error on new hardware

A fresh install on a very new consumer card throws no kernel image is available for execution on the device the first time a flash-attention-style kernel needs to compile. Using this part’s toolchain section, name the specific missing piece most likely responsible, and the two concrete defenses this part recommends for a first run.

See the worked solution (opens in a new tab)

Pick a cloud instance for a full fine tune

You need to full fine tune Qwen2.5-1.5B on a rented instance rather than local hardware. Using this part’s own static state and activation figures for that model, and the VRAM column of its cloud GPU table, compute whether a g6.xlarge, one L4, 24 GB, can hold the run, then do the same for a g6e.xlarge, one L40S, 48 GB, and state which instance you would actually provision.

See the worked solution (opens in a new tab)

Resolve a three-way tradeoff for a model-risk review

A model-risk review needs a base model with fully documented training data provenance. The same project also needs strong multilingual coverage and the deepest possible tooling ecosystem, since the team is small. Using this part’s three-way comparison, explain why no single candidate satisfies all three requirements at once, then justify which requirement you would refuse to compromise on and why.

See the worked solution (opens in a new tab)

Frequently asked questions

Which base model should I use to learn fine-tuning?

A Base checkpoint of 1 to 2 billion parameters, with a conventional decoder shape and a permissive licence. Qwen2.5-1.5B base is a strong default. It fully fine-tunes on one 32 GB card at about 24.6 GB of static state, it is Apache-2.0 at that size, and every tool in the ecosystem already supports it.

Should I fine-tune a Base or an Instruct model?

For learning, take Base. You impose the chat format yourself and watch instruction following appear, which makes cause and effect visible. For shipping, an Instruct checkpoint is often the better start, since it already follows instructions and you are only adjusting behaviour on top of that.

How much VRAM do I need to fully fine-tune a 1.5B model?

About 24.6 GB of static state under standard mixed-precision AdamW, plus activations that land near 8 GB at a batch of 4 and a sequence of 1,024. So a 32 GB card fits it only with the memory tricks on: bf16 compute, gradient checkpointing, and a small per-device batch made up with gradient accumulation.

What does “no kernel image is available for execution on the device” mean?

It means the library you installed contains no GPU code compiled for your chip. Every NVIDIA chip has a compute capability number, and compiled GPU code is built ahead of time for a specific list of them. Your card’s number is not on the list in that build. Editing your script cannot fix it. Install a build that covers your card.

Does the licence of the base model matter?

For private learning, barely. For anything a client or a review board will see, a lot. Apache-2.0 asks you to keep a notice and nothing more. A community licence adds an acceptable-use policy, a naming requirement and a large-deployer clause, and your dataset carries its own separate licence on top.

Why does vocabulary size affect my fine-tune?

Because the lookup table that turns token IDs into numbers has one row per vocabulary entry, and every number in it is a parameter. A large vocabulary therefore spends more of a small model’s fixed budget on that table and less on the layers. A small vocabulary reverses the split, at the cost of splitting non-English text and code into more tokens, which eats sequence length.

Sources and further reading

Previous