A language model has never seen your text. It receives a list of integers and nothing else. Byte-pair encoding is the algorithm that produces that list, and it finishes its work before a single model weight is touched. It is also responsible for a surprising share of the failures people file under “the model cannot reason”.
This is Part 2 of Inside the Inference Stack. Part 1 followed a prompt through the whole system, from characters typed in a box to a token streamed back. This part opens the first box on that path and shows what is inside it.
By the end you will have run the training algorithm by hand on a four-word corpus, small enough that you can check every number yourself. You will be able to say exactly why a model miscounts the letters in “strawberry”, why the same sentence costs three times more in Urdu than in English, and why nobody swaps a tokenizer after training. There is a short snippet near the end for inspecting your own strings.
- Part 1. How an LLM Answers a Question: The Complete Inference Path
- Part 2. Byte-Pair Encoding Explained: How LLMs Turn Text Into Tokens (you are here)
- Part 3. Inside One Transformer Block: The Residual Stream and Its Seven Matrices
- Part 4. Multi-Head Attention Explained: Why 64 Heads Instead of One
- Part 5. The Output Projection: How 64 Attention Heads Become One Thought
- Part 6. The Feed-Forward Network: Where a Transformer Keeps What It Knows
- Part 7. Mixture of Experts Explained: Conditional Computation From Zero
- Part 8. Attention Is All You Need, Dissected: The 2017 Figure, Box by Box
- Part 9. The Roofline Model: Why LLM Decode Is Memory-Bound
- Part 10. Continuous Batching and PagedAttention: How vLLM Keeps a GPU Busy
- Part 11. Benchmark Your Own LLM Serving Stack: Two Measurements, One Afternoon
What a tokenizer is, and the one choice inside it
A tokenizer is two data files and one procedure. The vocabulary is a fixed list of byte strings, each owning exactly one integer ID. The merge list is an ordered set of rules that decides how an arbitrary input string gets cut into entries from that vocabulary. Both files are produced once, before pretraining begins, and neither one changes for the life of the model.
Definition
Token, and token ID
A token is one entry in the vocabulary: a specific sequence of bytes such as the, " the" with its leading space attached, est, or a single fragment of an emoji. Its token ID is just its position in that list, an integer with no numeric meaning.
Origin: splitting text into tokens is older than neural networks and comes from lexical analysis in compilers and from corpus linguistics. The integer ID is what neural language modelling added in the early 2000s: a network multiplies matrices and cannot read text, so every vocabulary entry was given an index into a learned table of feature vectors, and all of a token’s meaning was pushed into that table.
Why it matters: the ID is also a row index. Token 8,921 selects row 8,921 of the embedding matrix, which is the only place a token acquires meaning. The integer itself carries none.
The design question is what a single vocabulary entry should be. Three answers are available: a character, a whole word, or a fragment in between. The first two are the ones people reach for first, and each fails in a way that surfaces as an infrastructure bill rather than as an error message.
Characters give you a tiny vocabulary: about a hundred symbols for English, or 256 if you work in raw bytes. Nothing is ever unrepresentable. The cost lands on sequence length instead. Attention work grows with the square of the sequence, so spending seven tokens on “unhappy” instead of one costs 49 times the attention arithmetic across that span. The KV cache, the per-token memory a model keeps for every position it has already processed, grows linearly with the same sequence, and on a real GPU that cache is what runs out first. Longer sequences also push you deeper into the regime where decode is memory-bound, so you pay again on every token generated.
Words give you short sequences and the mirror-image problem. The vocabulary is unbounded, because new product names, typos, URLs, function names and inflections arrive forever. Anything outside the fixed list becomes an out-of-vocabulary token, usually a single <unk> placeholder, which destroys the information before the model gets a chance to use it. Word units also break the relationship between related strings: “happy” and “unhappy” receive unrelated integers and therefore unrelated embedding rows, sharing nothing at all.
Subword units are the compromise that wins. A subword is a frequent fragment of text, which may be a whole common word, a prefix, a suffix, or a single byte for anything rare. Two properties follow. Vocabulary size becomes a dial you set rather than a fact about the language, and fragments are shared across words that look alike, so morphology survives into the integer representation.
Definition
Byte-pair encoding (BPE)
A procedure that builds a subword vocabulary from raw text by repeatedly finding the most frequent adjacent pair of symbols in a corpus and fusing that pair into a single new symbol.
Origin: published in 1994 as a data compression scheme, where fusing the most frequent byte pair into one unused byte value made the file smaller. Sennrich and colleagues brought it into language modelling in 2015 to fix the rare word problem in machine translation: a fixed word list turned everything it had never seen into a single unknown token, and subword pieces meant any word could at least be spelled out.
Why it matters: there is no linguistics inside it. Every token in a modern model exists because two symbols happened to co-occur often in a training corpus.
Vocabulary size is not a free parameter, because the vocabulary appears twice inside the model as a matrix of shape vocab by d_model. Once as the input embedding table, which turns each token ID into a vector, and once as the output projection, which turns the final hidden state back into one score per vocabulary entry. Many models tie those two matrices to share the weights. At our running d_model of 8,192, the arithmetic settles the argument on its own.
| Unit | Vocabulary | Tokens for a 1,000-word English page | Unseen input | Embedding matrix at d_model 8,192 |
|---|---|---|---|---|
| Bytes or characters | 256 | about 5,700 | always representable | 2.1M parameters |
| Words | unbounded, so you pick a cutoff (say 500,000) | about 1,200 | unknown token, information lost | 4.1B parameters |
| Subword (BPE) | a dial, typically 32,000 to 256,000 | about 1,300 | always representable | 1.05B parameters at 128,256 |
Read the last column first. A word-level vocabulary would spend more parameters on the embedding table alone than a 7B model spends on everything it has. A character-level vocabulary is almost free in parameters and ruinous in sequence length, which is the resource that actually constrains serving. Subword sits between the two: 128,256 by 8,192 comes to just over a billion parameters, a real and knowingly accepted cost that buys short sequences and universal coverage at the same time.
Section takeaways
- A tokenizer is a fixed vocabulary plus an ordered merge list, both frozen before pretraining starts.
- A token ID is a row index into the embedding matrix. It has no numeric meaning of its own.
- Characters fail on sequence length, which costs attention work and KV cache. Words fail on unbounded vocabulary, which costs embedding parameters and loses unseen input.
- Vocabulary size is a cost dial. At
d_model8,192 a 128,256-entry vocabulary costs 1.05B parameters, and a 500,000-entry one would cost 4.1B.
Training a tokenizer: a four-word corpus, merge by merge
Training a tokenizer has nothing in common with training a model. There is no gradient, no loss function and no randomness. It is a counting loop over a corpus that runs a fixed number of rounds and writes out a file. The whole algorithm fits on one line: count every adjacent pair of symbols, fuse the most frequent one into a new symbol, repeat.
Start: every symbol on its own
Here is the entire training corpus, expressed as word types with frequencies: “low” appears 5 times, “lower” 2 times, “newest” 6 times, “widest” 3 times. That is 16 word occurrences over four distinct words. BPE works on this frequency table rather than on the running text, which is why a corpus of billions of words still trains in minutes. Every word starts split into its base symbols.
Definition
Base vocabulary
The set of atomic symbols that exist before any merge happens, and the only symbols that can never be split further. In this toy example it is the ten letters l o w e r n s t i d. In GPT-2 and its descendants it is the 256 possible byte values.
Origin: the byte-level base vocabulary was shipped with GPT-2 in 2019. Tokenizers before it built their base from characters, which for Unicode means either an enormous starting alphabet or a shortlist that will eventually meet a character it does not hold and emit an unknown token. Dropping to the 256 byte values removed that failure mode by construction.
Why it matters: choosing bytes as the base makes the out-of-vocabulary problem disappear by construction. Every possible input is a sequence of bytes, and every byte is already in the vocabulary, so there is always a valid encoding. Emoji, Cyrillic, a mangled paste out of a PDF, a binary blob dropped in the chat box: all of it encodes.
Nothing has been learned at this point. The vocabulary is ten symbols, “low” is three tokens, and “newest” is six. Every improvement from here comes from the counting loop.
Round 1: count every pair, merge the winner
Count every adjacent pair across the corpus, weighting each occurrence by its word frequency. The pair e + s appears inside “newest” (6) and “widest” (3), so it scores nine. The pair s + t scores nine as well, from the same two words for the same reason. The pair l + o and the pair o + w each reach seven, five from “low” and two from “lower”. Counting the rest by hand takes a minute, and that minute is the entire algorithm.
Definition
Merge rule
One entry of the form “symbol A followed by symbol B becomes the single symbol AB”, stamped with the round number in which it was learned. Applying it to a word replaces every adjacent occurrence of that exact pair.
Origin: the rule is the unit of the original 1994 compression algorithm, where each learned pair had to be stored alongside the compressed data or the file could never be expanded again. Subword tokenisers kept the same object for the same reason. The rules are the only record of how a vocabulary was built, so there is no way to encode new text without them.
Why it matters: each merge rule creates exactly one new vocabulary entry, so vocabulary size and merge count move together. GPT-2’s 50,257 entries are 256 base bytes plus 50,000 merges plus one end-of-text marker.
The winning pair gets fused everywhere it occurs. “newest” becomes n e w es t and “widest” becomes w i d es t, dropping both from six symbols to five. The vocabulary is now eleven entries, and es is a token that will exist for the entire life of any model trained on this tokenizer.
Analogy
Think of a copy editor with no knowledge of English, working through a stack of pages with a tally sheet. They notice that the shape “es” keeps appearing, so they invent a shorthand squiggle for it and rewrite every occurrence. Next pass, the squiggle itself is now a shape that keeps appearing next to “t”, so they invent a shorthand for that pair too.
Where it breaks: a real editor would recognise “est” as an English suffix and generalise. The algorithm has no such concept. It would fuse an arbitrary pair of hex digits just as happily, if the counts came out that way.
That indifference is the point worth holding onto. Nobody told the algorithm that es is a meaningful English fragment. It won on frequency alone, and it would have won just as decisively on a corpus of chemical formulae or Base64 blobs if the counts had come out the same way. Every property people later attribute to the tokenizer, good or bad, traces back to what was in the corpus.
The tie at nine also has to be broken by a rule, and every implementation has one, usually first-seen order or a lexicographic comparison. The specific rule matters less than the fact that it is deterministic, because whichever pair wins gets baked into an artifact that a model will then train against for months.
Rounds 2 to 4: merges build on merges
This is the property that makes byte-pair encoding more than a compression trick: a merged symbol is itself a candidate for further merging. Round 1 created es. Round 2 can therefore merge es + t into est at count 9. Round 3 takes l + o at count 7. Round 4 takes the new symbol lo plus w at count 7 and produces low.
Four rounds in, the algorithm has bootstrapped from single letters to a complete English word, with no dictionary involved at any point. “low” is now one token where it was three. “newest” is n e w est, four tokens where it was six. That recursion is what lets 50,000 rounds produce tokens as long as ” unfortunately” from a base of individual bytes.
You stop whenever you like, and the stopping point is the whole vocabulary-size decision. GPT-2 ran 50,000 rounds. More merges buys shorter sequences and a fatter embedding table. Fewer merges does the reverse. There is no correct answer here, only a tradeoff picked for a particular deployment, which is why a code-heavy model and a multilingual model land on different numbers.
The merge list is the trained artifact
Training produced two files, and the second one is the interesting half. A vocabulary lists what exists. A merge list records in what order things came to exist, and encoding depends entirely on that order.
Definition
Merge list
The ordered log of every merge rule learned during training, stored with its rank: rule 1 first, rule 50,000 last. Encoding a new string means splitting it into base symbols and replaying that log from the top, applying each rule wherever it fits.
Origin: the ordering is inherited from the compression algorithm, where replaying the rules out of sequence gives back a different string. Subword tokenisers ship the list with the model instead of rebuilding it from the corpus, because a tokenizer that produced different IDs on different machines would send the same word to the wrong embedding rows.
Why it matters: rank is data. Two tokenizers with an identical vocabulary but a different merge order will produce different token IDs for the same string.
Order is also what separates byte-pair encoding from greedy longest-match, which people frequently assume is the same algorithm. It is not, and the two disagree on real inputs.
Definition
Greedy longest-match
An alternative encoding strategy that scans left to right and, at each position, takes the longest string that exists in the vocabulary. It consults only the vocabulary and ignores merge ranks entirely. WordPiece encodes this way.
Origin: longest-match scanning is an old technique from word segmentation for languages written without spaces, where a segmenter walked a dictionary looking for the longest entry that fitted at the current position. It reached subword tokenisation through WordPiece, first described in 2012 in work on Japanese and Korean voice search, and it survived because it needs only a vocabulary. There is no ordered rule list to store, ship or replay.
Take a merge list that learned n + i as rule 5 and u + n as rule 12, and encode the string “unit”. BPE replays in rank order, so rule 5 fires first and gives u ni t. Rule 12 can then never fire, because the n it needs is already sealed inside ni. The result is u, ni, t. Greedy longest-match starts at position 0, finds un in the vocabulary and uni absent from it, and returns un, i, t. Same token count, completely different integers, and the model has only ever been trained on one of the two.
Analogy
A merge list is a replay log, in the same sense as a database write-ahead log or a sequence of git commits. The final state is not the artifact. The ordered operations are, and you reproduce a state by replaying them from the beginning in exactly the recorded order.
Where it breaks: a database log can be rewritten or compacted while preserving the outcome. Reordering a merge list changes which integers come out, which silently invalidates every embedding row in a trained model.
This is why a tokenizer is frozen the moment a model starts pretraining on it. Token 8,921 carries meaning only because the model saw that integer a billion times in a consistent role. Change the merge order and a different string now maps to that integer, so every embedding row is attached to the wrong thing. There is no migration path and no adapter you can bolt on. You retrain.
Encoding a word that was never in the corpus
The training corpus contained “low”, “lower”, “newest” and “widest”. It never contained “lowest”. Encoding it shows the payoff for everything above: split into base symbols, then replay the merge list in rank order.
- start:
l o w e s t - rule 1 (
e+s):l o w es t - rule 2 (
es+t):l o w est - rule 3 (
l+o):lo w est - rule 4 (
lo+w):low est
The unseen word resolves into two tokens, both learned from other words, both carrying meaning the model has practised with millions of times. Nothing was unrepresentable, no unknown-token fallback was involved, and the model can generalise from “low” and from the superlative marker “est” to a word it never encountered. That generalisation is the entire reason the subword compromise won.
Production implementations add two things to this picture. They do not rescan the whole text once per rule, because that would be 50,000 passes; they keep a priority queue of candidate merges keyed by rank, which yields an identical answer far faster. And before any merging happens, they run a regular expression over the raw text to cut it into pre-tokens.
Definition
Pre-tokenizer
A regular expression applied before BPE that splits raw text into chunks, roughly at word boundaries, and attaches each leading space to the word that follows it. Merges are then only allowed inside a chunk, never across a boundary.
Origin: the 2015 subword work ran BPE inside whitespace-separated words, so word boundaries were respected without anyone having to ask for it. Byte-level BPE gave that up, and GPT-2 in 2019 added an explicit regular expression to stop merges running across character categories. Without it the learner spent vocabulary slots on near-duplicates of the same common word, one for each punctuation mark it happened to sit beside.
Why it matters: this is the single design choice behind most of the strange behaviour in the next section. It is why " the" and "the" are different tokens, and why digit handling can be capped by editing one regex.
Section takeaways
- Tokenizer training is a counting loop over a word-frequency table, with no gradient and no randomness anywhere in it.
- Round 1 of the toy corpus fuses
e+sat count nine, six occurrences from “newest” and three from “widest”. - Merged symbols are themselves merge candidates, which is how four rounds climb from single letters to the whole word “low”.
- The trained artifact is a rank-ordered merge list. Encoding replays it in order, which is a different algorithm from greedy longest-match and returns different integers.
- Rank order is why tokenizers are frozen before pretraining. Reordering merges detaches every embedding row from what it learned.
- An unseen word such as “lowest” encodes cleanly into pieces learned from other words, with no unknown token involved.
Where byte-pair encoding leaks into model behaviour
The four behaviours below get filed as reasoning failures and are nothing of the kind. Each one is a direct consequence of a boundary decision made during tokenizer training, before the model saw a single training example. Recognising them saves you from trying to fix them with prompting.
Arithmetic
Numbers are chunked by frequency like everything else, so “380” may be one token while “381” is two, 38 followed by 1. There is no positional digit system anywhere in the token IDs, which means the model cannot line up columns the way you were taught to. It has to learn addition over a chunking of numbers that is inconsistent by construction. Several newer tokenizers cap digit runs at three characters in the pre-tokenizer regex, and some force every digit onto its own token. Both help a great deal, and neither makes arithmetic free.
Spelling and counting characters
“strawberry” comes apart into roughly str, aw, berry. The model never receives letters. It receives three opaque integers with learned vectors attached, and the internal structure of each chunk is not visible to any layer. Asking how many r’s the word contains asks the model to report on structure it has never directly observed. It often answers correctly anyway, because a great deal of internet text spells words out letter by letter and large models absorb those descriptions, though that is recall rather than reading.
Analogy
Imagine reading a language printed on tiles, where each tile shows one whole word-fragment as a single glyph. You can read fluently, because you know what each tile means and which tiles follow which. Now someone asks you how many times a particular letter appears on a tile. You never saw letters; you saw a tile.
Where it breaks: a model can partially recover spelling from text that describes spelling, so the failure is soft and gets softer with scale. The same root cause also produces failures at reversing strings, finding words that fit a letter pattern, and rhyming reliably.
Whitespace and the leading space
Whitespace is text, so it costs tokens like text does. GPT-2 turns four spaces into four separate tokens, which makes one level of Python indentation cost four times what it should. Later tokenizers added merges for runs of spaces, and that single change is a large part of why code became cheaper to process. It is worth checking before you build a product that ships a lot of indented text.
The leading space matters more than people expect, and it follows directly from the pre-tokenizer definition above. Because a space is attached to the word after it, " the" and "the" are two different tokens with two different IDs and two unrelated embedding rows. The model spends nearly all of its training seeing the space-attached variant in mid-sentence positions.
In practice
Do not end a prompt with a trailing space. The model expects the next token to arrive carrying its own leading space, and a trailing space takes that space away and hands the model a continuation shape it saw rarely during training. The effect is usually small and occasionally is not, and it costs nothing to avoid.
Fertility outside English
The merge list was trained on a corpus, and for most widely deployed tokenizers that corpus was overwhelmingly English. English words therefore earned long, efficient tokens. Other scripts did not. Characters in Devanagari, Arabic or Thai also occupy multiple bytes each in UTF-8, so thin merge coverage over those byte sequences means paying several tokens per single character.
Definition
Fertility
The average number of tokens a tokenizer produces per unit of text, usually quoted per word or per 1,000 characters. Low fertility means efficient packing. High fertility means the same meaning consumes more tokens.
Origin: the word is borrowed from statistical machine translation of the early 1990s, where a source word’s fertility was the number of target words it produced. It was picked up as a tokenizer metric once multilingual models needed one number to compare vocabularies, after it became obvious that the same sentence could cost several times more tokens in one language than in another.
Why it matters: fertility is an infrastructure number rather than a linguistic curiosity. It multiplies your bill, divides your usable context window, inflates your KV cache and, because attention is quadratic, raises prefill compute faster than linearly.
Section takeaways
- Arithmetic errors come from frequency-based digit chunking, not from a missing capacity to add.
- Letter-level tasks fail because tokens are opaque integers. The model never observes the characters inside a chunk.
- Whitespace costs tokens. Old tokenizers charge four tokens for one Python indent, newer ones merge space runs.
- The pre-tokenizer attaches leading spaces to words, which makes a trailing space in a prompt a genuine, if minor, distribution shift.
- Fertility outside English is a property of the merge list, and it lands on cost, context, memory and compute at the same time.
What fertility does to your bill and your context window
Fertility is the number you actually budget against, so it helps to carry rough ratios in your head before you size anything. The figures below are the typical spread for a modern English-trained tokenizer such as cl100k_base.
| Content | Typical tokens per 1,000 characters | Roughly, versus English prose | Why |
|---|---|---|---|
| English prose | about 250 | 1.0x baseline | common words are whole tokens, around 4 characters each |
| European Latin-script prose (German, French) | about 280 to 330 | 1.1x to 1.3x | fewer whole words in the merge list, compounds split |
| Python or JavaScript source | about 300 to 400 | 1.2x to 1.6x | indentation, punctuation, identifiers split at case boundaries |
| Pretty-printed JSON | about 350 to 500 | 1.4x to 2.0x | braces, quotes and colons each cost, and every repeated key pays again |
| Non-Latin script (Hindi, Urdu, Thai) | about 500 to 1,000 or more | 2x to 4x | multi-byte characters with thin merge coverage |
Treat those as approximate ratios rather than measurements. They move with the tokenizer version, the domain and the specific text in front of you. The shape of the spread is the durable part, and your own corpus is the only source for a number you plan to put in a spreadsheet.
Analogy
Fertility works like a courier that bills per box rather than per kilogram. English prose ships in large boxes, so a page takes few of them. High-fertility text ships the same contents in small boxes, so the identical meaning arrives as four times the number of parcels and four times the charge.
Where it breaks: boxes are only a billing unit, whereas tokens are also real compute and real memory. Every extra token is another position in the KV cache and another row of attention arithmetic on every subsequent step.
Three consequences follow directly, and the first is the bill. Providers charge per token in both directions, so a 3x fertility ratio is a 3x input bill for the same sentence. Prompt caching removes a chunk of that for repeated prefixes and routing cheap requests to smaller models removes another, though neither one changes the underlying ratio.
The second is the context window, which is measured in tokens and therefore holds a variable amount of meaning. A 32,000-token window holds roughly 24,000 words of English and perhaps 8,000 words of a high-fertility language. If you are chunking documents for retrieval, a chunk size expressed in words is not portable across languages, and a config tuned on English will silently truncate elsewhere.
The third is GPU memory, which is where the tokenizer quietly makes a hardware decision for you. Take the running Llama-3-70B class model: 80 layers, 8 KV heads, head_dim 128, bf16. Per token that is 2 (for K and V) times 8 times 128 times 2 bytes times 80 layers, which comes to 327,680 bytes, about 320 KB. A 1,000-token conversation therefore holds roughly 320 MB of cache. Serve the same conversation in a language with 3x fertility and you are holding close to 1 GB for identical meaning. Multiply that by concurrent users and the merge list has just set your GPU count, which is a number worth having before you move a demo into production.
Section takeaways
- English prose runs about 250 tokens per 1,000 characters. JSON runs 1.4x to 2x that, and non-Latin scripts 2x to 4x.
- Fertility multiplies the bill directly, because providers charge per token in both directions.
- A context window is a token budget, so its capacity in words changes with the language and the format.
- At 320 KB of KV cache per token, a 3x fertility ratio turns a 320 MB conversation into nearly 1 GB of GPU memory.
- Published ratios are a starting point. Measure your own prompts, since one verbose template often dominates the total.
The neighbours: WordPiece, Unigram, SentencePiece and tiktoken
BPE is the most common approach and not the only one. Four names come up constantly in tokenizer configs, and knowing which is an algorithm and which is a library saves a lot of confusion when reading model cards.
| Approach | What it is | How it differs from BPE |
|---|---|---|
| WordPiece | the tokenizer behind BERT and its family | picks the merge that most improves training-data likelihood rather than the most frequent pair, and encodes by greedy longest-match with a continuation marker rather than by replaying merges |
| Unigram | a probabilistic subword model, common in T5 and many multilingual models | starts from a large candidate vocabulary and prunes it down, then scores several possible segmentations of a word at encode time and takes the most probable |
| SentencePiece | a library rather than an algorithm, running either BPE or Unigram underneath | consumes the raw string including spaces and marks word boundaries with a visible character, which is what lets it work on languages that do not use spaces |
| tiktoken | OpenAI’s tokenizer implementation | the same BPE algorithm engineered for speed in Rust, shipping fixed named encodings such as r50k_base, cl100k_base and o200k_base |
The practical distinction hides in the last two rows. WordPiece and Unigram are genuinely different training and encoding algorithms, so their outputs differ from BPE on the same text. SentencePiece and tiktoken are implementations, so a config naming either of them still leaves the algorithm question open, and you have to read further to find out which one is running underneath.
Section takeaways
- WordPiece maximises likelihood rather than raw frequency, and encodes by greedy longest-match.
- Unigram prunes a large vocabulary down and picks the most probable segmentation at encode time.
- SentencePiece and tiktoken are implementations, not algorithms. Either can be running BPE underneath.
- A model card that names only the library has told you nothing about the segmentation behaviour you will see.
Inspect your own strings
Everything above is checkable in about a minute, and the checking is worth more than the reading. Install tiktoken and look at what your actual production text becomes.
import tiktoken
enc = tiktoken.get_encoding("cl100k_base") # GPT-4 class encoding
text = "What is a quadratic equation?"
ids = enc.encode(text)
print(len(ids), "tokens")
for i in ids:
print(i, enc.decode_single_token_bytes(i))
The running prompt from Part 1 comes to seven tokens. Two details in the output are worth pausing on. The tokens print as byte strings rather than as text, because byte strings are what they are underneath. And the leading spaces appear attached to the words that follow them, which is the pre-tokenizer behaviour showing up in your own terminal instead of in a diagram.
In practice
Run three things through that snippet: a real production prompt, a paragraph of your product’s non-English content, and one of your JSON payloads. The fertility table stops being abstract in about thirty seconds, and you will usually find one field or one template quietly consuming a third of your input budget.
From here, Part 3 takes those integers into the first transformer block and follows what happens to them once they become vectors.
Section takeaways
- Measured fertility beats quoted fertility. The tokenizer is a twenty-line script away from being observable.
- Token output prints as byte strings because tokens are byte strings, and leading spaces belong to the following word.
- Templates and JSON keys are the usual hidden cost, since every repeated key pays its tokens on every request.
Key takeaways
- Characters and words both fail, in opposite directions. Subword units turn vocabulary size into a dial you set rather than a property of the language.
- Training is one counting loop: count every adjacent pair weighted by frequency, fuse the winner, repeat. In the toy corpus
e+swins round 1 with nine occurrences, six from “newest” and three from “widest”. - Merged tokens are candidates for further merging, which is how four rounds get from single letters to the whole word “low”.
- The ordered merge list is the trained artifact. Encoding replays it in rank order, which is a different algorithm from greedy longest-match and gives different answers.
- Starting from the 256 byte values removes the out-of-vocabulary problem completely. Any input encodes, always.
- Arithmetic errors, letter-counting errors, whitespace waste and leading-space sensitivity are tokenization artifacts rather than reasoning failures.
- Fertility is an infrastructure number. Non-Latin scripts can cost 2x to 4x the tokens for the same meaning, which lands on your bill, your context window and your KV cache at once.
Frequently asked questions
Why do LLMs get “how many r’s are in strawberry” wrong?
The model never sees letters. Byte-pair encoding splits the word into a few opaque chunks such as str, aw and berry, and each chunk arrives as a single integer with no internal structure. Any answer about spelling is recalled from text that described the spelling rather than read off the word itself.
What is the difference between byte-pair encoding and tiktoken?
Byte-pair encoding is the algorithm and tiktoken is one fast implementation of it, written in Rust with Python bindings. tiktoken also ships specific pre-trained encodings such as cl100k_base and o200k_base, which are merge lists rather than algorithms.
What exactly is a merge list?
It is the ordered log of every merge rule the tokenizer learned during training, stored with its rank from first to last. Encoding a string means splitting it into base symbols and replaying that log in order, so two tokenizers with the same vocabulary but a different merge order produce different token IDs.
Why do non-English prompts cost more tokens?
The merge list was trained on a corpus that was mostly English, so English words became long efficient tokens and other scripts did not. Characters in scripts like Devanagari or Arabic also take multiple bytes each in UTF-8, so the same meaning can take two to four times as many tokens.
Can you change a model’s tokenizer after training?
Not without retraining. Every token ID carries meaning only because the model saw that integer in a consistent role billions of times, so changing the merge list detaches every embedding row from what it learned. Tokenizers are frozen before pretraining starts for exactly this reason.
How many characters are in a token?
For English prose the usual rule of thumb is about 4 characters per token, or around 1.3 tokens per word. That ratio gets worse for code and JSON, and considerably worse for non-Latin scripts, so measure your own text rather than relying on the average.
Does a trailing space at the end of a prompt matter?
It can. Most tokenizers attach a leading space to the word that follows it, so a word with its space and the same word without are different tokens with different embeddings. Ending a prompt with a space leaves the model predicting a continuation shape it saw rarely during training.
Sources and further reading
- Neural Machine Translation of Rare Words with Subword Units, Sennrich et al., the paper that brought BPE into NLP.
- Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates, Kudo, which introduces the Unigram language model tokenizer.
- SentencePiece: A simple and language independent subword tokenizer and detokenizer, Kudo and Richardson.
- Google’s Neural Machine Translation System, which describes WordPiece as deployed.
- openai/tiktoken, the reference fast BPE implementation and its named encodings.
- Hugging Face tokenizers documentation, for training and inspecting your own merge lists.
Check your understanding
Take the 15 question quiz on this article
The same 15 questions every time, picked to cover every idea this article teaches, out of 30 written for it. The order of the questions and the order of the answer choices reshuffle on every attempt, so a retake never looks identical to the last one. It runs here on the page and keeps your place, so you can go and read a section and come back without losing anything. You get a full report at the end: your score, the correct answer to anything you missed, why it is correct, and a link straight back to the section it came from.
- 15 questions
- same set every time
- 8 knowledge areas
- hint on every question
- timed, no limit
