Understanding Tokenization
A language model cannot operate on raw characters or on abstract words; it needs a finite vocabulary of discrete symbols mapped to integer ids, which then index into an embedding table. Tokenization is the step that performs this segmentation, and its design has consequences that reach through the whole system.
Splitting on whitespace into whole words fails in two directions. The vocabulary becomes enormous and still incomplete, since new words, names, typos, and morphological variants appear endlessly; and any word absent from the vocabulary has to be replaced by an unknown-token placeholder, destroying information. Splitting into individual characters avoids that but produces very long sequences with little meaning per token.
Byte pair encoding takes the middle path. Starting from individual characters it repeatedly finds the most frequent adjacent pair of symbols and merges it into a new symbol, continuing until the vocabulary reaches a target size. Frequent words end up as single tokens, while rare words are decomposed into meaningful fragments. Raschka notes that GPT models use exactly this scheme, and that because of it they need no unknown-word token at all: any string can be represented, falling back to smaller pieces when necessary.
This has practical consequences worth internalizing. Token count differs from word count, often substantially, and it is tokens that context windows and pricing are measured in. Tokenizers trained predominantly on English also fragment other languages more aggressively, so the same sentence can cost several times more tokens in one language than another. And because the model sees tokens rather than letters, character-level tasks such as counting letters in a word are genuinely awkward for it.
Example of Tokenization
A common word such as "the" is frequent enough to survive as a single token. A rarer word such as "tokenization" may be split into pieces like "token" and "ization", so the model still sees the recognizable stem rather than an opaque unknown symbol.
The merge procedure is mechanical. Given a corpus of characters, the pair that co-occurs most often, perhaps "t" followed by "h", is merged into "th". The counts are recomputed and the next most frequent pair merged, and so on for a fixed number of merges. The learned merge list is the tokenizer.
Because the fallback is to ever-smaller units, and ultimately to raw bytes, no input is unrepresentable. An unseen proper noun or a novel technical term is simply encoded as several subword pieces rather than triggering an out-of-vocabulary failure.
Frequently Asked Questions
Why not just tokenize into words?
The vocabulary would be huge and still incomplete, since language continually produces new words, and every unseen word would collapse to an unknown-token placeholder. Subword tokenization keeps the vocabulary bounded while remaining able to encode anything.
Why do token counts differ so much between languages?
Because the merge rules are learned from a training corpus. A tokenizer built mainly on English learns merges that compress English efficiently, and languages under-represented in that corpus are split into more, smaller pieces, raising their token cost for identical content.
Why do language models struggle to count letters in a word?
They never see the letters. A word may arrive as one or two subword tokens with no explicit character structure, so questions about individual characters ask about information the model’s input representation has largely discarded.
The Bottom Line
Tokenization converts text into the bounded set of subword units a model can consume. Byte pair encoding is the standard method, removing the out-of-vocabulary problem entirely, and its consequences show up in context limits, pricing, cross-language cost, and the model’s blind spot for individual characters.