In short
When you send a message to a language model, it does not receive text. It receives a sequence of integers — identifiers. Each integer corresponds to a token: a unit of text carved out by an algorithm before the model ever starts processing anything.
A token is not a word. It is not a character either. It is something in between: a word fragment, a whole word, sometimes a punctuation mark or a space. Everything the model knows about human language was learned through this lens. Tokenization is therefore the first layer — invisible but fundamental — between the text and the model.
Explanation
What is a token, concretely?
Take the word “tokenization”. Depending on the tokenizer used, this word can be split into two or three fragments: token, ization — or token, iza, tion. Each fragment is a token. An integer identifier is assigned to it in the model’s vocabulary.
In English, the rule of thumb often cited is: 1 word ≈ 1.5 tokens. In French, this ratio is slightly less favorable. For languages such as Japanese, Arabic or Ukrainian, it can reach 3 to 5 tokens per word — with direct consequences on costs and processing limits (see next section).
The model never receives the raw text. It only receives the sequence of integer identifiers corresponding to the tokens. Decoding into readable text only happens at the output.
The BPE algorithm — how the vocabulary is built
The dominant technique in large language models is Byte Pair Encoding (BPE), or one of its variants. The idea is simple: build a subword vocabulary from statistical frequencies on a training corpus.
The algorithm proceeds in four steps:
- Initialization: the starting vocabulary contains all the individual characters from the corpus (or all bytes, in the byte-level BPE variant).
- Counting: the most frequent pair of adjacent tokens in the corpus is identified.
- Merging: this pair is merged into a new token, added to the vocabulary.
- Repetition: steps 2 and 3 are repeated until the target vocabulary size is reached (for example 100,000 tokens).
The result is a vocabulary that encodes the most common fragments of the language. Frequent words (“the”, “les”, “de”) often become single tokens. Rare words are broken down into recognizable sub-parts.
The advantage over whole-word tokenization: unknown words are no longer blocking. Instead of being replaced by an [UNK] (unknown) marker, they are decomposed into known parts. The advantage over character-by-character tokenization: sequences are much shorter, which mechanically reduces the size of contexts to process.
In short: BPE is a compromise between two extremes. Whole words = few tokens but lots of unknowns. Characters = no unknowns but sequences too long. BPE takes the middle: a vocabulary of “word pieces” that snap together like Lego to reconstruct any word, even a never-seen neologism.
Not all tokenizers are alike
Models do not all use the same tokenizer, and this choice has practical implications:
| Model | Tokenizer | Vocabulary size |
|---|---|---|
| GPT-2 | Byte-level BPE (tiktoken) | 50,257 |
| GPT-4 | Byte-level BPE (tiktoken, cl100k_base) | ~100,000 |
| Claude | BPE | ~65,000 |
| Llama 2 | SentencePiece BPE | 32,000 |
| Llama 3 | Extended BPE | ~128,000 |
The difference between tiktoken (used by OpenAI) and SentencePiece (used by Meta for Llama) lies in the initial level of splitting: tiktoken first encodes text in UTF-8 bytes then merges bytes; SentencePiece works directly at the Unicode code point level. The visible result: for a given text, the number of tokens produced may differ from one model to another.
A larger vocabulary tends to produce shorter sequences (fewer tokens per text), but requires a bigger embedding and a longer training. A smaller vocabulary is more compact but splits more finely — and therefore less efficiently for rare words.
In short: sending the same text to GPT-4 and to Llama 2 does not consume the same number of tokens. The vocabularies differ. The rule: to estimate cost or length, use the actual model’s tokenizer, not an “average” tokenizer. OpenAI provides a public counter (tiktoken), HuggingFace exposes the Llama and Mistral tokenizers.
Concrete examples
Cost is computed in tokens, not words
Language model APIs charge by consumption, per token — on input (the prompt) and on output (the response). A 1,000-word English text represents about 1,300 to 1,500 tokens. The same volume of text in French will be slightly longer in tokens; in Arabic or Japanese, it can reach 3,000 to 5,000 tokens.
Knowing the word/token ratio of the model you are using is therefore a directly economic skill. The GPT-4 tokenizer (cl100k_base) is publicly accessible and testable through online tools — which makes it possible to precisely estimate the cost of an API call before running it.
Non-Latin languages are penalized on two levels
The penalty for languages under-represented in training data is twofold. First, tokens are more numerous per word: a Ukrainian word written in 6 characters may consume 4 tokens where its English equivalent consumes only one. Second, the model has been exposed to fewer examples in that language, which also degrades the quality of understanding and generation.
A study published in Frontiers in Artificial Intelligence (2025) documents this inefficiency for Ukrainian: encoding is significantly less efficient than for English, with a direct impact on the effective length of the context window available for that language.
Numbers pose a specific problem
Tokenization of integers is not systematic in BPE. The number “1234” can be split into “12” and “34”, or into “1”, “234”, or even into “1234” as a single token — depending on the vocabulary and frequencies seen at training time. This irregularity is not innocuous: research has shown that the tokenization order (left-to-right or right-to-left) influences the model’s arithmetic performance on multi-digit additions. This is not an intentional bug — it is a limitation inherent to the statistical BPE approach on numeric symbols.
An example to anchor the intuition
The sentence “LLMs do not read words.” will be tokenized differently depending on the model. In cl100k_base (GPT-4), one can expect something like: ["LL", "Ms", " do", " not", " read", " words", "."] — i.e. 7 tokens for 5 words. The token ” read” includes the preceding space, which is characteristic of byte-level BPE. In a SentencePiece tokenizer, the split may differ slightly. These differences are invisible to the user but influence the way the model “sees” the text.
In short: remember three orders of magnitude. English: ~1.3 tokens per word. French: ~1.8. Arabic or Japanese: ~3-5. Direct consequences: a 1,000-word French text costs about 1,800 tokens to GPT-4; the same content in Arabic can reach 4,000-5,000 tokens. This inequality does not vanish with a better prompt — it is built into the tokenizer.
What to remember
- An LLM does not process raw text but a sequence of integers corresponding to tokens — word fragments, whole words, or characters.
- The dominant algorithm, BPE, builds the vocabulary by iteratively merging the most frequent pairs in a corpus. It allows handling unknown words without an
[UNK]marker. - Tokenizers differ between models: vocabulary, size, implementation. A same text does not produce the same number of tokens from one model to another.
- Non-Latin languages consume more tokens than English — sometimes 3 to 5 times more per word — which increases costs and reduces the effective context space.
- API billing is per token, not per word. Understanding this ratio is a practical skill for controlling costs.
- Tokenization of numbers is irregular in BPE, which can affect performance on arithmetic tasks.