In short

Tokenization splits text into units that the model can process. BPE is the best-known algorithm — but not the only one. WordPiece, Unigram LM, and SentencePiece follow different logics, with consequences for supported languages, usage costs, and model robustness. The choice of tokenizer is not neutral: in Arabic or Yoruba, the same text can consume ten times more tokens than in English.


BPE as a starting point

If you haven’t read the article on tokens and tokenization, that is the right place to start. It explains BPE step by step, the word-to-token ratio, and the problem with numbers.

In short: BPE (Byte Pair Encoding) iteratively merges the most frequent character pairs in a training corpus. It starts from individual characters and builds up in complexity. GPT-2, GPT-4, and Mistral use it in its original form or under the name tiktoken.

What follows assumes you are familiar with that mechanism. The goal here is to explore what lies beyond.


WordPiece: maximizing likelihood rather than frequency

WordPiece is the algorithm used by BERT, DistilBERT, and ALBERT. It resembles BPE on the surface — but its merging criterion is different.

BPE merges the most frequent pairs. WordPiece merges the pairs that maximize the likelihood of the training corpus. Concretely, it favors merges that make the corpus more probable according to a language model. This produces a slightly different vocabulary, more sensitive to morphological regularities.

A distinguishing sign: non-initial subwords carry the prefix ##. The word playing becomes play + ##ing. This convention makes segmentation readable and helps the model identify word boundaries.

BERT’s vocabulary contains 30,000 tokens. That is compact compared to recent tokenizers — which has implications for representation granularity.

In short: BPE and WordPiece look like two versions of the same cutter. BPE counts the most frequent pairs; WordPiece weighs how much each merge makes the entire corpus more probable. The difference is subtle but visible: WordPiece better captures morphological regularities (prefixes, suffixes) at the cost of slightly heavier training computation.


Unigram LM: the inverse logic of BPE

Kudo (2018) proposes a radically different approach with the Unigram Language Model. Where BPE starts from the bottom (characters) and builds up through merges, Unigram LM starts from the top (a large initial vocabulary) and prunes iteratively.

The algorithm begins with a large vocabulary — often several hundred thousand candidate subwords. It evaluates each token’s contribution to the corpus likelihood. Tokens whose removal causes the smallest drop in likelihood are eliminated first. The process iterates until the target vocabulary size is reached.

The result: multiple segmentations of the same word are possible, with different probabilities. During training, these segmentations can be sampled — which acts as a form of regularization. During inference, the most probable segmentation is selected via the Viterbi algorithm.

This probabilism is a fundamental difference from BPE: segmentation is not deterministic — it reflects a distribution.


SentencePiece: erasing spaces from the problem

BPE and WordPiece assume that text has already been pre-tokenized by spaces. In English, this is natural. In Japanese, Chinese, or Thai, there are no spaces between words. This assumption becomes a structural bias.

SentencePiece (Kudo 2018) solves this problem by treating raw text as a sequence of bytes. Spaces are not delimiters — they become an ordinary character, represented by the symbol ▁. The sentence “hello world” becomes ▁hello + ▁world.

This seemingly minor change has significant consequences:

  • The tokenizer is language-agnostic: Japanese, Arabic, Hindi, and Swahili all pass through the same pipeline without language-specific preprocessing.
  • Segmentation is reproducible: no dependency on language-specific tokenization tools.
  • The same algorithms (BPE or Unigram LM) run underneath — SentencePiece is an abstraction layer, not a standalone algorithm.

T5, LLaMA, Mistral, and Gemma all use SentencePiece. It has become the de facto standard for multilingual models.

In short: SentencePiece is not a new algorithm — it is a wrapper that makes BPE or Unigram language-agnostic. Its key contribution: treating space as an ordinary character, not a delimiter. Practical consequence: the same training pipeline works on Japanese, Chinese, Thai, French — without language-specific preprocessing.


Multilingual fairness: the ten-times problem

Here is a surprising fact: asking a question in Yoruba costs about ten times more than in English, for equivalent content. This is not a pricing policy — it is a direct consequence of tokenization.

Ahia et al. (2023, ACL) measured the fertility of tokenizers on translated texts: the number of tokens produced for the same text across different languages. For English, a word often corresponds to 1–2 tokens. For Yoruba or Amharic, the same word can generate 5 to 10.

Petrov et al. (2023, NeurIPS) extended this measurement to more than 1,000 languages and showed a correlation between high fertility and degraded performance. High-fertility languages are not only more expensive — they are also less well understood.

The mechanisms are several:

  • Proportional usage cost: providers bill per token, not per word.
  • Reduced context window: an Arabic text occupies more space in the context window.
  • Under-learning: models see few examples in these languages during training, and when they do, the representation is fragmented.

A partial response consists in increasing vocabulary size. LLaMA 2 had 32,000 tokens. LLaMA 3 has 128,000. Gemma goes up to 256,000. A larger vocabulary reduces fragmentation for rare languages — but increases the weight of the embedding layer, which must store one vector per token.

In short: “fertility” measures how many tokens an average word consumes in a given language. English: ~1.3 tokens per word. French: ~1.8. Arabic: ~3.5. Yoruba: ~5-10. This inequality is not a technical detail — it translates into bills (APIs charge per token) and performance (fragmented text = less rich representation).

Which tokenizer for which need

ContextRecommendationWhy
English-centric model, large tooling communityByte-level BPE (tiktoken GPT-4 style)De facto standard, HuggingFace compatible, performs in English.
Multilingual model with varied scripts (CJK, Arabic, Indic)SentencePiece with BPE or UnigramLanguage-agnostic, no specific preprocessing, better multilingual fairness.
Research in regularization / robustnessUnigram LM with segmentation samplingAllows varying segmentations during training, acts as data augmentation.
Tokenizer-free research, extreme typo robustnessByT5 or MegaByte (raw bytes)No vocabulary to build, maximum robustness, but 3-7× higher compute cost.
BERT / encoder model for specific NLP tasksClassical WordPieceInherited convention, compact vocabulary (~30k), sufficient for short tasks.

Tokenizer-free: the open question

If tokenization causes problems, why not remove it?

That is the idea behind tokenizer-free approaches. ByT5 (Xue et al., Google 2022) operates directly at the byte level. Each byte is a unit — no pre-built vocabulary needed. Advantages: robustness to spelling errors, underrepresented languages, and rare characters. The model sees raw data without preprocessing.

Major drawback: sequences are 3 to 7 times longer than with BPE. Attention mechanisms have a quadratic cost in sequence length — this computational overhead is significant.

MegaByte (Yu et al., Meta 2023) proposes a hierarchical architecture to work around this problem: a global model processes byte “patches,” while a local model generates bytes one by one. SpaceByte (2024) refines this idea to approach BPE performance without a tokenizer.

These approaches remain experimental. No major production model uses them yet. But they show that tokenization is not a physical constraint — it is an engineering choice, and that choice can evolve.


Key takeaways

  • WordPiece (BERT) merges pairs that maximize corpus likelihood, not raw frequency. The prefix ## marks non-initial subwords.
  • Unigram LM starts from a large vocabulary and prunes iteratively, unlike BPE which starts from characters and builds up. Segmentation is probabilistic.
  • SentencePiece treats spaces as ordinary characters, making the tokenizer language-agnostic. It is the standard for recent multilingual models.
  • Fertility is unequal: Yoruba, Amharic, and other languages with non-Latin scripts generate up to 10 times more tokens than English, with consequences for cost, context window, and performance.
  • Tokenizer-free approaches (ByT5, MegaByte) show that operating directly on bytes is possible, but the computational overhead remains an unresolved obstacle in production.