SentencePiece Unigram vs BPE · model_type, Character Coverage and Split Rules
What model_type actually selects
model_type determines how SentencePiece constructs and selects its subword vocabulary. The official options documentation sets its default to "unigram" and recommends Unigram, but that recommendation does not imply a universal performance result.
The two principal choices use different mechanisms:
unigramfits a probabilistic model and prunes the resulting vocabulary. Kudo’s 2018 subword regularization method defines likelihoods for multiple possible segmentations of the same input, allowing segmentations to be sampled probabilistically during training.bpebegins with characters and repeatedly merges frequent pairs. Its learned vocabulary therefore reflects a merge procedure rather than a probabilistic segmentation model.
This mechanism explains why the models can produce different pieces for the same string even when trained from the same corpus with the same requested vocabulary size. A Unigram candidate must survive probabilistic modeling and vocabulary pruning. A BPE piece must be produced or retained by the pair-merging procedure. Equal vocabulary sizes do not imply equal pieces, equal numbers of pieces per input, or equal downstream behavior.
The other training options in this article operate at different stages. character_coverage controls which characters enter the alphabet. The split options constrain where learned pieces may cross boundaries. byte_fallback controls how out-of-vocabulary characters are represented. None of these options replaces the distinction between probabilistic vocabulary pruning and BPE merging.
How character_coverage changes the alphabet
character_coverage is a training option with a documented default of 0.9995. It specifies the ratio of corpus characters that the vocabulary must cover. SentencePiece sorts characters by frequency and accumulates them until it reaches the target ratio.
Characters not reached by that cutoff are excluded from the alphabet. With byte fallback disabled, these characters map to <unk>. With byte fallback enabled, they are decomposed into byte tokens instead. Consequently, changing character_coverage can change both the set of ordinary character pieces available to the model and the representation of rare characters at encoding time.
The official documentation gives language-dependent recommendations:
- Use
1.0for languages with small alphabets, including English and German. - Use the default
0.9995for languages with large character sets, including Chinese, Japanese, and Korean, where it can prune rare noise characters and emojis.
These are official recommendations, not general performance claims. A higher coverage target retains a larger character alphabet. A lower target permits low-frequency characters to remain outside it. The setting does not itself impose whitespace, script, number, or delimiter boundaries.
required_chars provides an explicit exception. Its default is "". Any UTF-8 characters placed in this string must enter the alphabet regardless of character_coverage, so they do not map to <unk> through coverage-based exclusion. This is useful when a small set of characters must remain ordinary vocabulary units even under a lower coverage target.
Split rules and their boundaries
The split settings are independent constraints applied before the learned model determines its final pieces. Their documented defaults are:
| Option | Default | Boundary enforced | Effect when disabled or empty |
|---|---|---|---|
split_by_whitespace |
true |
A subword piece cannot cross a whitespace boundary. | The restriction is removed, so a learned piece may cross whitespace. |
split_by_unicode_script |
true |
A piece cannot cross Unicode script boundaries, such as a Latin-to-Kanji boundary. | A learned piece may mix characters from different scripts. |
split_by_number |
true |
Characters 0-9 are separated from other characters, preventing pieces such as abc123. |
Digits are no longer required to be separated from adjacent non-digit characters. |
split_digits |
false |
When set to true, every digit is emitted separately: 123 is guaranteed to become 1, 2, and 3. |
Digits remain subject to the other boundary rules, but consecutive digits may form multi-digit pieces. |
pretokenization_delimiter |
"" |
A nonempty static delimiter separates input during training, and the delimiter itself is removed. Pieces cannot cross that boundary. | Restoring "" removes this custom boundary rule. |
The official documentation describes pretokenization_delimiter as the most general way to introduce an arbitrary static segmentation boundary. It works with both Unigram and BPE. Unlike the other split options, its boundary is supplied explicitly rather than inferred from whitespace, scripts, or numbers.
These constraints compose. For example, split_by_number=true separates abc from 123, while split_digits=false allows the digits to remain together if the model has learned a suitable piece. Setting split_digits=true adds a stronger constraint that makes every digit an individual piece.
Turning a split option off removes a restriction; it does not require the model to create a particular cross-boundary piece. Vocabulary pruning, BPE merges, frequency, and other training constraints still determine whether such a piece exists.
Two related options affect whitespace representation or vocabulary contents rather than the whitespace boundary itself:
treat_whitespace_as_suffixdefaults tofalse. Setting it totrueplaces▁as a suffix, producing notation such asworld▁instead of▁world. It does not replacesplit_by_whitespace.allow_whitespace_only_piecesdefaults tofalse. Setting it totruepermits vocabulary pieces made entirely of whitespace characters.
What byte_fallback changes
byte_fallback controls the representation of out-of-vocabulary characters. Its default is false. When disabled, an out-of-vocabulary character uses <unk>.
When byte_fallback=true, SentencePiece decomposes the character into UTF-8 byte tokens such as <0xE3>. The official documentation states that this completely avoids <unk> tokens and calls byte fallback highly recommended for modern language models.
Coverage and byte fallback solve different parts of the problem:
character_coveragedetermines which characters are included in the learned alphabet.byte_fallbackdetermines how a character outside that alphabet is serialized.
Thus, if a low-frequency character is excluded under character_coverage=0.9995, disabling byte fallback sends it to <unk>, while enabling byte fallback represents it with byte tokens. Byte fallback does not add the excluded character to the ordinary alphabet; it supplies an alternative output representation.
A controlled Unigram-versus-BPE procedure
A fair model comparison must change the model construction mechanism while holding the alphabet, boundaries, vocabulary target, and downstream evaluation constant. Start with the same training corpus and held-out data. Set the same character_coverage, byte_fallback, split options, required_chars, and maximum piece length for both models. Change only model_type.
The documented vocab_size default is 8000. The documented hard_vocab_limit default is true, which makes training fail if the corpus does not contain enough unique subwords to reach the requested size. This is preferable to a silent comparison in which one model reaches the target and the other produces a smaller vocabulary. If hard_vocab_limit=false, SentencePiece can shrink the output vocabulary to the maximum available size, introducing an additional difference between the models.
A minimal paired training configuration is:
import sentencepiece as spm
common = {
"input": "corpus.txt",
"vocab_size": 8000,
"hard_vocab_limit": True,
"max_sentencepiece_length": 16,
"character_coverage": 0.9995,
"required_chars": "",
"byte_fallback": False,
"split_by_whitespace": True,
"split_by_unicode_script": True,
"split_by_number": True,
"split_digits": False,
"pretokenization_delimiter": "",
"treat_whitespace_as_suffix": False,
"allow_whitespace_only_pieces": False,
}
for model_type in ("unigram", "bpe"):
spm.SentencePieceTrainer.train(
model_prefix=model_type,
model_type=model_type,
**common
)
text = "abc123 漢字 café"
for model_type in ("unigram", "bpe"):
processor = spm.SentencePieceProcessor(
model_file=f"{model_type}.model"
)
print(model_type, processor.encode(text, out_type=str))
The printed segmentation is determined by the learned models; it cannot be inferred from model_type alone. The code also assumes that corpus.txt contains enough unique subwords to satisfy hard_vocab_limit=true.
For downstream evaluation, submit the two tokenizations to the same downstream training configuration, decoding procedure, and metric definitions. Use the same primary and secondary metrics for both systems. Segmentation examples and actual vocabulary sizes are useful diagnostics, but they do not replace the shared downstream measurement. Any preprocessing, normalization, or delimiter policy that differs between the two systems must be reported as part of the experiment rather than attributed to model_type.
The Unigram TSV frequency caveat
The official options documentation records an important Unigram-specific limitation for training from a raw word-list or TSV representation. During initial seed-vocabulary extraction, Unigram uses a Suffix Array and does not use frequency information at that stage; each unique word is treated as appearing once.
This behavior can prevent short, high-frequency words such as am or on from becoming candidate pieces. They may then be segmented as individual characters, such as a and m. A result from such data can therefore reflect the seed-extraction path rather than only the difference between probabilistic Unigram pruning and BPE merging.
The documented workarounds are to use model_type=bpe, repeat high-frequency words in the TSV, or add them to user_defined_symbols. Each changes the experimental conditions. Switching only the Unigram run to repeated entries or forced symbols no longer holds the representation constant, while selecting BPE as a workaround does not isolate why the two trained models differ. Keep the original paired comparison unchanged, disclose the limitation, and report any frequency-corrected condition separately.