BPE Dropout · Making Subword Segmentation More Robust for Unseen Words
The deterministic segmentation bottleneck
Provilkov et al. identify a limitation in conventional Byte Pair Encoding (BPE): although a vocabulary can admit several possible segmentations, the BPE procedure selects one unique sequence for each word. Frequent words may remain intact, while rare words are split into multiple subword tokens, but the segmentation decision itself is deterministic.
This creates a representation bottleneck during training. If the same word always appears with the same token boundaries, the model does not receive examples of its alternative decompositions. It therefore cannot learn as directly from segmentation ambiguity as a model trained on multiple valid segmentations of the same input. The originating paper states that deterministic BPE segmentation “may prevent a model from better learning the compositionality of words and being robust to segmentation errors.”
The diagnosis is not that alternative segmentations do not exist. They may exist under the same vocabulary. The problem is that ordinary BPE exposes only one of them. For words associated with the open-vocabulary problem in machine translation, this removes a useful source of training variation precisely where the model may need to reason about unfamiliar compositions.
What BPE-dropout changes
BPE-dropout retains conventional BPE and stochastically corrupts its segmentation procedure. In this terminology, “corruption” refers to variation in the selected token boundaries, not modification of the input characters. The same word can consequently receive different segmentations within the same fixed BPE framework.
This is the central distinction from deterministic BPE:
- Conventional BPE provides one segmentation for a given word.
- BPE-dropout provides multiple possible segmentations during training.
- The underlying framework remains BPE rather than switching to a separate subword segmentation model.
The training objective can therefore encounter alternative compositions instead of seeing only the segmentation selected by the ordinary BPE procedure. Provilkov et al. evaluate this training-time variation with standard BPE at inference: BPE-dropout is used to regularize training, while the inference procedure remains deterministic.
The distinction between the training and inference procedures matters. BPE-dropout is not presented as a requirement for stochastic deployment. Its role is to expose the model to segmentation ambiguity while it is learning, after which the experiment uses the standard BPE segmentation procedure.
BPE-dropout and the Unigram sampling model
The model family determines the name and provenance of the sampling method. With model_type set to BPE, SentencePiece calls the technique BPE-Dropout. With Unigram, it calls the corresponding technique Subword Regularization. Changing between these model families changes both the base segmentation model and the sampling approach; it is not merely a different value for a BPE sampling control.
The other construction options are separate from the sampling switch demonstrated here. The official on-the-fly example does not vary character_coverage, split_by_whitespace, byte_fallback, or a related split rule. It therefore does not establish which characters or token boundaries would change if those options were flipped. In particular, the alternative segmentations shown below should not be attributed to a change in character coverage, whitespace handling, byte fallback, or another BPE construction rule. They arise from sampling the segmentation under the same BPE framework.
This separation prevents a common experimental confound: changing tokenizer construction while also enabling BPE-dropout. A controlled comparison should keep the fixed BPE framework unchanged and vary the sampling behavior.
Enabling on-the-fly sampling
SentencePiece supports on-the-fly subword sampling during training. Given a loaded processor, the official example can be expressed as a small Python helper:
import sentencepiece as spm
def sample_new_york(sp: spm.SentencePieceProcessor):
for _ in range(3):
print(
sp.encode(
"New York",
out_type=str,
enable_sampling=True,
alpha=0.1,
nbest_size=-1,
)
)
The important controls in this call are:
enable_sampling=Trueexplicitly enables on-the-fly sampling.alpha=0.1is thealphasetting used in the official example.nbest_size=-1is the accompanyingnbest_sizesetting in that example.out_type=strmakes the token boundaries visible as strings.
The documentation presents alpha=0.1 and nbest_size=-1 as example values, not as documented defaults or universal recommendations. The example also does not define alpha as a percentage or identify it as a dropout probability. It should be reproduced as the documented sampling configuration rather than reinterpreted through an unsupported equivalence.
The loop performs three separate encoding calls. Its count of printed results comes from range(3), not from nbest_size. The encoder also returns a new sampled segmentation on each enabled call.
For the single input string New York, the official documentation gives this illustrative output:
['▁', 'N', 'e', 'w', '▁York']
['▁New', '▁York']
['▁New', '▁Y', 'o', 'r', 'k']
The README labels this output with “May output,” so the exact sequences and their order are not guaranteed. The block is an illustration of segmentation diversity, not a fixed expected transcript.
What the example demonstrates concretely is the movement of token boundaries:
- In the first result,
Newis divided into character-level pieces whileYorkremains a single piece. - In the second, both words remain intact.
- In the third,
Newremains intact whileYorkis divided into character-level pieces.
The input string is unchanged. The returned decomposition changes.
This is the implementation-level form of the paper’s claim that BPE-dropout leads to multiple segmentations within one fixed BPE framework. SentencePiece describes the effect as virtually augmenting the training data: the model receives different segmentations of the same input rather than storing each resegmentation as a separate corpus example. The official documentation connects this variation with greater resilience to spelling variations and noise.
Relationship to Kudo’s subword regularization
Kudo’s 2018 subword regularization work addresses the same broader opportunity from a different direction. It treats subword segmentation ambiguity as noise that can be sampled during training, allowing a sentence to be converted into multiple subword sequences. For improved sampling, Kudo proposes a segmentation algorithm based on a unigram language model.
The BPE-dropout paper characterizes Kudo’s method as the previous solution to deterministic BPE’s limitation. Its alternative is to keep conventional BPE and randomize that procedure directly. This preserves compatibility with BPE rather than replacing it with the Unigram-based approach.
Provilkov et al. report the following results for their machine translation experiments:
- Up to 3 BLEU compared with BPE.
- Up to 0.9 BLEU compared with the previous subword regularization method.
These are the paper’s translation results, using BPE-dropout during training and standard BPE during inference. They are not general tokenizer benchmark scores and should not be transferred to unrelated tasks. The evidence supports a translation result, not a universal performance claim for classification, question answering, language modeling, or arbitrary multilingual workloads.
Kudo’s paper separately reports consistent improvements especially in low-resource and out-of-domain settings, but provides no comparable numerical figures in the material associated with this comparison. The measurable BPE-dropout claims here should therefore remain tied to the translation experiment in which Provilkov et al. reported them.
The conceptual relationship can be stated precisely:
- Both methods use segmentation ambiguity as training-time noise.
- Kudo’s method samples through a Unigram-language-model-based procedure.
- BPE-dropout samples by corrupting the segmentation procedure of conventional BPE.
- Both seek improved robustness without treating deterministic BPE as the only possible training representation.
Reproducibility and random_seed
BPE-dropout uses a random number generator, and SentencePiece exposes a training option named random_seed. The official option is a uint32 with the documented default:
4294967295
This value is the maximum uint32. It is an official default, not a recommendation to use that value for every experiment.
The documentation uses the random generator for EM initialization in Unigram and for BPE-dropout. A fixed random_seed is therefore relevant when reproducing the stochastic behavior of these components, but it does not create an unconditional guarantee of identical output.
SentencePiece uses Abseil Random internally, specifically absl::BitGen. Abseil does not guarantee the stability of generated random sequences across different library versions or platforms. Consequently, passing a fixed random_seed does not guarantee permanent reproducibility across updates.
This limitation applies directly to sampled segmentation output. Two executions intended to represent the same experiment can be distinguished by the seed, but matching that seed does not by itself establish a cross-version or cross-platform guarantee. The New York example should therefore be read as a possible set of boundaries, not as a golden output that must recur after library or platform changes.
The practical boundary is straightforward: enable_sampling=True activates the documented sampling behavior; random_seed controls the random generator; and the documented Abseil limitation prevents the latter from being treated as permanent, environment-independent reproducibility.