Neural Machine Translation of Rare Words with Subword Units
Rico Sennrich, Barry Haddow, Alexandra Birch · 2015 · arXiv:1508.07909
Summary
Adapts byte pair encoding to text segmentation, building a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair, so that rare words decompose into known fragments.
Why it matters
It removed the out-of-vocabulary problem from neural sequence models. Raschka identifies this as the paper describing the byte pair encoding used for tokenization, and the scheme it introduced is what GPT-family models still use to segment their input.