Objective vs. Search: Decomposing What Makes a Good Tokeniser
AuthorsAhmetcan Yavuz, Clara Meister, Tiago Pimentel
Resources
This study shows that how a tokeniser searches may matter more for language-model efficiency than what objective it optimises.
Key results
FineWeb-Edu tokens used to train the English tokenisers.
Five-language corpus used for multilingual tokeniser experiments.
Vocabulary size used in the main large-vocabulary comparisons.
Largest language-model scale evaluated in the English experiments.
Mean BPB for BPE with a 1B model and 128k vocabulary.
Mean BPB for UnigramLM with a 1B model and 128k vocabulary.
What the paper found
Tokenisers used by systems such as OpenAI models and Llama-style decoders differ in two entangled ways: the objective they optimise—compression or unigram log-likelihood—and the search strategy—bottom-up merging or top-down pruning. This paper isolates those factors in a 2×2 design by introducing BottomUpLL, which combines bottom-up search with likelihood optimisation, and TopDownComp, which combines top-down pruning with compression. The experiments train tokenisers on 2B tokens from FineWeb-Edu for English and on a 20B-token, five-language corpus, then train Llama-style language models from 100M to 1B parameters with vocabularies of 8k, 32k, and 128k. Across nearly every condition, search procedure matters more than objective for bits-per-byte: bottom-up methods outperform their top-down counterparts. In the English 1B, 128k-vocabulary setting, BPE reaches 0.7790 BPB versus 0.7872 for UnigramLM, while multilingual results show the same bottom-up advantage. However, this ordering does not consistently transfer to grammaticality on BLiMP, MultiBLiMP, or ZhoBLiMP; at small vocabularies, likelihood objectives can help. The analysis also finds that learned vocabularies cluster by search strategy, with BPE and BottomUpLL sharing 80.5% of tokens at 128k, and that top-down deletion scores require approximations: a local replacement method recovers 81.5% of the vocabulary selected by exact scoring. The central conclusion is that tokeniser search is an independent architectural choice, not merely an implementation detail behind the objective.
Original abstract
Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these axes, making it unclear whether their observed differences stem from what is being optimised vs. how it is being optimised. We disentangle the two by introducing two new tokenisation algorithms that complete this 2x2 design space: BottomUpLL, a bottom-up likelihood-based tokeniser, and TopDownComp, a top-down compression-based tokeniser. We train language models with tokenisers produced by each algorithm, varying: model size, vocabulary sizes, and domain (English-only vs. multilingual). Evaluating models on bits-per-byte, we find that the search procedure -- not the objective -- is the dominant factor: bottom-up tokenisers consistently achieve lower bits-per-byte in most settings. Evaluating models on the BLiMP task, however, shows no consistent relationship between design choice and performance. Overall, our results disentangle the effect of tokeniser design choices on language modelling performance, offering concrete guidance for their more principled construction.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.