Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation
AuthorsM. Ali Bayram, Banu Diri, Savaş Yıldırım
Resources
This paper shows how to cheaply adapt a multilingual embedding model to Turkish by redesigning its tokenizer and distilling it offline, achieving strong retrieval-quality gains at low training cost.
Key results
The hybrid tokenizer is built to a 128K vocabulary, reported as 2^17 = 131,072 tokens.
Offline distillation uses approximately 580K balanced examples from a 40-language Wikipedia corpus.
The final student embeddingmagibu-200m is reported as an approximately 200M-parameter model.
embeddingmagibu-200m achieves 77.55 Pearson on the STSbTR test set, exceeding the teacher's 73.84.
On TR-MTEB, the model reaches a 63.9 average score and ranks 7th out of 26 models.
The paper reports a total training cost of $5–$20 for the offline distillation run.
What the paper found
This paper adapts Google’s EmbeddingGemma-300M into a Turkish-focused sentence encoder, embeddingmagibu-200m, by combining cross-lingual tokenizer surgery with offline distillation rather than full pretraining. The key novelty is a hybrid 131,072-token tokenizer built by selecting 64K high-frequency Turkish subwords from the Cosmos Turkish Corpus, pruning redundant teacher tokens, and then adding frequency-filtered multilingual tokens from a 40-language Wikipedia corpus to preserve broad coverage. Because the tokenizer changes vocabulary identities, the authors clone the EmbeddingGemma backbone and initialize the new embedding table with mean-composition mapping from each target token to its teacher-token decomposition, then distill only precomputed teacher vectors using a cosine loss on about 580K balanced examples. The result is a 200M-parameter model with 768-dimensional L2-normalized embeddings and an 8,192-token context window, trained in roughly four hours on a single NVIDIA A100 for an estimated $5–$20. On STSbTR, it reaches 77.55 Pearson and 77.45 Spearman, outperforming the 300M-parameter teacher’s 73.84 and 72.92. On TR-MTEB, it scores 63.9 on average, ranking 7th of 26 models and retaining 98% of the teacher’s score while using 33% fewer parameters. The largest gains appear in semantically sensitive tasks: +5.7 points on STS, +12.3 on NLI, and +6.9 on bitext mining versus the smaller 64K-vocabulary predecessor, showing that tokenizer design is a primary lever for Turkish embedding quality and long-context retrieval.
Original abstract
Sentence embeddings are a foundational component for semantic search, clustering, classification, and retrieval-augmented generation. This paper presents embeddingmagibu-200m, a Turkish-focused sentence embedding model that produces 768-dimensional L2-normalized vectors and supports an 8,192-token context window, far exceeding the 512-token limit of earlier BERT-based Turkish encoders. Instead of full pretraining, an efficient three-stage adaptation pipeline is introduced: (1) construct a Turkish-optimized multilingual tokenizer with a 131,072 vocabulary by pruning redundant tokens from the teacher's vocabulary and incorporating multilingual tokens via frequency analysis on a 40-language corpus, (2) clone a teacher embedding model while preserving transformer backbone weights and initializing a compatible embedding table for the new vocabulary via mean-composition token mapping, and (3) perform offline embedding distillation from precomputed teacher vectors using a cosine similarity objective over a balanced 40-language Wikipedia corpus. The resulting student model contains approximately 200M parameters and trains in roughly four hours on a single GPU by avoiding online teacher inference during training, at a total cost of $5-$20. Empirically, Pearson/Spearman correlations of 77.55%/77.45% are obtained on STSbTR, surpassing the 300M-parameter teacher model (73.84%/72.92%). On TR-MTEB (26 tasks), a mean score of 63.9% is achieved (7th out of 26 models), providing a competitive cost-quality trade-off with 33% fewer parameters than the teacher. To facilitate reproducibility and downstream use, all artifacts are released including model weights, tokenizer files, precomputed embedding datasets, and open-source cloning and distillation tooling.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.