NTH

Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation

AuthorsM. Ali Bayram, Banu Diri, Savaş Yıldırım

June 7, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows how to cheaply adapt a multilingual embedding model to Turkish by redesigning its tokenizer and distilling it offline, achieving strong retrieval-quality gains at low training cost.

Key results

131072
Tokenizer vocabulary

The hybrid tokenizer is built to a 128K vocabulary, reported as 2^17 = 131,072 tokens.

580K
Training corpus size

Offline distillation uses approximately 580K balanced examples from a 40-language Wikipedia corpus.

200M
Model size

The final student embeddingmagibu-200m is reported as an approximately 200M-parameter model.

77.55
STSbTR Pearson

embeddingmagibu-200m achieves 77.55 Pearson on the STSbTR test set, exceeding the teacher's 73.84.

63.9
TR-MTEB average

On TR-MTEB, the model reaches a 63.9 average score and ranks 7th out of 26 models.

5-20
Training cost

The paper reports a total training cost of $5–$20 for the offline distillation run.

What the paper found

This paper adapts Google’s EmbeddingGemma-300M into a Turkish-focused sentence encoder, embeddingmagibu-200m, by combining cross-lingual tokenizer surgery with offline distillation rather than full pretraining. The key novelty is a hybrid 131,072-token tokenizer built by selecting 64K high-frequency Turkish subwords from the Cosmos Turkish Corpus, pruning redundant teacher tokens, and then adding frequency-filtered multilingual tokens from a 40-language Wikipedia corpus to preserve broad coverage. Because the tokenizer changes vocabulary identities, the authors clone the EmbeddingGemma backbone and initialize the new embedding table with mean-composition mapping from each target token to its teacher-token decomposition, then distill only precomputed teacher vectors using a cosine loss on about 580K balanced examples. The result is a 200M-parameter model with 768-dimensional L2-normalized embeddings and an 8,192-token context window, trained in roughly four hours on a single NVIDIA A100 for an estimated $5–$20. On STSbTR, it reaches 77.55 Pearson and 77.45 Spearman, outperforming the 300M-parameter teacher’s 73.84 and 72.92. On TR-MTEB, it scores 63.9 on average, ranking 7th of 26 models and retaining 98% of the teacher’s score while using 33% fewer parameters. The largest gains appear in semantically sensitive tasks: +5.7 points on STS, +12.3 on NLI, and +6.9 on bitext mining versus the smaller 64K-vocabulary predecessor, showing that tokenizer design is a primary lever for Turkish embedding quality and long-context retrieval.

Original abstract

Sentence embeddings are a foundational component for semantic search, clustering, classification, and retrieval-augmented generation. This paper presents embeddingmagibu-200m, a Turkish-focused sentence embedding model that produces 768-dimensional L2-normalized vectors and supports an 8,192-token context window, far exceeding the 512-token limit of earlier BERT-based Turkish encoders. Instead of full pretraining, an efficient three-stage adaptation pipeline is introduced: (1) construct a Turkish-optimized multilingual tokenizer with a 131,072 vocabulary by pruning redundant tokens from the teacher's vocabulary and incorporating multilingual tokens via frequency analysis on a 40-language corpus, (2) clone a teacher embedding model while preserving transformer backbone weights and initializing a compatible embedding table for the new vocabulary via mean-composition token mapping, and (3) perform offline embedding distillation from precomputed teacher vectors using a cosine similarity objective over a balanced 40-language Wikipedia corpus. The resulting student model contains approximately 200M parameters and trains in roughly four hours on a single GPU by avoiding online teacher inference during training, at a total cost of $5-$20. Empirically, Pearson/Spearman correlations of 77.55%/77.45% are obtained on STSbTR, surpassing the 300M-parameter teacher model (73.84%/72.92%). On TR-MTEB (26 tasks), a mean score of 63.9% is achieved (7th out of 26 models), providing a competitive cost-quality trade-off with 33% fewer parameters than the teacher. To facilitate reproducibility and downstream use, all artifacts are released including model weights, tokenizer files, precomputed embedding datasets, and open-source cloning and distillation tooling.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →