NTH

Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?

AuthorsEj Zhou, Suchir Salhan, Catherine Arnett, Anna Korhonen

September 19, 2026 2 min read
Watch on YouTube
The one-line take

Independently trained language models may spontaneously learn representations that can be rotated and transferred across languages, suggesting multilingual capabilities without joint training.

Key results

0.78
FLORES-200 matched SGPT CKA

Matched last-layer CKA on parallel sentences, compared with 0.17 for shuffled pairs.

88.7%
Procrustes retrieval

Mean cosine-retrieval P@1 across 36 Goldfish language pairs.

0.71
Independent model matched CKA

Average matched CKA across five independently developed approximately 1B monolingual models, versus 0.18 shuffled.

85%
Cross-lingual factual transfer

Directional success rate for rotated English residuals patched into German on factual cloze tasks.

What the paper found

This paper tests whether cross-lingual alignment requires multilingual joint training by comparing strictly monolingual Transformer models trained on disjoint corpora, vocabularies, and parameters. Across Goldfish models evaluated with FLORES-200, Tatoeba, OPUS, and BouQUET, position-weighted SGPT representations produced matched last-layer CKA of 0.78 on FLORES-200 versus 0.17 for shuffled sentence pairs, showing semantic geometry beyond architectural or token-level overlap. Alignment increased with training-data scale, model scale, and linguistic proximity, and remained detectable across five independently developed approximately 1B monolingual models spanning Pythia, Llama, Qwen2.5, and Mistral-based systems, where matched CKA averaged 0.71 versus 0.18 shuffled. For construction, a single orthogonal Procrustes rotation learned from parallel sentences achieved 88.7% cosine-retrieval accuracy, outperforming affine and one-layer MLP mappings because it preserves angular geometry. For causation, cross-model activation patching injected rotated English residuals into German, French, Spanish, Japanese, and Chinese models; on country-to-capital factual cloze prompts, the transferred representation achieved up to 85% directional success, compared with 76–98% within-model ceilings and approximately 50% chance controls. The findings support the Platonic Representation Hypothesis: language models can independently converge toward partially universal conceptual representations, enabling post-hoc model stitching, merging, and modular multilingual systems without joint pretraining.

Original abstract

Cross-lingual alignment in multilingual language models is typically attributed to joint training: shared parameters, mixed-language batches, or explicit alignment objectives. We ask whether monolingual models trained on non-parallel data learn alignable representations without joint training. By testing on strictly monolingual language models, such as the Goldfish model families and independently developed models from different research labs, we find three results. Correlation: these models develop alignable representational geometry across layers, with alignment strengthening as data scale, model scale, or linguistic proximity increases. Construction: a single Procrustes rotation fit on parallel sentences maps hidden states between models. Causation: the same rotation transfers functional content; patching a rotated English residual into a German model on a factual cloze flips the prediction to the donor's capital in most cases. We confirm that cross-lingual alignment can emerge from the structure of language and the information it carries rather than from joint training, and this points to practical future directions including model stitching, merging, and modular multilingual systems built from monolingual components.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →