NTH

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

AuthorsRaphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior, Amélie Chatelain

August 25, 2026 2 min read
Watch on YouTube
The one-line take

This paper openly trains strong dense and late-interaction search models and shows that token-level matching can improve multilingual and unseen-script retrieval.

Key results

665M
English pre-training pairs

Curated from 1.4B query-document pairs across 34 public sources.

56.20
English BEIR DenseOn

Average nDCG@10 for the 149M-parameter dense model.

57.22
English BEIR LateOn

Average nDCG@10 for the 149M-parameter late-interaction model.

2.8B
Multilingual pre-training corpus

Query-document pairs produced through translate-train and cross-lingual pairing.

67.04
Full MIRACL LateOn

Average nDCG@10 across all evaluated MIRACL languages.

77.92
Full MLDR LateOn

Average nDCG@10 for multilingual long-document retrieval.

What the paper found

DenseOn and LateOn present a fully open retrieval recipe designed to close the reproducibility gap behind closed systems such as Qwen3-Embedding and Mistral AI’s embedding ecosystem. The pipeline reconstructs 665M English contrastive pairs from 1.4B candidates across 34 public sources, adds 1.88M supervised pairs with NV-Retriever hard negatives, and trains matched 149M-parameter ModernBERT models: DenseOn uses single-vector CLS embeddings, while LateOn uses ColBERT-style token-level late interaction. On BEIR, they reach 56.20 and 57.22 average nDCG@10, respectively. For multilingual, long-context, and code search, the authors translate training data into eight languages using Mistral-Small-3.1-24B-Instruct and Qwen3-32B, add cross-lingual, MIRACL, MLDR, and LateOn-Code supervision, and create a 2.8B-pair corpus plus a 16.3M-sample fine-tuning set. The resulting 307M-parameter mmBERT-base models reveal the central finding: LateOn generalizes beyond the translated languages and scripts far better than DenseOn. LateOn scores 67.04 on full MIRACL and 77.92 on full MLDR, compared with DenseOn’s 58.02 and 51.59, while reaching 73.48 on MTEB Code without code-specific pre-training. The models, datasets, and training code are released, enabling audits and controlled comparisons against systems such as BGE-M3, Qwen3-Embedding, and pplx-embed.

Original abstract

State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scripts. This suggests that token-level matching turns translate-train from a target-language expansion strategy into a multilingual generalization recipe. We publicly release the models, datasets, and training code.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →