DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
AuthorsRaphaël Sourty, Antoine Chaffin, Paulo Roberto Moura Junior, Amélie Chatelain
Resources
This paper openly trains strong dense and late-interaction search models and shows that token-level matching can improve multilingual and unseen-script retrieval.
Key results
Curated from 1.4B query-document pairs across 34 public sources.
Average nDCG@10 for the 149M-parameter dense model.
Average nDCG@10 for the 149M-parameter late-interaction model.
Query-document pairs produced through translate-train and cross-lingual pairing.
Average nDCG@10 across all evaluated MIRACL languages.
Average nDCG@10 for multilingual long-document retrieval.
What the paper found
DenseOn and LateOn present a fully open retrieval recipe designed to close the reproducibility gap behind closed systems such as Qwen3-Embedding and Mistral AI’s embedding ecosystem. The pipeline reconstructs 665M English contrastive pairs from 1.4B candidates across 34 public sources, adds 1.88M supervised pairs with NV-Retriever hard negatives, and trains matched 149M-parameter ModernBERT models: DenseOn uses single-vector CLS embeddings, while LateOn uses ColBERT-style token-level late interaction. On BEIR, they reach 56.20 and 57.22 average nDCG@10, respectively. For multilingual, long-context, and code search, the authors translate training data into eight languages using Mistral-Small-3.1-24B-Instruct and Qwen3-32B, add cross-lingual, MIRACL, MLDR, and LateOn-Code supervision, and create a 2.8B-pair corpus plus a 16.3M-sample fine-tuning set. The resulting 307M-parameter mmBERT-base models reveal the central finding: LateOn generalizes beyond the translated languages and scripts far better than DenseOn. LateOn scores 67.04 on full MIRACL and 77.92 on full MLDR, compared with DenseOn’s 58.02 and 51.59, while reaching 73.48 on MTEB Code without code-specific pre-training. The models, datasets, and training code are released, enabling audits and controlled comparisons against systems such as BGE-M3, Qwen3-Embedding, and pplx-embed.
Original abstract
State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scripts. This suggests that token-level matching turns translate-train from a target-language expansion strategy into a multilingual generalization recipe. We publicly release the models, datasets, and training code.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.