KletterMix: Climbing Toward High-Quality German Pretraining Data
AuthorsMaurice Kraus, Ruben Härle, Sebastian Sztwiertnia, Abbas Goher Khan, Mehdi Ali, Michael Fromm, Kristian Kersting
Resources
This paper builds a large, carefully translated German pretraining corpus and shows that better curated data can noticeably improve German language models.
Key results
KletterMix is a German pretraining corpus containing 725B GPT-2 tokens.
The full translation campaign ran on 1,008 NVIDIA B200 GPUs.
The full corpus translation campaign took about 10 days.
The deployed target-only proxy reached 0.725 Pearson against COMETKiwi on the disjoint validation split.
The deployed target-only proxy reached 0.719 Spearman against COMETKiwi on the disjoint validation split.
In matched 12B-token pretraining, KletterMix-Filt0.60 achieved the best Core Avg. at 40.2.
What the paper found
KletterMix, from TU Darmstadt, the Lamarr Institute, Fraunhofer IAIS, hessian.AI, and DFKI, introduces a 725B-token German pretraining corpus built by translating the ClimbMix English mixture into German while preserving document boundaries, source clusters, metadata, and topical diversity. The pipeline uses Qwen3.5-397B-A17B-FP8 with length-aware routing, contextual chunk translation for long documents, dynamic generation budgets, and shard-wise execution on 1,008 NVIDIA B200 GPUs over about 10 days. To score quality at scale, the authors use COMETKiwi on a pilot set, then train a target-only gradient-boosted proxy on cheap German-side features; this proxy achieves 0.725 Pearson, 0.719 Spearman, and 0.0486 MAE against COMETKiwi on a disjoint 18,275-document validation split. In controlled 12B-token pretraining of Qwen3-0.6B, KletterMix lowers validation loss versus FineWeb2-DE and GermanWeb and improves downstream German 5-shot Core Avg. from 38.3 for FineWeb2-DE and 36.8 for GermanWeb to 38.7 unfiltered, with a proxy-filtered variant reaching 40.2. The biggest gains appear on HellaSwag and ARC-Challenge, and annealing a FineWeb2-DE checkpoint on KletterMix beats annealing on GermanWeb, reaching 39.4 Core Avg. versus 37.6. The paper’s main claim is that careful translation can transfer not just German surface text, but the richer mixture structure of a strong English corpus.
Original abstract
High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are often smaller, less carefully curated, weakly documented, and rarely validated through controlled training experiments. We introduce KletterMix, a high-quality German corpus for language model pretraining and annealing, designed as a reusable dataset artifact for the natural language processing and modeling community. KletterMix is built by translating a state-of-the-art English pretraining corpus into German while preserving document boundaries, metadata, source structure, and topical diversity. This construction yields a German corpus with the scale and diversity of a modern pretraining dataset, while enabling direct comparison to its English source. We document the dataset through a broad set of corpus-level analyses, including translation quality, document length distributions, topic coverage, source composition, and geographic metadata. Using COMETKiwi, we show that the translated documents achieve strong quality across diverse domains, suggesting that careful translation can preserve much of the semantic and stylistic richness of the original corpus. Beyond dataset construction, we evaluate KletterMix as training data. Through controlled pretraining and annealing ablations against established German corpora, we show that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations. These results demonstrate that carefully curated translated data can substantially strengthen the German pretraining data ecosystem.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.