NTH

A Sovereign, Open-Source Foundation Model for German and English

AuthorsThe Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben Härle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom Röhr, Sebastian von Rohrscheidt, Jörg Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim Köhler, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max Lübbering

July 17, 2026 2 min read
Watch on YouTube
The one-line take

Soofi S is a new open bilingual foundation model for German and English that combines MoE and Mamba-Transformer ideas to deliver strong quality with efficient long-context inference.

Key results

31.6B
Total parameters

Soofi S model capacity in the hybrid Mamba–Transformer MoE architecture

3.2B
Active parameters per token

Parameters activated for each forward pass

27T
Pretraining tokens

Approximate tokens consumed across the three-phase training curriculum

15.3%
German annealing share

German portion of the high-quality annealing mixture

4.82k
Long-context decode throughput

Aggregate decode TPS per GPU at 40K context and batch size 32

9.6
GPQA-Diamond improvement

Percentage-point improvement over the architecture-identical Nemotron 3 Nano baseline

What the paper found

The Soofi-Team introduces Soofi S 30B-A3B, a sovereign German–English foundation model developed by researchers from Fraunhofer IAIS, DFKI, Technische Universität Darmstadt, KI Bundesverband, and other German institutions, with funding from BMWE. Its architecture follows NVIDIA’s Nemotron 3 Nano reference design: a hybrid Mamba-2, Grouped-Query Attention, and Mixture-of-Experts network with 31.6B total parameters but only 3.2B active per token, reducing attention-cache growth for long-context serving. Trained on Deutsche Telekom’s Industrial AI Cloud in Munich using NVIDIA infrastructure, the model consumed 27T tokens in a Warmup–Stable–Decay curriculum built from Nemotron-CC, HPLT, FinePDFs, Dolma, code, mathematics, reasoning, and synthetic SFT data. German was deliberately increased to 15.3% of the annealing mixture. At 40K-token context and batch size 32, Soofi S reached 4.82k decode TPS per GPU, 9.2 times the throughput of Ministral 3 14B, while extending usable context to 1M tokens. In capability evaluations, it scored 70.1 on the English aggregate and 79.1 on the German aggregate, outperforming open European baselines such as Apertus 70B and remaining competitive with Qwen3.5, Gemma 3, Ministral 3, and Olmo 3. Against the architecture-identical Nemotron 3 Nano, the data recipe improved GPQA-Diamond by 9.6 points and German language proficiency by 15.1 points. The planned release includes weights, selected checkpoints, training and evaluation code, hyperparameters, and detailed per-source data accounting, although the commercially licensed Genios corpus cannot be redistributed.

Original abstract

We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis