A Sovereign, Open-Source Foundation Model for German and English
AuthorsThe Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben Härle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom Röhr, Sebastian von Rohrscheidt, Jörg Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim Köhler, Alexander Löser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max Lübbering
Resources
Soofi S is a new open bilingual foundation model for German and English that combines MoE and Mamba-Transformer ideas to deliver strong quality with efficient long-context inference.
Key results
Soofi S model capacity in the hybrid Mamba–Transformer MoE architecture
Parameters activated for each forward pass
Approximate tokens consumed across the three-phase training curriculum
German portion of the high-quality annealing mixture
Aggregate decode TPS per GPU at 40K context and batch size 32
Percentage-point improvement over the architecture-identical Nemotron 3 Nano baseline
What the paper found
The Soofi-Team introduces Soofi S 30B-A3B, a sovereign German–English foundation model developed by researchers from Fraunhofer IAIS, DFKI, Technische Universität Darmstadt, KI Bundesverband, and other German institutions, with funding from BMWE. Its architecture follows NVIDIA’s Nemotron 3 Nano reference design: a hybrid Mamba-2, Grouped-Query Attention, and Mixture-of-Experts network with 31.6B total parameters but only 3.2B active per token, reducing attention-cache growth for long-context serving. Trained on Deutsche Telekom’s Industrial AI Cloud in Munich using NVIDIA infrastructure, the model consumed 27T tokens in a Warmup–Stable–Decay curriculum built from Nemotron-CC, HPLT, FinePDFs, Dolma, code, mathematics, reasoning, and synthetic SFT data. German was deliberately increased to 15.3% of the annealing mixture. At 40K-token context and batch size 32, Soofi S reached 4.82k decode TPS per GPU, 9.2 times the throughput of Ministral 3 14B, while extending usable context to 1M tokens. In capability evaluations, it scored 70.1 on the English aggregate and 79.1 on the German aggregate, outperforming open European baselines such as Apertus 70B and remaining competitive with Qwen3.5, Gemma 3, Ministral 3, and Olmo 3. Against the architecture-identical Nemotron 3 Nano, the data recipe improved GPQA-Diamond by 9.6 points and German language proficiency by 15.1 points. The planned release includes weights, selected checkpoints, training and evaluation code, hyperparameters, and detailed per-source data accounting, although the commercially licensed Genios corpus cannot be redistributed.
Original abstract
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English. Its hybrid design activates only 3B of 30B parameters per token and keeps the inference cache near-constant as context grows, giving it a decisive throughput advantage over dense models for long-context, high-concurrency deployment. Pretrained on roughly 27 trillion tokens with deliberately up-weighted German, Soofi S matches dense 14 to 27B models on aggregate English and German benchmarks while achieving the best code aggregates in both languages among 17 open base models, and outperforms every European sovereign baseline in our comparison, including ones far larger in active parameters. Among fully open models, Soofi S obtains the highest English and German evaluation scores, ahead of Olmo 3 32B and Apertus 70B. Soofi S was built end-to-end on the German Industrial AI Cloud, a sovereign HPC scale AI infrastructure operated by Deutsche Telekom in Munich. Soofi S will be released under highly permissive, open-access terms: weights, selected intermediate checkpoints, full per-source data accounting, hyperparameters, and training and evaluation code. Where source licenses permit, data-construction artifacts are released under permissive licenses; commercially licensed sources are documented with aggregate statistics and exact mixture accounting.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.