NTH

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

AuthorsJiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren

September 16, 2026 2 min read
Watch on YouTube
The one-line take

ZGCM-1 shows how a fully open 7B model can combine efficient training, long-context reasoning, and tool-using agents to compete with much larger systems.

Key results

7.39B
Model parameters

Dense foundation-model scale.

256K
Maximum context

Supported by the hybrid sliding-window/global attention architecture.

3.94
Long-context throughput speedup

Relative throughput over full attention at 256K context.

6.4
KV-cache reduction

Reduction relative to full attention at 256K context.

4.2
Pre-training time-to-loss improvement

Speedup over a BF16/AdamW baseline.

75.0%
AIME 2026

ZGCM-1-7B thinking-mode score.

What the paper found

ZGCM-1 is a fully open 7.39B dense foundation model designed to combine deliberate reasoning with external tool use for mathematics, web research, and binary analysis. Its 256K-token architecture interleaves gated 128-token sliding-window attention with global attention in a 5:1 ratio, using Grouped-Query Attention, Partial RoPE, FP8 computation, the Muon optimizer, and TWEO outlier regularization. Progressive mid-training expands context from 16K to 64K to 256K while reformulating agent trajectories as Markov Decision Process state-action supervision, and mixed think/no-think supervised fine-tuning balances long reasoning with concise responses. Compared with full attention, the hybrid design delivers a 3.94× throughput gain at 256K and cuts KV-cache memory by 6.4×; combined with FP8, Muon, TWEO, and system co-design, pre-training time-to-loss improves by 4.2× over a BF16/AdamW baseline. ZGCM-1-7B scores 75.0% on AIME 2026 and remains competitive with much larger systems including Qwen3-235B-A22B, GLM-5.1, Kimi-K2, DeepSeek-R1, Claude 4 Sonnet, and GPT-4o, while the paper reports 63.1% on WebWalkerQA and 62.0% on Binary Function Search. The project releases weights, checkpoints, training code, data recipes, logs, and evaluation harnesses, positioning reproducibility and AI-agent-assisted model development as core contributions rather than afterthoughts.

Original abstract

In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis