NTH

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs

AuthorsJiarong Zhao, Zhikai Lei, Zhiheng Xi, Rui Zheng, Hang Yan, Jie Zhou, Qin Chen, Liang He

July 27, 2026 2 min read
Watch on YouTube
The one-line take

NexForge turns high-level capability requirements into large-scale executable training tasks, substantially improving open-source LLM agent performance.

Key results

3.6K
Terminal-3.6K task corpus

Executable terminal tasks used for Qwen3.5-35B-A3B post-training.

52.0%
Terminal-Bench 2.0 after Terminal-3.6K

Qwen3.5-35B-A3B accuracy after training, up from 22.5%.

1338
GDPval after Office-2K

Qwen3.5-35B-A3B Elo after office-task post-training, up from 813.

58.4%
Scaled Terminal-Bench 2.0

Accuracy achieved after scaling NexForge data to 43.2K terminal tasks.

75.3%
Nex-N2 Terminal-Bench 2.1

Performance of the publicly available Nex-N2 model family.

1585
Nex-N2 GDPval

GDPval Elo reached by Nex-N2 after larger-scale NexForge training.

What the paper found

NexForge, from researchers at East China Normal University, Fudan University, and Shanghai Qiji Zhifeng, addresses a central bottleneck in LLM-agent post-training: existing synthesis pipelines are bound to specific tools, repositories, or skill graphs. Its requirement-driven pipeline begins with a high-level capability request, researches real-world demand, builds weighted task profiles and a diverse scenario reservoir, compiles compatible task directives, materializes CPU-only executable workspaces, and distills teacher-agent trajectories. GPT-5.5 drives synthesis, while DeepSeek-V4-Pro supplies teacher rollouts. Without domain-specific infrastructure, 3.6K terminal tasks raise Qwen3.5-35B-A3B Base from 22.5% to 52.0% on Terminal-Bench 2.0, and 2K office tasks raise GDPval from 813 to 1338 Elo, bringing the open model near systems such as Gemini 3.1 Pro and ahead of several specialized Qwen-based pipelines. Scaling to 43.2K terminal tasks reaches 58.4%, comparable to Claude Opus 4.6 with Claude Code. The resulting Nex-N2 family, trained on larger NexForge corpora, reaches 75.3% on Terminal-Bench 2.1 and 1585 GDPval Elo, challenging proprietary systems from Anthropic, Google, and OpenAI without changing the base architecture. Ablations show that distribution-aware compatibility filtering is essential: it improves directive-scenario match rate from 7.0% to 81.0%. The authors note that the generated tasks are not yet reliable benchmarks or reinforcement-learning environments because they lack universal verifiers.

Original abstract

Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate biases rather than real-world demand. We introduce NexForge, a requirement-driven framework that takes high-level capability requirements as input and synthesizes diverse, executable agent tasks and expert trajectories for SFT. NexForge first investigates real-world demand to construct representative scenarios and task profiles, then performs distribution-aware compilation to generate task directives. For each directive, NexForge automatically retrieves or constructs the required files, dependencies, and runtime configurations, and finally synthesizes expert rollouts and produces training trajectories. Without domain-specific infrastructure, NexForge produces 3.6K terminal and 2K office tasks, improving Qwen3.5-35B-A3B Base from 22.5\% to 52.0\% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval; scaling further to 43.2K terminal tasks yields 58.4\%, on par with Claude Opus 4.6 equipped with Claude Code. Scaled further, NexForge-synthesized data contributes to the training of Nex-N2, a family of publicly available agent models that lift Qwen3.5-35B-A3B to 75.3\% on Terminal-Bench 2.1 and to 1585 Elo on GDPval -- achieving state-of-the-art open-source performance and surpassing several frontier proprietary systems. Nex-N2 models are available at https://nex.sii.edu.cn/.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis