NTH

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

AuthorsYihao Chen, Shi Chang, Khaled Chawa, Feng Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan

August 5, 2026 2 min read
Watch on YouTube
The one-line take

MindForge trains smaller language models to build complete software from scratch by learning through source-free, end-to-end programming environments.

Key results

562
Source-free environments

MindForge environments spanning six compiled programming languages.

1001
Teacher trajectories

Whole-life-cycle program-development trajectories generated with GLM-5.2.

973
SFT trajectories

Trajectories retained after the 256K-token training-length filter.

49.51%
ProgramBench pass rate

MindForge-27B score, increased from the Qwen3.6-27B base score of 37.98%.

78.00%
RepoZero-C2Rust gain

MindForge-27B all-pass rate, compared with 47.00% for the base model.

15.92%
DeepSWE gain

MindForge-27B score, compared with 1.76% for the base model.

What the paper found

MindForge, developed by researchers at Huawei Canada, Queen’s University, and the University of Manitoba, targets a gap in coding-agent training: building complete software from behavior rather than modifying visible source code. Its pipeline converts open-source command-line programs into source-free environments exposing only a compiled reference executable and documentation, then uses GLM-5.2 through Mini-SWE-Agent to generate and refine whole-life-cycle trajectories covering specification discovery, architecture, implementation, debugging, testing, and refinement. The resulting corpus contains 562 environments across six compiled languages and 1,001 trajectories, with 973 retained for supervised fine-tuning. Distilling these trajectories into Qwen3.6-27B raises average ProgramBench test pass rate from 37.98% to 49.51%, a gain of 11.53 percentage points, while generalizing to seven unseen software-engineering benchmarks. The largest transfer occurs on RepoZero-C2Rust, where performance rises from 47.00% to 78.00%, and on DeepSWE, from 1.76% to 15.92%. MindForge combines infrastructure-noise recovery with localized reasoning rewrites that preserve tool calls and environment outputs, producing cleaner supervision without source leakage. Behavioral analysis shows the tuned model sustains longer workflows while reducing command-failure rate from 10.98% to 9.35%, suggesting that whole-life-cycle trajectory distillation teaches small language models not only to reason about software, but to convert diagnoses into concrete implementation changes.

Original abstract

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier models. Moreover, the fine-tuned model consistently improves over the base model across all seven unseen software engineering benchmarks, spanning long-horizon repository generation and translation, bug fixing, feature implementation, and cross-language issue resolution, with absolute gains of 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, 10.70/4.56 on NL2Repo-Bench (with/without tests), 5.04 on SWE-bench Verified, 5.93 on SWE-bench Pro, 5.22 on SWE-bench Multilingual, and 4.94 on FeatBench.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis