MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
AuthorsYihao Chen, Shi Chang, Khaled Chawa, Feng Lin, Boyuan Chen, Shaowei Wang, Ahmed E. Hassan
Resources
MindForge trains smaller language models to build complete software from scratch by learning through source-free, end-to-end programming environments.
Key results
MindForge environments spanning six compiled programming languages.
Whole-life-cycle program-development trajectories generated with GLM-5.2.
Trajectories retained after the 256K-token training-length filter.
MindForge-27B score, increased from the Qwen3.6-27B base score of 37.98%.
MindForge-27B all-pass rate, compared with 47.00% for the base model.
MindForge-27B score, compared with 1.76% for the base model.
What the paper found
MindForge, developed by researchers at Huawei Canada, Queen’s University, and the University of Manitoba, targets a gap in coding-agent training: building complete software from behavior rather than modifying visible source code. Its pipeline converts open-source command-line programs into source-free environments exposing only a compiled reference executable and documentation, then uses GLM-5.2 through Mini-SWE-Agent to generate and refine whole-life-cycle trajectories covering specification discovery, architecture, implementation, debugging, testing, and refinement. The resulting corpus contains 562 environments across six compiled languages and 1,001 trajectories, with 973 retained for supervised fine-tuning. Distilling these trajectories into Qwen3.6-27B raises average ProgramBench test pass rate from 37.98% to 49.51%, a gain of 11.53 percentage points, while generalizing to seven unseen software-engineering benchmarks. The largest transfer occurs on RepoZero-C2Rust, where performance rises from 47.00% to 78.00%, and on DeepSWE, from 1.76% to 15.92%. MindForge combines infrastructure-noise recovery with localized reasoning rewrites that preserve tool calls and environment outputs, producing cleaner supervision without source leakage. Behavioral analysis shows the tuned model sustains longer workflows while reducing command-failure rate from 10.98% to 9.35%, suggesting that whole-life-cycle trajectory distillation teaches small language models not only to reason about software, but to convert diagnoses into concrete implementation changes.
Original abstract
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier models. Moreover, the fine-tuned model consistently improves over the base model across all seven unseen software engineering benchmarks, spanning long-horizon repository generation and translation, bug fixing, feature implementation, and cross-language issue resolution, with absolute gains of 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, 10.70/4.56 on NL2Repo-Bench (with/without tests), 5.04 on SWE-bench Verified, 5.93 on SWE-bench Pro, 5.22 on SWE-bench Multilingual, and 4.94 on FeatBench.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.