Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
AuthorsYubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen
Resources
By training code models to fill in dependency-aware function gaps, the paper improves coding-agent performance and preserves broader tool-use abilities after agentic post-training.
Key results
Decontaminated Python corpus drawn from GitHub repositories.
Function-aware fill-in-the-middle training examples.
Token budget for the released FIM corpus.
Point improvement after FIM mid-training plus SWE-Lego.
Point improvement after FIM mid-training plus SWE-Lego.
Point recovery over agentic post-training alone on the 14B model.
What the paper found
This paper, from researchers at the University of Waterloo, the University of British Columbia, NVIDIA, Verdent AI, and the Vector Institute, proposes function-aware fill-in-the-middle, or FIM, as a dedicated mid-training stage for coding agents. The key idea is that a function call mirrors an agent step: context leads to an action, an externally computed return, and a continuation. The method parses Python files into program dependency graphs, selects single functions or connected function groups using complexity and inferability scores, and trains models to reconstruct the masked code together with a reasoning trace. Using Gemini-3-Flash to generate and filter rationales, the team built a decontaminated corpus from 968 GitHub repositories containing 400K FIM samples and 2.6B tokens. They mid-trained Qwen2.5-Coder-Instruct 7B and 14B, plus Qwen3-8B, before applying R2E-Gym, SWE-Smith, or SWE-Lego post-training. On Qwen3-8B with SWE-Lego, FIM mid-training raised SWE-Bench-Verified by 3.2 points and SWE-Bench-Lite by 5.4 points; similar gains appeared across the Qwen2.5-Coder models and post-training pipelines. The intervention also countered capability erosion: on the 14B model, LiveCodeBench recovered 11.1 points after agentic post-training, while τ-bench and BFCL improved by 3.9 and 2.4 points. Behavioral analysis suggests the gain comes from better recovery after negative tool observations and more iterative editing, especially on patches modifying multiple functions within one file. The main limitations are Python-only data, reliance on Gemini-3-Flash for the strongest rationales, and limited validation beyond Qwen model families.
Original abstract
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.