NTH

Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

AuthorsYubo Wang, Jiarong Liang, Yuxuan Zhang, Xuye Liu, Cong Wei, Yuyu Zhang, Ping Nie, Wenhu Chen

July 22, 2026 2 min read
Watch on YouTube
The one-line take

By training code models to fill in dependency-aware function gaps, the paper improves coding-agent performance and preserves broader tool-use abilities after agentic post-training.

Key results

968
Source repositories

Decontaminated Python corpus drawn from GitHub repositories.

400K
FIM samples

Function-aware fill-in-the-middle training examples.

2.6B
Mid-training tokens

Token budget for the released FIM corpus.

3.2
Qwen3-8B SWE-Bench-Verified gain

Point improvement after FIM mid-training plus SWE-Lego.

5.4
Qwen3-8B SWE-Bench-Lite gain

Point improvement after FIM mid-training plus SWE-Lego.

11.1
LiveCodeBench recovery

Point recovery over agentic post-training alone on the 14B model.

What the paper found

This paper, from researchers at the University of Waterloo, the University of British Columbia, NVIDIA, Verdent AI, and the Vector Institute, proposes function-aware fill-in-the-middle, or FIM, as a dedicated mid-training stage for coding agents. The key idea is that a function call mirrors an agent step: context leads to an action, an externally computed return, and a continuation. The method parses Python files into program dependency graphs, selects single functions or connected function groups using complexity and inferability scores, and trains models to reconstruct the masked code together with a reasoning trace. Using Gemini-3-Flash to generate and filter rationales, the team built a decontaminated corpus from 968 GitHub repositories containing 400K FIM samples and 2.6B tokens. They mid-trained Qwen2.5-Coder-Instruct 7B and 14B, plus Qwen3-8B, before applying R2E-Gym, SWE-Smith, or SWE-Lego post-training. On Qwen3-8B with SWE-Lego, FIM mid-training raised SWE-Bench-Verified by 3.2 points and SWE-Bench-Lite by 5.4 points; similar gains appeared across the Qwen2.5-Coder models and post-training pipelines. The intervention also countered capability erosion: on the 14B model, LiveCodeBench recovered 11.1 points after agentic post-training, while τ-bench and BFCL improved by 3.9 and 2.4 points. Behavioral analysis suggests the gain comes from better recovery after negative tool observations and more iterative editing, especially on patches modifying multiple functions within one file. The main limitations are Python-only data, reliance on Gemini-3-Flash for the strongest rationales, and limited validation beyond Qwen model families.

Original abstract

Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the action-observation-continuation loop of a coding agent is structurally isomorphic to a function call site, where a caller binds arguments, a callee returns a value computed elsewhere, and downstream code consumes that value. This conditioning structure exists at internet scale in ordinary code. We exploit it through function-aware fill-in-the-middle (FIM) mid-training: a self-supervised objective that masks functions selected via program dependency graph analysis and a complexity-inferability double criterion. We mid-train Qwen2.5-Coder-Instruct (7B/14B) and Qwen3-8B on a 2.6B-token decontaminated corpus drawn from 968 GitHub repositories, then apply existing agentic post-training pipelines. Mid-training improves SWE-Bench-Verified by +2.8/+3.0 at 7B/14B and by +3.2 on Qwen3-8B; SWE-Bench-Lite gains are +3.7/+4.0/+5.4 on the same models. The improvement holds across two post-training pipelines (R2E-Gym, SWE-Smith) and on a non-Qwen2.5 base (Qwen3-8B with SWE-Lego). Beyond in-domain gains, mid-training also mitigates the capability erosion that agentic post-training otherwise inflicts on non-agent coding (e.g., LiveCodeBench) and non-coding tool-use benchmarks (tau-bench, BFCL): although the mid-training corpus contains Python code only, the function-call inductive bias survives post-training and yields consistent gains.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis