FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents
AuthorsDongyi Lv, Fushun E, Aichen Cai, Liang Huang, Ya Zhang, Qiuyu Ding, Canhui Wu, Zhi Wang, Yuesong Zhang, Jiaqi Wang, Nan Duan
Resources
FailForge teaches code agents to learn from their failures, turning past mistakes into reusable skills that substantially improve software-engineering task performance.
Key results
Number of executable software-engineering training instances.
Tasks where all five initial rollouts failed.
Share of persistently failed instances recovered by guided rerollouts.
Qwen3.5-4B result with FailForge, compared with 59.6 for standard RFT.
Improvement over the standard RFT baseline on SWE-bench Verified.
Teacher tokens required per pass@1 improvement point for FailForge.
What the paper found
FailForge targets a weakness in rejection-sampling fine-tuning: tasks on which every rollout fails are discarded, even though they represent the model’s capability frontier. Using 2,401 SWE-Gym tasks, the framework identifies 832 persistently failed instances, then has a diagnosing agent analyze complete trajectories, test outputs, reasoning traces, and a reference patch to produce a leakage-filtered, repository-agnostic procedural skill. That skill guides new rollouts from a Kimi-K2.6 teacher running OpenHands; successful traces are stripped of the injected skill before training, forcing Qwen3.5-4B or Qwen3.5-9B to internalize the investigation strategy rather than depend on inference-time hints. FailForge recovers 26.2% of previously failed instances and raises Qwen3.5-4B’s SWE-bench Verified pass@1 from 59.6 to 66.2, a 6.6-point improvement, while outperforming both extra sampling and instance-specific hints. At matched training size, recovered traces increase pass@1 from 59.6 to 63.8, indicating higher supervision value per sample. The method transfers across repositories, languages, and harnesses, including Claude Code, and costs 2.12B teacher tokens per pass@1 improvement point, versus 3.01B for hints and 4.32B for more sampling. GPT-5.4 performs leakage filtering, achieving 96.7% agreement with human consensus, helping preserve transferable procedure instead of solution-specific details.
Original abstract
Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering tasks, retaining those that pass the tests, and fine-tuning on the successful rollouts. However, even strong code agents repeatedly fail on a substantial fraction of such tasks, and standard RFT simply discards these failures. The discarded samples are precisely the hardest and most informative ones, drawn from verifiable instances that are costly to curate. Stronger base models may reduce the number of failures, but the remaining hard cases still define the frontier for further improvement. We propose FailForge, an agentic framework that converts failed rollouts into training signal. For each failed instance, an agent diagnoses the failure from error feedback and execution traces, distills the diagnosis into a concise and actionable skill, and injects the skill into the agent context for a guided second attempt. Trajectories that succeed under skill guidance are folded back into the RFT corpus. Crucially, the skill is removed at training time, so the model internalizes the recovered behavior rather than relying on external hints at inference. FailForge recovers over 26% of previously failed instances at marginal additional cost, and training Qwen3.5-4B on the augmented corpus improves the SWE-bench Verified resolve rate by 6.6 points over a strong RFT baseline, with gains concentrated on the hardest problems.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.