CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
AuthorsBowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo
AffiliationsLLM Core, Xiaomi · Peking University · University of Hong Kong · Renmin University of China
Resources
CodeMidas turns existing codebases into scalable, automatically verified RL environments that train coding agents to perform better across diverse software tasks.
Key results
Verifiable tasks retained for reinforcement-learning training
Open-source codebases used to construct the environments
Languages represented in the training corpus
Post-RL score, up from 10.0%
Post-RL score, up from 63.7%
Higher mean pass rate for rollouts with agent-written and executed checks
What the paper found
CodeMidas presents an agentic pipeline that converts implemented functionality in open-source codebases into executable reinforcement-learning environments using source code as the only task-specific input. Agents identify public interfaces, write behavioral specifications, remove the target implementation, generate execution-grounded tests from the original program, and apply leakage checks, verifier audits, execution-consistency checks, and rollout filtering. The resulting corpus contains 5,545 verifiable tasks from 3,185 codebases across 23 programming languages and 15 technical domains. Training Xiaomi’s MiMo-V2.5 with GRPO and binary execution rewards improved every external benchmark: DeepSWE rose from 10.0% to 21.7%, ProgramBench’s Almost Solved score increased from 4.5 to 21.5, and Terminal-Bench v2.1 pass rate climbed from 63.7% to 72.2%; gains also appeared on SWE-bench Pro and RepoZero C2Rust. Scaling high-quality task pools from 1k to 3k to 5,545 produced progressively stronger results, while a filtered 3k subset outperformed an unfiltered 8k sample, showing that verifier reliability matters more than raw task count. Trajectory analysis found that RL-trained agents explored more before editing and diversified self-verification, with agent-written checks associated with a 4.2 percentage-point higher pass rate. The central result is that source code itself can supply scalable, reliable RL environments for coding agents across issue repair, program construction, translation, and terminal work.
Original abstract
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.