KAT-Coder-V2.5 Technical Report
AuthorsBo Huang, Fengxiang Li, Hao Xu, Haoyang Huang, Hongyi Fu, Jinhua Hao, Kun Yuan, Minglei Zhang, Pengcheng Xu, Shiyang Liu, Wenhao Zhuang, Yuze Shi, Zongxian Feng, Chao Wang, Cheng He, Chongling Rao, Deyu Cao, Fan Yang, Gang Xiong, Haochen Liu, Jiabao Li, Jian Liang, Jinghui Jia, Jingwen Chang, Jun Du, Junyu Shi, Min Li, Mingqi Wu, Qiang Gao, Shangpeng Yan, Shaotong Qi, Shu Xu, Shuo Zhou, Tiankuo Xu, Tong Zheng, Weilun Zhao, Xiancheng Meng, Xianda Sun, Xiaoyu Jiang, Xunhao Jia, Yao Xia, Yimeng Xu, Yinghan Cui, Yingpeng Chen, Yiwen Ning, Yong Wang, Yuxuan Sun, Zhongsheng Liu, Ming Sun, Cheng Luo, Chen Yang, Han Li, Kun Gai
Resources
KAT-Coder-V2.5 trains coding agents to autonomously modify and test real software repositories using scalable environments, tool-use trajectories, and reinforcement learning.
Key results
AutoBuilder produced 100K executable repository environments.
The environments span 12 programming languages.
AutoBuilder raised construction success from 16.5% to 57.2%.
KAT-Coder-V2.5 scored 65.2 on SWE-Bench Pro.
KAT-Coder-V2.5 achieved 94.9 on PinchBench, the best evaluated result.
Sandbox feedback errors were reduced from roughly 16% to below 2%.
What the paper found
The KwaiKAT Team’s KAT-Coder-V2.5 argues that autonomous coding agents are constrained less by parameter count than by training infrastructure: reproducible repositories, verifiable rewards, and high-quality trajectories. Its AutoBuilder pipeline reconstructs multilingual repositories in sandboxes, generates precise tasks from code and test patches, and verifies both fail-to-pass and pass-to-pass behavior. AutoBuilder produced 100K executable environments across 12 languages, raising environment-construction success from 16.5% to 57.2%. For general tool use, KwaiClawEnv combines executable services, controllable multi-step tasks, parallel rollouts, and hard-rule plus LLM-as-Judge filtering. Reinforcement learning adds randomized white-box and black-box harnesses, a reliability-hardened sandbox, asymmetric actor–critic PPO with hindsight-augmented value estimation, and process-aware rewards for search, testing, tool discipline, and partial progress. Sandbox feedback errors fell from roughly 16% to below 2%. Finally, Multi-Teacher On-Policy Distillation fuses five experts spanning software engineering, Claw tool use, terminal work, web coding, and general knowledge. Under a unified Claude Code harness, KAT-Coder-V2.5 scored 65.2 on SWE-Bench Pro and 94.9 on PinchBench, ranking second to Opus 4.8 on repository-level software engineering while achieving the best result on PinchBench. The report also compares GLM-5.1, GLM-5.2, Kimi-K2.6, Codex, OpenHands, and NVIDIA Polar-related infrastructure, positioning KAT-Coder-V2.5 as a systems-engineered alternative to frontier models from Anthropic and other labs.
Original abstract
We present KAT-Coder-V2.5, a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a single-turn code generator. Its capability is bottlenecked less by model scale than by the scarcity of reproducible environments, verifiable rewards, and high-value trajectories, which we address with an end-to-end agentic post-training framework. AutoBuilder reconstructs multilingual repositories into sandboxed environments with fail-to-pass and pass-to-pass verification at scale, from which we regenerate self-contained task specifications, recover near-miss trajectories, and distill supervision through process-aware filtering, while KwaiClawEnv synthesizes large-scale tool-use trajectories from executable services and real task seeds. We further scale reinforcement learning with harness randomization, a reliability-hardened sandbox, an asymmetric actor--critic PPO with hindsight-augmented value estimation, and a harness-oriented reward framework, and unify SWE, Agent-Claw, and WebCoding experts via Multi-Teacher On-Policy Distillation. Across six software-engineering and agentic benchmarks, KAT-Coder-V2.5 delivers the best agentic tool-use result on PinchBench and ranks second only to the frontier Opus 4.8 on repository-level software engineering. Our service is available at https://streamlake.com/product/kat-coder.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.