NTH
Research collection

Code Generation research

Research on models that write, repair, and reason about software. Compare programming benchmarks, execution feedback, and developer-facing capabilities.

43 papers · Latest edition October 4, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Code Generation papers

Newest editions first.

02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis
04Code Generation

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo

CodeMidas turns existing codebases into scalable, automatically verified RL environments that train coding agents to perform better across diverse software tasks.

Read analysis
06Code Generation

Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering

Alexander Krentsel, Shubham Agarwal, Mert Cemri, Shu Liu, Sidharth Sankhe, Ziming Mao, Matei Zaharia, Ion Stoica

AI coding agents can pass tests yet fail in the real world, so dependable development requires continuously narrowing the gaps between human intent, evaluation models, and deployment reality.

Read analysis
07Code Generation

ExecCritic: Learn to Test, Test to Improve for Coding Agents

Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao

ExecCritic trains one coding agent to create trustworthy tests and another to repair code from them, substantially improving repository-level bug fixing.

Read analysis
08Code Generation

When LLM Decompilers Recompile More and Preserve Less

Chang Liu, Edward Raff, Kristopher Micinski

This work shows that LLM decompilers can produce code that compiles and passes known tests while silently changing program behavior or erasing vulnerabilities.

Read analysis
10Code Generation

CogEvol: Towards Efficient and Reliable Learning Environment Generation

Shangqing Tu, Daniel Zhang-Li, Yucheng Wang, Shiyu Gan, Yanpeng Wang, Huiqiang Rong, Mofei Chen, Shen Yang, Yini Chen, Yinuo Duan, Haoxuan Li, Binglin Liu, Ye He, Danqi Zheng, Zhanxin Hao, Yuxuan Wu, Mengting Tao, Yuqiu Liu, Jifan Yu, Juanzi Li, Bin Xu, Lei Hou, Huiqin Liu, Yu Zhang

CogEvol is a compact, production-trained model that turns course briefs into reliable slides and interactive educational web pages much faster and more cheaply than larger coding models.

Read analysis
13Code Generation

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, Haoli Bai

LEGO-RL turns real coding-agent harnesses into trainable reinforcement-learning environments and improves SWE-bench performance across three popular platforms.

Read analysis
14Code Generation

Repo0: Design-Driven Zero-to-All Code Generation

Silin Chen, Haoyi Teng, Xiaodong Gu, Yuling Shi, Jiale Huang, Yongpan Wang, Hongyu Zhang, Haibing Guan

Repo0 helps coding agents build entire software repositories by first evolving and stabilizing an explicit modular architecture before writing the code.

Read analysis
15Code Generation

Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation

Tianyi Xiong, Zhengyuan Yang, Xiaofei Wang, Chung-Ching Lin, Ruichun Ma, Kevin Lin, Zhendong Wang, Linjie Li, Chenxi Liu, Ruibo Chen, Ramani Duraiswami, Heng Huang, Lijuan Wang

RubSE helps vision-language models iteratively repair generated interfaces by turning visual feedback into prioritized, structured rubrics instead of letting each code edit disrupt the whole page.

Read analysis
20Code Generation

Characterizing the Quality Profile of AI-Generated C++ in Production

Michael Tran, Fred Lewis, Kun Yang, Saksham Thakur, Aditya Kini, Aditya Patil, Milad Hashemi, Parthasarathy Ranganathan

A massive real-world study finds that AI-written C++ can increase maintenance and compute costs, but targeted feedback helps make the code more efficient.

Read analysis
26Code Generation

KAT-Coder-V2.5 Technical Report

Bo Huang, Fengxiang Li, Hao Xu, Haoyang Huang, Hongyi Fu, Jinhua Hao, Kun Yuan, Minglei Zhang, Pengcheng Xu, Shiyang Liu, Wenhao Zhuang, Yuze Shi, Zongxian Feng, Chao Wang, Cheng He, Chongling Rao, Deyu Cao, Fan Yang, Gang Xiong, Haochen Liu, Jiabao Li, Jian Liang, Jinghui Jia, Jingwen Chang, Jun Du, Junyu Shi, Min Li, Mingqi Wu, Qiang Gao, Shangpeng Yan, Shaotong Qi, Shu Xu, Shuo Zhou, Tiankuo Xu, Tong Zheng, Weilun Zhao, Xiancheng Meng, Xianda Sun, Xiaoyu Jiang, Xunhao Jia, Yao Xia, Yimeng Xu, Yinghan Cui, Yingpeng Chen, Yiwen Ning, Yong Wang, Yuxuan Sun, Zhongsheng Liu, Ming Sun, Cheng Luo, Chen Yang, Han Li, Kun Gai

KAT-Coder-V2.5 trains coding agents to autonomously modify and test real software repositories using scalable environments, tool-use trajectories, and reinforcement learning.

Read analysis
33Code Generation

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Ruoyu Wang, Jierun Chen, Shaowei Wang, Chaofan Tao, Sidi Yang, Yuxin Jiang, Kim-Hui Yap, Lifeng Shang, Xiaohui Li, Haoli Bai

This paper teaches coding agents not just to write pull requests, but to review, critique, and revise them in a loop that more closely matches real software development.

Read analysis
35Code Generation

LLM Agents Can See Code Repositories

Dongjian Ma, Silin Chen, Yufei Yang, Yulin Shi, Yanfu yan, Xiaodong Gu

This paper shows that coding agents work better when they can 'see' a repository’s structure, and that combining visual graphs with text can cut token use while preserving or improving bug-fixing performance.

Read analysis
42Code Generation

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

Chuan Xiao, Zhengbo Jiao, Shaobo Wang, Wei Wang, Bing Zhao, Hu Wei, Linfeng Zhang, Lin Qu

This paper shows how coding agents can learn from their own past mistakes by turning old solving traces into new training tasks, steadily improving software-engineering performance over multiple iterations.

Read analysis