Repo0: Design-Driven Zero-to-All Code Generation
AuthorsSilin Chen, Haoyi Teng, Xiaodong Gu, Yuling Shi, Jiale Huang, Yongpan Wang, Hongyu Zhang, Haibing Guan
Resources
Repo0 helps coding agents build entire software repositories by first evolving and stabilizing an explicit modular architecture before writing the code.
Key results
Percentage-point improvement over RPG.
Percentage-point improvement over RPG.
Voting Rate achieved with metrics-guided structural convergence.
What the paper found
Repo0 reframes zero-to-all code generation as continuous software-architecture evolution rather than one-shot planning. Starting from natural-language requirements, it builds a Dual-DAG: one directed acyclic graph represents functional coordination among requirements, another represents implementation dependencies among components, and an alignment relation preserves traceability between them. The agent iteratively applies split, merge, revise, save, and add actions, using cohesion, Jaccard-based coupling, connectivity, and requirement-coverage checks to refine component boundaries until structural convergence. Only then does it generate packages and files through dependency-aware, test-driven development, creating importable skeletons, synthesizing tests, implementing code, and using validation failures for localized repair. On the RepoCraft benchmark’s six real-world Python repositories, evaluated with OpenAI’s GPT-5 mini and DeepSeek V3.2, Repo0 achieved the best Functionality Coverage and Pass Rate in every reported setting. Against RPG, the strongest static-graph baseline, it improved Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. In a statsmodels setting, metrics-guided evolution raised Functionality Coverage from 75.90% to 80.68%, Pass Rate from 81.90% to 85.51%, and Voting Rate to 98.65%. Ablations show structural evolution is the largest contributor, while removing the Dual-DAG, requirement context, or dependency-aware ordering also reduces performance; unconstrained LLM-directed refinement instead tends to over-decompose repositories and lower correctness.
Original abstract
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.