NTH

Repo0: Design-Driven Zero-to-All Code Generation

AuthorsSilin Chen, Haoyi Teng, Xiaodong Gu, Yuling Shi, Jiale Huang, Yongpan Wang, Hongyu Zhang, Haibing Guan

August 28, 2026 2 min read
Watch on YouTube
The one-line take

Repo0 helps coding agents build entire software repositories by first evolving and stabilizing an explicit modular architecture before writing the code.

Key results

20.08
Maximum Functionality Coverage improvement

Percentage-point improvement over RPG.

29.74
Maximum Pass Rate improvement

Percentage-point improvement over RPG.

98.65%
Evolved statsmodels Voting Rate

Voting Rate achieved with metrics-guided structural convergence.

What the paper found

Repo0 reframes zero-to-all code generation as continuous software-architecture evolution rather than one-shot planning. Starting from natural-language requirements, it builds a Dual-DAG: one directed acyclic graph represents functional coordination among requirements, another represents implementation dependencies among components, and an alignment relation preserves traceability between them. The agent iteratively applies split, merge, revise, save, and add actions, using cohesion, Jaccard-based coupling, connectivity, and requirement-coverage checks to refine component boundaries until structural convergence. Only then does it generate packages and files through dependency-aware, test-driven development, creating importable skeletons, synthesizing tests, implementing code, and using validation failures for localized repair. On the RepoCraft benchmark’s six real-world Python repositories, evaluated with OpenAI’s GPT-5 mini and DeepSeek V3.2, Repo0 achieved the best Functionality Coverage and Pass Rate in every reported setting. Against RPG, the strongest static-graph baseline, it improved Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. In a statsmodels setting, metrics-guided evolution raised Functionality Coverage from 75.90% to 80.68%, Pass Rate from 81.90% to 85.51%, and Voting Rate to 98.65%. Ablations show structural evolution is the largest contributor, while removing the Dual-DAG, requirement context, or dependency-aware ordering also reduces performance; unconstrained LLM-directed refinement instead tends to over-decompose repositories and lower correctness.

Original abstract

Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly from natural-language requirements while maintaining a modular repository architecture throughout development. We present Repo0, a continuous structural evolution framework for zero-to-all code generation. Repo0 maintains an explicit architectural state instantiated as a Dual-Directed-Acyclic-Graph (Dual-DAG), consisting of a requirement-level DAG, a component-level DAG, and their alignment relation. Starting from natural-language requirements, it iteratively evolves component boundaries through structural actions guided by modularity metrics until structural convergence, after which the converged architecture guides test-driven development code generation. We evaluate Repo0 on six real-world repositories from RepoCraft using GPT-5 mini and DeepSeek V3.2. Repo0 achieves the highest Functionality Coverage and Pass Rate across all settings. Compared with RPG, the strongest repository-planning baseline, Repo0 improves Functionality Coverage by up to 20.08 percentage points and Pass Rate by up to 29.74 percentage points. Ablation and structural-evolution analyses further demonstrate the importance of the Dual-DAG architectural state, modularity-guided structural evolution, and explicit structural convergence.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis