NTH

Characterizing the Quality Profile of AI-Generated C++ in Production

AuthorsMichael Tran, Fred Lewis, Kun Yang, Saksham Thakur, Aditya Kini, Aditya Patil, Milad Hashemi, Parthasarathy Ranganathan

August 11, 2026 3 min read
Watch on YouTube
The one-line take

A massive real-world study finds that AI-written C++ can increase maintenance and compute costs, but targeted feedback helps make the code more efficient.

Key results

10.46M
C++ lines analyzed

C++ lines with informative provenance used for static-analysis comparisons.

5%
Compute growth gap

Approximate higher compute growth for AI-heavy functions.

8%
Memory growth gap

Approximate higher memory growth for AI-heavy functions.

31%
Feedback efficiency improvement

Improvement in the efficiency score from taxonomy-informed feedback.

What the paper found

This study follows AI-generated C++ through authoring, review, static analysis, deployment, and runtime monitoring, addressing a gap left by benchmark evaluations of systems such as GitHub Copilot and ChatGPT. Across 3.52M submitted changes and 10.46M provenance-labeled C++ lines in a large enterprise monorepo, AI-generated code developed a distinct quality profile: its static-analysis burden concentrated in Interface and Coupling Burden and Copy and Allocation Overhead, while source code used explicit loops about 2.0x as often and standard-library or API calls about 0.4x as often as human-written code. These patterns were associated with higher review friction and production resource costs: AI-heavy functions reached approximately 5% higher compute growth and 8% higher memory growth than human-heavy functions. The analysis also found that generated code was more likely to trigger initial build instability but less likely to be reverted, suggesting that the dominant penalty is chronic maintenance and efficiency overhead rather than catastrophic failure. A controlled intervention reimplemented 50 C++ functions under progressively richer prompts; taxonomy-informed feedback raised the efficiency score to 0.385, a 31% improvement over untargeted prompting, and reduced targeted static-analysis warnings by 11.1%. The results argue that production evaluation must combine authoring-time provenance, category-level static findings, code-review outcomes, and deployed compute measurements, while also showing that targeted feedback can mitigate recurring weaknesses in AI-generated C++.

Original abstract

The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent studies point to a growing trade-off, revealing persistent challenges with code quality and maintainability. Industry leaders, including frontier AI labs, echo these concerns. As large language models are increasingly relied upon to author production code, understanding their impact on shipped software quality has become a critical priority. However, assessing these effects in industrial workflows remains difficult due to observability barriers. We study the impact of AI-generated code on production quality within a large enterprise operating global products relied upon by billions of users daily. Driven by this scale and user trust, the organization values code quality and has built thorough observability for every line of code deployed into production, enabling us to overcome measurement barriers to assess these effects. This study presents a large-scale empirical analysis of AI-generated C++ code from April 2025 to April 2026, tracking 3.52 million code changes across this enterprise's brownfield codebase. The core purpose is to understand the quality, performance, and maintenance characteristics of AI-generated code compared to human-written code in a production environment at scale. We find that AI-generated C++ code has a distinct quality profile, showing higher rates of interface and coupling burdens, copy and allocation overheads, and a reliance on explicit loops over optimized standard APIs. These issues translate into tangible downstream costs, including increased review effort and a 5-8% increase in compute resource consumption. However, we demonstrate that providing models with targeted, taxonomy-informed feedback can mitigate these effects, leading to an 11.1% reduction in targeted static analysis warnings and improved computational efficiency.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis