NTH

Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents

AuthorsA H M Nazmus Sakib, Dipayan Banik, Murtuza Jadliwala

July 17, 2026 2 min read
Watch on YouTube
The one-line take

This study finds that autonomous coding workflows frequently contain overlooked security problems, especially supply-chain issues and leaked credentials, calling for stronger safeguards at the human-AI collaboration boundary.

Key results

38.9%
PRs with security smells

Share of 4,022 analyzed agent-generated pull requests containing at least one security smell.

82.3%
Supply-chain integrity share

Proportion of all detected security smells attributed to supply-chain integrity issues.

99.6%
Critical smells from hard-coded credentials

Share of critical-severity smells classified as hard-coded credentials.

67.6%
Genuine secrets introduced by humans

Share of 74 confirmed genuine credentials committed by human collaborators.

81.1%
Genuine secrets missed before integration

Share of confirmed credentials lacking a bot or human review comment before integration.

0.836
LLM judge F1

F1 score against the manually annotated validation gold standard.

What the paper found

The paper “Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents,” by researchers at the University of Texas at San Antonio and Danovo Energy Solutions, examines security smells introduced by autonomous coding systems including GitHub Copilot, OpenAI Codex, Claude Code, Cursor, and Devin. Using the AIDev dataset, the authors analyzed 16,112 added file changes across 4,022 pull requests, applying a six-category taxonomy grounded in OWASP guidance, Center for Internet Security benchmarks, and GitHub hardening practices. An LLM-as-a-judge pipeline based on quantized Qwen3.6-35B-A3B-FP8 and Google DeepMind’s Gemma-4-26B-A4B-IT-FP8 was validated against manual annotations. The central finding is that 38.9% of agent-generated pull requests contained at least one security smell. Supply-chain integrity problems, such as mutable action or container-image tags and unpinned installations, represented 82.3% of all detected smells, while hard-coded credentials accounted for 99.6% of critical-severity findings. Manual investigation confirmed 74 genuine credentials; human collaborators introduced 67.6% of them, compared with agents introducing the remainder, and reviewers or security bots failed to comment on 81.1% before integration. The judge achieved 0.836 F1 on the validation sample, but its 0.775 recall indicates that reported prevalence is likely an underestimate. The authors conclude that agent-generated code requires context-aware, point-of-collaboration guardrails rather than reliance on conventional review pipelines alone.

Original abstract

The increasing adoption of autonomous coding agents accelerates software development but also introduces scoped security risks within high-impact file paths that can outpace traditional human review capacity. While prior research has primarily evaluated these systems in terms of functional correctness and productivity, this paper presents a large-scale empirical study using the AIDev dataset to systematically characterize security code smells in agent-generated pull requests (PRs). Through a combination of a validated LLM-as-a-judge framework and manual qualitative analysis, we identify and classify security misconfigurations across 16,112 file changes spanning 4,022 pull requests. Our results reveal that 38.9% of agent-generated PRs contain at least one security smell, with supply chain integrity issues accounting for 82.3% of all detected security smells. Furthermore, hard-coded credentials constitute 99.6% of all critical-severity security smells. Crucially, we find that human collaborators are responsible for introducing 67.6% of genuine leaked secrets within these agent-assisted workflows, while existing automated and human review processes fail to detect 81.1% of these credentials prior to integration. These findings highlight substantial security risks in agent-assisted software development workflows and suggest a potential reduction in developer vigilance. They also underscore the urgent need for context-aware security guardrails implemented directly at the point of human-AI collaboration.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis