Security Tests as Executable Specifications for LLM Code Generation: Benefits, Trade-offs, and Coverage Limits
AuthorsYunhao Liang, Chengguang Gan, Ruixuan Ying
Resources
Security tests can help LLMs repair vulnerable code, but their success depends heavily on test coverage, feedback design, and the model-task combination.
Key results
Primary study scale across benchmarks and model conditions
Average increase in hidden functional-and-security joint success
Initially unsuccessful shared candidates repaired
Hidden joint regressions among structured-feedback repairs
Initially unsuccessful shared candidates repaired with fixed raw logs
Visible-joint candidates failing hidden joint evaluation
What the paper found
This paper introduces SecTDD, a controlled scaffold that uses security tests as executable specifications before code generation and as feedback during repair, while separating test visibility, revision triggers, and failure representation. Across 2705 trajectories covering 31 task instances, three benchmarks, 16 CWE categories, and Qwen models from Qwen2.5-Coder-7B-Instruct through Qwen3.6-27B plus DeepSeek-V4-Flash, showing all visible tests upfront increased hidden functional-and-security joint success by 19.3 percentage points on average, but helped only seven of nine benchmark–model conditions and harmed two. Shared-candidate comparisons provide stronger causal evidence: structured feedback repaired 80 initially unsuccessful candidates with 0 joint regressions, while fixed raw logs repaired 83 with 3 regressions. Structured feedback did not consistently outperform raw logs; head-to-head, it won 6 cells, lost 6, and tied 453. The central limitation is oracle coverage: candidates passing every visible test still failed hidden behavior families, with the highest visible-to-hidden failure rate reaching 18.8 percent. SecTDD therefore is not presented as a universally superior prompting method. Its contribution is mechanistic: executable feedback can repair secure code across model families, but benefits depend on the model, task, entry point, and whether visible tests expose the hidden vulnerability behavior.
Original abstract
Large language models (LLMs) can generate functionally useful code that remains vulnerable, while security-focused interventions may break intended behavior. We investigate security tests as executable specifications both before generation and during iterative repair. We develop SecTDD, a controlled test-feedback scaffold that separates three factors: whether tests are shown upfront, whether failed executions trigger revision, and how failures are selected and represented. The evaluation uses behavior-partitioned visible and hidden tests and byte-identical initial candidates for repair comparisons. Across 2,705 trajectories, 31 task instances, three secure-code benchmarks, 16 CWE categories, and two model families, showing all visible tests upfront increases hidden functional-and-security joint success by 19.3 percentage points on average, but improves only seven of nine benchmark-model conditions and harms two. In shared-candidate comparisons, structured feedback repairs 80 initially unsuccessful candidates with no joint regressions; fixed raw feedback repairs 83 but causes three regressions. Structured and raw feedback are otherwise nearly indistinguishable head-to-head (six wins, six losses, and 453 ties). Candidates that pass all visible tests still fail hidden behavior families under every common regime. These results show that executable feedback can repair secure-code generation, but its benefits depend on the model, task, feedback entry point, and especially test coverage.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.