NTH

Security Tests as Executable Specifications for LLM Code Generation: Benefits, Trade-offs, and Coverage Limits

AuthorsYunhao Liang, Chengguang Gan, Ruixuan Ying

August 19, 2026 2 min read
Watch on YouTube
The one-line take

Security tests can help LLMs repair vulnerable code, but their success depends heavily on test coverage, feedback design, and the model-task combination.

Key results

2705
Evaluation trajectories

Primary study scale across benchmarks and model conditions

19.3%
Upfront-test improvement

Average increase in hidden functional-and-security joint success

80
Structured-feedback repairs

Initially unsuccessful shared candidates repaired

0
Structured-feedback regressions

Hidden joint regressions among structured-feedback repairs

83
Raw-feedback repairs

Initially unsuccessful shared candidates repaired with fixed raw logs

18.8%
Maximum visible-to-hidden failure rate

Visible-joint candidates failing hidden joint evaluation

What the paper found

This paper introduces SecTDD, a controlled scaffold that uses security tests as executable specifications before code generation and as feedback during repair, while separating test visibility, revision triggers, and failure representation. Across 2705 trajectories covering 31 task instances, three benchmarks, 16 CWE categories, and Qwen models from Qwen2.5-Coder-7B-Instruct through Qwen3.6-27B plus DeepSeek-V4-Flash, showing all visible tests upfront increased hidden functional-and-security joint success by 19.3 percentage points on average, but helped only seven of nine benchmark–model conditions and harmed two. Shared-candidate comparisons provide stronger causal evidence: structured feedback repaired 80 initially unsuccessful candidates with 0 joint regressions, while fixed raw logs repaired 83 with 3 regressions. Structured feedback did not consistently outperform raw logs; head-to-head, it won 6 cells, lost 6, and tied 453. The central limitation is oracle coverage: candidates passing every visible test still failed hidden behavior families, with the highest visible-to-hidden failure rate reaching 18.8 percent. SecTDD therefore is not presented as a universally superior prompting method. Its contribution is mechanistic: executable feedback can repair secure code across model families, but benefits depend on the model, task, entry point, and whether visible tests expose the hidden vulnerability behavior.

Original abstract

Large language models (LLMs) can generate functionally useful code that remains vulnerable, while security-focused interventions may break intended behavior. We investigate security tests as executable specifications both before generation and during iterative repair. We develop SecTDD, a controlled test-feedback scaffold that separates three factors: whether tests are shown upfront, whether failed executions trigger revision, and how failures are selected and represented. The evaluation uses behavior-partitioned visible and hidden tests and byte-identical initial candidates for repair comparisons. Across 2,705 trajectories, 31 task instances, three secure-code benchmarks, 16 CWE categories, and two model families, showing all visible tests upfront increases hidden functional-and-security joint success by 19.3 percentage points on average, but improves only seven of nine benchmark-model conditions and harms two. In shared-candidate comparisons, structured feedback repairs 80 initially unsuccessful candidates with no joint regressions; fixed raw feedback repairs 83 but causes three regressions. Structured and raw feedback are otherwise nearly indistinguishable head-to-head (six wins, six losses, and 453 ties). Candidates that pass all visible tests still fail hidden behavior families under every common regime. These results show that executable feedback can repair secure-code generation, but its benefits depend on the model, task, feedback entry point, and especially test coverage.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis