NTH

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

AuthorsYitong Zhang, Shiteng Lu, Jia Li

July 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that grammar constraints meant to make code generation safer can actually be exploited to jailbreak LLMs into producing malicious code, and it introduces a defense that resists this attack.

Key results

10
models evaluated

CodeSpear and CodeShield were tested across 10 LLMs including GPT-5, MiniMax, Qwen2.5, and LLaMA3 variants.

54.92%
average ASR baseline

Average attack success rate for Vanilla across RMCBench and MalwareBench.

81.82%
average ASR CodeSpear

Average attack success rate achieved by CodeSpear across 10 models.

36.61%
average MR baseline

Average malicious rate for Vanilla across RMCBench and MalwareBench.

54.58%
average MR CodeSpear

Average malicious rate achieved by CodeSpear across 10 models.

What the paper found

This paper shows that grammar-constrained decoding, a reliability feature now built into mainstream inference stacks such as vLLM, SGLang, OpenAI, and Fireworks AI, can become a jailbreak channel for code generation LLMs. The attack, CodeSpear, works by forcing an aligned model to answer malicious prompts under a benign programming-language grammar, which removes natural-language refusal from the valid output space and pushes the model into producing grammar-valid malicious code. Across 10 models, including GPT-5, GPT-5-mini, GPT-OSS-120B, MiniMax-M2.5, MiniMax-M2.7, Qwen2.5-Coder-7B, Qwen2.5-Coder-32B, Qwen2.5-7B, Qwen2.5-32B, and LLaMA3-8B, CodeSpear raises average attack success rate from 54.92% to 81.82% and malicious rate from 36.61% to 54.58% on RMCBench and MalwareBench. The defense, CodeShield, adapts DPO to the code modality by training on preference triples that rank natural-language refusals above structurally diverse honeypot code, and honeypot code above harmful code. On Qwen2.5-Coder-7B, Qwen2.5-7B, and LLaMA3-8B, CodeShield cuts average ASR under CodeSpear from 83.11% to 5.57% on Qwen2.5-Coder-7B and from 66.74% to 7.87% on LLaMA3-8B, while preserving benign code utility on HumanEval and MBPP with only minor pass@3 drops such as 78.00% to 77.00% on MBPP. The key takeaway is that safety alignment confined to natural language is brittle when decoding is grammar-restricted, and code-modality alignment is required.

Original abstract

Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code. Meanwhile, Grammar-Constrained Decoding (GCD) has been widely adopted to improve the reliability of LLM-generated code by enforcing syntactic validity. In this paper, we reveal a counterintuitive risk: this reliability-oriented technique can itself become an attack surface. We uncover a new jailbreak attack, termed CodeSpear, that exploits GCD to induce LLMs into generating malicious code. Our experiments show that simply applying a benign code grammar constraint can effectively jailbreak LLMs. To address this vulnerability, we propose CodeShield, a safety alignment approach that robustly preserves safe behavior even under attacker-controlled grammar constraints. CodeShield aligns the model in the code modality by teaching it to generate honeypot code under GCD. Such code is semantically harmless, so it does not implement the malicious request, and structurally diverse, so it is difficult to suppress through grammar tightening. At the same time, CodeShield still preserves natural-language refusals when natural language is available. Experiments on 10 popular LLMs across 4 benchmarks show that CodeSpear outperforms representative jailbreak baselines and increases the attack success rate by more than 30 percentage points on average. CodeShield also restores safety under CodeSpear while preserving benign utility. Our findings reveal a fundamental risk of GCD and call for greater attention to its potential security implications.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis