What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
AuthorsSaisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth
Resources
The paper shows that many AI compliance detectors may ignore the actual rules they are supposed to enforce, and offers a lightweight way to audit them.
Key results
Labelled pairs used to calibrate the training-free activation direction
In-domain AUROC on outcome-ablated OmniCompliance
Mean AUROC on distributions withheld from calibration
Null performance used to assess calibrate-once transfer
Step-by-step judge performance on four-cell rule–scenario quadruples
Percentage-point gain in mechanically verified pass rate from ICS-guided ranking
What the paper found
This paper audits whether compliance detectors evaluate a written rule or merely recognize that a scenario looks risky. It identifies “rule blindness”: deleting, shuffling, or replacing the governing rule leaves detection performance effectively unchanged across activation probes and guards, including Llama Guard 3, Qwen3Guard, WildGuard, and Latent Policy Guard. The proposed Internal Compliance Score, or ICS, is a training-free difference-of-means direction over the monitored model’s residual-stream activations, calibrated with 10 labelled pairs and scored with one projection. On outcome-ablated OmniCompliance, ICS reaches 0.952 AUROC, but its pre-registered advantage over trivial baselines passes in only 11 of 20 domains, while four of seven public compliance benchmarks are lexically degenerate. In a crossed benchmark with 200 rule–scenario templates, cheap detectors remain near chance because neither the rule nor scenario alone predicts the label; a step-by-step judge reaches 0.849 AUROC and 74.4% quadruple exact match, showing that explicit composition is possible but absent from efficient one-pass detectors. ICS transfers across held-out distributions at 0.728 AUROC against a 0.557 null, yet pooled TF-IDF matches that mean, limiting claims of representation-level understanding. Its practical use is candidate ranking: on IFEval, ICS improves mechanically verified pass rates by 5.2 percentage points, though a white-box adaptive attack defeats the mechanism. The results suggest that systems such as OpenAI Moderation, GPT-4o, Claude Haiku, and Gemini should be evaluated with counterfactual rule tests rather than accuracy alone.
Original abstract
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart. A purpose-built benchmark crossing two rules with two scenarios, so that neither alone predicts the label, confirms the failure under a design no prior benchmark rules out, and shows that step by step reasoning, not any fast detector we test, is what escapes it. Auditing at scale requires a retraining-free detector, so we introduce the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten labelled pairs and scored by a single projection. We hold ICS to the same scrutiny as the guards it audits: a pre-registered criterion for beating trivial baselines is not met, and a bag-of-words model matches its pooled generalisation exactly. It remains useful because it is inexpensive, letting us audit four deployed guard models, an 8B zero-shot judge, and thirteen benchmarks, and it raises the mechanically verified pass rate when used to rank candidate responses, though an adaptive white-box attack removes this gain. We release the counterfactual protocol and crossed-rule benchmark so rule blindness can be tested in future probe and guard claims.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.