NTH

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

AuthorsSaisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu, Pratinav Seth

August 22, 2026 3 min read
Watch on YouTube
The one-line take

The paper shows that many AI compliance detectors may ignore the actual rules they are supposed to enforce, and offers a lightweight way to audit them.

Key results

10
ICS calibration size

Labelled pairs used to calibrate the training-free activation direction

0.952
ICS OmniCompliance AUROC

In-domain AUROC on outcome-ablated OmniCompliance

0.728
Leave-one-distribution-out ICS AUROC

Mean AUROC on distributions withheld from calibration

0.557
Budget-matched null AUROC

Null performance used to assess calibrate-once transfer

74.4%
Crossed benchmark reasoning exact match

Step-by-step judge performance on four-cell rule–scenario quadruples

5.2
IFEval selection improvement

Percentage-point gain in mechanically verified pass rate from ICS-guided ranking

What the paper found

This paper audits whether compliance detectors evaluate a written rule or merely recognize that a scenario looks risky. It identifies “rule blindness”: deleting, shuffling, or replacing the governing rule leaves detection performance effectively unchanged across activation probes and guards, including Llama Guard 3, Qwen3Guard, WildGuard, and Latent Policy Guard. The proposed Internal Compliance Score, or ICS, is a training-free difference-of-means direction over the monitored model’s residual-stream activations, calibrated with 10 labelled pairs and scored with one projection. On outcome-ablated OmniCompliance, ICS reaches 0.952 AUROC, but its pre-registered advantage over trivial baselines passes in only 11 of 20 domains, while four of seven public compliance benchmarks are lexically degenerate. In a crossed benchmark with 200 rule–scenario templates, cheap detectors remain near chance because neither the rule nor scenario alone predicts the label; a step-by-step judge reaches 0.849 AUROC and 74.4% quadruple exact match, showing that explicit composition is possible but absent from efficient one-pass detectors. ICS transfers across held-out distributions at 0.728 AUROC against a 0.557 null, yet pooled TF-IDF matches that mean, limiting claims of representation-level understanding. Its practical use is candidate ranking: on IFEval, ICS improves mechanically verified pass rates by 5.2 percentage points, though a white-box adaptive attack defeats the mechanism. The results suggest that systems such as OpenAI Moderation, GPT-4o, Claude Haiku, and Gemini should be evaluated with counterfactual rule tests rather than accuracy alone.

Original abstract

Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard and activation probe we test, including a policy-conditioned guard that correctly cites the governing clause yet barely changes its verdict when that clause is swapped for its permissive counterpart. A purpose-built benchmark crossing two rules with two scenarios, so that neither alone predicts the label, confirms the failure under a design no prior benchmark rules out, and shows that step by step reasoning, not any fast detector we test, is what escapes it. Auditing at scale requires a retraining-free detector, so we introduce the Internal Compliance Score (ICS): a training-free activation readout calibrated from ten labelled pairs and scored by a single projection. We hold ICS to the same scrutiny as the guards it audits: a pre-registered criterion for beating trivial baselines is not met, and a bag-of-words model matches its pooled generalisation exactly. It remains useful because it is inexpensive, letting us audit four deployed guard models, an 8B zero-shot judge, and thirteen benchmarks, and it raises the mechanically verified pass rate when used to rank candidate responses, though an adaptive white-box attack removes this gain. We release the counterfactual protocol and crossed-rule benchmark so rule blindness can be tested in future probe and guard claims.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis