NTH

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

AuthorsTaebong Kim, Youngsik Hong, Minsik Kim, Sunyoung Choi, Jaewon Jang, Minseo Kim

September 1, 2026 2 min read
Watch on YouTube
The one-line take

A fast two-pass test reveals when sequence models secretly use future information, even when their attention masks look correct.

Key results

192
Injected-fault localization

The audit localized all 192 injected faults to the exact offending layer.

2
Forward passes

The audit requires two forward passes and no backward pass.

256
Zamba2 chunk boundary

Zamba2-1.2B leakage begins at chunk boundary 256.

128
Nemotron-H chunk boundary

NVIDIA’s Nemotron-H-8B leakage begins at chunk boundary 128.

What the paper found

The paper introduces prefix invariance as the defining causality condition for autoregressive sequence models: changing a future token must leave every earlier layer representation unchanged. Its audit uses per-layer forward hooks, a single-token suffix perturbation, and two forward passes with caching disabled; it requires no gradients, labels, training, or accelerator, and reports the first offending layer. This is broader than inspecting an attention mask, which can miss leakage caused by scans, aggregation, normalization, or unmasked state-space and recurrent mixers such as Mamba and RWKV. Across 192 injected-fault trials spanning attention, state-space, recurrent, and hybrid architectures, the audit exactly localized all 192 faults, while static mask inspection detected 0 of 192. The study also found a real implementation defect in the PyTorch chunked-scan paths for Zamba2-1.2B and NVIDIA’s Nemotron-H-8B: an inter-chunk recurrence reduced over the output-chunk axis instead of the input-chunk axis, allowing future information into earlier positions at chunk boundaries of 256 and 128, respectively. A reference-axis patch removed the leak. The failure was invisible when sequences stayed below the relevant chunk size, motivating a rule that audit lengths must exceed architectural chunk, window, and kernel parameters. The authors also require a positive-control liveness test on the same loaded checkpoint, after inert Falcon-H1 loads produced apparently clean but unauditable outputs; Microsoft’s Phi-4-mini-flash-reasoning was likewise among models blocked by unavailable custom kernels.

Original abstract

Hybrid sequence models must satisfy prefix invariance: representations at position t must not depend on future inputs, yet this is rarely verified. We formalize prefix invariance and give a lightweight audit, two forward passes, no training or gradients, yielding a per-layer score localizing where causality breaks. Attention-mask inspection, the field's default check, is incomplete: causality is a graph-level property, and leaks can occur via scans, aggregations, or normalization despite correct masks. Across 192 injected-fault trials on eight checkpoints, mask inspection detected none, while our audit localized all 192/192 to the exact layer. Static/dynamic analysis of chunked-scan code in transformers found the same defect in Zamba2 and Nemotron-H, an inter-chunk axis error fixed via the reference implementation. The method fits on one page and runs in seconds.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis