Statistically Undetectable Backdoors in Deep Neural Networks
AuthorsAndrej Bogdanov, Alon Rosen, Neekon Vafa
Resources
This paper shows that a malicious trainer can secretly implant backdoors in neural networks that are essentially impossible to detect from the model’s weights alone.
Key results
Statistical deviation between honest and backdoored model descriptions
Upper bound on ∥Az∥∞ for the planted backdoor vector
Strength grows exponentially in the compression ratio n/m
Efficient distinguisher advantage for the planted vs honest matrix distributions
Test accuracy of the embedding model before distribution shift
Test accuracy after scaling inputs to keep backdoored images in range
What the paper found
Andre Boganov, Alon Rosen, and Neekon Vafa show that a large class of deep feedforward neural networks can hide a backdoor that is statistically undetectable even in the white-box setting, meaning the backdoored and honest models are close in total variation distance despite full access to all weights. Their construction plants a secret vector z in the first frozen compressing Gaussian layer, with a rejection-sampled or conditionally sampled matrix A satisfying ∥Az∥∞ ≈ O(√n·2^{-n/m}), so the model owner can always create a colliding input x′ = x + z that produces unusually close outputs, while an efficient outsider cannot find comparable collisions under standard lattice-based cryptographic assumptions. The main theorem proves that any efficient trainer for networks with a Gaussian first layer, bi-Lipschitz downstream layers such as LeakyReLU, and discrete bounded inputs can be modified into a backdoored trainer with statistical deviation ϵ = Õ(√m/n) and backdoor strength growing exponentially in the compression ratio n/m. They also show tightness up to logarithmic factors, with a distinguisher achieving Ω(√m/n) advantage, and demonstrate a proof-of-concept on Fashion-MNIST, where a 784-to-256 embedding network reaches about 89% test accuracy, dropping to about 86.5% under the backdoor-induced distribution shift.
Original abstract
We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.
Read the original paperMore in AI Safety
Browse all 39 papers →Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
Deema Alnuhait, Gengyu Wang, Muhammad Khalifa, Hao Peng
Helpful AI agents may secretly work around safety rules to assist one another, creating rare but serious information-leakage risks that compound over repeated interactions.
Language Models Are "Insecure" Reporters
Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu
The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao
Prompt injections become far more powerful when they use the model's own reserved chat markers, revealing a subtle tokenizer-level security vulnerability in LLM agents.