NTH

Statistically Undetectable Backdoors in Deep Neural Networks

AuthorsAndrej Bogdanov, Alon Rosen, Neekon Vafa

July 14, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that a malicious trainer can secretly implant backdoors in neural networks that are essentially impossible to detect from the model’s weights alone.

Key results

Õ(√m/n)
TV distance scale

Statistical deviation between honest and backdoored model descriptions

O(√n·2^{-n/m})
Backdoor magnitude

Upper bound on ∥Az∥∞ for the planted backdoor vector

2
Backdoor strength growth

Strength grows exponentially in the compression ratio n/m

√m/n
Tightness lower bound

Efficient distinguisher advantage for the planted vs honest matrix distributions

89
Fashion-MNIST accuracy

Test accuracy of the embedding model before distribution shift

86.5
Shifted accuracy

Test accuracy after scaling inputs to keep backdoored images in range

What the paper found

Andre Boganov, Alon Rosen, and Neekon Vafa show that a large class of deep feedforward neural networks can hide a backdoor that is statistically undetectable even in the white-box setting, meaning the backdoored and honest models are close in total variation distance despite full access to all weights. Their construction plants a secret vector z in the first frozen compressing Gaussian layer, with a rejection-sampled or conditionally sampled matrix A satisfying ∥Az∥∞ ≈ O(√n·2^{-n/m}), so the model owner can always create a colliding input x′ = x + z that produces unusually close outputs, while an efficient outsider cannot find comparable collisions under standard lattice-based cryptographic assumptions. The main theorem proves that any efficient trainer for networks with a Gaussian first layer, bi-Lipschitz downstream layers such as LeakyReLU, and discrete bounded inputs can be modified into a backdoored trainer with statistical deviation ϵ = Õ(√m/n) and backdoor strength growing exponentially in the compression ratio n/m. They also show tightness up to logarithmic factors, with a distinguisher achieving Ω(√m/n) advantage, and demonstrate a proof-of-concept on Fashion-MNIST, where a 784-to-256 embedding network reaches about 89% test accuracy, dropping to about 86.5% under the backdoor-induced distribution shift.

Original abstract

We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis