NTH

Invisible Shortcuts: Why Vision Encoders Know Your Camera

AuthorsVladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos, Noa Garcia, Giorgos Tolias

August 8, 2026 2 min read
Watch on YouTube
The one-line take

Vision models may recognize not just what is in an image, but also clues about which camera or processing pipeline produced it.

Key results

0.067
ImageNet21k JPEG-semantic correlation

Cramér’s V, compared with 0.047 for ImageNet1k.

40M
Re-LAION-2B Exif subset

Images used to study acquisition-metadata correlations.

6238
Caption topic count

Topics extracted from Re-LAION-2B captions.

0.396
Strong acquisition correlation

Cramér’s V for the stronger 6.4M-image Re-LAION-2B subset.

60.4
Generated-image detection average

Accuracy for the maximally correlated model, versus 57.3 for the uncorrelated model.

30.0
ImageNet-Sketch after mitigation

Accuracy after mitigation, up from 26.4.

What the paper found

Researchers from Czech Technical University in Prague and the University of Osaka show that vision encoders exploit invisible shortcuts left by cameras and image-processing pipelines. JPEG quality, resizing, sharpening, focal length, aperture, ISO, and camera model leave pixel-level traces that become predictive when correlated with semantic supervision. In ImageNet1k and ImageNet21k, JPEG–class correlations measured by Cramér’s V are 0.047 and 0.067, while models trained on the more correlated ImageNet21k encode metadata more strongly and suffer greater semantic distraction. Controlled ResNet50 experiments confirm a causal relationship: increasing JPEG–label correlation raises metadata prediction and harms semantic generalization, and the effect transfers to unseen metadata attributes. The same pattern appears in caption-based vision-language training: the authors analyze 40M Exif-tagged images from Re-LAION-2B, organize captions into 6238 topics, and train CLIP-loss models on 6.4M-image subsets whose topic–camera correlations reach 0.396, compared with 0.255 and 0.166 for baseline and weaker settings. Mitigation works during and after pretraining: DINOv2-style color jitter, blur, and grayscale reduce sensitivity, while a post-hoc adversarial linear layer suppresses JPEG traces and generalizes to other metadata without sacrificing semantic utility. The trade-off is revealing: metadata-sensitive models are better generated-image detectors, with average accuracy rising from 57.3 for the uncorrelated model to 60.4 at maximal correlation, but removing these shortcuts improves out-of-distribution robustness, including ImageNet-Sketch accuracy from 26.4 to 30.0. The findings also suggest Stable Diffusion images can inherit semantic metadata traces from their training data.

Original abstract

Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: https://github.com/ryan-caesar-ramos/visual-encoder-traces

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis