Invisible Shortcuts: Why Vision Encoders Know Your Camera
AuthorsVladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos, Noa Garcia, Giorgos Tolias
Vision models may recognize not just what is in an image, but also clues about which camera or processing pipeline produced it.
Key results
Cramér’s V, compared with 0.047 for ImageNet1k.
Images used to study acquisition-metadata correlations.
Topics extracted from Re-LAION-2B captions.
Cramér’s V for the stronger 6.4M-image Re-LAION-2B subset.
Accuracy for the maximally correlated model, versus 57.3 for the uncorrelated model.
Accuracy after mitigation, up from 26.4.
What the paper found
Researchers from Czech Technical University in Prague and the University of Osaka show that vision encoders exploit invisible shortcuts left by cameras and image-processing pipelines. JPEG quality, resizing, sharpening, focal length, aperture, ISO, and camera model leave pixel-level traces that become predictive when correlated with semantic supervision. In ImageNet1k and ImageNet21k, JPEG–class correlations measured by Cramér’s V are 0.047 and 0.067, while models trained on the more correlated ImageNet21k encode metadata more strongly and suffer greater semantic distraction. Controlled ResNet50 experiments confirm a causal relationship: increasing JPEG–label correlation raises metadata prediction and harms semantic generalization, and the effect transfers to unseen metadata attributes. The same pattern appears in caption-based vision-language training: the authors analyze 40M Exif-tagged images from Re-LAION-2B, organize captions into 6238 topics, and train CLIP-loss models on 6.4M-image subsets whose topic–camera correlations reach 0.396, compared with 0.255 and 0.166 for baseline and weaker settings. Mitigation works during and after pretraining: DINOv2-style color jitter, blur, and grayscale reduce sensitivity, while a post-hoc adversarial linear layer suppresses JPEG traces and generalizes to other metadata without sacrificing semantic utility. The trade-off is revealing: metadata-sensitive models are better generated-image detectors, with average accuracy rising from 57.3 for the uncorrelated model to 60.4 at maximal correlation, but removing these shortcuts improves out-of-distribution robustness, including ImageNet-Sketch accuracy from 26.4 to 30.0. The findings also suggest Stable Diffusion images can inherit semantic metadata traces from their training data.
Original abstract
Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: https://github.com/ryan-caesar-ramos/visual-encoder-traces
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.