Explaining AI-Image Detection: What the Heatmap Actually Shows
AuthorsLeonid Kuturin, Ilya Sotnikov, Mark Khusnutdinov, Mikhail Potemkin, Pavel Baranas, Aleksandra Korepanova, Alexander Kalashnikov
Resources
It shows that AI-image detectors may learn file-compression fingerprints instead of synthesis—and that attractive heatmaps are not automatically faithful explanations.
Key results
Total images in the marketplace review-photo corpus.
PE-Core-336 detector performance before compression-history correction.
Performance after synthetic images are re-encoded to the real class’s WebP format.
Native PR-AUC gain attributed to symmetric encoding in the three-seed factorial.
Local edits detected out of 368, versus 19 for the first-fix detector.
Conditional pixel average precision against edited-region masks.
What the paper found
In “Explaining AI-Image Detection: What the Heatmap Actually Shows,” researchers from Sirius Educational Centre and HSE University audit both an AI-image detector and the heatmaps meant to justify its decisions on marketplace review photos. Their corpus contains 186,527 images: 169,751 real photographs and 16,776 synthetic images generated through OpenRouter using models including OpenAI’s gpt_image_2, Black Forest Labs’ FLUX, Google Gemini, Microsoft MAI, ByteDance Seedream, and xAI Grok. A frozen Perception Encoder, PE-Core-336, combined with 44 FFT and spatial-rich-model forensic features reaches 0.9999 PR-AUC on the collected data, but falls to 0.7254 when synthetic images are re-encoded into the real class’s WebP format, revealing that compression history—not synthesis—drives much of the apparent performance. An asymmetric repair merely relocates that shortcut; a three-seed factorial shows that symmetric encoding supplies a +0.176 native PR-AUC gain, while removing forensic features contributes essentially nothing. The selected detector identifies 142 of 368 local edits, compared with 19 for the first-fix model that mostly classified edits as real. For explanations, the authors require causal intervention: deleting highlighted regions must reduce the detector logit more than detector-blind controls, with bootstrap intervals excluding zero. Across 17 attribution maps, perturbation methods lead, attention rollout ranks third, and no gradient-CAM variant shows a positive advantage. Their SLIC-based region_ensemble reaches 0.466 pixel AP, roughly level with the center prior at 0.462, while costing 12.4 seconds per map; the authors therefore present it as a localizer, not proof of faithful explanation.
Original abstract
A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study that detection problem, so we build a detector and attach an attribution map as its evidence, then measure what that pair delivers on 186,527 images under controls designed to change our conclusions when something is wrong. Compression history, not synthesis, drives naive evaluation: our strongest model reaches 0.9999 PR-AUC (area under the precision-recall curve) on a product-disjoint split, yet falls to 0.7254 once we re-encode synthetics into the real class's format, while five public detectors move by at most 0.07. Aligning one class relocates the cue rather than removing it, and the repaired model then assigns native files a median probability of synthesis of 0.0004. One identical final encode for both classes repairs that, and a three-seed factorial credits the encoding change with the whole gain (+0.176 +- 0.009 PR-AUC). That encode equalises the last stage only: forensic features alone still separate the classes at 0.7145 against a base rate of 0.254. For evidence we test maps causally, against controls that never consult the detector. Whether an attribution ranking exists at all depends on whether the detector reacts to the image. On our first-fix detector, which calls 96 of 100 edited frames real, no map beats a random one. On the detector we selected, twelve of seventeen maps clear that control on edited images and eight on generated ones; perturbation leads both axes and no gradient-CAM variant shows a positive advantage. The trivial controls never clear it, and on generated images the centre prior is worse than random. Our ensembled regional map clears both axes and takes the top pixel AP at 12.4 s per map against 44.9 for occlusion. Clearing a detector-blind control is not yet a faithful explanation, and we demonstrate none.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.