NTH

Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection

AuthorsGuray Ozgur, Fadi Boutros, Naser Damer

August 26, 2026 3 min read
Watch on YouTube
The one-line take

This study shows that tomatoes, potatoes, and onions can help train face anti-spoofing systems by teaching detectors to recognize presentation artifacts rather than facial identity.

Key results

12,480
TPO presentations

Total face-free bona fide, print, and replay presentations in the TPO dataset.

78
TPO vegetable identities

Distinct tomatoes, potatoes, and onions used in the dataset.

92.70%
TPO cross-dataset AUC

Average video-level AUC across MSU-MFSD, CASIA-FASD, Idiap Replay-Attack, and OULU-NPU.

14.15%
TPO cross-dataset HTER

Average HTER corresponding to the TPO-only FoundPAD model.

81.02%
SynthASpoof cross-dataset AUC

Average AUC for FoundPAD trained on synthetic faces.

92.55%
Face-plus-TPO AUC

Average AUC after replacing part of single-source face training with TPO under a fixed optimization budget.

What the paper found

This paper tests whether face presentation attack detection, or PAD, actually needs faces during training. It introduces TPO, a controlled face-free dataset of tomatoes, potatoes, and onions containing 12,480 presentations from 78 vegetable identities, with bona fide captures plus print and replay attacks recorded using Microsoft Surface and Samsung Galaxy devices. The central hypothesis is that PAD models learn recapture artifacts—such as moiré, halftoning, display structure, reflections, gamma distortion, and sensor noise—rather than facial identity or appearance. Using FoundPAD, a CLIP ViT-B/16 encoder adapted with rank-stabilized LoRA, the TPO-only model reaches 92.70% average video-level AUC and 14.15% HTER across MSU-MFSD, CASIA-FASD, Idiap Replay-Attack, and OULU-NPU, despite never using faces in downstream PAD training. This exceeds the 81.02% AUC achieved by the synthetic-face dataset SynthASpoof. Under a fixed optimization budget, replacing part of single-source face training with TPO raises average AUC from 89.33% to 92.55% and lowers HTER from 17.54% to 13.97%, showing that face-free data contributes complementary presentation-process information rather than merely more samples. Ablations indicate that diversity across print and replay instruments matters more than redundant frames, while frequency analysis rejects a single universal spectral shortcut. The evidence supports privacy-preserving, identity-independent PAD development, although the conclusions are limited to visible-spectrum print and replay attacks, not three-dimensional masks, depth, or infrared sensing.

Original abstract

Face presentation attack detection (PAD) is traditionally formulated as a face-specific problem, although many of the visual artifacts introduced by print, replay, and recapture processes are not inherently tied to facial appearance. In this work, we investigate whether transferable PAD representations can be learned without using faces during downstream PAD training. To this end, we introduce TPO, a controlled face-free presentation attack dataset consisting of bona fide, print, and replay recordings of, almost randomly chosen, tomatoes, potatoes, and onions acquired under protocols that closely mirror conventional face PAD datasets. Using a foundation-model-based PAD architecture, we demonstrate that a detector trained on TPO achieves an average AUC of 92.70% across four standard cross-dataset face PAD benchmarks, outperforming training on synthetic faces and remaining competitive with models trained on real face datasets. Conversely, models trained on face PAD datasets transfer consistently above chance to TPO, suggesting that the learned representations capture characteristics of the presentation process rather than object semantics. Furthermore, incorporating TPO into conventional face PAD training consistently improves cross-dataset performance under fixed optimization budgets, indicating that face-free data provides complementary information rather than simply additional training samples. Finally, representation and frequency analyses provide further evidence that transferable PAD representations cannot be explained by a single spectral artifact but instead encode richer presentation cues shared across object categories. Together, these results provide empirical evidence that transferable presentation attack representations can be learned independently of facial content, opening new opportunities for privacy-preserving and identity-independent PAD development.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis