NTH

EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment

AuthorsGuray Ozgur, Mustafa Efe Tamyapar, Naser Damer, Fadi Boutros

August 26, 2026 3 min read
Watch on YouTube
The one-line take

EXPL-FR turns opaque face-recognition embeddings into prompt-queryable semantic explanations using a lightweight vision-language alignment adapter.

Key results

978
Semantic vocabulary

Candidate text prompts spanning 22 attribute categories.

71.66%
FR-space vocabulary verification

Mean verification accuracy using the 978 CLIP prompt anchors mapped into AdaFace’s face-recognition space.

19.68
Cross-modal transfer gain

Percentage-point improvement over the 51.98% CLIP-native vocabulary projection.

0.931
Top-100 signature AUC

Identity-separation AUC for the selected 100 prompts, compared with 0.910 for all 978 prompts.

0.92
Prompt-driven RFW audit correlation

Mean Kendall correlation between prompt-only model rankings and measured ethnicity-related verification errors.

What the paper found

EXPL-FR makes face-recognition embeddings explainable without retraining the recognizer or using attribute labels. It freezes a face-recognition model and a vision-language model such as OpenAI’s CLIP or Google’s SigLIP, then trains only a lightweight four-layer MLP adapter on face images to map visual embeddings into the recognizer’s space; the same adapter maps text prompts into that space, creating 978 semantic anchors across 22 categories. A label-free detectability test selects the 100 concepts that remain separable after alignment, producing identity-level, per-image, and differential explanations for genuine, impostor, and morph comparisons. With CLIP and AdaFace, the full prompt vocabulary reaches 71.66% mean verification accuracy in the face-recognition space, compared with 51.98% in CLIP’s native image-text space, isolating a 19.68-point gain from cross-modal transfer. The filtered top-100 signature reaches 0.931 identity-separation AUC versus 0.910 for all 978 prompts. On RFW, the fully prompt-driven audit ranks four face-recognition models by ethnicity-related error with Kendall correlation 0.92, matching supervised and VLM-pseudo-label audits while requiring no annotations. Validation on WebFace4M, CelebA, GAN-Control, LFW, and DCMorph shows that models preserve cues such as eyewear and hair color but suppress many capture-condition factors. The method remains limited by CLIP or SigLIP blind spots, especially pose and illumination, and by semantic cues that cannot be represented through a small textual vocabulary.

Original abstract

Deep face recognition (FR) models reach near-saturated accuracy but remain opaque: a practitioner cannot ask which semantic attributes a similarity score relied upon. EXPL-FR answers this inside the FR model's own embedding space. A lightweight adapter aligns a vision-language model's (VLM) image encoder with the frozen FR space, trained on face images alone and never on text. Because the VLM's encoders share one space, the same adapter applies to the text encoder, turning 978 attribute prompts in 22 categories, also extendable, into FR-space anchors at no extra cost. We do not assume this transfer works: a face-verification protocol measures it, and an ablation changing only the adapter isolates its contribution. Not every concept survives, because an FR model earns its invariances by discarding the factors it must verify identities across. A label-free detectability measure compares each concept's separability in FR space against the VLM space, and the 100 most detectable form the model's readable semantic signature, which separates identities better than the full vocabulary. We cover four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations. We benchmark attribute-level auditing under three supervision settings, human labels (current practice), VLM pseudo-labels, and our fully prompt-driven audit, against real verification behavior. With no labels, the prompt-driven audit ranks four FR models by their measured per-ethnicity RFW errors and ranks controlled attribute changes by their true verification cost.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis