CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection


Abstract

Presentation Attack Detection (PAD) serves as a crucial safeguard for face recognition systems against presentation attacks such as printed photos, replayed videos, and 3D masks. Despite significant progress, existing PAD models still struggle to generalize across unseen domains due to variations in sensors, lighting, and attack materials. Recent Vision-Language Models (VLMs) have shown strong generalization ability, yet their applications in PAD remain limited because learned prompts, typically optimized under class-label supervision, fail to explicitly align with fine-grained attack-relevant visual semantics. As a result, the learned representations often overfit domain-specific artifacts instead of capturing transferable attack cues. To address this, we propose Concept-Informed Prompts Guided Presentation Attack Detection (CPG-PAD), a framework that introduces model-level concept guidance into the prompt learning process. Specifically, we design a Visual Concept-driven Enhancement (VCE) module that employs eXplainable AI (XAI) techniques to automatically discover PAD-relevant visual concepts and generate concept-associated heatmaps providing localized fine-grained guidance. Guided by these heatmaps, a Prompt-based Concept Injection (PCI) mechanism integrates these concepts into the prompt space through a Visual-Prompt Decoder (VPD) and a concept-mapping loss, enabling prompts to align with the model’s internal concept space. This design enables CPG-PAD to capture generalizable and domain-invariant attack cues while effectively suppressing dataset-specific biases. Extensive experiments across nine benchmark datasets demonstrate that CPG-PAD consistently achieves state-of-the-art cross-domain performance under multi-source, limited-source, and single-source settings.

Presentation Attack Detection, Explainable AI, Domain Generalization

1 Introduction↩︎

With the widespread adoption of face recognition in applications like identity verification and online payments, security concerns have become increasingly prominent. Presentation Attack Detection (PAD) is essential for detecting presentation attacks (PAs) using printed images [1], video replays [2], and 3D masks [3]. While existing PAD methods, ranging from hand-crafted features [4][6] to deep-learning-based approaches [7], [8], have shown promising results in controlled settings, they struggle to generalize across different domains due to significant domain shifts. Mitigating domain shifts between seen source samples and unseen target samples remains a considerable challenge for PAD models.

Figure 1: Comparison with Existing CLIP-like PAD Methods. (a) Previous methods aim to learn carefully designed prompts through class label supervision, which limits their performance. (b) In contrast, PAD-relevant visual concepts discovered from VLMs teach model to learn concept-informed prompts, which enhance PAD performance.

Traditional PAD methods [9][11] seek to bridge the distribution gap between source and target domains using adversarial adaptation, meta-learning, or feature alignment, etc. However, their performance is constrained by the uni-modal nature of these approaches. With the emergence of pretrained vision-language models (VLMs), the generalization of downstream tasks [12] improves through their rich multi-modal knowledge. Recent studies [7], [8] further show that CLIP-like frameworks significantly enhance PAD performance. These approaches outperform traditional PAD methods by providing additional context. However, human-level prompts struggle to convey the critical visual clues for PAD classification. For instance, clues like photo edges and moiré patterns can be described using natural language, while subtle geometric distortions and illumination-reflection discrepancies are difficult to express. As a result, constructing and learning prompts relying on human-level understanding limit the performance of existing CLIP-like PAD methods.

Surprisingly, we find that leveraging eXplainable AI (XAI) methods [13][15] enables us to transfer general knowledge embedded in pretrained VLMs into PAD process from a model-level perspective and feature space. This finding motivates us to construct and learn model-oriented prompts informed by these explanatory model-level signals. Specifically, with the aid of concept-based XAI techniques, we successfully discover visual concepts within the model’s feature space and enhance domain data using corresponding feature heatmaps. These concept-associated heatmaps enable PAD models to learn multiple model-oriented prompts that are aligned with the discovered visual concepts, which guide PAD model to capture critical clues for classification beyond human-level understanding. A comparison with existing CLIP-like PAD methods is illustrated in 1.

Based on the above, we propose Concept-informed Prompts Guided Presentation Attack Detection (CPG-PAD), which exploits and incorporates visual concepts from pretrained VLMs into PAD process. Specifically, we introduce Visual Concept-driven Enhancement (VCE), a mechanism that leverages general pretrained VLMs to discover PAD-relevant visual concepts, while enhancing the domain data by generating fine-grained feature heatmaps corresponding to each discovered concept. The discovered concepts are closely linked to the objectives of attack detection, from a model-level perspective and within its feature space. With the aid of these concept-associated heatmaps, we introduce Prompt-based Concept Injection (PCI) to learn multiple concept-informed prompts by aligning textual prompts with the discovered visual concepts. In particular, learnable prompts interact with visual features through the Visual-Prompt Decoder (VPD) and are supervised by concept-associated heatmaps to inject concept information into the prompts. In this way, CPG-PAD accurately captures generalizable visual features, suppressing interference from domain-specific information when encountering unseen target domains, thereby enhancing the generalizability of PAD task.

Our contributions can be summarized as follows:

  • We propose Concept-informed Prompts Guided PAD (CPG-PAD), a novel framework which enhances domain generalization performance by incorporating visual concepts discovered by employing XAI techniques on pretrained VLMs.

  • CPG-PAD introduces Visual Concept-driven Enhancement (VCE), and Prompt-based Concept Injection (PCI). VCE automatically discovers PAD-relevant visual concepts, and enhances domain data by generating corresponding feature heatmaps. With the concept-associated heatmaps, PCI learns multiple concept-informed prompts through Visual-Prompt Decoder (VPD).

  • Extensive experiments and analysis demonstrate the superiority of CPG-PAD over state-of-the-art competitors on widely-used benchmark datasets.

2 Related Works↩︎

2.1 Presentation Attack Detection↩︎

The earliest attempts at Presentation Attack Detection (PAD) focused on handcrafted feature engineering. Techniques such as Local Binary Patterns (LBP) [16], [17], Histogram of Oriented Gradients (HOG) [18], and Scale-Invariant Feature Transform (SIFT) [19] were employed to capture texture irregularities caused by presentation attacks. Although these methods provided a foundation for early PAD, they were easily affected by changes in environmental conditions, including lighting variations, scaling factors, and head pose, which hindered robustness in practice.

Deep learning advanced PAD research by enabling robust and discriminative feature extraction. Convolution Neural Networks (CNNs) quickly became the mainstream choice [20][22]. Yet, the inherently local nature of convolutions limited their capacity to jointly capture fine-grained details and long-range dependencies, leading to vulnerabilities in challenging attack scenarios [23]. To overcome these shortcomings, transformer-based models [24], [25] have been explored. By leveraging self-attention mechanisms, they model global relationships across image patches and extend beyond CNNs’ receptive field constraints. While these architectures have achieved strong intra-domain evaluation, they often fail to generalize well in unseen environments.

To address the domain shift problem, researchers have explored Domain Adaptation (DA) and Domain Generalization (DG) strategies. DA [26][28] aims to bridge distribution gaps between source and target domains using unlabeled target data. Existing methods employ adversarial adaptation [26], meta-learning [27], or extract generalizable features from large-scale unlabeled data [28]. However, obtaining target domain data during training remains a significant challenge in real-world scenarios. In contrast, DG [9], [10], [29][31] tackles this issue by enhancing model robustness across multiple source domains without relying on target data. Traditional DG approaches learn domain-invariant features through feature alignment [9] or hierarchical relationships [32] to improve generalizability in PAD. Despite these improvements, single-modality DG-PAD still faces difficulties when domain discrepancies are large or training data limited.

Unlike DA and DG, which focus on improving feature generalization, auxiliary-supervised approaches provide additional supervisory signals to guide feature learning. Some works augment the classification task with physical or physiological auxiliary signals such as depth maps or remote photoplethysmography (rPPG). For example, Liu et al. [20] jointly estimate pixel-wise facial depth and rPPG rhythm along with the binary real/fake decision, showing that these auxiliary losses help guide the network to focus on fundamental attack-discriminative cues rather than overfit to texture patterns. Subsequent works [22] refine this supervision by introducing spatial gradient features and contrastive depth losses to sharpen the depth supervision signal. Other methods [33][35] extend depth supervision across multiple frames or design adaptive fusion with rPPG, e.g. exploiting temporal and depth information using multiframe depth estimation. Related efforts also introduce explanation-based or attention-based guidance as auxiliary supervision, using saliency maps to regularize model attention and improve interpretability [36].

Recent progress in large-scale pretrained vision and vision-language models, particularly CLIP [12], has opened new avenues for PAD. A first line of research [37], [38] investigates how foundation models can be adapted to PAD through simple transfer strategies, such as feature extraction, lightweight adaptation, or zero-shot inference. These studies suggest that pretrained models provide strong general visual priors, but directly transferring such knowledge to PAD remains challenging without task-specific guidance. Building on this, CLIP-based PAD methods further exploit cross-modal supervision by aligning facial images with textual prompts. FLIP [7] fine-tunes pretrained vision-language models by associating facial images with textual prompts such as “a photo of a real face” or “a photo of a fake face”, achieving substantial improvements in DG-PAD. Building on this, subsequent works [8], [39][43] propose frameworks that more effectively leverage multi-modal knowledge by disentangling features or constructing fine-grained prompts. Although existing CLIP-like methods achieve better performance than traditional PAD methods, their generalizability is still constrained by domain-specific biases.

2.2 Visual Concept Discovery in XAI↩︎

Explainable AI (XAI) aims to improve the interpretability and transparency of deep learning models by providing explanations for their decisions. In computer vision, gradient-based methods [13], [44][46] leverage activation and gradient information to attribute decisions to specific image regions. Perturbation-based methods [47] perturb the input image and observe the resulting changes in the model’s output to infer decision attribution. However, these methods focus on attributing individual decisions locally and lack the ability to explain the model from a global perspective. In contrast, concept-based approaches aim to generate multiple global explanations by analyzing groups of data that share common characteristics. Kim et al. [48] first introduced a method that goes beyond attribution-based approaches by measuring the influence of pre-selected concepts on a model’s outputs. ACE [14] and CRAFT [15] further advance this idea by automatically discovering concepts through clustering and factorization of intermediate neural activations. Unfortunately, existing concept-based XAI methods are primarily designed for general classification tasks and are not well suited to specific downstream applications. Building on these ideas, we seek to adapt concept-based XAI techniques for PAD, using visual concepts as a novel source of supervision.

2.3 Motivation↩︎

In the history of PAD research, auxiliary supervision from signals such as rPPG, depth maps, and NIR images has been shown to improve domain generalization by providing complementary cues beyond RGB data. At the same time, explainability techniques enable us to uncover machine-level concepts and heatmaps from powerful pretrained models, offering a new form of supervision that does not rely on additional sensors. This perspective differs from prior attention-guided PAD methods [36] that provide sample-level saliency supervision, as well as recent CLIP-based methods such as TF-FAS [41] and CCPE [40] that mainly transfer pretrained knowledge through richer semantic guidance or prompt engineering. Motivated by this, we propose a Visual Concept-driven Enhancement (VCE) process to distill PAD-relevant concepts and heatmaps from the CLIP visual encoder. These concepts and heatmaps are incorporated into a novel learning framework as auxiliary supervision signals, which are injected into learnable prompts to strengthen the generalization ability of PAD models.

Figure 2: Detailed information of Concept-informed Prompts Guided PAD (CPG-PAD). We first generate concept-associated heatmaps \mathcal{H}^{gt} through Visual concept-driven Enhancement (VCE) (details in 3). Next, multiple concept-informed prompts are learned via Prompt-based Concept Injection (PCI). Specifically, domain visual data is encoded into a CLS token V_{cls} and image tokens V_{img} through a learnable visual encoder. Learnable prompt embeddings \{\mathcal{T}^{r/f}_j\}_{j=1}^J and fixed class-wise prompts “A photo of a {CLS} face" is encoded into Prompt Features \{F^{r/f}_j\}_{j=1}^J through a fixed text encoder. Multiple feature heatmaps \mathcal{H}^{dec} are decoded through Visual-Prompt Decoder (VPD) by interaction between \{F^{r/f}_j\}_{j=1}^J and V_{img}. The training process is supervised by the Cross Entropy loss \mathcal{L}_{cls} and a concept mapping loss \mathcal{L}_M, which injects visual concept into learnable prompts by aligning \mathcal{H}^{dec} with \mathcal{H}^{gt}.

3 Method↩︎

The proposed Concept-informed Prompts Guided PAD (CPG-PAD) can be decomposed into two steps. We first enhance each domain through Visual Concept-driven Enhancement (VCE). Then we learn multiple concept-informed prompts via Prompt-based Concept Injection (PCI) process. The total framework of CPG-PAD is shown in 2.

3.1 Preliminaries↩︎

We adopt a general domain generalization PAD training setting. \(\mathcal{D}\)=\(\{D_i, i = 1, 2, \cdots, I\}\) denotes the domain data where the \(i-th\) domain \(\mathcal{D}_i = \{(\boldsymbol{x}_j, y_j), j=1, 2, \cdots, N_i\}\) contains \(N_i\) image and label pairs. \(\mathcal{C}(\cdot)\) denotes CLIP model with pretrained parameters. CLIP model \(\mathcal{C}(\cdot)\) contains two parts, the visual encoder \(VisEnc(\cdot)\) extracts CLS token \(V_{cls} \in \mathbb{R}^{d_v}\) and image tokens \(V_{img} \in \mathbb{R}^{H \times W \times d_v}\) and the text encoder \(TextEnc(\cdot)\) generates text feature \(F \in \mathbb{R}^{d_t}\) for each category: \[\begin{align} V_{cls}, V_{img} = VisEnc(\mathcal{X}); \;\;\;F = TextEnc(\mathcal{T}) \nonumber \end{align}\] where \(\mathcal{X}\) and \(\mathcal{T}\) indicate input image and prompts respectively. Finally, visual tokens \(V_{cls}, V_{img}\) and text feature \(F\) are projected to the same dimension \(C\) to calculate similarity scores. Notably, \(d_t=C=512\), \(d_v=768\) in ViT-B/16.

3.2 Visual Concept-driven Enhancement↩︎

In this section, we introduce Visual Concept-driven Enhancement (VCE), which operates in two stages: concept discovery and concept-associated heatmap generation. Notably, VCE utilizes a pretrained CLIP model without requiring any parameter updates. The detailed process is illustrated in 3. For clear illustration, we use the symbol \(N\) to represent the number of domain inputs for all domains.

Figure 3: Detailed information of Visual Concept-driven Enhancement (VCE). VCE consists of two stages, (a) concept discovery and (b) concept-associated heatmap generation.

3.2.1 Concept Discovery↩︎

The process of concept discovery is illustrated in 3 (a) with dark yellow box. We generally consider a single domain \(\mathcal{D}_i\), the data can be simply represented as input images \((\boldsymbol{x}_1, \cdots, \boldsymbol{x}_{N}) \in \mathcal{X}^{N}\) where \(\boldsymbol{x_j} \in \mathbb{R}^{D}\) and their associated labels \((y_1, \cdots, y_{N}) \in \mathcal{Y}^{N}\). Firstly, we select subsets of images \(\mathcal{X}^{fake/real}_{sub} \in \mathbb{R}^{N_s \times D}\) from the original dataset \(\mathcal{X}^{N}\) for both real and fake class to discover concepts. Here we randomly select \(r\) frames from each video to form the subset where \(r\) is a hyperparameter (e.g. select a single frame from each live video to form \(\mathcal{X}^{real}_{sub}\)). We define \(\pi(\cdot)\) as a straightforward filter function to create candidate concepts. Specifically, \(\pi(\cdot)\) crops the image into \(64 \times 64\) patches at regular intervals in both vertical and horizontal directions and resize each patch to dimension \(D\). Feed \(\mathcal{X}_{sub}\) into \(\pi(\cdot)\) to obtain an auxiliary dataset \(\mathcal{X}^{fake/real}_{patch} \in \mathbb{R}^{N_a \times D}\) which contains all candidate concepts. The following part will focus on the auxiliary dataset of a single class \(\mathcal{X}_{patch} \in \mathbb{R}^{N_a \times D}\), as the same operations apply to both the fake and real classes.

To automatically discover visual concepts from domain auxiliary dataset \(\mathcal{X}_{patch}\), we feed it to the pretrained visual encoder to obtain activation \(\mathbf{A} = VisEnc (\mathcal{X}_{patch}) \in \mathbb{R}^{N_a \times H \times W \times C}\) where \(HW\) indicates the shape of activation map and \(C\) indicates the number of channels. We apply Semi-NMF (Semi Non-negative Matrix Factorization) [49] to factorize activation maps, as it can naturally handle mixed-sign activations in CLIP features while encouraging interpretable, parts-based representations, which are difficult to achieve with other factorization methods. For instance, standard NMF is not applicable due to its non-negativity constraint on inputs, while PCA tends to produce less semantically interpretable components [50]. Semi-NMF decomposes the average pooled activations \(\mathbf{\bar{A}} = AvgPool(\mathbf{A}) \in \mathbb{R}^{N_a \times C}\) into a product of concept coefficients \(\mathbf{U} \in \mathbb{R}^{N_a \times K}\) and concept basis \(\mathbf{W} \in \mathbb{R}^{C \times K}\) by solving: \[\begin{align} (\mathbf{U}, \mathbf{W}) = \mathop{\arg\min}\limits_{\mathbf{U} \geq 0, \mathbf{W}} \Vert \mathbf{\bar{A}} - \mathbf{U} \mathbf{W}^\intercal \Vert^2_F, \label{eq:seminmf} \end{align}\tag{1}\] where \(K\) indicates the number of concepts one wishes to discover and \(\Vert \cdot \Vert^2_F\) denotes the Frobenius norm.

According to Ding et al. [49], the objective can be solved by iteratively updating \(\mathbf{U}\) and \(\mathbf{W}\) as follows: \[\begin{align} & \mathbf{W} = \boldsymbol{\bar{A}}^\intercal \mathbf{U} (\mathbf{U}^\intercal \mathbf{U})^{-1}, \\ & \mathbf{U} = \mathbf{U} \cdot \sqrt{\frac{ (\boldsymbol{\bar{A}} \mathbf{W})^{+} + \mathbf{U} (\mathbf{W}^\intercal \mathbf{W})^{-} }{ (\boldsymbol{\bar{A}} \mathbf{W})^{-} + \mathbf{U} (\mathbf{W}^\intercal \mathbf{W})^{+} } }, \end{align}\] where the separated parts of matrix \(\mathbf{M}\) are: \[\begin{align} \mathbf{M}^{+} = (\vert \mathbf{M} \vert + \mathbf{M}) / 2, \;\;\;\mathbf{M}^{-} = (\vert \mathbf{M} \vert - \mathbf{M}) / 2 \end{align}\]

\(\mathbf{W}\) is the discovered concepts where each column \(\mathbf{W}_k \in \mathbb{R}^{C}\) corresponds to a single concept basis. These concepts will be utilized in the subsequent heatmap generation process.

3.2.2 Concept-Associated Heatmap Generation↩︎

The process of concept-associated heatmap generation is illustrated in Figure 3 (b) with green box. After concept discovery, we get \(K\) concept basis \(\mathbf{W} \in \mathbb{R}^{C \times K}\) from the selected subset \(\mathcal{X}_{sub}\) with \(N_s\) images, which serves as an approximation of the concepts present in the original domain containing \(N\) images. Given an image \(\boldsymbol{x}\) (from the original domain), we factorize the activations \(\mathbf{A} = VisEnc(\boldsymbol{x}) \in \mathbb{R}^{H \times W \times C}\) in each position into several concept coefficients \(\mathbf{U} \in \mathbb{R}^{H \times W \times K}\) corresponding to \(\mathbf{W}\) by solving similar objective in Equation 1 : \[\begin{align} (\mathbf{U}(s, t), \mathbf{W}) = \mathop{\arg\min}\limits_{\mathbf{U}(s, t) \geq 0, \mathbf{W}} \Vert \mathbf{A}(s, t) - \mathbf{U}(s, t) \mathbf{W}^\intercal \Vert^2_F, \label{eq:seminmf2} \end{align}\tag{2}\] where \(\mathbf{W}\) is fixed based on the values obtained during the concept discovery process, and \(\mathbf{A}(s, t) \in \mathbb{R}^{C}\) represents the activation vector at position \((s, t)\) within the spatial dimensions \((H, W)\). The result \(\mathbf{U}(s, t) \in \mathbb{R}^{K}\) indicates the concept coefficients in position \((s, t)\).

Specifically, we can regard \(\mathbf{U}(s, t, k)\) as the influence of concept \(k\) at \((s, t)\) position of the activation map and the \(\mathbf{U}(\cdot, \cdot, k)\) can be seen as an attention map of concept \(k\). In this way, we factorize a single input image into several concepts and their corresponding attention maps. With the help of \(\mathbf{U}\), we can get the fine-grained feature heatmaps by standardizing them to have zero mean and unit variance: \[\begin{align} \mathcal{H}^{gt}(k) = \frac{\mathbf{U}(\cdot, \cdot, k)- mean(\mathbf{U}(\cdot, \cdot, k))}{std(\mathbf{U}(\cdot, \cdot, k))}, \end{align}\] where \(mean\) and \(std\) are operators to calculate the mean and variance of activation map \(\mathbf{U}(\cdot, \cdot, k)\).

3.3 Prompt-based Concept Injection↩︎

In this section, we propose Prompt-based Concept Injection (PCI), which learns concept-informed prompts by integrating visual concept information into textual prompts. PCI constructs multiple learnable prompts \(\mathcal{S}\) for both fake and real classes. For concept injection, \(\mathcal{S}\) is decoded into heatmaps \(\mathcal{H}^{dec}\) via a proposed Visual-Prompt Decoder (VPD) and supervised by concept-associated heatmaps \(\mathcal{H}^{gt}\) generated by the former VCE to align with discovered visual concepts. For PAD classification, the well-learned concept-informed prompts \(\mathcal{S}\) guide CPG-PAD to identify generalizable features and project to the final classification logits \(l\).

3.3.1 Learnable Prompts↩︎

The learnable prompts \(\mathcal{S}\) consist of two components: fixed class-wise text descriptions \(\mathcal{P}^{text}\) and multiple learnable prompt embeddings \(\{ \mathcal{T}^{prompt}_j \}^{J}_{j=1}\), where \(J\) is a hyperparameter. Following the setup in [7], we construct the fixed class-wise text description \(\mathcal{P}^{text}\) as {A photo of a fake face} and {A photo of a real face}. These texts are tokenized and embedded using CLIP’s tokenizer and embedding layer, resulting in fixed class-wise embeddings \(\mathcal{T}^{text} \in \mathbb{R}^{n_t \times d_t}\). The learnable prompt embeddings \(\{ \mathcal{T}^{prompt}_j \}^{J}_{j=1}\) where \(\mathcal{T}^{prompt}_j \in \mathbb{R}^{n_p \times d_t}\) are concatenated with the fixed class-wise embeddings \(\mathcal{T}^{text}\) and the necessary \([SOS]\) and \([EOS]\) embeddings to form complete learnable prompts \(\mathcal{S}\). The process can be formulated as follows: \[\begin{align} & [SOS], \mathcal{T}^{text}, [EOS] = \mathcal{C}_{Embed}( \mathcal{C}_{tokenize}( \mathcal{P}^{text} ) ), \nonumber \\ & \;\;\;\;\;\;\;where \;\; \mathcal{T}^{text} \in \mathbb{R}^{n_t \times d_t},\;\;[EOS], [SOS] \in \mathbb{R}^{d_t},\nonumber \\ & \mathcal{S}_j = concat([SOS], \mathcal{T}^{prompt}_j, \mathcal{T}^{text}, [EOS]) \in \mathbb{R}^{77 \times d_t}, \nonumber \\ & \mathcal{S} = \{ \mathcal{S}_j \}^{J}_{j = 1} \in \mathbb{R}^{J \times 77\times d_t}, \nonumber \end{align}\] where \(\mathcal{C}\) indicates CLIP model, \(d_t\) indicates embedding dimension of the text encoder in CLIP. The parameters \(n_t\) and \(n_p\) correspond to the embedding lengths of the fixed class-wise embeddings and the learnable prompt embeddings, respectively, ensuring that their sum satisfies \(n_t + n_p = 75\). The complete learnable prompts \(\mathcal{S}\) are subsequently aligned with visual concept information through concept injection. Notably, the same learnable prompts \(\mathcal{S}^{r/f} =\mathcal{S}\) are constructed for both fake and real classes.

Figure 4: Detailed structure of Visual-Prompt Decoder (VPD) module. MSA denotes multi-head self attention and MDA denotes Multi-head dual attention.

3.3.2 Learning Prompts via Concept Injection↩︎

The prompt learning process via concept injection is illustrated in Figure 2 with blue box. With the constructed learnable prompts \(\mathcal{S}^{r/f}\), we propose a Visual-Prompt Decoder (VPD) to better learn multiple concept-informed prompts for both fake and real classes. VPD can effectively decode multiple feature heatmaps with the interaction between multiple prompts \(\mathcal{S}^{r/f}\) and image tokens \(V_{img}\), which is later aligned with discovered concepts. For clear description, we use \(\mathcal{S}\) to represent \(\mathcal{S}^f\) or \(\mathcal{S}^r\).

Feed the prompts \(\mathcal{S}\) into the text encoder to obtain the prompt features \(F = TextEnc(\mathcal{S}) \in \mathbb{R}^{J \times C}\). Similarly, process the input image through the image encoder to derive a CLS token and image tokens \(V_{cls}, V_{img} = VisEnc(X)\) where \(V_{cls} \in \mathbb{R}^{C}\) and \(V_{img} \in \mathbb{R}^{H \times W \times C}\). VPD takes the prompt features \(F\) and image tokens \(V_{img}\) as inputs and decodes into heatmaps \(\mathcal{H}^{dec} = VPD(F, V_{img})\) through their interactions. The decoded heatmaps are later supervised by concept-associated heatmaps \(\mathcal{H}^{gt}\).

Specifically, VPD is a module with several layers of multi-head self-attention mechanism \(MSA(\cdot)\) (identical to the traditional Transformer [51]) and a newly proposed multi-head dual-attention module \(MDA(\cdot)\). Multi-head dual-attention calculates a single attention map for the two inputs and updates both through matrix multiplication to enable effective interaction, illustrated in Figure 4. For each layer in VPD, it first applies multi-head self-attention to both the image tokens \(V_{img}\) and the prompt features \(F\), producing \(\hat{V}_{img}\) and \(\hat{F}\), respectively. Then, the processed image tokens \(\hat{V}_{img}\) and prompt features \(\hat{F}\) are projected into query \(Q_*\) and key \(K_*\) representations using separate linear layers with layer normalization, resulting in \(Q_V, K_V\) for image tokens and \(Q_F, K_F\) for prompt features where \(Q_V, K_V \in \mathbb{R}^{H \times W \times C}\) and \(Q_F, K_F \in \mathbb{R}^{J \times C}\). Next, the attention matrix \(Attn\) is computed using a scaled dot-product between \(Q_V\) and \(Q_F\), followed by a softmax operation forming the shape \(\mathbb{R}^{H \times W \times J}\). The final representations are obtained by applying two separate Multi-Layer Perceptrons (MLPs) to the attended features. The image tokens are updated using the attended prompt keys \(K_F\), while the prompt features are refined through the attended image keys \(K_V\). Notably, the MLPs consist of two linear layers with GELU activation.

After passing through the Depth layer of the VPD, the decoded heatmap \(\mathcal{H}^{dec} \in \mathbb{R}^{H \times W \times J}\) is obtained via \(\mathrm{Conv2d} \left(F @ {V_{img}}^\intercal\right)\). With the generated heatmaps \(\mathcal{H}^{gt}\), we inject visual concept information into learnable prompts through concept mapping loss \(\mathcal{L}_M\) which supervises the class-specific heatmaps conditioned on the ground-truth label by minimizing the optimal Mean Square Error (MSE) through Hungarian Matching [52]: \[\begin{gather} \mathcal{H}^{dec} = (1-\delta)\mathcal{H}^{dec}_{fake} + \delta\mathcal{H}^{dec}_{real}, \text{where } \delta = \mathbb{I}(\text{label=real}), \nonumber \\ \mathcal{L}_M = \frac{1}{J} \min_{\pi \in S_J} \sum_{j=1}^{J} MSE(\mathcal{H}^{dec} (j), \mathcal{H}^{gt}(\pi(j))). \end{gather}\] where \(\mathbb{I}(\cdot)\) is the indicator function, \(S_J\) is the set of all permutations of \(\{1, \dots, J\}\), \(\pi\) is the optimal permutation found via the Hungarian algorithm.

Thus, we learn the concept-informed prompts \(\mathcal{S}^{r/f}\) through visual concept injection with the help of the VPD.

3.3.3 PAD Classification↩︎

With the learned concept-informed prompts, CPG-PAD can identify significant features in PAD classification process, illustrated in 2 with pink dotted box. Specifically, prompt features \(F^{r/f}\) are used to calculate the similarity score \(s = concat(F^r @ V_{cls}, F^f @ V_{cls})\) where \(s \in \mathbb{R}^{J*2}\). A Linear layer is used to project the similarity score to logits along with a softmax operation \(l=Softmax(Linear(s))\). We supervise the logits using traditional cross entropy loss \(\mathcal{L}_{cls} = CE(l, label)\).

The total loss \(\mathcal{L}\) can be formulated in two parts, classification loss \(\mathcal{L}_{cls}\) and concept mapping loss \(\mathcal{L}_M\): \[\begin{align} &\mathcal{L} = \mathcal{L}_{cls} + \alpha \mathcal{L}_M. \end{align}\]

Table 1: Stability analysis of Concept Discovery. We use \(r=1,2,3\) (r frames per video) across four dataset settings. The results demonstrate the stability of concept discovery process.
Dataset Class Top-5 Top-10 Top-15 Random
C live 0.9795 0.9413 0.8863 0.0552
fake 0.9617 0.9232 0.8599
1-5 live 0.9705 0.9302 0.8099
fake 0.9660 0.9313 0.8212
1-5 live 0.9581 0.9116 0.8078
fake 0.9887 0.9796 0.9010
1-5 live 0.9659 0.9308 0.8195
fake 0.9773 0.9501 0.8447
1-5 Avg. All 0.9710 0.9373 0.8438

4 Visual Concept Discovery Analysis↩︎

In this section, we analyze the automatically discovered visual concepts and the generated heatmaps through Visual Concept-driven Enhancement (VCE) process. We choose to analyze CelebA-Spoof [53] dataset which is a large-scale PAD dataset containing rich attributes on face, illumination, environment and attack types.

Figure 5: Visualization of discovered visual concepts in CelebA-Spoof dataset.
Figure 6: Visualization of concept-associated heatmaps in CelebA-Spoof dataset.

4.1 Concept Discovery Stability↩︎

To evaluate the stability of our concept discovery algorithm, we conduct experiments with three different settings (\(r=1,2,3\)) on each dataset-class combination. For each setting, we repeat the concept discovery process with three different random seeds. For each pair of runs, we compute the cosine similarity between the extracted concept basis vectors and apply maximum bipartite matching (Hungarian algorithm) to find the optimal one-to-one correspondence. Note that Semi-NMF uses a deterministic K-means initialization, and thus different random seeds do not introduce additional variation.

1 reports the average cosine similarity of the Top-\(K\) matched concepts across all pairwise comparisons. The Top-5 concepts achieve 97.1% average similarity, Top-10 concepts achieve 93.7%, and all 15 concepts achieve 84.4%, all significantly higher than the random baseline (5.5%). We also observe that the most discriminative concepts (e.g., paper edges, cut holes) consistently appear within the Top-10 concepts. These results indicate that the concept discovery process is highly stable across different \(r\) values.

4.2 Visual Concept Visualization↩︎

We employ ActMax [54], a technique used in the XAI domain for concept visualization, to project the concepts, originally represented as features, into the RGB space. Specifically, ActMax selects the top-k most activated patches for each concept and treats them as representatives.

The discovered concepts in CelebA-Spoof dataset are presented in 5. In 5 (a), VCE identifies semantic attack concepts, such as the ‘cut holes’ made for the eyes or mouth and the ‘fold lines’ created by bending the paper, whereas others may be more subtle, such as the ambiguity introduced by print attacks. Notably, although some concept visualization may appear similar for real and fake classes, they are distinguishable in the feature space. The results demonstrate that our method can uncover meaningful visual concepts.

4.3 Concept-Associated Heatmap Visualization↩︎

With the discovered visual concepts, VCE generates concept-associated heatmaps for each sample. After normalization, the heatmaps are resized to 224 × 224 for clearer visualization. The visualizations on the CelebA-Spoof dataset are presented in 6, where red denotes stronger attention and blue denotes weaker attention. As shown, the generated heatmaps align well with the discovered visual concepts. For instance, in 5, the heatmaps of concept 1 and concept 4 highlight the eyes and mouth, while those of concept 2 and concept 3 capture clear attack cues such as ‘fold lines’ and ‘cut holes’.

Table 2: The results of P1 evaluation on ICMO datasets. Note that * indicates the corresponding method using CelebA-Spoof [53] as the supplementary source dataset. \(\downarrow\) / \(\uparrow\) represents that the smaller/larger value, the better performance. We BOLD the best results and UNDERLINE the second best results of CLIP-based methods with and without the CelebA-Spoof dataset.
Method OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
2-3 (lr)4-5 (lr)6-7 (lr)8-9 (lr)10-10 HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\)
SDA (AAAI’21) [27] 15.40 91.80 24.50 84.40 15.60 90.10 23.10 84.30 19.65
DRDG (IJCAI’21) [29] 12.43 95.81 19.05 88.79 15.56 91.79 16.63 91.75 15.66
FGHV (AAAI’22) [55] 9.17 96.92 12.47 93.47 16.29 90.11 13.58 93.55 12.88
GDA (ECCV’22) [56] 9.20 98.00 12.20 93.00 10.00 96.00 14.40 92.60 11.45
SSAN-R (CVPR’22) [57] 6.67 98.75 10.00 96.67 8.88 96.79 13.72 93.63 9.82
SA-FAS (CVPR’23) [58] 5.95 96.55 8.78 95.37 6.58 97.54 10.00 96.23 7.82
IADG (CVPR’23) [9] 5.41 98.19 8.70 96.40 10.62 94.50 8.86 97.14 8.40
UDG-FAS (ICCV’23) [28] 5.95 98.47 9.82 96.76 5.86 98.62 10.97 95.36 8.15
DiVT-M (WACV’23) [59] 2.86 99.14 8.67 96.92 3.71 99.29 13.06 94.04 7.08
TTDG (CVPR’24) [31] 4.16 98.48 7.59 98.18 9.62 98.18 10.00 96.15 7.84
GAC-FAS (CVPR’24) [10] 5.00 97.56 8.20 95.16 4.29 98.87 8.60 97.16 6.52
CLIP (ICML’21) [12] 4.04 99.13 5.00 98.89 6.57 98.45 6.09 98.12 5.43
CoOp (IJCV’22) [60] 3.86 99.08 2.33 98.92 6.07 98.52 5.83 98.97 4.37
CoCoOp (CVPR’22) [61] 4.16 99.01 5.17 98.19 6.21 98.50 6.00 98.49 5.39
CFPL-FAS (CVPR’24) [8] 3.09 99.45 2.56 99.10 5.43 98.41 3.33 99.05 3.60
S-CPTL (ACM MM’24) [39] 1.43 99.17 0.89 99.00 6.86 98.63 4.12 99.02 3.33
CCPE (TMM’25) [40] 3.10 99.21 1.33 99.36 6.08 94.36 5.57 98.49 4.02
CPG-PAD (Ours) 2.86 99.41 2.67 99.27 1.29 99.90 1.49 99.80 2.08
ViT* (ECCV’22) [62] 1.58 99.68 5.70 98.91 9.25 97.15 7.47 98.42 6.00
FLIP-MCL* (ICCV’23) [7] 4.95 98.11 0.54 99.98 4.25 99.07 2.31 99.63 3.01
CFPL-FAS* (CVPR’24) [8] 1.43 99.28 2.56 99.10 5.43 98.41 2.50 99.42 2.98
FGPL* (ACM MM’24) [63] 2.86 98.12 3.89 98.19 3.50 99.54 1.77 99.23 3.01
S-CPTL* (ACM MM’24) [39] 1.25 99.35 0.52 99.90 5.84 98.78 3.95 99.72 2.89
CCPE* (TMM’25) [40] 2.86 99.24 1.30 99.98 4.15 99.31 3.64 99.67 2.99
CPG-PAD* (Ours) 1.67 99.74 0.78 99.86 0.71 99.80 1.58 99.63 1.19
Table 3: The results of P1 evaluation on ICMO datasets. Note that * indicates using CelebA-Spoof [53] as the supplementary source dataset and indicates methods using LLMs. \(\downarrow\) / \(\uparrow\) represents that the smaller/larger value, the better performance. We BOLD the best average HTER(%).
Method OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
2-3 (lr)4-5 (lr)6-7 (lr)8-9 (lr)10-10 HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\) AUC \(\uparrow\) HTER(%) \(\downarrow\)
TF-FAS(ECCV’24) [41] 3.44 99.42 0.81 99.92 2.24 99.67 2.26 99.48 2.19
CPG-PAD (Ours) 2.86 99.41 2.67 99.27 1.29 99.90 1.49 99.80 2.08
TF-FAS*(ECCV’24) [41] 1.49 99.80 0.58 99.99 1.56 99.89 1.43 99.93 1.27
I-FAS*(AAAI’25) [64] 0.32 99.88 0.04 99.99 3.22 98.48 1.74 99.66 1.33
CPG-PAD* (Ours) 1.67 99.74 0.78 99.86 0.71 99.80 1.58 99.63 1.19
Table 4: Extra metric result of P1 evaluation. We report BPCER10, BPCER20, BPCER100 metric for fair comparison following standard ISO/IEC 30107-3 [65]. We also report results using EER threshold on validation set (@val-thr).
Method Metric OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
FLIP [7] APCER@val-thr 7.14 6.00 1.43 8.75 5.83
BPCER@val-thr 4.29 0.22 6.00 1.01 2.88
HTER@val-thr 5.71 3.11 3.71 4.88 4.35
BPCER10 3.81 0.22 2.14 0.80 1.74
BPCER20 5.24 0.22 2.86 1.98 2.57
BPCER100 20.48 0.67 6.00 4.69 7.96
CPG-PAD APCER@val-thr 1.90 2.00 3.00 1.64 2.14
BPCER@val-thr 2.86 5.33 0.00 1.46 2.41
HTER@val-thr 2.38 3.67 1.50 1.55 2.28
BPCER10 1.43 1.33 0.00 0.00 0.69
BPCER20 1.43 2.00 0.00 0.49 0.98
BPCER100 5.71 8.00 1.43 4.88 5.01

5 Experiments↩︎

5.1 Experimental Setups↩︎

5.1.1 Datasets↩︎

We employ nine datasets to evaluate the performance of our method, including MSU-MFSD (M) [66], CASIA-FASD (C) [67], Idiap Replay-Attack (I) [68], OULU-NPU (O) [69], CASIA-SURF (S) [70], CASIA-CeFA (C) [71], WMCA (W) [72], CelebA-Spoof [53] and SiW-Mv2 [73]. Idiap Replay-Attack, CASIA-FASD, MSU-MFSD and OULU-NPU datasets (ICMO) differ in material, lighting, background, and resolution. CASIA-SURF, CASIA-CeFA, and WMCA datasets (CSW) encompass a broader range of topics, various types of attacks, and diverse sampling environments. In all experimental protocols, each dataset is treated as an individual domain, and we use the notation \(XY \rightarrow Z\) to indicate that datasets X and Y serve as the source domain for training, while dataset Z is used as an unseen target domain for evaluation.

Table 5: The results of P2 evaluation on CSW datasets. Note that * indicates the corresponding method using CelebA-Spoof [53] as the supplementary source dataset. \(\downarrow\) / \(\uparrow\) represents that the smaller/larger value, the better performance. We BOLD the best results and UNDERLINED the second best results of CLIP-based methods with and without the CelebA-Spoof dataset.
Methods CS \(\rightarrow\) W SW \(\rightarrow\) C CW \(\rightarrow\) S Avg.
2-4 (lr)5-7 (lr)8-10 (lr)11-11 HTER(%) \(\downarrow\) AUC \(\uparrow\) TPR@ FPR=1% HTER(%) \(\downarrow\) AUC \(\uparrow\) TPR@ FPR=1% HTER(%) \(\downarrow\) AUC \(\uparrow\) TPR@ FPR=1% HTER(%) \(\downarrow\)
ViT (ECCV’22) [62] 21.04 89.12 30.09 17.12 89.05 22.71 17.16 90.25 30.23 18.44
CLIP-V (ICML’21) [12] 20.00 87.72 16.44 17.67 89.67 20.70 8.32 97.23 57.28 15.33
CLIP (ICML’21) [12] 17.05 89.37 8.17 15.22 91.99 17.08 9.34 96.62 60.75 13.87
CoOp (IJCV’22) [60] 9.52 90.49 10.68 18.30 87.47 11.50 11.37 95.46 40.40 13.06
CoCoOp (CVPR’22) [61] 13.89 90.74 - 15.49 89.40 - 13.76 95.59 - 14.38
S-CPTL (ACM MM’24) [39] 8.99 94.01 - 12.78 91.64 - 9.48 95.83 - 10.42
FGPL (ACM MM’24) [63] 14.05 92.65 33.33 19.00 88.53 13.33 11.00 94.72 34.00 14.68
CFPL-FAS (CVPR’24) [8] 9.04 96.48 25.84 14.83 90.36 8.33 8.77 96.83 53.34 10.88
CPG-PAD (Ours) 8.65 97.77 67.42 13.37 93.76 12.23 7.71 98.09 71.01 9.91
ViT* (ECCV’22) [62] 7.98 97.97 73.61 11.13 95.46 47.59 13.35 94.13 49.97 10.82
FLIP-MCL* (ICCV’23) [7] 4.46 99.16 83.86 9.66 96.69 59.00 11.71 95.21 57.98 8.61
S-CPTL* (ACM MM’24) [39] 4.22 99.30 - 9.59 96.73 - 10.97 97.40 - 8.33
CFPL-FAS* (CVPR’24) [8] 4.40 99.11 85.23 8.13 96.70 62.41 8.50 97.00 55.66 7.01
CPG-PAD* (Ours) 6.99 98.20 76.48 4.19 99.16 74.09 8.59 97.37 61.15 6.59
Table 6: The results of P4 evaluation. We report the HTER(%) results on the different 12 scenarios on ICMO datasets. We only BOLD the best results and Avg. denotes the mean HTER(%).
Methods C → I C → M C → O I → C I → M I → O M → C M → I M → O O → C O → I O → M Avg.
ADDA (CVPR’17) [74] 41.8 36.6 - 49.8 35.1 - 39.0 35.2 - - - - 39.6
DRCN (ECCV’16) [11] 44.4 27.6 - 48.9 42.0 - 28.9 36.8 - - - - 38.1
DupGAN (CVPR’18) [75] 42.4 33.4 - 46.5 36.2 - 27.1 35.4 - - - - 36.8
KSA (TIFS’18) [76] 39.3 15.1 - 12.3 33.3 - 9.1 34.9 - - - - 24.0
DR-UDA (TIFS’20) [26] 15.6 9.0 28.7 34.2 29.0 38.5 16.8 3.0 30.2 19.5 25.4 27.4 23.1
USDAN-Un (PR’21) [77] 16.0 9.2 - 30.2 25.8 - 13.3 3.4 - - - - 16.3
GDA (ECCV’22) [56] 15.1 5.8 - 29.7 20.8 - 12.2 2.5 - - - - 14.4
CDFTN-L (AAAI’23) [78] 1.7 8.1 29.9 11.9 9.6 29.9 8.8 1.3 25.6 19.1 5.8 6.3 13.2
CPG-PAD (Ours) 7.14 4.29 2.99 14.67 7.38 6.61 4.00 2.29 2.48 4.78 5.71 1.67 5.33
FLIP-MCL* (ICCV’23) [7] 10.57 7.15 3.19 0.68 7.22 4.22 0.19 5.88 3.95 0.19 5.69 8.40 4.84
CPG-PAD* (Ours) 4.29 8.33 3.02 3.33 5.71 3.88 0.67 0.71 2.68 1.22 1.00 5.48 3.36

5.1.2 Protocols↩︎

Following previous works [7], [8], we adopt 4 evaluation protocols to assess the generalizability of our method: multi-source evaluation (P1), limited-source evaluation (P2 and P3), single-source evaluation (P4), unknown PAI (Presentation Attack Instruments) evaluation (P5). For each protocol, we train the model on the training sets of the source domains and evaluate it on the test set of the target domain. Following previous works, evaluation is conducted at the video level, where two frames are sampled from each video and their scores are averaged to obtain the final video-level prediction.

5.1.3 Evaluation Metrics↩︎

Following standard evaluation principles, we employed three key metrics, HTER, AUC, and TPR, to assess the model’s performance. (1) HTER (Half Total Error Rate) quantifies the trade-off between false rejection and false acceptance errors, computed as the average of the False Rejection Rate (FRR) and the False Acceptance Rate (FAR). (2) AUC (Area Under the Curve) measures the overall classification performance by representing the area under the Receiver Operating Characteristic (ROC) curve. (3) TPR (True Positive Rate) evaluates the model’s ability to correctly identify attack samples, with the classification threshold adjustable based on specific application requirements to optimize performance. Besides, following standard ISO/IEC 30107-3 [65] for biometric PAD, we report (1) the BPCERs (Bona Fide Presentation Classification Error Rate) observed at APCER (Attack Presentation Classification Error Rate) values or security thresholds of 1% (BPCER100), 5% (BPCER20), and 10% (BPCER10) and draw DET curves for P1 protocol.

Thresholding protocol and limitation: Following common cross-database protocols, HTER is computed using the EER threshold estimated from the target test scores for direct comparison with prior works. We acknowledge that such a posteriori thresholding may introduce optimistic bias, and therefore additionally report results under a stricter setting where the threshold is determined on the validation set and then evaluated on the test set (P1). We also provide BPCER/APCER metrics and DET curves to complement HTER.

5.1.4 Implementation Details↩︎

For the PAD task, we use a pretrained CLIP [12] model from OpenAI6 with ViT-B/16 as the image encoder, the same as previous methods. We keep the parameters in the text encoder unchanged, and finetune the image encoder with a simple Convpass adapter. We resize the image to 224×224 with a batch size of 24, and train all models for 50 epochs. We use the Adam optimizer and set the learning rate to 1e-4 for training. In all experiments we set \(r=2\), \(K=15\) and set the number of learnable concept-informed prompts \(J=K=15\). We set the depth of VPD \(Depth=4\), the coefficient of concept mapping loss \(\alpha=0.01\).

Table 7: The results of P5 evaluation on SiW-Mv2 dataset. We follow [37] and use leave-one-out strategy to evaluate the generalizability of unknown PAIs. The results of ViT-B/16 come from [37].
Approaches Metrics Covering Make-up 3D Attack 2D Attack Avg.\(\pm\)Std.
FunE. PEye PMouth PaperG. Ob. Impers. Cosmetic HalfM. Silicone TransM. Paper Mann. Replay Print
SiW-Mv2 baseline[73] HTER(%) \(\downarrow\) 29.50 2.70 1.10 11.90 1.30 24.50 10.90 8.00 9.20 0.00 0.60 4.00 17.90 9.60 9.40\(\pm\)8.80
BPCER100(%) \(\downarrow\) 91.10 63.00 11.60 96.00 1.70 76.20 60.80 38.60 52.50 0.00 0.00 33.4 60.70 21.10 43.34\(\pm\)33.19
ViT-B/16 HTER(%) \(\downarrow\) 13.46 0.19 0.39 1.62 8.60 0.19 11.37 3.51 0.58 3.21 0.19 0.19 16.27 11.72 5.11\(\pm\)5.87
BPCER10(%) \(\downarrow\) 14.67 0.39 0.39 0.39 8.11 0.39 15.44 2.70 0.39 2.32 0.00 0.39 17.76 15.44 5.63\(\pm\)7.04
BPCER20(%) \(\downarrow\) 33.20 0.39 0.77 1.16 22.01 0.39 20.46 4.25 0.77 3.09 0.39 0.39 23.55 26.25 9.79\(\pm\)12.21
BPCER100(%) \(\downarrow\) 50.19 0.39 0.77 1.93 22.39 0.39 26.25 8.49 1.16 7.34 0.39 0.39 32.05 42.47 13.90\(\pm\)17.47
CLIP HTER(%) \(\downarrow\) 10.64 0.19 0.19 0.38 14.79 0.19 8.19 0.19 0.19 2.77 0.00 0.19 2.62 1.85 3.03\(\pm\)4.54
BPCER10(%) \(\downarrow\) 11.11 0.38 0.38 0.38 9.58 0.38 6.90 0.38 0.38 1.53 0.00 0.38 2.30 0.77 2.49\(\pm\)3.64
BPCER20(%) \(\downarrow\) 27.59 0.38 0.38 0.38 10.34 0.38 9.58 0.38 0.38 2.68 0.00 0.38 2.30 1.15 4.02\(\pm\)7.31
BPCER100(%) \(\downarrow\) 63.98 0.38 0.38 0.77 10.34 0.38 61.69 0.38 0.38 4.60 0.00 0.38 6.13 5.36 11.08\(\pm\)21.34
CPG-PAD HTER(%) \(\downarrow\) 4.45 0.19 0.19 0.19 0.77 0.19 4.19 0.19 0.19 0.57 0.00 0.19 2.62 0.38 1.02\(\pm\)1.49
BPCER10(%) \(\downarrow\) 2.30 0.38 0.38 0.38 1.53 0.38 0.77 0.38 0.38 0.38 0.00 0.38 1.92 0.77 0.74\(\pm\)0.66
BPCER20(%) \(\downarrow\) 3.83 0.38 0.38 0.38 1.53 0.38 2.68 0.38 0.38 0.77 0.00 0.38 2.68 0.77 1.07\(\pm\)1.12
BPCER100(%) \(\downarrow\) 57.47 0.38 0.38 0.38 1.53 0.38 22.61 0.38 0.38 1.15 0.00 0.38 6.90 0.77 6.65\(\pm\)15.23
Table 8: The results of P3 evaluation. We BOLD the best results.
MI \(\rightarrow\) C MI \(\rightarrow\) O Avg.
2-3 HTER AUC HTER AUC HTER
SSDG-R (CVPR’20) [79] 19.86 86.46 27.92 78.72 23.89
SSAN-R (CVPR’22) [57] 25.56 83.89 24.44 82.86 25.00
HFN+MP (TIFS’22) [80] 30.89 72.48 20.94 85.71 25.92
CIFAS (ICME’22) [81] 22.67 83.89 24.63 81.48 23.65
DiVT-M (WACV’23) [59] 20.11 86.71 23.61 85.73 21.86
DGUA-FAS (ICIP’23) [82] 19.22 86.81 20.05 88.75 19.64
BUDoPT (ECCV’24) [83] 5.33 98.92 5.94 98.37 5.64
CPG-PAD (Ours) 2.67 99.25 2.25 99.81 2.46

5.2 PAD Performance↩︎

5.2.1 Multi-source evaluation (P1)↩︎

Multi-source evaluation (P1). In P1 protocol, we construct a domain generalization benchmark using four datasets, ICMO, resulting in 4 scenarios where three datasets are seen as source domains and the remaining one serves as the target domain. We include traditional PAD methods [9], [10], [27][29], [31], [55][59] and current proposed CLIP-based methods [8], [12], [39], [40], [60], [61] as baseline. As shown in 2, we can draw the following conclusions: (1) CLIP-based methods outperform the traditional DG methods, even the finetuned CLIP model has a lower average HTER (5.43%) compared to the most competitive traditional DG method GAC-FAS (6.52%), supporting the motivation that general knowledge from CLIP model can benefit PAD task. (2) Without supplementary training data, the proposed CPG-PAD outperforms SOTA methods with an average HTER of 2.08%, surpassing S-CPTL [39] by 1.25% and even surpassing S-CPTL* which is trained with additional CelebA-Spoof dataset. On all 8 settings, CPG-PAD achieves top-two performance on 7 of the settings. (3) With supplementary training data (CelebA-Spoof), CPG-PAD also surpasses the recently proposed methods with an average HTER of 1.19%. The above results and conclusions prove that CPG-PAD can better leverage the general knowledge from pretrained CLIP model and effectively enhance PAD generalizability in multi-source scenarios.

Recent MLLM-based FAS methods [41], [64], [84] introduce large language models or multimodal reasoning into PAD, but usually require substantially larger model capacity than CLIP-based frameworks. As we can see in 3, our method CPG-PAD surpasses TF-FAS [41] and I-FAS [64] on average HTER metric with a small margin. However, achieving a performance level that is nearly on par is already sufficient to highlight the effectiveness of CPG-PAD since CPG-PAD is a vision-only model that does not rely on any extra annotations or LLMs.

Figure 7: DET curve comparison of P1 evaluation.

Following the standard ISO/IEC 30107-3 [65], we further compare CPG-PAD with the reproducible FLIP baseline using extra metrics as reported in 4, together with the DET curves in 7. Extra metrics include BPCER10, BPCER20, BPCER100 and BPCER/APCER/HTER results using EER threshold on validation set. Although FLIP achieves better performance in one protocol, CPG-PAD obtains better results in most cases and shows stronger overall operating-point performance. The DET curves further demonstrate that CPG-PAD yields a more favorable APCER-BPCER trade-off.

5.2.2 Limited-source evaluation (P2 and P3)↩︎

In this evaluation experiment, we test on two protocols (P2 and P3) with two datasets as source domains and one dataset as target domain. For P2, we establish a domain generalization benchmark focused on multimodal sensors, incorporating the CSW datasets, leading to 3 sub-experiments where two datasets are seen as source domain and the remaining one serves as target domain. As 5 shows, without additional training data, the proposed CPG-PAD outperforms the previous SOTA method S-CPTL [39] with an average HTER of 9.91% and achieves top-two performance across eight of nine metrics. With CelebA-Spoof as additional training data, the proposed CPG-PAD also achieves SOTA performance with an average HTER of 6.59%. Notably, as the results show, the SW \(\rightarrow\) C scenario is difficult for all previous methods since the domain gap is larger than the other two scenarios. However, CPG-PAD* achieves nearly half the HTER of the second-best CFPL-FAS*. For P3, we follow [83] and construct 2 scenarios with ICMO datasets which set M and I as source domains and test on C and O separately. In all settings, the proposed CPG-PAD surpasses previous methods, demonstrating its strong generalizability in limited-source PAD scenarios.

Table 9: The effect of MLP, CML and VPD. Improvement/ Degradation indicates the comparison against the baseline CLIP [12]. “MLP” indicates Multiple Learnable Prompts, “CML” indicates Concept Mapping Loss, “VPD” indicates Visual-Prompt Decoder.
Components OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
5-6 (lr)7-8 (lr)9-10 (lr)11-12 (lr)13-13 CLIP MLP CML VPD HTER (%) \(\downarrow\) AUC \(\uparrow\) HTER (%) \(\downarrow\) AUC \(\uparrow\) HTER (%) \(\downarrow\) AUC \(\uparrow\) HTER (%) \(\downarrow\) AUC \(\uparrow\) HTER (%) \(\downarrow\)
4.04 99.13 5.00 98.89 6.57 98.45 6.09 98.12 5.43
2.86 (-1.18) 99.57 (+0.44) 3.89 (-1.11) 99.30 (+0.41) 5.00 (-1.57) 98.73 (+0.28) 2.22 (-3.87) 99.63 (+1.51) 3.49 (-1.94)
4.29 (-0.25) 98.66 (-0.47) 3.44 (-1.56) 99.01 (+0.12) 3.71 (-2.86) 99.17 (+0.72) 1.59 (-4.50) 99.78 (+1.66) 3.26 (-2.17)
2.86 (-1.18) 99.41 (+0.28) 2.67 (-2.33) 99.27 (+0.38) 1.29 (-5.28) 99.90 (+1.45) 1.49 (-4.60) 99.80 (+1.68) 2.08 (-3.35)

5.2.3 Single-source evaluation (P4)↩︎

For P4, we strictly follow the setup in [7] to construct 12 scenarios using a single-source-to-single-target approach, leveraging the ICMO datasets. As shown in 6, the proposed CPG-PAD surpasses suboptimal method CDFTN-L [78] on nine of twelve settings with an average HTER of 5.33%, lower than half of CDFTN-L. Since it’s not fair to leverage pretrained models, we trained CPG-PAD with an additional CelebA-Spoof dataset to compare with FLIP-MCL* (both leverage pretrained models). The results indicate that CPG-PAD can outperform FLIP-MCL* with an average HTER of 3.36% demonstrating its strong generalizability on single source scenarios.

5.2.4 Unknown PAI evaluation (P5)↩︎

We evaluate the generalizability of the CPG-PAD for the challenging scenario of unknown PAI, including 3D masks (i.e. silicone masks, transparent masks and mannequin head) and make-up (obfuscation, impersonation and cosmetic). For this purpose, we follow the leave-one-out protocol in [37] using SiW-Mv2 [73] database: thirteen PAI species are used for training and the remaining PAI species is tested. 7 reports in compliance with ISO/IEC 30107-3 and benchmarks against the SiW-Mv2 baseline in terms of HTER and BPCER.

Compared with our CLIP baseline, CPG-PAD achieves clear improvements under the unknown-PAI protocol on SiW-Mv2, reducing the average HTER from 3.03% to 1.02% and the average BPCER100 from 11.08% to 6.65%. The gains are particularly evident on attack instruments with stronger appearance variation or localized attack cues, such as obfuscation make-up, cosmetic make-up, Funny Eye covering, transparent mask, and print attacks. For example, the HTER on obfuscation decreases from 14.79% to 0.77%, while the BPCER100 on cosmetic make-up drops substantially from 61.69% to 22.61%. These results indicate that concept guidance is especially beneficial for unseen PAIs involving facial region modification, occlusion boundaries, and material inconsistencies.

5.3 Ablation Studies↩︎

We conduct ablation studies on the same P1 benchmark and report results across all scenarios to demonstrate the effectiveness of each component of CPG-PAD as well as the influence of key hyperparameters.

5.3.1 Effect of CPG-PAD components↩︎

To investigate the effect of each component in CPG-PAD, we gradually add them to the baseline CLIP and report the results of P1 in 9. The differences of each component are reported in parentheses, with green denoting improvements and red denoting degradations. The CLIP baseline in 9 refers to a CLIP-based PAD classification baseline with fixed class prompts, a frozen text encoder, and visual-side adaptation through a lightweight Convpass adapter, rather than a fully finetuned CLIP trained with contrastive loss. On the base of CLIP, we first add Multiple Learnable Prompts (MLP) to encourage the model to freely discover different sub-direction features that limitedly improve the performance. Adding MLP improves average HTER from 5.43% to 3.49% since it provides the model with greater flexibility. Then we add the Concept Mapping Loss (CML) with a simple 2D convolution layer after the visual encoder. This makes a further step to improve the performance since the auxiliary supervision of concept-associated heatmaps helps model to learn more generalizable features. However, it can not fully leverage the visual concept information since the decoder is only a simple 2D convolution layer. Finally, we change the simple decoder with the proposed Visual-Prompt Decoder (VPD) to make full use of the concept-associated heatmaps, which improve the average HTER to 2.08%. This allows the model to effectively align the textual prompts with visual concepts, thereby enhancing the generalizability of PAD task. While adding different components of CPG-PAD to baseline CLIP, HTER and AUC is improved on almost all scenarios (only one except), indicating the effectiveness of each component proposed in CPG-PAD.

5.3.2 Effect of Learnable Prompt Number↩︎

We report the result of P1 with different \(J\) in 10. Since the concept number \(K\) is set to 15 empirically, there are only 15 heatmaps for each sample. Thus, the number of learnable prompt \(J\) can’t exceed \(K\). We uniformly select values in the range of 3 to 15 to analyze the impact. As \(J\) increases, the model is able to learn and exploit a greater number of concepts concurrently. In 10, we come up with the conclusion that the optimal value of \(J\) is 15 which utilizes all visual concepts. Moreover, larger values of \(J\) correspond to better performance, indicating that the more concepts discovered by VCE are utilized, the better the results. This demonstrates the effectiveness of VCE.

Table 10: The effect of learnable prompt number \(J\). We report the HTER(%) of different \(J\). Since \(J\) can not exceed the discovered concept number \(K=15\), we uniformly select values in the range of 3 to 15 and analyze their impact. We BOLD the best result in each column.
\(J\) OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
3 7.14 5.33 3.71 4.39 5.14
6 7.14 6.00 3.57 3.94 5.16
9 4.29 4.00 1.43 3.50 3.31
12 4.29 3.33 2.86 2.16 3.16
15 2.86 2.67 1.29 1.49 2.08
Table 11: The effect of Visual-Prompt Decoder layer depth \(Depth\). We report HTER(%) setting \(Depth\) to 0, 1, 2, 4, 8 and BOLD the best result in each column.
\(Depth\) OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
0 4.29 3.44 3.71 1.59 3.26
1 2.61 5.89 2.93 1.62 3.26
2 2.86 3.44 2.86 1.94 2.78
4 2.86 2.67 1.29 1.49 2.08
8 2.86 4.00 0.71 2.50 2.52
Table 12: The effect of Concept Mapping Loss coefficient \(\alpha\). We report HTER(%) setting \(\alpha\) to 0.1, 0.05, 0.01, 0.005, 0.001. We BOLD the best result in each column.
\(\alpha\) OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
0.1 4.29 6.67 5.71 3.81 5.12
0.05 2.86 4.56 2.29 1.50 2.80
0.01 2.86 2.67 1.29 1.49 2.08
0.005 4.05 4.56 3.00 2.00 3.40
0.001 4.29 4.00 5.00 2.44 3.93
Table 13: The effect of different backbone. We report HTER(%) using different ViT version including ViT-B/16, ViT-B/32 and ViT-L/14.
Backbone OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
ViT-B/16 2.86 2.67 1.29 1.49 2.08
ViT-B/32 4.52 3.89 9.29 2.68 5.10
ViT-L/14 1.43 1.33 1.50 1.50 1.44
Table 14: The effect of frame number \(r\) in VCE. We reproduce VCE process and report HTER(%) by using different settings (r=1,r=2,r=3) on P1 protocol.
r OCI \(\rightarrow\) M OMI \(\rightarrow\) C OCM \(\rightarrow\) I ICM \(\rightarrow\) O Avg.
1 2.86 2.67 1.43 1.69 2.16
2 2.86 2.67 1.29 1.49 2.08
3 2.86 2.78 1.36 1.69 2.17

5.3.3 Effect of Visual-Prompt Decoder Depth↩︎

Visual-Prompt Decoder (VPD) plays a crucial role in the Prompt-based Concept Injection (PCI) process (shown in 9). To investigate its impact on cross-domain performance, we experimented with different \(Depth\) of VPD layers, where 0 layers indicate replacing VPD with a simple 2D convolution layer. We observed that increasing the number of layers improves performance when \(Depth\) is below 4; however, when \(Depth\) is set to 8, the performance drops. Our analysis suggests that an overly complex network may make the decoder too strong, preventing useful information in the heatmaps from being effectively back-propagated to the learnable prompts, thereby leading to degraded performance. Moreover, an 8-layer network also introduces a significant increase in parameters and computational costs. Therefore, setting \(Depth\) to 4 offers the best trade-off, achieving optimal performance without excessive computational burden.

5.3.4 Effect of Concept Mapping Loss Coefficient↩︎

The coefficient of Concept Mapping Loss (CML) is also an important hyper-parameter of CPG-PAD method. The result of different coefficient \(\alpha\) is reported in 12. A small \(\alpha\) (0.005 and 0.001) dilutes the effect of CML with the basic cross-entropy loss, leading to performance degradation. On the other hand, a large \(\alpha\) (0.1) also results in degraded performance. This is because the concept-associated heatmaps are extracted from the visual feature space of the pretrained CLIP, whereas in the PAD task, the optimal visual space may not be perfectly aligned with that of pretrained CLIP. A large \(\alpha\) forces the model to over-align with the pretrained CLIP, which ultimately harms performance. Therefore, an intermediate value of \(\alpha=0.01\) proves to be the optimal choice.

5.3.5 Effect of Backbone↩︎

To study the effect of different backbone version, we use different ViT version as backbone including ViT-B/16, ViT-B/32 and ViT-L/14. The results are shown in 13. We can draw two conclusions from the results: (1) The best performance is achieved when using ViT-L/14 as the backbone, which is consistent with expectations since ViT-L/14 (\(\sim\)​304M parameters) has significantly more capacity than ViT-B/16 (\(\sim\)​86M). (2) Using ViT-B/32 leads to a noticeable performance degradation. This is mainly because, given a 224×224 input, it produces only 7×7 patches, which is much coarser compared to the 14×14 patches in ViT-B/16, resulting in a loss of fine-grained information.

5.3.6 Effect of VCE Frame Number↩︎

In the concept discovery stage of VCE, we sample \(r\) frames per video to discover visual concepts. To investigate the impact of \(r\), we conduct experiments with \(r=1,2,3\), and report the results in 14. The results show that the best value is \(r=2\). Using fewer frames may lead to information loss, while incorporating more frames can introduce redundant information and noise interference. Across different values of \(r\), CPG-PAD demonstrates stable performance with only a marginal variation.

Figure 8: T-SNE visualization of OCM \rightarrow I in P1 benchmark. We show T-SNE visualization of CLIP+MLP baseline and our proposed method CPG-PAD.

5.4 Visualization and Analysis↩︎

We compare CPG-PAD against the degraded model (CLIP+MLP) through T-SNE [85] visualization on OCM \(\rightarrow\) I scenarios in P1. As shown in 8, the CLIP+MLP baseline produces feature distributions where red (real) and blue (fake) samples are intermixed, and the dark-colored target samples are scattered widely within each class, reflecting large intra-class variation and unclear decision boundaries. By contrast, CPG-PAD forms two visibly distinct clusters of real and fake samples, and the target samples (dark points) are tightly grouped around their source counterparts (light points), showing improved compactness and cross-domain alignment.

Table 15: Computational overhead. We compare CPG-PAD with CLIP.
Model Training Time (batch=24) Inference Time (batch=64) Training Memory (batch=24) Inference Memory (batch=64)
CLIP 0.10s 0.07s 3196M 1490M
CPG-PAD 0.12s 0.08s 6708M 1632M

5.5 Computational Overhead↩︎

Table 15 shows the computational time for training and inference on a single RTX 3090 GPU, the results show that CPG-PAD is not a heavy burden for both training and inference. The comparison between CPG-PAD and CLIP demonstrates the applicability of CPG-PAD in real-world scenarios.

6 Conclusion↩︎

In this paper, we introduce Concept-informed Prompts Guided PAD (CPG-PAD), a novel framework that enhances the generalizability of PAD models by incorporating visual concepts from pretrained VLMs into PAD process. Concretely, we propose Visual Concept-driven Enhancement (VCE) which can discover PAD-relevant visual concepts and enhance domain data with fine-grained feature heatmaps corresponding to each discovered concept. With the help of these concept-associated heatmaps, Prompt-based Concept Injection (PCI) encourages the model to learn multiple concept-informed prompts via the cooperation of Visual-Prompt Decoder (VPD) and concept mapping loss, allowing CPG-PAD to precisely identify generalizable features when encountering unseen target domains. Extensive cross-domain quantitative experiments and in-depth analyses demonstrate the effectiveness of the proposed CPG-PAD method.

References↩︎

[1]
Z. Zhang, J. Yan, S. Liu, Z. Lei, D. Yi, and S. Z. Li, “A face antispoofing database with diverse attacks,” in 2012 5th IAPR international conference on biometrics (ICB), 2012, pp. 26–31.
[2]
I. Chingovska, A. Anjos, and S. Marcel, “On the effectiveness of local binary patterns in face anti-spoofing,” in 2012 BIOSIG-proceedings of the international conference of biometrics special interest group (BIOSIG), 2012, pp. 1–7.
[3]
A. Liu et al., “Contrastive context-aware learning for 3d high-fidelity mask face presentation attack detection,” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 2497–2507, 2022.
[4]
G. Kim, S. Eum, J. K. Suhr, D. I. Kim, K. R. Park, and J. Kim, “Face liveness detection based on texture and frequency analyses,” in 2012 5th IAPR international conference on biometrics (ICB), 2012, pp. 67–72.
[5]
J. Yang, Z. Lei, S. Liao, and S. Z. Li, “Face liveness detection with component dependent descriptor,” in 2013 international conference on biometrics (ICB), 2013, pp. 1–6.
[6]
Z. Zhang, D. Yi, Z. Lei, and S. Z. Li, “Face liveness detection by learning multispectral reflectance distributions,” in 2011 IEEE international conference on automatic face & gesture recognition (FG), 2011, pp. 436–441.
[7]
K. Srivatsan, M. Naseer, and K. Nandakumar, “FLIP: Cross-domain face anti-spoofing with language guidance,” in Proceedings of the IEEE/CVF international conference on computer vision (ICCV), 2023, pp. 19685–19696.
[8]
A. Liu et al., “Cfpl-fas: Class free prompt learning for generalizable face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 222–232.
[9]
Q. Zhou et al., “Instance-aware domain generalization for face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 20453–20463.
[10]
B. M. Le and S. S. Woo, “Gradient alignment for cross-domain face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 188–199.
[11]
M. Ghifary, W. B. Kleijn, M. Zhang, D. Balduzzi, and W. Li, “Deep reconstruction-classification networks for unsupervised domain adaptation,” in Computer vision–ECCV 2016: 14th european conference, amsterdam, the netherlands, october 11–14, 2016, proceedings, part IV 14, 2016, pp. 597–613.
[12]
A. Radford et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning, 2021, pp. 8748–8763.
[13]
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 618–626.
[14]
A. Ghorbani, J. Wexler, J. Y. Zou, and B. Kim, “Towards automatic concept-based explanations,” Advances in neural information processing systems, vol. 32, 2019.
[15]
T. Fel et al., “Craft: Concept recursive activation factorization for explainability,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 2711–2721.
[16]
T. de Freitas Pereira, A. Anjos, J. M. De Martino, and S. Marcel, “LBP- TOP based countermeasure against face spoofing attacks,” in Asian conference on computer vision, 2012, pp. 121–132.
[17]
Z. Boulkenafet, J. Komulainen, and A. Hadid, “Face spoofing detection using colour texture analysis,” IEEE Transactions on Information Forensics and Security, vol. 11, no. 8, pp. 1818–1830, 2016.
[18]
J. Komulainen, A. Hadid, and M. Pietikäinen, “Context based face anti-spoofing,” in 2013 IEEE sixth international conference on biometrics: Theory, applications and systems (BTAS), 2013, pp. 1–8.
[19]
K. Patel, H. Han, and A. K. Jain, “Secure face unlock: Spoof detection on smartphones,” IEEE transactions on information forensics and security, vol. 11, no. 10, pp. 2268–2283, 2016.
[20]
Y. Liu, A. Jourabloo, and X. Liu, “Learning deep models for face anti-spoofing: Binary or auxiliary supervision,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 389–398.
[21]
Z. Yu et al., “Searching central difference convolutional networks for face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 5295–5305.
[22]
Z. Wang et al., “Deep spatial gradient and temporal depth learning for face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 5042–5051.
[23]
Z. Wang, Q. Wang, W. Deng, and G. Guo, “Face anti-spoofing using transformers with relation-aware mechanism,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 4, no. 3, pp. 439–450, 2022.
[24]
A. George and S. Marcel, “On the effectiveness of vision transformers for zero-shot face anti-spoofing,” in 2021 IEEE international joint conference on biometrics (IJCB), 2021, pp. 1–8.
[25]
H.-P. Huang et al., “Adaptive transformers for robust few-shot cross-domain face anti-spoofing,” in European conference on computer vision, 2022, pp. 37–54.
[26]
G. Wang, H. Han, S. Shan, and X. Chen, “Unsupervised adversarial domain adaptation for cross-domain face presentation attack detection,” IEEE Transactions on Information Forensics and Security, vol. 16, pp. 56–69, 2020.
[27]
J. Wang, J. Zhang, Y. Bian, Y. Cai, C. Wang, and S. Pu, “Self-domain adaptation for face anti-spoofing,” in Proceedings of the AAAI conference on artificial intelligence, 2021, vol. 35, pp. 2746–2754.
[28]
Y. Liu et al., “Towards unsupervised domain generalization for face anti-spoofing,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 20654–20664.
[29]
S. Liu et al., “Dual reweighting domain generalization for face presentation attack detection,” arXiv preprint arXiv:2106.16128, 2021.
[30]
R. Shao, X. Lan, and P. C. Yuen, “Regularized fine-grained meta face anti-spoofing,” in Proceedings of the AAAI conference on artificial intelligence, 2020, vol. 34, pp. 11974–11981.
[31]
Q. Zhou, K.-Y. Zhang, T. Yao, X. Lu, S. Ding, and L. Ma, “Test-time domain generalization for face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 175–187.
[32]
C. Hu, K.-Y. Zhang, T. Yao, S. Ding, and L. Ma, “Rethinking generalizable face anti-spoofing via hierarchical prototype-guided distribution refinement in hyperbolic space,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 1032–1041.
[33]
Z. Wang et al., “Exploiting temporal and depth information for multi-frame face anti-spoofing,” arXiv preprint arXiv:1811.05118, 2018.
[34]
P.-K. Huang, M.-C. Chin, and C.-T. Hsu, “Face anti-spoofing via robust auxiliary estimation and discriminative feature learning,” in Asian conference on pattern recognition, 2021, pp. 443–458.
[35]
Y. Bian, P. Zhang, J. Wang, C. Wang, and S. Pu, “Learning multiple explainable and generalizable cues for face anti-spoofing,” in ICASSP 2022-2022 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2022, pp. 2310–2314.
[36]
S. Pan, S. Hoque, and F. Deravi, “An attention-guided framework for explainable biometric presentation attack detection,” Sensors, vol. 22, no. 9, p. 3365, 2022.
[37]
L. J. Gonzalez-Soler, J. E. Tapia, and C. Busch, “Are foundation models all you need for zero-shot face presentation attack detection?” in 2025 IEEE 19th international conference on automatic face and gesture recognition (FG), 2025, pp. 1–10.
[38]
G. Ozgur, E. Caldeira, T. Chettaoui, F. Boutros, R. Ramachandra, and N. Damer, “FoundPAD: Foundation models reloaded for face presentation attack detection,” in Proceedings of the winter conference on applications of computer vision, 2025, pp. 745–755.
[39]
J. Guo et al., “Style-conditional prompt token learning for generalizable face anti-spoofing,” in Proceedings of the 32nd ACM international conference on multimedia, 2024, pp. 994–1003.
[40]
J. Guo et al., “Domain generalization for face anti-spoofing via content-aware composite prompt engineering,” IEEE Transactions on Multimedia, 2025.
[41]
X. Wang et al., “TF-FAS: Twofold-element fine-grained semantic guidance for generalizable face anti-spoofing,” in European conference on computer vision, 2024, pp. 148–168.
[42]
A. Liu et al., “DGPDL: Domain-guided prompt distribution learning for generalizable face anti-spoofing,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026.
[43]
A. Liu et al., “ICPE-FAS: Instance and category prompts engineering for generalizable face anti-spoofing: A. Liu et al.” International Journal of Computer Vision, vol. 134, no. 6, p. 292, 2026.
[44]
H. G. Ramaswamy et al., “Ablation-cam: Visual explanations for deep convolutional network via gradient-free localization,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2020, pp. 983–991.
[45]
A. Chattopadhay, A. Sarkar, P. Howlader, and V. N. Balasubramanian, “Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks,” in 2018 IEEE winter conference on applications of computer vision (WACV), 2018, pp. 839–847.
[46]
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek, “On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation,” PloS one, vol. 10, no. 7, p. e0130140, 2015.
[47]
V. Petsiuk, “Rise: Randomized input sampling for explanation of black-box models,” arXiv preprint arXiv:1806.07421, 2018.
[48]
B. Kim et al., “Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),” in International conference on machine learning, 2018, pp. 2668–2677.
[49]
C. H. Ding, T. Li, and M. I. Jordan, “Convex and semi-nonnegative matrix factorizations,” IEEE transactions on pattern analysis and machine intelligence, vol. 32, no. 1, pp. 45–55, 2008.
[50]
T. Fel et al., “A holistic approach to unifying automatic concept extraction and concept importance estimation,” Advances in Neural Information Processing Systems, vol. 36, pp. 54805–54818, 2023.
[51]
A. Vaswani et al., “Attention is all you need,” in Proceedings of the 31st international conference on neural information processing systems, 2017, pp. 6000–6010.
[52]
H. W. Kuhn, “The hungarian method for the assignment problem,” Naval research logistics quarterly, vol. 2, no. 1–2, pp. 83–97, 1955.
[53]
Y. Zhang et al., “Celeba-spoof: Large-scale face anti-spoofing dataset with rich annotations,” in Computer vision–ECCV 2020: 16th european conference, glasgow, UK, august 23–28, 2020, proceedings, part XII 16, 2020, pp. 70–85.
[54]
A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune, “Synthesizing the preferred inputs for neurons in neural networks via deep generator networks,” Advances in neural information processing systems, vol. 29, 2016.
[55]
S. Liu, S. Lu, H. Xu, J. Yang, S. Ding, and L. Ma, “Feature generation and hypothesis verification for reliable face anti-spoofing,” in Proceedings of the AAAI conference on artificial intelligence, 2022, vol. 36, pp. 1782–1791.
[56]
Q. Zhou et al., “Generative domain adaptation for face anti-spoofing,” in European conference on computer vision, 2022, pp. 335–356.
[57]
Z. Wang et al., “Domain generalization via shuffled style assembly for face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 4123–4133.
[58]
Y. Sun, Y. Liu, X. Liu, Y. Li, and W.-S. Chu, “Rethinking domain generalization for face anti-spoofing: Separability and alignment,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 24563–24574.
[59]
C.-H. Liao, W.-C. Chen, H.-T. Liu, Y.-R. Yeh, M.-C. Hu, and C.-S. Chen, “Domain invariant vision transformer learning for face anti-spoofing,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, pp. 6098–6107.
[60]
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” Int. J. Comput. Vision, vol. 130, no. 9, pp. 2337–2348, Sep. 2022, doi: 10.1007/s11263-022-01653-1.
[61]
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learning for vision-language models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16816–16825.
[62]
H.-P. Huang et al., “Adaptive transformers for robust few-shot cross-domain face anti-spoofing,” in European conference on computer vision, 2022, pp. 37–54.
[63]
X. Hu et al., “Fine-grained prompt learning for face anti-spoofing,” in Proceedings of the 32nd ACM international conference on multimedia, 2024, pp. 7619–7628.
[64]
G. Zhang et al., “Interpretable face anti-spoofing: Enhancing generalization with multimodal large language models,” arXiv preprint arXiv:2501.01720, 2025.
[65]
ISO/IEC JTC1 SC37 Biometrics, ISO/IEC 30107-3. Information technology - biometric presentation attack detection - part 3: Testing and reporting. International Organization for Standardization, 2023.
[66]
D. Wen, H. Han, and A. K. Jain, “Face spoof detection with image distortion analysis,” IEEE Transactions on Information Forensics and Security, vol. 10, no. 4, pp. 746–761, 2015.
[67]
Z. Zhang, J. Yan, S. Liu, Z. Lei, D. Yi, and S. Z. Li, “A face antispoofing database with diverse attacks,” in 2012 5th IAPR international conference on biometrics (ICB), 2012, pp. 26–31.
[68]
I. Chingovska, A. Anjos, and S. Marcel, “On the effectiveness of local binary patterns in face anti-spoofing,” in 2012 BIOSIG-proceedings of the international conference of biometrics special interest group (BIOSIG), 2012, pp. 1–7.
[69]
Z. Boulkenafet, J. Komulainen, L. Li, X. Feng, and A. Hadid, “OULU-NPU: A mobile face presentation attack database with real-world variations,” in 2017 12th IEEE international conference on automatic face & gesture recognition (FG 2017), 2017, pp. 612–618.
[70]
S. Zhang et al., “Casia-surf: A large-scale multi-modal benchmark for face anti-spoofing,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 2, no. 2, pp. 182–193, 2020.
[71]
A. Liu, Z. Tan, J. Wan, S. Escalera, G. Guo, and S. Z. Li, “Casia-surf cefa: A benchmark for multi-modal cross-ethnicity face anti-spoofing,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2021, pp. 1179–1187.
[72]
A. George, Z. Mostaani, D. Geissenbuhler, O. Nikisins, A. Anjos, and S. Marcel, “Biometric face presentation attack detection with multi-channel convolutional neural network,” IEEE transactions on information forensics and security, vol. 15, pp. 42–55, 2019.
[73]
X. Guo, Y. Liu, A. Jain, and X. Liu, “Multi-domain learning for updating face anti-spoofing models,” in European conference on computer vision, 2022, pp. 230–249.
[74]
E. Tzeng, J. Hoffman, K. Saenko, and T. Darrell, “Adversarial discriminative domain adaptation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7167–7176.
[75]
L. Hu, M. Kan, S. Shan, and X. Chen, “Duplex generative adversarial network for unsupervised domain adaptation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 1498–1507.
[76]
H. Li, W. Li, H. Cao, S. Wang, F. Huang, and A. C. Kot, “Unsupervised domain adaptation for face anti-spoofing,” IEEE Transactions on Information Forensics and Security, vol. 13, no. 7, pp. 1794–1809, 2018.
[77]
Y. Jia, J. Zhang, S. Shan, and X. Chen, “Unified unsupervised and semi-supervised domain adaptation network for cross-scenario face anti-spoofing,” Pattern Recognition, vol. 115, p. 107888, 2021.
[78]
H. Yue et al., “Cyclically disentangled feature translation for face anti-spoofing,” in Proceedings of the AAAI conference on artificial intelligence, 2023, vol. 37, pp. 3358–3366.
[79]
Y. Jia, J. Zhang, S. Shan, and X. Chen, “Single-side domain generalization for face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 8484–8493.
[80]
R. Cai, Z. Li, R. Wan, H. Li, Y. Hu, and A. C. Kot, “Learning meta pattern for face anti-spoofing,” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 1201–1213, 2022.
[81]
Y. Liu, Y. Chen, W. Dai, C. Li, J. Zou, and H. Xiong, “Causal intervention for generalizable face anti-spoofing,” in 2022 IEEE international conference on multimedia and expo (ICME), 2022, pp. 01–06.
[82]
Z.-W. Hong, Y.-C. Lin, H.-T. Liu, Y.-R. Yeh, and C.-S. Chen, “Domain-generalized face anti-spoofing with unknown attacks,” in 2023 IEEE international conference on image processing (ICIP), 2023, pp. 820–824.
[83]
S.-Q. Liu, Q. Wang, and P. C. Yuen, “Bottom-up domain prompt tuning for generalized face anti-spoofing,” in European conference on computer vision, 2024, pp. 170–187.
[84]
H. Zhang et al., “From intuition to investigation: A tool-augmented reasoning MLLM framework for generalizable face anti-spoofing,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2026, pp. 40855–40865.
[85]
L. Van der Maaten and G. Hinton, “Visualizing data using t-SNE.” Journal of machine learning research, vol. 9, no. 11, 2008.

  1. Haoyuan Zhang is with the School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China; the State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China (e-mail: zhanghaoyuan2023@ia.ac.cn).↩︎

  2. Xiangyu Zhu, Ajian Liu and Siran Peng are with the State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; the School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China (e-mails: xiangyu.zhu@ia.ac.cn, ajian.liu@ia.ac.cn, pengsiran2023@ia.ac.cn).↩︎

  3. Li Gao is with the China Mobile Financial Technology Co., Ltd., Beijing 100032, China (e-mail: gaolids@chinamobile.com).↩︎

  4. Zhen Lei is with the State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; the School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China; the Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences, Hong Kong, China; the School of Computer Science and Engineering, the Faculty of Innovation Engineering, Macau University of Science and Technology, Macau, China (e-mail: zhen.lei@ia.ac.cn).↩︎

  5. Corresponding author: Zhen Lei.↩︎

  6. https://github.com/OpenAI/CLIP↩︎