Quantifying Training Membership Information in the
Hyperspherical Embedding Geometry of Face Recognition Models

nsal Öztürk\(^{1}\)Sébastien Marcel\(^{1,2}\)
\(^{1}\)Idiap Research Institute, Martigny, Switzerland
\(^{2}\)Université de Lausanne (UNIL), Lausanne, Switzerland
{unsal.ozturk, sebastien.marcel}@idiap.ch


Abstract

Face recognition models represent each face as an embedding vector on the unit hypersphere by clustering embeddings of the same identity while pushing different identities apart through angular-margin losses. Because these losses act only on training identities, non-member identities may form clusters with different geometric properties. In this paper, we quantify the magnitude of this difference and what training-time factors control it. We compute four statistics based on cluster geometry across 180 face recognition models in a factorial design over IResNet backbone size, loss head, training duration, and the number of training identities, and evaluate each configuration on nine benchmarks. Our results indicate that the number of training identities has the largest effect on member/non-member separability, while backbone and loss head contribute far less, and that, on a same-domain held-out reference, the geometric membership signal decreases monotonically as more identities are added to training. We provide an analysis of cross-domain (pose, age, quality, ethnicity) non-member benchmarks and report that these inflate the apparent membership signal. Finally, we fuse all four statistics with a learned classifier to reveal additional membership information beyond the best individual statistic.

1

1 Introduction↩︎

Figure 1: Measurement pipeline. Per-identity statistics (s_1, s_2, \ldots) are computed from the embedding clusters of a face recognition model. If the distributions of these statistics differ between member identities (present in the training set) and non-member identities (absent from it) as shown on the left, the explicit cluster geometry carries membership information on these statistics; if they coincide (right), it does not.

Modern face recognition (FR) systems are trained on large-scale datasets of face crops, typically collected from the internet [1]. To represent a person, the system takes multiple crops of that person’s face and passes each through a deep neural network that produces a 512-dimensional embedding vector, normalised to lie on the unit hypersphere. The training of these models is primarily driven by angular-margin losses [2][4], which pull the embeddings of crops belonging to the same identity into a cluster on the hypersphere while pushing apart the clusters of different identities. During training, a classification head maps embeddings to identity logits and this head is discarded after training. When deployed, the system typically uses only the backbone neural network. Enrolled faces are stored as embeddings, and the similarity between two faces is measured by the inner product of their corresponding embeddings.

As angular-margin training directly optimises the compactness of each identity’s cluster, we pose the following question: how much information about training-set membership is retained in the geometry of these embedding clusters? If training identities form systematically tighter or better-separated clusters than unseen identities, then the embedding geometry carries a residual membership signal.

Quantifying this residual signal is related to membership inference attacks (MIA), a line of work investigating whether a given sample or identity was part of a model’s training data [5], [6]. In practice, membership inference serves as an auditing tool for data privacy by allowing individuals or regulators to verify whether a model was trained on data belonging to a specific person in the context of data protection rights such as the right to erasure under the GDPR [7]. MIA literature primarily focuses on closed-set classifiers that expose prediction confidences, loss values, or the model weights directly [5], [6], [8], [9]. In our setting, we consider open-set face recognition where neither logits nor losses nor model weights are available, so the membership signal must be sought entirely in the geometry of the embeddings themselves.

To measure this signal, we compute four different statistics capturing different aspects of cluster geometry on the hypersphere and measure how much they differ between training and unseen identities (Figure 1). We train 180 face recognition models from scratch (Section 4.2) and evaluate each model on nine benchmarks, including a same-domain held-out split that controls for the pose, age, quality, and ethnicity shifts present in the other benchmarks. To our knowledge, this constitutes the largest systematic study of membership information in face embedding geometry. Our contributions are: (1) we quantify the membership information retained in four geometric statistics computed from face embeddings; (2) we identify the primary factors that control how much of this information persists; (3) we show that cross-domain non-member benchmarks (e.g.differing in pose or age) inflate the apparent membership signal by conflating domain shift with training-attributable differences; (4) we demonstrate that fusing all statistics captures membership information beyond the best individual statistic.

2 Related Work↩︎

Membership inference on classifiers. Membership inference literature targets closed-set classifiers, where the attacker can observe prediction confidences, logits, or loss values of the target model. In the open-set face recognition setting, however, none of these signals are available and this has motivated a line of work operating directly on embeddings. Shokri et al. [5] introduce the shadow-model paradigm and train substitute classifiers to learn a binary meta-classifier on output distributions. Salem et al. [8] relax many of these assumptions and show that a single shadow model and only the top-\(k\) predictions often suffice. Carlini et al. [6] recast MIA as a likelihood-ratio attack (LiRA) and derive near-optimal attacks. They show that per-example hardness varies by orders of magnitude. Hu et al. [9] provide a comprehensive survey.

3pt

@ l l c c c @ Work & Domain & Level & # Mod. & Fact.

Li  [10] & FR / re-identification & Identity & 6 &
FACE-AUD. [11] & FR / few-shot & Identity & 4 &
SLMIA-SR [12] & Speaker recognition & Identity & 5 &
EncoderMI [13] & Contrastive learning & Sample & 4 &

MINT [14] & FR & Sample & 3 &
DeAlcala [15] & FR & Sample & 9 & \(^\ast\)
Mancera [16] & Object classification & Sample & 4 &
Huang [17] & FR & Sample & 2 &
Ours & FR & Identity & 180 &


\(^\ast\)Varies factors individually but not in a fully crossed design.

Membership inference on embedding models. Existing embedding-based approaches differ in the granularity of the membership question: some ask whether a specific image was used for training (sample level), while others ask whether any image of a given person was included (identity level). Li et al. [10] formalise identity-level inference on metric-embedding models and use intra-similarity and inter-dissimilarity of probe embeddings as attack features on person re-identification and face recognition. Chen et al. [11] propose FACE-AUDITOR, a system that queries few-shot face recognition models with designed probing sets and uses the returned similarity scores as auditing features. They report accuracies up to 99%. SLMIA-SR [12] addresses the analogous problem in speaker recognition by training shadow models with controlled membership mixing ratios. All three operate at the identity level. EncoderMI [13], by contrast, operates at the sample level on contrastive-learning encoders including CLIP and exploits the observation that an overfitted encoder produces more similar representations for augmented views of a training sample than for an unseen one.

White-box approaches on face recognition. A parallel line of work allows access to model internals. The motivation is that intermediate representations such as activation maps, batch-normalisation statistics, or gradient signals may carry a stronger membership trace than the final embedding alone, at the cost of requiring more privileged access to the model. These methods also typically operate at the sample level, determining whether a specific image (rather than a person) was included in training. The Membership Inference Test (MINT) of DeAlcala et al. [14] trains an MLP or CNN on intermediate activation maps to classify individual face images as member or non-member and achieves up to 90% accuracy on three ResNet-100 models. A companion study [15] analyses factors that affect MINT performance, including loss function, dropout, and number of training epochs. Mancera et al. [16] extend MINT to object classification and report 70–80% precision depending on layer depth. Huang et al. [17] exploit distances between intermediate features and batch-normalisation parameters for sample-level membership and model-inversion attacks [18]. Even with this level of access, Rezaei and Liu [19] show that the advantage over random guessing can be modest once the model generalises well.

Memorisation and embedding geometry. The question of why a geometric membership signal might exist at all connects to a discussion on memorisation in deep networks. If a model must memorise certain training examples to achieve low error, particularly on the tail of the data distribution, then its internal representations and output embeddings may carry traces of those examples. Feldman [20] formalises this argument, showing theoretically that long-tailed distributions necessitate memorisation for near-optimal generalisation. Angular-margin losses [2][4] amplify this effect, potentially making training identities geometrically distinguishable from unseen ones. Unlearning methods [21], [22] and differential privacy [23] aim to erase or bound such signals.

Comparison with this work. Our contribution is a controlled measurement study of how much membership information geometric statistics carry, not a stronger attack feature. Table ¿tbl:tab:related-summary? summarises the literature. Prior embedding-level work proposes specific attack features and demonstrates that membership inference is feasible through evaluating a small number of models trained under a single configuration. Our study differs in three respects. First, the goal: we measure how much membership information is retained in the explicit cluster geometry and identify which training factors control the amount that persists. Second, the experimental scale: 180 models, whereas the largest prior study evaluates nine. Third, domain control: we add a same-domain held-out reference, which aims to separate training-attributable signal from domain-shift effects. Several of the per-identity statistics we evaluate have already been used as features in prior attacks. For instance, the intra-similarity score proposed by Li et al. [10] is equivalent to the mean pairwise cosine similarity that we include as a baseline statistic (defined in Section 3.2).

3 Methodology↩︎

3.1 Access Model and Measurement Procedure↩︎

We assume that an auditor can query an FR model \(f\) with \(k \geq 2\) probe images of a target identity \(y\) and obtain the corresponding \(\ell_2\)-normalised embeddings. From the resulting cluster, the auditor computes four per-identity statistics (Section 3.2) and uses them to assess if the cluster geometry carries information about \(y\)’s membership in the training set. For each model we train, we compute these statistics on member identities (from WebFace4M) and on non-member identities drawn from the benchmarks in Table 1.

Table 1: Dataset overview for membership evaluation. “Eval.subj.”: identities with \({\geq}2\) crops (required for per-identity cluster statistics). \(\dagger\)For WebFace4M this is the held-out split we sample for evaluation; the dataset itself contains far more identities with \({\geq}2\) crops.
Dataset Primary variation Subjects Eval.subj. Images Covariate metadata
LFW [24] Unconstrained (baseline) 5,749 1,680 13,233 (subject name only)
CPLFW [25] Cross-pose (subset of LFW identities) 3,929 3,909 11,652 yaw (estimated from landmarks)
CFP-FF [26] Frontal-only probes 500 500 7,000 pose \(\in\)
CFP-FP [26] Frontal+profile probes 500 7,000 pose \(\in\)
AgeDB-30 [27] Large intra-identity age gap (\(\geq\)30 yr) 440 440 12,240 age (3–100 yr), gender
XQLFW [28] Low image quality (paired with LFW) 5,749 1,672 13,233 (quality implicit by design)
RFW [29] Ethnicity-balanced (4 groups) 11,429 11,429 40,607 ethnicity, gender, nationality
IJB-C [30] Unconstrained, mixed image+video 3,531 3,531 469,376 ethnicity, gender, nationality
WebFace4M [1] Large-scale training source 205,990 \({\leq}\,5{,}000\)\(\dagger\) 4,000,000 (subject ID only)

3pt

For each statistic, we obtain a scalar score per identity. To quantify how much membership information a statistic carries, we sweep a decision threshold \(\tau\) across the range of observed values and classify an identity as a member if its score exceeds \(\tau\) (or falls below \(\tau\) for statistics where lower values indicate membership). We report the area under the resulting ROC curve (AUC), hereafter threshold-classifier AUC, as a threshold-free summary, which is equal to the probability that a randomly chosen member scores higher than a randomly chosen non-member. This serves as our primary measure of the membership information carried by a given statistic. We use AUC rather than the true positive rate at a fixed false positive rate, because our framing concerns separability across all decision thresholds rather than the operating point most relevant for a deployable attack. Of the four statistics defined in the next section, two require only the embeddings and two additionally require the training-time scale and margin hyperparameters. We refer to these two access levels as black-box (embeddings only) and grey-box (embeddings plus \(s\) and \(m\)). The grey-box setting is realistic when models are released alongside their training configurations.

3.2 Per-Identity Geometric Statistics↩︎

Given an identity \(y\) with \(k \geq 2\) probe images, we extract embeddings \(\{e_1, \ldots, e_k\}\) via the target model and compute four statistics from the resulting cluster. Each statistic captures a different aspect of cluster shape. The underlying intuition is that an angular-margin loss explicitly optimises the representation of training identities, so members are expected to form tighter clusters on the embedding hypersphere than non-members. If that expectation holds, the statistic will carry membership information; the question is how much, and what factors control it.

Pairwise cosine similarity (PairCos). The mean cosine similarity over pairs of embeddings within the identity: \[\bar{s} = \binom{k}{2}^{-1} \sum_{i < j} \cos(e_i, e_j).\] Higher values indicate a more coherent cluster.

vMF concentration \(\kappa\). We model the per-identity embedding distribution as a von Mises-Fisher (vMF) distribution on the unit hypersphere and estimate the concentration parameter \(\kappa\) via maximum likelihood. We use the closed-form approximation of Banerjee et al. [31]: \[\hat{\kappa}_0 = \frac{\bar{R}\,(d - \bar{R}^2)}{1 - \bar{R}^2}, \qquad \bar{R} = \left\| \frac{1}{k}\sum_{i=1}^{k} e_i \right\|,\] where \(d = 512\) is the embedding dimensionality and \(\bar{R}\) is the mean resultant length. Higher \(\hat{\kappa}\) corresponds to greater directional concentration. Note that ordering identities by \(\hat{\kappa}\) is equivalent to ordering them by cluster inertia \(I = k^{-1}\sum_i \|e_i - \mu\|^2\), where \(\mu = k^{-1}\sum_i e_i\) is the cluster mean. For \(\ell_2\)-normalised embeddings, \(I\) reduces to \(1 - \bar{R}^2\), and \(\hat{\kappa}_0\) is strictly decreasing in \(1-\bar{R}^2\), so the two quantities produce identical identity rankings.

Penalised logit (PenLogit). To approximate the margin-penalised logit that the training loss would compute for the cluster, we use a leave-one-out scheme. For each embedding \(e_i\), we compute a prototype \(\hat{w}_{\backslash i} = \bar{e}_{\backslash i}/\|\bar{e}_{\backslash i}\|\) as the \(\ell_2\)-normalised mean of the remaining \(k{-}1\) embeddings, and the margin-penalised logit is \[\ell_i = s \cdot g\left(\cos\angle\left(e_i, \hat{w}_{\backslash i}\right),\; m\right),\] where \(s\) and \(m\) are the scale and margin of the loss head, and \(g\) is the head-specific penalisation function (e.g. \(g(\cos\theta, m) = \cos(\theta + m)\) for ArcFace). The statistic value is the mean \(\bar{\ell} = k^{-1}\sum_i \ell_i\). A higher penalised logit indicates that the embeddings are tightly aligned to their own prototype.

Prototype softmax CE (ProtoCE). We construct a surrogate of the training loss. For each identity \(y\), we compute a class prototype \(\hat{w}_y = \mu / \|\mu\|\) and treat all other identities in the evaluation set as negative classes. The cross-entropy loss with the margin-modified logit for the target class and standard cosine logits for the negatives gives \[\mathcal{L}_y = -\log \frac{\exp\bigl(s\cdot g(\cos\theta_y,\, m)\bigr)}{\exp\bigl(s\cdot g(\cos\theta_y,\, m)\bigr) + \sum_{j \neq y} \exp(s\cdot\cos\theta_j)}\] We report \(-\mathcal{L}_y\) so that higher values indicate membership, consistent with the sign convention of the other statistics. Because the softmax denominator depends on the composition and size of the evaluation set, ProtoCE values are not directly comparable across benchmarks with different identity pools. Throughout, all AUC values are oriented so that \(\text{AUC} > 0.5\) indicates better-than-chance separation.

4 Experimental Setup↩︎

4.1 Training Data↩︎

All models are trained on WebFace4M [1], a large-scale face dataset containing approximately 4 million images of 205,990 identities. WebFace4M is well-suited to this study because its size permits creating training subsets that span several orders of magnitude in identity count while still providing sufficient images per identity for meaningful embedding statistics. Every probe image is aligned to a canonical \(112\!\times\!112\) crop with the standard InsightFace five-landmark pipeline.

4.2 Ablation Design↩︎

We vary four factors in a fully crossed design.

Backbone capacity. We train three IResNet architectures: IResNet-34, IResNet-50, and IResNet-100 [2], [32].

Angular-margin loss head. We pair each backbone with three angular-margin loss heads: ArcFace [2] (\(s\!=\!64\), \(m\!=\!0.5\)), CosFace [3] (\(s\!=\!64\), \(m\!=\!0.4\)), and MagFace [4] (\(s\!=\!64\), per-sample \(m \in [0.45, 0.8]\) depending on the un-normalised feature magnitude).

Training-set size (\(n_{\text{ids}}\)). To study the effect of identity coverage, we create five partitions by randomly sampling \(n_{\text{ids}}\in \{1\text{K}, 10\text{K}, 50\text{K}, 100\text{K}\}\) identities (with a fixed seed), plus one partition that uses all available identities. The identities not selected for training serve as same-domain non-members in the held-out WebFace4M benchmark.

Training duration. We train the backbones for 5, 10, 20, and 40 epochs to track the evolution of cluster-statistic separability. Longer training exposes each identity to more gradient updates and may therefore affect the gap between member and non-member statistic distributions.

The full grid spans \(3~\text{backbones} \times 3~\text{losses} \times 5~n_{\text{ids}}\text{ levels} \times 4~\text{epochs} = 180\) model configurations. All models are trained with SGD (momentum 0.9, weight decay \(5\!\times\!10^{-4}\)), a polynomial-decay learning-rate schedule starting at \(0.1\) with one warmup epoch, batch size 128, mixed-precision (FP16), and random horizontal flips.

4.3 Evaluation Benchmarks↩︎

The nine benchmarks (Table 1) supply standard verification protocols for measuring model quality and provide non-member pools with diverse domain characteristics so that we can study how pose, quality, and demographics affect the apparent cluster-statistic separability.

Eight of the nine benchmarks are drawn from sources disjoint from WebFace4M (we verified that no identity overlap exists by cross-checking subject identifiers); Table 1 lists their sizes and available covariate metadata. Two dataset pairs form natural experiments (matched comparisons where a single attribute varies). CFP-FF vs.CFP-FP [26] contrasts frontal-only against frontal-plus-profile probes for the same 500 identities, so any statistic difference is attributable to pose. LFW [24] vs.XQLFW [28] shares the same identity set and pair protocol but degrades one image per pair to isolate image quality. CPLFW [25] (cross-pose LFW), AgeDB-30 [27] (age gap \(\geq\)​30 yr), RFW [29] (ethnicity-balanced, four groups), and IJB-C [30] complete the cross-domain set.

Held-out WebFace4M split. The ninth benchmark is constructed from WebFace4M itself: for each model trained on \(n_{\text{ids}}\) identities, we sample \(\min(n_{\text{ids}},\, 2{,}500)\) members and 2,500 non-members from the remaining subjects; at \(n_{\text{ids}}{=}\)all no non-members remain. As both groups share the same acquisition pipeline, AUCs are arguably attributable to the cluster-geometry difference between seen and unseen identities rather than to domain shift.

5 Results and Discussion↩︎

This section presents experimental results and analysis for the 180-model grid described in Section 4. We examine verification performance, the effect of training-set size and duration on membership separability, distributional confounds, and statistic fusion. Full per-epoch, per-model results and ROC/DET curves are provided in the supplementary material.

Table 2: What controls membership information in cluster geometry? Top: partial \(\hat{\eta}^2_p\) from a Type-II ANOVA (each factor tested after adjusting for all others) on the AUC of a threshold classifier separating member from non-member identities (\(N\!=\!180\) models); values close to 1 indicate that the factor explains nearly all variance in this separability. Bottom: distributional-shift effect sizes (Cohen’s \(d'\), Spearman \(\rho\), Kruskal–Wallis \(\epsilon^2\)) averaged over all models; larger absolute values indicate a stronger confound.
PairCos \(\kappa\) PenLogit ProtoCE
\(n_\mathrm{ids}\)
Epoch
Backbone
Loss
\(n_\mathrm{ids}{\times}\)Ep
Residual
\(n_\mathrm{ids}\):1K \(\cdot\) 10K \(\cdot\) 50K \(\cdot\) 100K \(\cdot\) all
Epoch:ep5 \(\cdot\) ep10 \(\cdot\) ep20 \(\cdot\) ep40
Cohen’s \(d'\)
XQLFW \(-\) LFW \(-1.64\) \(-0.48\) \(-1.81\) \(+1.13\)
CFP-FP \(-\) CFP-FF \(-2.27\) \(-1.13\) \(-2.98\) \(+1.14\)
CPLFW \(-\) LFW \(-1.72\) \(-0.72\) \(-2.06\) \(+0.31\)
1-5 Spearman \(\rho\)
AgeDB (age \(\sigma\)) \(+0.01\) \(-0.26\) \(+0.17\) \(+0.33\)
1-5 Kruskal-Wallis \(\epsilon^2\)
RFW (4 groups) \(0.017\) \(0.017\) \(0.018\) \(0.123\)

4pt

5.1 Verification Performance↩︎

We first establish that our model grid spans a meaningful range of verification quality. Table ¿tbl:tab:eer? reports the verification equal error rate (EER) at epoch 40 for every combination of backbone, loss head, and \(n_{\text{ids}}\) on the cross-domain benchmarks. The table is dense and provided primarily as a per-cell reference, and the discussion below focuses on patterns and points back to specific cells only as needed. Performance generally improves with \(n_{\text{ids}}\), with diminishing gains beyond \(100\)K identities. Once \(n_{\text{ids}}\) is held fixed, backbone depth provides small reductions in verification EER on domain-shifted benchmarks, and the choice of loss head has an even smaller effect. Notably, the highest cluster-statistic separability occurs at \(n_{\text{ids}}{=}1\)K, where verification performance is worst. As \(n_{\text{ids}}\) grows, verification improves while separability decreases. However, non-trivial separability persists even at \(n_{\text{ids}}\) levels that achieve competitive verification EER, so the patterns reported below are not confined to poorly performing models.

Table ¿tbl:tab:eer? also reports the threshold-classifier AUC for all four cluster statistics. No single statistic achieves the highest AUC across all benchmarks. PenLogit ranks first most often at moderate-to-high \(n_{\text{ids}}\), in particular on the pose- and quality-degraded sets, whereas ProtoCE attains the highest AUC on the frontal and high-quality sets (CFP-FF, LFW, IJB-C), so the best statistic remains dataset-dependent. Because the eight non-WebFace4M benchmarks differ from the training source in pose, quality, and demographics, the observed AUC conflates training-attributable and domain-induced differences, which we discuss in Section 5.3. All four statistics exhibit similar qualitative trends across factors, so we show only PairCos distributions in Figure 2.

3.5pt

@ c c @ |ccccc|ccccc|ccccc|ccccc|ccccc|ccccc|ccccc|ccccc @ & & & & & & & & &
(lr)3-7(lr)8-12(lr)13-17(lr)18-22(lr)23-27(lr)28-32(lr)33-37(lr)38-42 & & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All
& E & 7.7 & 1.3 & 0.4 & 0.2 & 0.3 & 21.3 & 12.8 & 7.3 & 5.8 & 5.3 & 39.6 & 20.7 & 10.7 & 9.0 & 8.4 & 34.6 & 17.4 & 6.3 & 3.9 & 2.8 & 26.3 & 7.7 & 1.8 & 1.1 & 1.0 & 9.6 & 1.4 & 0.3 & 0.3 & 0.1 & 27.1 & 9.6 & 3.6 & 3.1 & 2.8 & 8.3 & 2.9 & 1.7 & 1.5 & 1.3
& C & .99 & .92 & .56 & .71 & .77 & .99 & 1. & .96 & .91 & .89 & 1. & 1. & .90 & .80 & .75 & 1. & .99 & .78 & .61 & .52 & 1. & 1. & .94 & .83 & .76 & 1. & .99 & .53 & .71 & .78 & 1. & 1. & .96 & .84 & .75 & .99 & .94 & .72 & .62 & .58
& \(\kappa\) & .78 & .60 & .84 & .89 & .90 & .83 & .94 & .70 & .58 & .54 & .96 & .90 & .61 & .71 & .74 & .98 & .93 & .52 & .67 & .72 & 1. & 1. & .82 & .63 & .55 & 1. & .99 & .51 & .68 & .73 & 1. & .99 & .92 & .81 & .73 & .99 & .96 & .78 & .69 & .65
& L & 1. & .98 & .74 & .59 & .51 & 1. & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 & 1. & 1. & .90 & .79 & .73 & 1. & 1. & .97 & .91 & .87 & 1. & .96 & .54 & .72 & .79 & 1. & 1. & .93 & .81 & .73 & .96 & .90 & .63 & .53 & .52
& S & .78 & .59 & .86 & .90 & .90 & .93 & .93 & .63 & .54 & .55 & .83 & .82 & .78 & .83 & .83 & .99 & .97 & .56 & .63 & .66 & 1. & 1. & .65 & .58 & .61 & 1. & .98 & .55 & .72 & .73 & .99 & .98 & .85 & .70 & .66 & 1. & 1. & .93 & .84 & .83
(l)2-42 & E & 9.3 & 1.3 & 0.4 & 0.2 & 0.3 & 22.8 & 13.3 & 7.5 & 5.7 & 5.5 & 39.4 & 21.0 & 10.6 & 9.0 & 8.1 & 34.2 & 17.6 & 6.1 & 4.0 & 2.6 & 27.8 & 8.7 & 1.8 & 1.2 & 0.9 & 10.2 & 1.2 & 0.3 & 0.3 & 0.1 & 27.2 & 10.4 & 3.8 & 3.0 & 2.6 & 8.7 & 3.1 & 1.7 & 1.5 & 1.4
& C & .97 & .89 & .58 & .71 & .77 & .97 & .99 & .94 & .91 & .88 & 1. & 1. & .89 & .80 & .75 & 1. & .98 & .77 & .61 & .53 & 1. & 1. & .93 & .83 & .75 & 1. & .96 & .58 & .72 & .78 & 1. & 1. & .96 & .85 & .76 & .97 & .92 & .69 & .60 & .56
& \(\kappa\) & .62 & .51 & .86 & .89 & .91 & .65 & .88 & .61 & .53 & .50 & .83 & .80 & .69 & .76 & .78 & .94 & .89 & .57 & .69 & .72 & 1. & .99 & .77 & .59 & .52 & 1. & .96 & .55 & .69 & .73 & .99 & .99 & .92 & .81 & .74 & .98 & .94 & .76 & .68 & .64
& L & .99 & .97 & .71 & .59 & .51 & 1. & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 & 1. & .99 & .89 & .80 & .73 & 1. & 1. & .96 & .91 & .87 & .99 & .92 & .59 & .73 & .79 & 1. & 1. & .92 & .82 & .73 & .93 & .87 & .60 & .52 & .53
& S & .67 & .53 & .87 & .91 & .90 & .93 & .90 & .62 & .54 & .56 & .57 & .67 & .83 & .86 & .84 & .97 & .95 & .56 & .62 & .65 & .99 & .99 & .57 & .62 & .62 & 1. & .96 & .59 & .73 & .73 & .98 & .97 & .84 & .69 & .65 & .99 & 1. & .94 & .87 & .85
(l)2-42 & E & 8.9 & 1.2 & 0.5 & 0.4 & 0.2 & 23.0 & 13.8 & 7.5 & 5.9 & 5.1 & 40.1 & 20.6 & 10.7 & 8.9 & 8.3 & 34.3 & 17.4 & 6.2 & 4.0 & 2.8 & 28.5 & 7.9 & 1.9 & 1.2 & 0.9 & 10.5 & 1.4 & 0.3 & 0.2 & 0.2 & 27.4 & 10.2 & 3.6 & 2.9 & 2.8 & 8.7 & 2.9 & 1.6 & 1.4 & 1.3
& C & .99 & .91 & .58 & .72 & .77 & .99 & 1. & .96 & .91 & .89 & 1. & 1. & .89 & .81 & .75 & 1. & .99 & .77 & .60 & .52 & 1. & 1. & .94 & .84 & .76 & 1. & .98 & .56 & .72 & .78 & 1. & 1. & .96 & .84 & .75 & .99 & .94 & .71 & .62 & .57
& \(\kappa\) & .78 & .57 & .84 & .89 & .90 & .80 & .93 & .69 & .59 & .54 & .95 & .88 & .62 & .71 & .74 & .98 & .93 & .53 & .67 & .72 & 1. & 1. & .82 & .63 & .55 & 1. & .98 & .53 & .69 & .74 & 1. & .99 & .92 & .81 & .73 & .99 & .96 & .78 & .69 & .65
& L & 1. & .98 & .72 & .59 & .51 & 1. & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 & 1. & 1. & .89 & .79 & .72 & 1. & 1. & .97 & .92 & .87 & 1. & .95 & .57 & .73 & .79 & 1. & 1. & .92 & .81 & .73 & .96 & .89 & .62 & .53 & .52
& S & .79 & .60 & .86 & .90 & .90 & .93 & .95 & .72 & .63 & .64 & .88 & .86 & .70 & .78 & .79 & .99 & .97 & .64 & .55 & .59 & 1. & 1. & .72 & .50 & .53 & 1. & .97 & .62 & .76 & .75 & .99 & .99 & .86 & .73 & .70 & .99 & .99 & .87 & .79 & .78
& E & 8.2 & 1.2 & 0.4 & 0.2 & 0.2 & 21.0 & 12.5 & 7.0 & 5.2 & 4.5 & 39.3 & 20.6 & 10.0 & 8.3 & 7.7 & 34.3 & 17.5 & 6.0 & 3.5 & 2.1 & 26.0 & 7.9 & 1.6 & 0.9 & 0.6 & 9.5 & 1.5 & 0.3 & 0.1 & 0.2 & 26.4 & 9.5 & 3.9 & 3.0 & 2.5 & 8.2 & 3.0 & 1.5 & 1.4 & 1.2
& C & .99 & .92 & .50 & .69 & .76 & .99 & 1. & .96 & .92 & .89 & 1. & 1. & .92 & .82 & .75 & 1. & .99 & .83 & .64 & .53 & 1. & 1. & .96 & .86 & .77 & 1. & .99 & .55 & .67 & .76 & 1. & 1. & .98 & .88 & .78 & .99 & .95 & .75 & .65 & .59
& \(\kappa\) & .80 & .62 & .81 & .88 & .90 & .86 & .95 & .74 & .62 & .57 & .98 & .90 & .57 & .70 & .74 & .98 & .94 & .56 & .64 & .71 & 1. & 1. & .88 & .67 & .56 & 1. & .99 & .58 & .64 & .72 & 1. & 1. & .95 & .85 & .76 & .99 & .97 & .81 & .71 & .66
& L & 1. & .98 & .77 & .62 & .53 & 1. & 1. & .99 & .98 & .97 & 1. & 1. & .99 & .97 & .95 & 1. & 1. & .92 & .82 & .74 & 1. & 1. & .98 & .93 & .88 & 1. & .97 & .53 & .67 & .77 & 1. & 1. & .95 & .85 & .75 & .97 & .91 & .67 & .56 & .50
& S & .80 & .61 & .84 & .89 & .89 & .94 & .94 & .67 & .57 & .56 & .86 & .83 & .76 & .83 & .83 & .99 & .97 & .62 & .60 & .66 & 1. & 1. & .69 & .57 & .62 & 1. & .98 & .51 & .70 & .72 & .99 & .99 & .88 & .74 & .67 & 1. & 1. & .95 & .87 & .85
(l)2-42 & E & 9.4 & 1.5 & 0.3 & 0.2 & 0.1 & 22.0 & 13.9 & 7.3 & 5.5 & 4.9 & 39.5 & 20.2 & 9.6 & 8.6 & 7.7 & 34.9 & 17.2 & 5.9 & 3.3 & 2.1 & 27.9 & 8.5 & 1.4 & 0.9 & 0.6 & 10.3 & 1.3 & 0.4 & 0.1 & 0.1 & 27.3 & 10.2 & 3.5 & 2.7 & 2.8 & 8.8 & 3.0 & 1.6 & 1.4 & 1.3
& C & .98 & .89 & .54 & .69 & .76 & .97 & .99 & .96 & .92 & .89 & 1. & 1. & .91 & .82 & .76 & 1. & .99 & .81 & .64 & .54 & 1. & 1. & .95 & .85 & .76 & 1. & .97 & .51 & .69 & .76 & 1. & 1. & .97 & .89 & .79 & .97 & .92 & .72 & .62 & .57
& \(\kappa\) & .65 & .51 & .84 & .89 & .90 & .66 & .89 & .65 & .56 & .52 & .86 & .83 & .66 & .74 & .77 & .95 & .90 & .51 & .67 & .72 & 1. & 1. & .83 & .63 & .54 & 1. & .97 & .52 & .66 & .71 & .99 & .99 & .94 & .85 & .76 & .98 & .95 & .78 & .70 & .65
& L & 1. & .97 & .75 & .61 & .52 & .99 & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 & 1. & .99 & .91 & .81 & .74 & 1. & 1. & .97 & .92 & .87 & .99 & .94 & .52 & .69 & .77 & 1. & 1. & .95 & .85 & .76 & .94 & .88 & .64 & .54 & .51
& S & .69 & .55 & .85 & .90 & .89 & .93 & .91 & .66 & .56 & .58 & .60 & .70 & .80 & .85 & .83 & .97 & .96 & .62 & .60 & .63 & 1. & .99 & .62 & .62 & .64 & 1. & .96 & .52 & .71 & .72 & .98 & .97 & .87 & .73 & .67 & .99 & 1. & .97 & .89 & .87
(l)2-42 & E & 8.6 & 1.1 & 0.3 & 0.3 & 0.1 & 21.3 & 13.4 & 7.0 & 5.3 & 4.3 & 39.6 & 20.6 & 9.8 & 8.5 & 7.6 & 34.0 & 17.2 & 5.8 & 3.5 & 2.3 & 27.4 & 7.2 & 1.5 & 0.9 & 0.8 & 10.4 & 1.5 & 0.4 & 0.2 & 0.1 & 27.5 & 9.4 & 3.4 & 2.8 & 2.7 & 8.6 & 2.8 & 1.5 & 1.3 & 1.3
& C & .99 & .92 & .53 & .70 & .76 & .99 & 1. & .96 & .92 & .89 & 1. & 1. & .91 & .82 & .75 & 1. & .99 & .81 & .63 & .53 & 1. & 1. & .96 & .86 & .78 & 1. & .99 & .52 & .69 & .76 & 1. & 1. & .97 & .87 & .78 & .99 & .95 & .74 & .64 & .59
& \(\kappa\) & .80 & .59 & .82 & .88 & .90 & .81 & .93 & .72 & .62 & .57 & .96 & .88 & .59 & .70 & .74 & .98 & .93 & .52 & .65 & .71 & 1. & 1. & .85 & .66 & .57 & 1. & .99 & .54 & .66 & .72 & 1. & .99 & .94 & .84 & .76 & .99 & .96 & .80 & .71 & .66
& L & 1. & .98 & .76 & .61 & .52 & 1. & 1. & .99 & .97 & .97 & 1. & 1. & .99 & .97 & .96 & 1. & 1. & .91 & .80 & .73 & 1. & 1. & .98 & .93 & .88 & 1. & .96 & .50 & .69 & .77 & 1. & 1. & .94 & .84 & .75 & .97 & .90 & .65 & .55 & .50
& S & .81 & .62 & .84 & .89 & .90 & .93 & .95 & .75 & .65 & .64 & .90 & .87 & .69 & .78 & .79 & .99 & .97 & .68 & .53 & .59 & 1. & 1. & .75 & .52 & .54 & 1. & .97 & .55 & .73 & .74 & .99 & .99 & .89 & .76 & .72 & 1. & .99 & .90 & .81 & .80
& E & 8.5 & 1.2 & 0.3 & 0.2 & 0.3 & 20.2 & 12.8 & 6.1 & 5.1 & 4.2 & 38.0 & 19.3 & 9.1 & 7.7 & 7.0 & 34.2 & 16.8 & 5.2 & 2.8 & 1.5 & 26.2 & 6.7 & 1.2 & 0.7 & 0.6 & 9.6 & 1.4 & 0.3 & 0.1 & 0.1 & 26.4 & 9.6 & 3.3 & 2.6 & 2.4 & 8.1 & 2.8 & 1.5 & 1.3 & 1.2
& C & .99 & .92 & .54 & .65 & .75 & .99 & .99 & .97 & .94 & .91 & 1. & 1. & .93 & .84 & .76 & 1. & .99 & .86 & .69 & .55 & 1. & 1. & .97 & .89 & .78 & 1. & .99 & .62 & .61 & .73 & 1. & 1. & .99 & .92 & .82 & .99 & .95 & .78 & .68 & .61
& \(\kappa\) & .83 & .63 & .78 & .87 & .90 & .88 & .96 & .78 & .68 & .61 & .98 & .91 & .52 & .67 & .73 & .98 & .95 & .61 & .60 & .69 & 1. & 1. & .91 & .72 & .58 & 1. & .99 & .65 & .59 & .69 & 1. & 1. & .96 & .89 & .79 & .99 & .97 & .83 & .75 & .68
& L & 1. & .98 & .80 & .67 & .55 & 1. & 1. & .99 & .98 & .97 & 1. & 1. & .99 & .98 & .96 & 1. & 1. & .93 & .84 & .75 & 1. & 1. & .98 & .94 & .88 & 1. & .97 & .59 & .61 & .75 & 1. & 1. & .96 & .89 & .78 & .97 & .91 & .70 & .60 & .52
& S & .82 & .62 & .82 & .89 & .88 & .94 & .94 & .72 & .61 & .58 & .88 & .84 & .72 & .81 & .82 & .99 & .97 & .66 & .57 & .64 & 1. & 1. & .71 & .55 & .63 & 1. & .98 & .57 & .66 & .70 & .99 & .99 & .90 & .78 & .70 & 1. & 1. & .97 & .91 & .87
(l)2-42 & E & 8.7 & 1.1 & 0.3 & 0.3 & 0.1 & 21.8 & 13.5 & 7.2 & 5.2 & 4.2 & 38.7 & 19.6 & 9.1 & 7.5 & 7.0 & 34.2 & 16.4 & 5.0 & 2.5 & 1.6 & 28.6 & 7.7 & 1.4 & 0.8 & 0.5 & 10.9 & 1.4 & 0.2 & 0.2 & 0.1 & 26.7 & 9.7 & 3.4 & 2.7 & 2.5 & 8.5 & 2.9 & 1.6 & 1.3 & 1.3
& C & .98 & .89 & .51 & .66 & .75 & .97 & .99 & .97 & .93 & .90 & 1. & 1. & .92 & .84 & .76 & 1. & .99 & .83 & .68 & .56 & 1. & 1. & .96 & .88 & .77 & 1. & .98 & .56 & .64 & .74 & 1. & 1. & .98 & .92 & .82 & .97 & .92 & .74 & .66 & .59
& \(\kappa\) & .69 & .53 & .82 & .88 & .90 & .71 & .90 & .71 & .62 & .57 & .90 & .84 & .62 & .72 & .76 & .96 & .91 & .54 & .63 & .71 & 1. & 1. & .87 & .67 & .55 & 1. & .98 & .58 & .61 & .70 & .99 & .99 & .95 & .88 & .79 & .98 & .95 & .80 & .73 & .67
& L & 1. & .97 & .77 & .65 & .54 & 1. & 1. & .99 & .98 & .97 & 1. & 1. & .99 & .97 & .96 & 1. & .99 & .92 & .84 & .75 & 1. & 1. & .97 & .94 & .88 & .99 & .94 & .54 & .64 & .75 & 1. & 1. & .96 & .88 & .78 & .94 & .88 & .66 & .58 & .51
& S & .72 & .56 & .84 & .89 & .88 & .95 & .92 & .72 & .60 & .60 & .66 & .72 & .77 & .83 & .82 & .98 & .96 & .66 & .56 & .61 & 1. & .99 & .62 & .61 & .65 & 1. & .95 & .51 & .69 & .71 & .98 & .97 & .88 & .77 & .69 & .99 & 1. & .98 & .92 & .90
(l)2-42 & E & 8.8 & 1.1 & 0.3 & 0.1 & 0.1 & 22.0 & 12.9 & 6.4 & 5.1 & 4.3 & 40.0 & 20.3 & 8.9 & 7.9 & 7.0 & 33.8 & 16.3 & 5.0 & 3.0 & 1.4 & 27.9 & 7.0 & 1.2 & 0.7 & 0.7 & 10.1 & 1.3 & 0.3 & 0.2 & 0.1 & 26.2 & 9.0 & 3.0 & 2.7 & 2.4 & 8.5 & 2.7 & 1.4 & 1.3 & 1.2
& C & .99 & .91 & .51 & .65 & .75 & .99 & 1. & .97 & .94 & .91 & 1. & 1. & .92 & .84 & .76 & 1. & .99 & .83 & .68 & .55 & 1. & 1. & .96 & .88 & .78 & 1. & .99 & .57 & .63 & .74 & 1. & 1. & .98 & .92 & .80 & .99 & .94 & .76 & .68 & .61
& \(\kappa\) & .81 & .58 & .81 & .87 & .90 & .86 & .94 & .76 & .67 & .60 & .96 & .89 & .55 & .67 & .73 & .98 & .93 & .57 & .61 & .70 & 1. & 1. & .89 & .71 & .58 & 1. & .98 & .59 & .61 & .70 & 1. & .99 & .95 & .88 & .78 & .99 & .96 & .82 & .74 & .68
& L & 1. & .98 & .78 & .66 & .54 & 1. & 1. & .99 & .98 & .97 & 1. & 1. & .99 & .98 & .96 & 1. & 1. & .92 & .83 & .74 & 1. & 1. & .98 & .94 & .88 & 1. & .96 & .54 & .63 & .75 & 1. & 1. & .96 & .88 & .78 & .97 & .89 & .67 & .59 & .52
& S & .83 & .61 & .82 & .88 & .89 & .94 & .95 & .78 & .69 & .65 & .91 & .87 & .67 & .77 & .79 & .99 & .97 & .71 & .51 & .59 & 1. & 1. & .76 & .52 & .56 & 1. & .97 & .51 & .68 & .71 & .99 & .99 & .90 & .81 & .74 & .99 & .99 & .93 & .86 & .83
& E & 9.7 & 1.8 & 0.3 & 0.3 & 0.3 & 19.2 & 10.0 & 5.2 & 5.2 & 4.2 & 38.1 & 24.7 & 11.8 & 9.8 & 8.8 & 34.1 & 19.8 & 6.6 & 4.5 & 3.1 & 28.9 & 13.9 & 2.7 & 1.9 & 1.3 & 11.7 & 2.3 & 0.4 & 0.3 & 0.2 & 29.2 & 13.3 & 4.3 & 3.6 & 3.0 & 8.7 & 3.9 & 1.7 & 1.6 & 1.4
& C & .98 & .80 & .65 & .70 & .76 & .98 & .98 & .90 & .88 & .85 & 1. & .99 & .88 & .84 & .77 & 1. & .97 & .72 & .65 & .55 & 1. & 1. & .95 & .90 & .82 & 1. & .92 & .57 & .63 & .71 & 1. & 1. & .95 & .91 & .83 & .96 & .88 & .65 & .62 & .57
& \(\kappa\) & .62 & .65 & .88 & .89 & .90 & .67 & .67 & .53 & .52 & .51 & .78 & .65 & .68 & .71 & .74 & .94 & .79 & .60 & .65 & .70 & 1. & 1. & .80 & .72 & .61 & 1. & .92 & .55 & .61 & .67 & .99 & .98 & .90 & .87 & .79 & .97 & .92 & .72 & .69 & .65
& L & 1. & .93 & .66 & .60 & .52 & 1. & .99 & .97 & .96 & .95 & 1. & 1. & .98 & .97 & .96 & 1. & .99 & .86 & .82 & .75 & 1. & 1. & .97 & .95 & .91 & 1. & .87 & .58 & .64 & .73 & 1. & 1. & .92 & .87 & .80 & .93 & .81 & .57 & .54 & .51
& S & .67 & .58 & .88 & .89 & .89 & .85 & .76 & .57 & .54 & .54 & .62 & .51 & .79 & .82 & .80 & .97 & .93 & .58 & .54 & .60 & .99 & .99 & .65 & .50 & .56 & .99 & .96 & .57 & .67 & .69 & .98 & .96 & .84 & .76 & .70 & .99 & 1. & .92 & .87 & .86

5.2 Effect of \(n_{\text{ids}}\) and Training Duration↩︎

a
b
c

d

Figure 2: PairCos (\(\bar{s}\)) kernel density estimates. The bold black curve (with shading) shows member identities (present in the training set). All other curves show non-member identities: the nine thin solid lines correspond to cross-distribution benchmarks whose domain differs from the training data (one colour per dataset; see legend), while the dashed line (“Within-dist”) shows non-members from the held-out WebFace4M split, which shares the training domain (no domain shift). (a, b) vary training epochs at fixed \(n_{\text{ids}}\); (c) varies \(n_{\text{ids}}\) at epoch 40.. a — IR-100 / MagFace / 100K identities., b — IR-100 / MagFace / 1K identities., c — IR-100 / MagFace / epoch 40.

Figure 2 presents the main finding of this study. Panel (c) shows that the PairCos distributions of members and non-members converge as \(n_{\text{ids}}\) increases from 1K to all identities. At \(n_{\text{ids}}{=}1\)K the distributions are well separated, whereas at \(n_{\text{ids}}{=}100\)K and beyond the separation narrows, though it remains above chance even on the held-out WebFace4M split (Table 3, discussed in Section 5.4).

Panels (a) and (b) of Figure 2 isolate the effect of training duration at two extremes of identity coverage. At \(n_{\text{ids}}{=}1\)K, the PairCos separation between members and non-members increases with the number of training epochs. At \(n_{\text{ids}}{=}100\)K the same trend is present but attenuated.

Table 2 reports the ANOVA results (\(\hat{\eta}^2_p\), the proportion of variance explained after accounting for all other factors). Training-set size explains the most variance (\(\hat{\eta}^2_p \geq .95\)), followed by training duration (\(\hat{\eta}^2_p \geq .89\)). The strong \(n_{\text{ids}}\!\times\!\text{Epoch}\) interaction (\(\hat{\eta}^2_p \geq .92\)) further shows that the epoch effect is moderated by identity coverage. The \(n_{\text{ids}}\!\times\!\text{Epoch}\) cell means on the same-domain WebFace4M split show that the mean PairCos AUC increase from epoch 5 to epoch 40 ranges from \(+.38\) at \(n_{\text{ids}}{=}10\)K down to \(+.16\) at \(n_{\text{ids}}{=}100\)K, with \(n_{\text{ids}}{=}1\)K reaching a ceiling near 1.0 already at epoch 10. Models with small identity pools thus show the largest separability gains over training, whereas models trained on \(100\)K identities have only a modest increase.

Backbone and loss head. The ANOVA also shows that backbone architecture and loss head explain far less variance in threshold-classifier AUC, the loss-head factor reaching at most \(\hat{\eta}^2_p = .04\) and backbone depth at most \(.09\) across all four statistics — an order of magnitude below the contributions of \(n_{\text{ids}}\) and number of epochs. Beyond the IResNet grid, a ViT-T [33], [34] backbone (CosFace) trained under the identical protocol [35] reproduces the same \(n_{\text{ids}}\) trend in both EER and membership AUC (Table ¿tbl:tab:eer?, bottom block), so the effect is not specific to convolutional backbones.

5.3 Effect of Input-Data Distribution↩︎

The benchmarks used in this study differ in pose, quality, age, and demographics, which affect embedding-cluster geometry independently of membership. The distributional-shift rows of Table 2 quantify these effects, and we discuss each factor below. For matched-pair benchmarks, we also translate each effect into the AUC change it would produce at the same-domain training-attributable baseline at \(n_{\text{ids}}{=}100\)K (where the AUC is \(\approx 0.71\)). We do so by passing each AUC through the inverse standard normal cumulative distribution function (a probit transform), on which equal shifts in the underlying non-member distribution correspond to equal numerical changes regardless of baseline. We take the matched-pair difference on this transformed scale and convert it back to AUC at the chosen baseline. This step is needed because the AUC scale saturates near 1, so a fixed shift in the underlying distribution produces a smaller AUC change near the ceiling than in the middle.

Pose. Profile-view faces produce large shifts in cluster-statistic distributions. The CFP-FP vs.CFP-FF comparison is the cleanest controlled experiment: both sets share the same 500 identities, so the observed Cohen’s \(d'\) (standardised mean difference) is attributable to pose alone (Table 2). CPLFW vs.LFW yields even larger \(|d'|\) for PairCos and PenLogit, but those two sets are not identity-matched, so other factors may contribute. In both comparisons, profile views widen the intra-identity distribution regardless of membership. This pushes statistic values towards the non-member range and inflates separability. At the \(n_{\text{ids}}{=}100\)K baseline, the pose contribution from CFP-FF to CFP-FP corresponds to an AUC change of \(+0.26\), matching or exceeding the training contribution above chance (\(+0.21\)).

Image quality. XQLFW shares LFW’s identity set and pair protocol but degrades one image per pair. For PairCos and PenLogit, the resulting Cohen’s \(d'\) is of the same order as that of cross-pose variation though somewhat smaller; \(\kappa\) shows a smaller quality effect (\(|d'|{=}0.48\) vs. \(1.13\) for pose). Lower image quality increases intra-identity spread in embedding space and shifts cluster-statistic distributions in the same direction as pose variation. At the same baseline, the quality contribution from LFW to XQLFW corresponds to an AUC change of \(+0.22\), comparable to the pose contribution.

Age. Within AgeDB-30, the Spearman rank correlation \(\rho\) (a non-parametric measure of monotonic association) between within-identity age spread and each cluster statistic (Table 2) is weak and varies in sign across statistics. Only \(\kappa\) exhibits the expected negative association (\(\rho = -0.26\)), under which identities spanning wider age ranges form less concentrated clusters, whereas PairCos is essentially uncorrelated (\(\rho = +0.01\)) and both PenLogit and ProtoCE carry small positive correlations (\(\rho = +0.17\) and \(\rho = +0.33\)). Age-range variation therefore exerts the weakest and least consistent influence on cluster geometry among the confounds examined here.

Ethnicity. A Kruskal–Wallis test (non-parametric one-way ANOVA) across the four RFW demographic groups (African, Asian, Caucasian, Indian) yields \(\epsilon^2 \approx 0.02\) (small) for PairCos, \(\kappa\), and PenLogit, and \(\epsilon^2 = 0.12\) (medium) for ProtoCE. Cluster-statistic values therefore vary across demographic groups within the non-member pool, with ProtoCE again being the most sensitive statistic. Although these effect sizes are modest relative to the \(n_{\text{ids}}\) factor, they indicate that the demographic composition of a non-member benchmark affects the baseline against which membership is measured.

Pose and quality produce the largest shifts in cluster-statistic distributions, with demographic composition and age contributing smaller and, for age, less consistent effects. The matched-pair pose and quality contributions show that the domain-induced component can match or exceed the training contribution itself at higher \(n_{\text{ids}}\), so the threshold-classifier AUC on any non-WebFace4M benchmark cannot be interpreted as membership signal alone.

5.4 Statistic Fusion↩︎

To test whether combining statistics recovers more membership information, we train a two-hidden-layer MLP (\(64{\to}32\) units, ReLU, BatchNorm, dropout 0.3; Adam with \(\eta=10^{-3}\)) on all four cluster statistics. For each epoch-40 configuration we perform five-fold stratified cross-validation with an 80/20 identity split on same-domain WebFace4M identities, reserving 15% of each training fold as a validation set for early stopping.

Table 3: Statistic fusion recovers more membership information than any single statistic. Same-domain held-out split (epoch 40, 5-fold stratified CV). Each cell shows best single-statistic AUC / fold-averaged MLP AUC. Bold highlights the MLP gain over the best single statistic.
\(\nids\)
3-6 Backbone Loss 1K 10K 50K 100K
IR-34 Arc 1.00 / 1.00 .98 / 1.00 .80 / .87 .65 / .69
Cos .99 / 1.00 .98 / .99 .79 / .88 .65 / .71
Mag 1.00 / 1.00 .98 / 1.00 .79 / .85 .64 / .68
1-6 Arc 1.00 / 1.00 .99 / 1.00 .84 / .91 .68 / .73
Cos .99 / 1.00 .98 / .99 .82 / .91 .68 / .75
Mag 1.00 / 1.00 .98 / 1.00 .82 / .89 .67 / .72
1-6 Arc 1.00 / 1.00 .99 / 1.00 .86 / .92 .73 / .79
Cos .99 / 1.00 .98 / .99 .84 / .92 .71 / .80
Mag 1.00 / 1.00 .98 / 1.00 .85 / .91 .72 / .79

3pt

Table 3 reports the best single-statistic AUC alongside the fold-averaged MLP AUC. The MLP consistently matches or outperforms the best individual statistic, with the largest gains at \(n_{\text{ids}}{=}50\text{K}\) and \(100\text{K}\) where individual statistics are weakest (one-sided Wilcoxon signed-rank: \(W{=}666\), \(p<10^{-10}\)).

5.5 Limitations↩︎

Generalisation across training data. All experiments use a single training source (WebFace4M). Whether the trends generalise to other training datasets is an open question, as comparably large datasets are not publicly available or have been retracted.

Experimental controls. We do not evaluate defences such as DP-SGD [23] or knowledge distillation [36]. The number of probe images \(k\) per identity varies across benchmarks and is not controlled for, which may affect estimates of the cluster statistics non-uniformly. The held-out split is constructed by random partitioning; in practice, membership boundaries may correlate with identity-level attributes.

Statistical assumptions and scope. Because our four statistics capture only cluster geometry, the measured separability is a lower bound on the membership information available from embeddings; a richer model operating on raw embedding vectors could extract more. The probit-scale translation in Section 5.3 assumes approximately Gaussian member and non-member statistic distributions and stable variances across matched pairs.

6 Conclusion and Future Work↩︎

We evaluated how much membership information is retained in the embedding geometry of 180 open-set face recognition models across nine benchmarks. The results lead to three practical conclusions. First, training-set size has the largest effect on member/non-member separability, while backbone architecture and loss head contribute far less. Training on more identities is therefore the only design choice we tested that substantially reduces the geometric membership signal. Second, cross-domain non-member benchmarks inflate the apparent signal by conflating domain shift with training-attributable differences. Privacy audits should therefore use same-domain held-out references, since cross-domain non-members overestimate the apparent membership signal. Third, even at high \(n_{\text{ids}}\) where individual statistics are weak, fusing all statistics with a learned classifier recovers additional membership information (Table 3). As our statistics capture only cluster geometry, the measured separability is a lower bound on the membership information available from embeddings. Future work includes extending these benchmarks to other backbones and evaluating whether defences such as [23] can reduce the geometric signal without degrading performance.

Acknowledgments↩︎

This work has received funding from the European Union’s Horizon Europe research and innovation programme under Grant Agreement No. 101189650 (CERTAIN: Certification for Ethical and Regulatory Transparency in Artificial Intelligence), and the Swiss State Secretariat for Education, Research and Innovation (SERI).

References↩︎

[1]
Z. Zhu et al., “WebFace260M: A benchmark for million-scale deep face recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 2, pp. 2627–2644, 2023, doi: 10.1109/TPAMI.2022.3169734.
[2]
J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “ArcFace: Additive angular margin loss for deep face recognition,” in IEEE conference on computer vision and pattern recognition, CVPR 2019, long beach, CA, USA, june 16-20, 2019, 2019, pp. 4690–4699, doi: 10.1109/CVPR.2019.00482.
[3]
H. Wang et al., “CosFace: Large margin cosine loss for deep face recognition,” in 2018 IEEE conference on computer vision and pattern recognition, CVPR 2018, salt lake city, UT, USA, june 18-22, 2018, 2018, pp. 5265–5274, doi: 10.1109/CVPR.2018.00552.
[4]
Q. Meng, S. Zhao, Z. Huang, and F. Zhou, “MagFace: A universal representation for face recognition and quality assessment,” in IEEE conference on computer vision and pattern recognition, CVPR 2021, virtual, june 19-25, 2021, 2021, pp. 14225–14234, doi: 10.1109/CVPR46437.2021.01400.
[5]
R. Shokri, M. Stronati, C. Song, and V. Shmatikov, “Membership inference attacks against machine learning models,” in 2017 IEEE symposium on security and privacy, SP 2017, san jose, CA, USA, may 22-26, 2017, 2017, pp. 3–18, doi: 10.1109/SP.2017.41.
[6]
N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramèr, “Membership inference attacks from first principles,” in 43rd IEEE symposium on security and privacy, SP 2022, san francisco, CA, USA, may 22-26, 2022, 2022, pp. 1897–1914, doi: 10.1109/SP46214.2022.9833649.
[7]
Art. 17 (Right to erasure)“Regulation (EU) 2016/679 of the European Parliament and of the Council (general data protection regulation).” Official Journal of the European Union, L 119, 1–88, 2016.
[8]
A. Salem, Y. Zhang, M. Humbert, P. Berrang, M. Fritz, and M. Backes, ML-Leaks: Model and data independent membership inference attacks and defenses on machine learning models,” in 26th annual network and distributed system security symposium, NDSS 2019, san diego, CA, USA, february 24-27, 2019, 2019, doi: 10.14722/ndss.2019.23119.
[9]
H. Hu, Z. Salcic, L. Sun, G. Dobbie, P. S. Yu, and X. Zhang, “Membership inference attacks on machine learning: A survey,” ACM Comput. Surv., vol. 54, no. 11s, pp. 1–37, 2022, doi: 10.1145/3523273.
[10]
G. Li, S. Rezaei, and X. Liu, “User-level membership inference attack against metric embedding learning,” CoRR, vol. abs/2203.02077, 2022, doi: 10.48550/arXiv.2203.02077.
[11]
M. Chen, Z. Zhang, T. Wang, M. Backes, and Y. Zhang, FACE-AUDITOR: Data auditing in facial recognition systems,” in 32nd USENIX security symposium, USENIX security 2023, anaheim, CA, USA, august 9-11, 2023, 2023, pp. 7195–7212, [Online]. Available: https://www.usenix.org/conference/usenixsecurity23/presentation/chen-min.
[12]
G. Chen, Y. Zhang, and F. Song, SLMIA-SR: Speaker-level membership inference attacks against speaker recognition systems,” in 31st annual network and distributed system security symposium, NDSS 2024, san diego, CA, USA, february 26 - march 1, 2024, 2024, doi: 10.14722/ndss.2024.241323.
[13]
H. Liu, J. Jia, W. Qu, and N. Z. Gong, EncoderMI: Membership inference against pre-trained encoders in contrastive learning,” in Proceedings of the 2021 ACM SIGSAC conference on computer and communications security, CCS 2021, virtual event, republic of korea, november 15-19, 2021, 2021, pp. 2081–2095, doi: 10.1145/3460120.3484749.
[14]
D. DeAlcala, A. Morales, J. Fierrez, G. Mancera, R. Tolosana, and J. Ortega-Garcia, “Is my data in your AI? Membership inference test (MINT) applied to face biometrics,” IEEE Access, vol. 13, pp. 163805–163819, 2025, doi: 10.1109/ACCESS.2025.3608951.
[15]
D. DeAlcala, G. Mancera, A. Morales, J. Fiérrez, R. Tolosana, and J. Ortega-Garcia, “A comprehensive analysis of factors impacting membership inference,” in IEEE/CVF conference on computer vision and pattern recognition workshops, CVPRW 2024, seattle, WA, USA, june 17-18, 2024, 2024, pp. 3585–3593, doi: 10.1109/CVPRW63382.2024.00362.
[16]
G. Mancera, D. DeAlcala, A. Morales, R. Tolosana, and J. Fierrez, “Membership inference test: Auditing training data in object classification models,” in Deployable AI workshop (DAI 2025), co-located with the 39th AAAI conference on artificial intelligence (AAAI-25), 2025, [Online]. Available: https://arxiv.org/abs/2601.12929.
[17]
Y. Huang, H. Chen, Y. Wang, and L. Wang, “Inference attacks against face recognition model without classification layers,” CoRR, vol. abs/2401.13719, 2024, doi: 10.48550/arXiv.2401.13719.
[18]
M. Fredrikson, S. Jha, and T. Ristenpart, “Model inversion attacks that exploit confidence information and basic countermeasures,” in Proceedings of the 22nd ACM SIGSAC conference on computer and communications security, CCS 2015, denver, CO, USA, october 12-16, 2015, 2015, pp. 1322–1333, doi: 10.1145/2810103.2813677.
[19]
S. Rezaei and X. Liu, “On the difficulty of membership inference attacks,” in IEEE/CVF conference on computer vision and pattern recognition, CVPR 2021, virtual, june 19-25, 2021, 2021, pp. 7892–7900, doi: 10.1109/CVPR46437.2021.00780.
[20]
V. Feldman, “Does learning require memorization? A short tale about a long tail,” in Proceedings of the 52nd annual ACM SIGACT symposium on theory of computing, STOC 2020, chicago, IL, USA, june 22-26, 2020, 2020, pp. 954–959, doi: 10.1145/3357713.3384290.
[21]
L. Bourtoule et al., “Machine unlearning,” in 42nd IEEE symposium on security and privacy, SP 2021, san francisco, CA, USA, may 23-27, 2021, 2021, pp. 141–159, doi: 10.1109/SP40001.2021.00019.
[22]
A. Golatkar, A. Achille, and S. Soatto, “Eternal sunshine of the spotless net: Selective forgetting in deep networks,” in 2020 IEEE/CVF conference on computer vision and pattern recognition, CVPR 2020, seattle, WA, USA, june 13-19, 2020, 2020, pp. 9301–9309, doi: 10.1109/CVPR42600.2020.00932.
[23]
M. Abadi et al., “Deep learning with differential privacy,” in Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, CCS 2016, vienna, austria, october 24-28, 2016, 2016, pp. 308–318, doi: 10.1145/2976749.2978318.
[24]
G. B. Huang, M. Ramesh, T. Berg, and E. Learned-Miller, “Labeled faces in the wild: A database for studying face recognition in unconstrained environments,” University of Massachusetts, Amherst, 07-49, 2007. [Online]. Available: https://vis-www.cs.umass.edu/lfw/lfw.pdf.
[25]
T. Zheng and W. Deng, “Cross-pose LFW: A database for studying cross-pose face recognition in unconstrained environments,” Beijing University of Posts; Telecommunications, 18-01, 2018. [Online]. Available: http://www.whdeng.cn/CPLFW/Cross-Pose-LFW.pdf.
[26]
S. Sengupta, J.-C. Chen, C. D. Castillo, V. M. Patel, R. Chellappa, and D. W. Jacobs, “Frontal to profile face verification in the wild,” in 2016 IEEE winter conference on applications of computer vision, WACV 2016, lake placid, NY, USA, march 7-10, 2016, 2016, pp. 1–9, doi: 10.1109/WACV.2016.7477558.
[27]
S. Moschoglou, A. Papaioannou, C. Sagonas, J. Deng, I. Kotsia, and S. Zafeiriou, “AgeDB: The first manually collected, in-the-wild age database,” in 2017 IEEE conference on computer vision and pattern recognition workshops, CVPR workshops 2017, honolulu, HI, USA, july 21-26, 2017, 2017, pp. 1997–2005, doi: 10.1109/CVPRW.2017.250.
[28]
M. Knoche, S. Hörmann, and G. Rigoll, “Cross-quality LFW: A database for analyzing cross- resolution image face recognition in unconstrained environments,” in 16th IEEE international conference on automatic face and gesture recognition, FG 2021, jodhpur, india, december 15-18, 2021, 2021, pp. 1–5, doi: 10.1109/FG52635.2021.9666960.
[29]
M. Wang, W. Deng, J. Hu, X. Tao, and Y. Huang, “Racial faces in the wild: Reducing racial bias by information maximization adaptation network,” in 2019 IEEE/CVF international conference on computer vision, ICCV 2019, seoul, korea (south), october 27 - november 2, 2019, 2019, pp. 692–702, doi: 10.1109/ICCV.2019.00078.
[30]
B. Maze et al., IARPA janus benchmark - C: Face dataset and protocol,” in 2018 international conference on biometrics, ICB 2018, gold coast, australia, february 20-23, 2018, 2018, pp. 158–165, doi: 10.1109/ICB2018.2018.00033.
[31]
A. Banerjee, I. S. Dhillon, J. Ghosh, and S. Sra, “Clustering on the unit hypersphere using von mises-fisher distributions,” J. Mach. Learn. Res., vol. 6, pp. 1345–1382, 2005, [Online]. Available: https://jmlr.org/papers/v6/banerjee05a.html.
[32]
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE conference on computer vision and pattern recognition, CVPR 2016, las vegas, NV, USA, june 27-30, 2016, 2016, pp. 770–778, doi: 10.1109/CVPR.2016.90.
[33]
A. Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in 9th international conference on learning representations, ICLR 2021, virtual event, austria, may 3-7, 2021, 2021, [Online]. Available: https://openreview.net/forum?id=YicbFdNTTy.
[34]
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Proceedings of the 38th international conference on machine learning, ICML 2021, 18-24 july 2021, virtual event, 2021, vol. 139, pp. 10347–10357, [Online]. Available: http://proceedings.mlr.press/v139/touvron21a.html.
[35]
X. An et al., “Killing two birds with one stone: Efficient and robust training of face recognition CNNs by partial FC,” in IEEE/CVF conference on computer vision and pattern recognition, CVPR 2022, new orleans, LA, USA, june 18-24, 2022, 2022, pp. 4032–4041, doi: 10.1109/CVPR52688.2022.00401.
[36]
G. E. Hinton, O. Vinyals, and J. Dean, NeurIPS Deep Learning Workshop“Distilling the knowledge in a neural network,” CoRR, vol. abs/1503.02531, 2015, [Online]. Available: https://arxiv.org/abs/1503.02531.

  1. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.↩︎