Quantifying Training Membership Information in the
Hyperspherical Embedding Geometry of Face Recognition Models
July 16, 2026
Face recognition models represent each face as an embedding vector on the unit hypersphere by clustering embeddings of the same identity while pushing different identities apart through angular-margin losses. Because these losses act only on training identities, non-member identities may form clusters with different geometric properties. In this paper, we quantify the magnitude of this difference and what training-time factors control it. We compute four statistics based on cluster geometry across 180 face recognition models in a factorial design over IResNet backbone size, loss head, training duration, and the number of training identities, and evaluate each configuration on nine benchmarks. Our results indicate that the number of training identities has the largest effect on member/non-member separability, while backbone and loss head contribute far less, and that, on a same-domain held-out reference, the geometric membership signal decreases monotonically as more identities are added to training. We provide an analysis of cross-domain (pose, age, quality, ethnicity) non-member benchmarks and report that these inflate the apparent membership signal. Finally, we fuse all four statistics with a learned classifier to reveal additional membership information beyond the best individual statistic.
Modern face recognition (FR) systems are trained on large-scale datasets of face crops, typically collected from the internet [1]. To represent a person, the system takes multiple crops of that person’s face and passes each through a deep neural network that produces a 512-dimensional embedding vector, normalised to lie on the unit hypersphere. The training of these models is primarily driven by angular-margin losses [2]–[4], which pull the embeddings of crops belonging to the same identity into a cluster on the hypersphere while pushing apart the clusters of different identities. During training, a classification head maps embeddings to identity logits and this head is discarded after training. When deployed, the system typically uses only the backbone neural network. Enrolled faces are stored as embeddings, and the similarity between two faces is measured by the inner product of their corresponding embeddings.
As angular-margin training directly optimises the compactness of each identity’s cluster, we pose the following question: how much information about training-set membership is retained in the geometry of these embedding clusters? If training identities form systematically tighter or better-separated clusters than unseen identities, then the embedding geometry carries a residual membership signal.
Quantifying this residual signal is related to membership inference attacks (MIA), a line of work investigating whether a given sample or identity was part of a model’s training data [5], [6]. In practice, membership inference serves as an auditing tool for data privacy by allowing individuals or regulators to verify whether a model was trained on data belonging to a specific person in the context of data protection rights such as the right to erasure under the GDPR [7]. MIA literature primarily focuses on closed-set classifiers that expose prediction confidences, loss values, or the model weights directly [5], [6], [8], [9]. In our setting, we consider open-set face recognition where neither logits nor losses nor model weights are available, so the membership signal must be sought entirely in the geometry of the embeddings themselves.
To measure this signal, we compute four different statistics capturing different aspects of cluster geometry on the hypersphere and measure how much they differ between training and unseen identities (Figure 1). We train 180 face recognition models from scratch (Section 4.2) and evaluate each model on nine benchmarks, including a same-domain held-out split that controls for the pose, age, quality, and ethnicity shifts present in the other benchmarks. To our knowledge, this constitutes the largest systematic study of membership information in face embedding geometry. Our contributions are: (1) we quantify the membership information retained in four geometric statistics computed from face embeddings; (2) we identify the primary factors that control how much of this information persists; (3) we show that cross-domain non-member benchmarks (e.g.differing in pose or age) inflate the apparent membership signal by conflating domain shift with training-attributable differences; (4) we demonstrate that fusing all statistics captures membership information beyond the best individual statistic.
Membership inference on classifiers. Membership inference literature targets closed-set classifiers, where the attacker can observe prediction confidences, logits, or loss values of the target model. In the open-set face recognition setting, however, none of these signals are available and this has motivated a line of work operating directly on embeddings. Shokri et al. [5] introduce the shadow-model paradigm and train substitute classifiers to learn a binary meta-classifier on output distributions. Salem et al. [8] relax many of these assumptions and show that a single shadow model and only the top-\(k\) predictions often suffice. Carlini et al. [6] recast MIA as a likelihood-ratio attack (LiRA) and derive near-optimal attacks. They show that per-example hardness varies by orders of magnitude. Hu et al. [9] provide a comprehensive survey.
3pt
@ l l c c c @ Work & Domain & Level & # Mod. & Fact.
Li [10] & FR / re-identification & Identity & 6 &
FACE-AUD. [11] & FR / few-shot & Identity & 4 &
SLMIA-SR [12] & Speaker recognition & Identity & 5 &
EncoderMI [13] & Contrastive learning & Sample & 4 &
MINT [14] & FR & Sample & 3 &
DeAlcala [15] & FR & Sample & 9 & \(^\ast\)
Mancera [16] & Object classification & Sample & 4 &
Huang [17] & FR & Sample & 2 &
Ours & FR & Identity & 180 &
\(^\ast\)Varies factors individually but not in a fully crossed design.
Membership inference on embedding models. Existing embedding-based approaches differ in the granularity of the membership question: some ask whether a specific image was used for training (sample level), while others ask whether any image of a given person was included (identity level). Li et al. [10] formalise identity-level inference on metric-embedding models and use intra-similarity and inter-dissimilarity of probe embeddings as attack features on person re-identification and face recognition. Chen et al. [11] propose FACE-AUDITOR, a system that queries few-shot face recognition models with designed probing sets and uses the returned similarity scores as auditing features. They report accuracies up to 99%. SLMIA-SR [12] addresses the analogous problem in speaker recognition by training shadow models with controlled membership mixing ratios. All three operate at the identity level. EncoderMI [13], by contrast, operates at the sample level on contrastive-learning encoders including CLIP and exploits the observation that an overfitted encoder produces more similar representations for augmented views of a training sample than for an unseen one.
White-box approaches on face recognition. A parallel line of work allows access to model internals. The motivation is that intermediate representations such as activation maps, batch-normalisation statistics, or gradient signals may carry a stronger membership trace than the final embedding alone, at the cost of requiring more privileged access to the model. These methods also typically operate at the sample level, determining whether a specific image (rather than a person) was included in training. The Membership Inference Test (MINT) of DeAlcala et al. [14] trains an MLP or CNN on intermediate activation maps to classify individual face images as member or non-member and achieves up to 90% accuracy on three ResNet-100 models. A companion study [15] analyses factors that affect MINT performance, including loss function, dropout, and number of training epochs. Mancera et al. [16] extend MINT to object classification and report 70–80% precision depending on layer depth. Huang et al. [17] exploit distances between intermediate features and batch-normalisation parameters for sample-level membership and model-inversion attacks [18]. Even with this level of access, Rezaei and Liu [19] show that the advantage over random guessing can be modest once the model generalises well.
Memorisation and embedding geometry. The question of why a geometric membership signal might exist at all connects to a discussion on memorisation in deep networks. If a model must memorise certain training examples to achieve low error, particularly on the tail of the data distribution, then its internal representations and output embeddings may carry traces of those examples. Feldman [20] formalises this argument, showing theoretically that long-tailed distributions necessitate memorisation for near-optimal generalisation. Angular-margin losses [2]–[4] amplify this effect, potentially making training identities geometrically distinguishable from unseen ones. Unlearning methods [21], [22] and differential privacy [23] aim to erase or bound such signals.
Comparison with this work. Our contribution is a controlled measurement study of how much membership information geometric statistics carry, not a stronger attack feature. Table ¿tbl:tab:related-summary? summarises the literature. Prior embedding-level work proposes specific attack features and demonstrates that membership inference is feasible through evaluating a small number of models trained under a single configuration. Our study differs in three respects. First, the goal: we measure how much membership information is retained in the explicit cluster geometry and identify which training factors control the amount that persists. Second, the experimental scale: 180 models, whereas the largest prior study evaluates nine. Third, domain control: we add a same-domain held-out reference, which aims to separate training-attributable signal from domain-shift effects. Several of the per-identity statistics we evaluate have already been used as features in prior attacks. For instance, the intra-similarity score proposed by Li et al. [10] is equivalent to the mean pairwise cosine similarity that we include as a baseline statistic (defined in Section 3.2).
We assume that an auditor can query an FR model \(f\) with \(k \geq 2\) probe images of a target identity \(y\) and obtain the corresponding \(\ell_2\)-normalised embeddings. From the resulting cluster, the auditor computes four per-identity statistics (Section 3.2) and uses them to assess if the cluster geometry carries information about \(y\)’s membership in the training set. For each model we train, we compute these statistics on member identities (from WebFace4M) and on non-member identities drawn from the benchmarks in Table 1.
| Dataset | Primary variation | Subjects | Eval.subj. | Images | Covariate metadata |
|---|---|---|---|---|---|
| LFW [24] | Unconstrained (baseline) | 5,749 | 1,680 | 13,233 | (subject name only) |
| CPLFW [25] | Cross-pose (subset of LFW identities) | 3,929 | 3,909 | 11,652 | yaw (estimated from landmarks) |
| CFP-FF [26] | Frontal-only probes | 500 | 500 | 7,000 | pose \(\in\) |
| CFP-FP [26] | Frontal+profile probes | 500 | 7,000 | pose \(\in\) | |
| AgeDB-30 [27] | Large intra-identity age gap (\(\geq\)30 yr) | 440 | 440 | 12,240 | age (3–100 yr), gender |
| XQLFW [28] | Low image quality (paired with LFW) | 5,749 | 1,672 | 13,233 | (quality implicit by design) |
| RFW [29] | Ethnicity-balanced (4 groups) | 11,429 | 11,429 | 40,607 | ethnicity, gender, nationality |
| IJB-C [30] | Unconstrained, mixed image+video | 3,531 | 3,531 | 469,376 | ethnicity, gender, nationality |
| WebFace4M [1] | Large-scale training source | 205,990 | \({\leq}\,5{,}000\)\(\dagger\) | 4,000,000 | (subject ID only) |
3pt
For each statistic, we obtain a scalar score per identity. To quantify how much membership information a statistic carries, we sweep a decision threshold \(\tau\) across the range of observed values and classify an identity as a member if its score exceeds \(\tau\) (or falls below \(\tau\) for statistics where lower values indicate membership). We report the area under the resulting ROC curve (AUC), hereafter threshold-classifier AUC, as a threshold-free summary, which is equal to the probability that a randomly chosen member scores higher than a randomly chosen non-member. This serves as our primary measure of the membership information carried by a given statistic. We use AUC rather than the true positive rate at a fixed false positive rate, because our framing concerns separability across all decision thresholds rather than the operating point most relevant for a deployable attack. Of the four statistics defined in the next section, two require only the embeddings and two additionally require the training-time scale and margin hyperparameters. We refer to these two access levels as black-box (embeddings only) and grey-box (embeddings plus \(s\) and \(m\)). The grey-box setting is realistic when models are released alongside their training configurations.
Given an identity \(y\) with \(k \geq 2\) probe images, we extract embeddings \(\{e_1, \ldots, e_k\}\) via the target model and compute four statistics from the resulting cluster. Each statistic captures a different aspect of cluster shape. The underlying intuition is that an angular-margin loss explicitly optimises the representation of training identities, so members are expected to form tighter clusters on the embedding hypersphere than non-members. If that expectation holds, the statistic will carry membership information; the question is how much, and what factors control it.
Pairwise cosine similarity (PairCos). The mean cosine similarity over pairs of embeddings within the identity: \[\bar{s} = \binom{k}{2}^{-1} \sum_{i < j} \cos(e_i, e_j).\] Higher values indicate a more coherent cluster.
vMF concentration \(\kappa\). We model the per-identity embedding distribution as a von Mises-Fisher (vMF) distribution on the unit hypersphere and estimate the concentration parameter \(\kappa\) via maximum likelihood. We use the closed-form approximation of Banerjee et al. [31]: \[\hat{\kappa}_0 = \frac{\bar{R}\,(d - \bar{R}^2)}{1 - \bar{R}^2}, \qquad \bar{R} = \left\| \frac{1}{k}\sum_{i=1}^{k} e_i \right\|,\] where \(d = 512\) is the embedding dimensionality and \(\bar{R}\) is the mean resultant length. Higher \(\hat{\kappa}\) corresponds to greater directional concentration. Note that ordering identities by \(\hat{\kappa}\) is equivalent to ordering them by cluster inertia \(I = k^{-1}\sum_i \|e_i - \mu\|^2\), where \(\mu = k^{-1}\sum_i e_i\) is the cluster mean. For \(\ell_2\)-normalised embeddings, \(I\) reduces to \(1 - \bar{R}^2\), and \(\hat{\kappa}_0\) is strictly decreasing in \(1-\bar{R}^2\), so the two quantities produce identical identity rankings.
Penalised logit (PenLogit). To approximate the margin-penalised logit that the training loss would compute for the cluster, we use a leave-one-out scheme. For each embedding \(e_i\), we compute a prototype \(\hat{w}_{\backslash i} = \bar{e}_{\backslash i}/\|\bar{e}_{\backslash i}\|\) as the \(\ell_2\)-normalised mean of the remaining \(k{-}1\) embeddings, and the margin-penalised logit is \[\ell_i = s \cdot g\left(\cos\angle\left(e_i, \hat{w}_{\backslash i}\right),\; m\right),\] where \(s\) and \(m\) are the scale and margin of the loss head, and \(g\) is the head-specific penalisation function (e.g. \(g(\cos\theta, m) = \cos(\theta + m)\) for ArcFace). The statistic value is the mean \(\bar{\ell} = k^{-1}\sum_i \ell_i\). A higher penalised logit indicates that the embeddings are tightly aligned to their own prototype.
Prototype softmax CE (ProtoCE). We construct a surrogate of the training loss. For each identity \(y\), we compute a class prototype \(\hat{w}_y = \mu / \|\mu\|\) and treat all other identities in the evaluation set as negative classes. The cross-entropy loss with the margin-modified logit for the target class and standard cosine logits for the negatives gives \[\mathcal{L}_y = -\log \frac{\exp\bigl(s\cdot g(\cos\theta_y,\, m)\bigr)}{\exp\bigl(s\cdot g(\cos\theta_y,\, m)\bigr) + \sum_{j \neq y} \exp(s\cdot\cos\theta_j)}\] We report \(-\mathcal{L}_y\) so that higher values indicate membership, consistent with the sign convention of the other statistics. Because the softmax denominator depends on the composition and size of the evaluation set, ProtoCE values are not directly comparable across benchmarks with different identity pools. Throughout, all AUC values are oriented so that \(\text{AUC} > 0.5\) indicates better-than-chance separation.
All models are trained on WebFace4M [1], a large-scale face dataset containing approximately 4 million images of 205,990 identities. WebFace4M is well-suited to this study because its size permits creating training subsets that span several orders of magnitude in identity count while still providing sufficient images per identity for meaningful embedding statistics. Every probe image is aligned to a canonical \(112\!\times\!112\) crop with the standard InsightFace five-landmark pipeline.
We vary four factors in a fully crossed design.
Backbone capacity. We train three IResNet architectures: IResNet-34, IResNet-50, and IResNet-100 [2], [32].
Angular-margin loss head. We pair each backbone with three angular-margin loss heads: ArcFace [2] (\(s\!=\!64\), \(m\!=\!0.5\)), CosFace [3] (\(s\!=\!64\), \(m\!=\!0.4\)), and MagFace [4] (\(s\!=\!64\), per-sample \(m \in [0.45, 0.8]\) depending on the un-normalised feature magnitude).
Training-set size (\(n_{\text{ids}}\)). To study the effect of identity coverage, we create five partitions by randomly sampling \(n_{\text{ids}}\in \{1\text{K}, 10\text{K}, 50\text{K}, 100\text{K}\}\) identities (with a fixed seed), plus one partition that uses all available identities. The identities not selected for training serve as same-domain non-members in the held-out WebFace4M benchmark.
Training duration. We train the backbones for 5, 10, 20, and 40 epochs to track the evolution of cluster-statistic separability. Longer training exposes each identity to more gradient updates and may therefore affect the gap between member and non-member statistic distributions.
The full grid spans \(3~\text{backbones} \times 3~\text{losses} \times 5~n_{\text{ids}}\text{ levels} \times 4~\text{epochs} = 180\) model configurations. All models are trained with SGD (momentum 0.9, weight decay \(5\!\times\!10^{-4}\)), a polynomial-decay learning-rate schedule starting at \(0.1\) with one warmup epoch, batch size 128, mixed-precision (FP16), and random horizontal flips.
The nine benchmarks (Table 1) supply standard verification protocols for measuring model quality and provide non-member pools with diverse domain characteristics so that we can study how pose, quality, and demographics affect the apparent cluster-statistic separability.
Eight of the nine benchmarks are drawn from sources disjoint from WebFace4M (we verified that no identity overlap exists by cross-checking subject identifiers); Table 1 lists their sizes and available covariate metadata. Two dataset pairs form natural experiments (matched comparisons where a single attribute varies). CFP-FF vs.CFP-FP [26] contrasts frontal-only against frontal-plus-profile probes for the same 500 identities, so any statistic difference is attributable to pose. LFW [24] vs.XQLFW [28] shares the same identity set and pair protocol but degrades one image per pair to isolate image quality. CPLFW [25] (cross-pose LFW), AgeDB-30 [27] (age gap \(\geq\)30 yr), RFW [29] (ethnicity-balanced, four groups), and IJB-C [30] complete the cross-domain set.
Held-out WebFace4M split. The ninth benchmark is constructed from WebFace4M itself: for each model trained on \(n_{\text{ids}}\) identities, we sample \(\min(n_{\text{ids}},\, 2{,}500)\) members and 2,500 non-members from the remaining subjects; at \(n_{\text{ids}}{=}\)all no non-members remain. As both groups share the same acquisition pipeline, AUCs are arguably attributable to the cluster-geometry difference between seen and unseen identities rather than to domain shift.
This section presents experimental results and analysis for the 180-model grid described in Section 4. We examine verification performance, the effect of training-set size and duration on membership separability, distributional confounds, and statistic fusion. Full per-epoch, per-model results and ROC/DET curves are provided in the supplementary material.
| PairCos | \(\kappa\) | PenLogit | ProtoCE | |
| \(n_\mathrm{ids}\) | ||||
| Epoch | ||||
| Backbone | ||||
| Loss | ||||
| \(n_\mathrm{ids}{\times}\)Ep | ||||
| Residual | ||||
| \(n_\mathrm{ids}\):1K \(\cdot\) 10K \(\cdot\) 50K \(\cdot\) 100K \(\cdot\) all | ||||
| Epoch:ep5 \(\cdot\) ep10 \(\cdot\) ep20 \(\cdot\) ep40 | ||||
| Cohen’s \(d'\) | ||||
| XQLFW \(-\) LFW | \(-1.64\) | \(-0.48\) | \(-1.81\) | \(+1.13\) |
| CFP-FP \(-\) CFP-FF | \(-2.27\) | \(-1.13\) | \(-2.98\) | \(+1.14\) |
| CPLFW \(-\) LFW | \(-1.72\) | \(-0.72\) | \(-2.06\) | \(+0.31\) |
| 1-5 Spearman \(\rho\) | ||||
| AgeDB (age \(\sigma\)) | \(+0.01\) | \(-0.26\) | \(+0.17\) | \(+0.33\) |
| 1-5 Kruskal-Wallis \(\epsilon^2\) | ||||
| RFW (4 groups) | \(0.017\) | \(0.017\) | \(0.018\) | \(0.123\) |
4pt
We first establish that our model grid spans a meaningful range of verification quality. Table ¿tbl:tab:eer? reports the verification equal error rate (EER) at epoch 40 for every combination of backbone, loss head, and \(n_{\text{ids}}\) on the cross-domain benchmarks. The table is dense and provided primarily as a per-cell reference, and the discussion below focuses on patterns and points back to specific cells only as needed. Performance generally improves with \(n_{\text{ids}}\), with diminishing gains beyond \(100\)K identities. Once \(n_{\text{ids}}\) is held fixed, backbone depth provides small reductions in verification EER on domain-shifted benchmarks, and the choice of loss head has an even smaller effect. Notably, the highest cluster-statistic separability occurs at \(n_{\text{ids}}{=}1\)K, where verification performance is worst. As \(n_{\text{ids}}\) grows, verification improves while separability decreases. However, non-trivial separability persists even at \(n_{\text{ids}}\) levels that achieve competitive verification EER, so the patterns reported below are not confined to poorly performing models.
Table ¿tbl:tab:eer? also reports the threshold-classifier AUC for all four cluster statistics. No single statistic achieves the highest AUC across all benchmarks. PenLogit ranks first most often at moderate-to-high \(n_{\text{ids}}\), in particular on the pose- and quality-degraded sets, whereas ProtoCE attains the highest AUC on the frontal and high-quality sets (CFP-FF, LFW, IJB-C), so the best statistic remains dataset-dependent. Because the eight non-WebFace4M benchmarks differ from the training source in pose, quality, and demographics, the observed AUC conflates training-attributable and domain-induced differences, which we discuss in Section 5.3. All four statistics exhibit similar qualitative trends across factors, so we show only PairCos distributions in Figure 2.
3.5pt
@ c c @ |ccccc|ccccc|ccccc|ccccc|ccccc|ccccc|ccccc|ccccc @ & & & & & & & & &
(lr)3-7(lr)8-12(lr)13-17(lr)18-22(lr)23-27(lr)28-32(lr)33-37(lr)38-42 & & 1K & 10K & 50K & 100K & All & 1K & 10K &
50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K & 100K & All & 1K & 10K & 50K
& 100K & All & 1K & 10K & 50K & 100K & All
& E & 7.7 & 1.3 & 0.4 & 0.2 & 0.3 & 21.3 & 12.8 & 7.3 & 5.8 & 5.3 & 39.6 & 20.7 & 10.7 & 9.0 & 8.4
& 34.6 & 17.4 & 6.3 & 3.9 & 2.8 & 26.3 & 7.7 & 1.8 & 1.1 & 1.0 & 9.6 & 1.4 & 0.3 & 0.3 &
0.1 & 27.1 & 9.6 & 3.6 & 3.1 & 2.8 & 8.3 & 2.9 & 1.7 & 1.5 & 1.3
& C & .99 & .92 & .56 & .71 & .77 & .99 & 1. & .96 & .91 & .89 & 1. & 1. &
.90 & .80 & .75 & 1. & .99 & .78 & .61 & .52 & 1. & 1. & .94 & .83 & .76 & 1. &
.99 & .53 & .71 & .78 & 1. & 1. & .96 & .84 & .75 & .99 & .94 & .72 & .62 & .58
& \(\kappa\) & .78 & .60 & .84 & .89 & .90 & .83 & .94 & .70 & .58 & .54 &
.96 & .90 & .61 & .71 & .74 & .98 & .93 & .52 & .67 & .72 & 1. & 1. & .82 & .63 &
.55 & 1. & .99 & .51 & .68 & .73 & 1. & .99 & .92 & .81 & .73 & .99 & .96 & .78 &
.69 & .65
& L & 1. & .98 & .74 & .59 & .51 & 1. & 1. & .99 & .97 & .96 & 1. & 1. &
.99 & .97 & .95 & 1. & 1. & .90 & .79 & .73 & 1. & 1. & .97 & .91 & .87 & 1. &
.96 & .54 & .72 & .79 & 1. & 1. & .93 & .81 & .73 & .96 & .90 & .63 & .53 & .52
& S & .78 & .59 & .86 & .90 & .90 & .93 & .93 & .63 & .54 & .55 & .83 & .82 & .78 & .83 & .83 & .99 & .97 & .56 & .63 &
.66 & 1. & 1. & .65 & .58 & .61 & 1. & .98 & .55 & .72 & .73 & .99 & .98 & .85 & .70 & .66 & 1. & 1.
& .93 & .84 & .83
(l)2-42 & E & 9.3 & 1.3 & 0.4 & 0.2 & 0.3 & 22.8 & 13.3 & 7.5 & 5.7 & 5.5 & 39.4 & 21.0 & 10.6 & 9.0 &
8.1 & 34.2 & 17.6 & 6.1 & 4.0 & 2.6 & 27.8 & 8.7 & 1.8 & 1.2 & 0.9 & 10.2 & 1.2 & 0.3 & 0.3 &
0.1 & 27.2 & 10.4 & 3.8 & 3.0 & 2.6 & 8.7 & 3.1 & 1.7 & 1.5 & 1.4
& C & .97 & .89 & .58 & .71 & .77 & .97 & .99 & .94 & .91 & .88 & 1. & 1. & .89 & .80 & .75 & 1. & .98 & .77 & .61
& .53 & 1. & 1. & .93 & .83 & .75 & 1. & .96 & .58 & .72 & .78 & 1. & 1. & .96 &
.85 & .76 & .97 & .92 & .69 & .60 & .56
& \(\kappa\) & .62 & .51 & .86 & .89 & .91 & .65 & .88 & .61 & .53 & .50 & .83 & .80 & .69 &
.76 & .78 & .94 & .89 & .57 & .69 & .72 & 1. & .99 & .77 & .59 & .52 & 1. & .96 & .55 & .69
& .73 & .99 & .99 & .92 & .81 & .74 & .98 & .94 & .76 & .68 & .64
& L & .99 & .97 & .71 & .59 & .51 & 1. & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 & 1. & .99 & .89 & .80 &
.73 & 1. & 1. & .96 & .91 & .87 & .99 & .92 & .59 & .73 & .79 & 1. & 1. & .92 & .82 & .73
& .93 & .87 & .60 & .52 & .53
& S & .67 & .53 & .87 & .91 & .90 & .93 & .90 & .62 & .54 & .56 & .57 & .67 & .83 & .86 & .84
& .97 & .95 & .56 & .62 & .65 & .99 & .99 & .57 & .62 & .62 & 1. & .96 & .59 & .73 & .73 & .98 & .97 & .84 & .69 & .65 & .99 & 1.
& .94 & .87 & .85
(l)2-42 & E & 8.9 & 1.2 & 0.5 & 0.4 & 0.2 & 23.0 & 13.8 & 7.5 & 5.9 & 5.1 & 40.1 & 20.6 & 10.7 &
8.9 & 8.3 & 34.3 & 17.4 & 6.2 & 4.0 & 2.8 & 28.5 & 7.9 & 1.9 & 1.2 & 0.9 & 10.5 & 1.4 & 0.3 & 0.2 & 0.2 & 27.4 &
10.2 & 3.6 & 2.9 & 2.8 & 8.7 & 2.9 & 1.6 & 1.4 & 1.3
& C & .99 & .91 & .58 & .72 & .77 & .99 & 1. & .96 & .91 & .89 & 1. & 1. & .89 & .81
& .75 & 1. & .99 & .77 & .60 & .52 & 1. & 1. & .94 & .84 & .76 & 1. & .98 & .56 & .72 & .78 &
1. & 1. & .96 & .84 & .75 & .99 & .94 & .71 & .62 & .57
& \(\kappa\) & .78 & .57 & .84 & .89 & .90 & .80 & .93 & .69 & .59 & .54 & .95 & .88 & .62 & .71 & .74 & .98 & .93
& .53 & .67 & .72 & 1. & 1. & .82 & .63 & .55 & 1. & .98 & .53 & .69 & .74 & 1. & .99 &
.92 & .81 & .73 & .99 & .96 & .78 & .69 & .65
& L & 1. & .98 & .72 & .59 & .51 & 1. & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 &
1. & 1. & .89 & .79 & .72 & 1. & 1. & .97 & .92 & .87 & 1. & .95 & .57 & .73 & .79 & 1. &
1. & .92 & .81 & .73 & .96 & .89 & .62 & .53 & .52
& S & .79 & .60 & .86 & .90 & .90 & .93 & .95 & .72 & .63 & .64
& .88 & .86 & .70 & .78 & .79 & .99 & .97 & .64 & .55 & .59 & 1. & 1. & .72 & .50 & .53 &
1. & .97 & .62 & .76 & .75 & .99 & .99 & .86 & .73 & .70 & .99
& .99 & .87 & .79 & .78
& E & 8.2 & 1.2 & 0.4 & 0.2 & 0.2 & 21.0 & 12.5 & 7.0 & 5.2 & 4.5 & 39.3 & 20.6 & 10.0
& 8.3 & 7.7 & 34.3 & 17.5 & 6.0 & 3.5 & 2.1 & 26.0 & 7.9 & 1.6 & 0.9 & 0.6 & 9.5 & 1.5 & 0.3 & 0.1 &
0.2 & 26.4 & 9.5 & 3.9 & 3.0 & 2.5 & 8.2 & 3.0 & 1.5 & 1.4 & 1.2
& C & .99 & .92 & .50 & .69 & .76 & .99 & 1. & .96 & .92 & .89 & 1. & 1. &
.92 & .82 & .75 & 1. & .99 & .83 & .64 & .53 & 1. & 1. & .96 & .86 & .77
& 1. & .99 & .55 & .67 & .76 & 1. & 1. & .98 & .88 & .78 & .99 & .95 & .75 & .65
& .59
& \(\kappa\) & .80 & .62 & .81 & .88 & .90 & .86 & .95 & .74 & .62 & .57 &
.98 & .90 & .57 & .70 & .74 & .98 & .94 & .56 & .64 & .71 & 1. & 1. &
.88 & .67 & .56 & 1. & .99 & .58 & .64 & .72 & 1. & 1. & .95 & .85 & .76 & .99 &
.97 & .81 & .71 & .66
& L & 1. & .98 & .77 & .62 & .53 & 1. & 1. & .99 & .98 & .97 & 1. & 1. &
.99 & .97 & .95 & 1. & 1. & .92 & .82 & .74 & 1. & 1. & .98 & .93 & .88 & 1. &
.97 & .53 & .67 & .77 & 1. & 1. & .95 & .85 & .75 & .97 & .91 & .67 &
.56 & .50
& S & .80 & .61 & .84 & .89 & .89 & .94 & .94 & .67 & .57 & .56 & .86 & .83 & .76 & .83 & .83 & .99 & .97 & .62 &
.60 & .66 & 1. & 1. & .69 & .57 & .62 & 1. & .98 & .51 & .70 & .72 & .99 & .99 & .88 & .74 & .67 & 1.
& 1. & .95 & .87 & .85
(l)2-42 & E & 9.4 & 1.5 & 0.3 & 0.2 & 0.1 & 22.0 & 13.9 & 7.3 & 5.5 & 4.9 & 39.5 & 20.2 & 9.6 & 8.6 & 7.7 & 34.9
& 17.2 & 5.9 & 3.3 & 2.1 & 27.9 & 8.5 & 1.4 & 0.9 & 0.6 & 10.3 & 1.3 & 0.4 & 0.1 & 0.1 & 27.3
& 10.2 & 3.5 & 2.7 & 2.8 & 8.8 & 3.0 & 1.6 & 1.4 & 1.3
& C & .98 & .89 & .54 & .69 & .76 & .97 & .99 & .96 & .92 & .89 & 1. & 1. & .91 & .82 & .76 & 1. & .99 & .81 & .64
& .54 & 1. & 1. & .95 & .85 & .76 & 1. & .97 & .51 & .69 & .76 & 1. & 1. & .97 & .89 & .79 & .97 &
.92 & .72 & .62 & .57
& \(\kappa\) & .65 & .51 & .84 & .89 & .90 & .66 & .89 & .65 & .56 & .52 & .86 & .83 & .66 &
.74 & .77 & .95 & .90 & .51 & .67 & .72 & 1. & 1. & .83 & .63 & .54 & 1. & .97 & .52 & .66 & .71 & .99 & .99 &
.94 & .85 & .76 & .98 & .95 & .78 & .70 & .65
& L & 1. & .97 & .75 & .61 & .52 & .99 & 1. & .99 & .97 & .96 & 1. & 1. & .99 & .97 & .95 & 1. & .99 & .91 & .81 &
.74 & 1. & 1. & .97 & .92 & .87 & .99 & .94 & .52 & .69 & .77 & 1. & 1. & .95 & .85 & .76 & .94 & .88 & .64 &
.54 & .51
& S & .69 & .55 & .85 & .90 & .89 & .93 & .91 & .66 & .56 & .58 & .60 & .70 & .80 & .85 & .83 & .97 & .96
& .62 & .60 & .63 & 1. & .99 & .62 & .62 & .64 & 1. & .96 & .52 & .71 & .72 & .98 & .97 & .87 & .73 & .67 & .99 & 1. &
.97 & .89 & .87
(l)2-42 & E & 8.6 & 1.1 & 0.3 & 0.3 & 0.1 & 21.3 & 13.4 & 7.0 & 5.3 & 4.3 & 39.6 & 20.6 & 9.8 & 8.5 &
7.6 & 34.0 & 17.2 & 5.8 & 3.5 & 2.3 & 27.4 & 7.2 & 1.5 & 0.9 & 0.8 & 10.4 & 1.5 & 0.4 & 0.2 &
0.1 & 27.5 & 9.4 & 3.4 & 2.8 & 2.7 & 8.6 & 2.8 & 1.5 & 1.3 & 1.3
& C & .99 & .92 & .53 & .70 & .76 & .99 & 1. & .96 & .92 & .89 & 1. & 1. & .91 & .82 & .75 &
1. & .99 & .81 & .63 & .53 & 1. & 1. & .96 & .86 & .78 & 1. & .99 & .52 & .69 & .76 &
1. & 1. & .97 & .87 & .78 & .99 & .95 & .74 & .64 & .59
& \(\kappa\) & .80 & .59 & .82 & .88 & .90 & .81 & .93 & .72 & .62 & .57 & .96 & .88 & .59 & .70 & .74 & .98 & .93
& .52 & .65 & .71 & 1. & 1. & .85 & .66 & .57 & 1. & .99 & .54 & .66 & .72 & 1. & .99 & .94 & .84 & .76 &
.99 & .96 & .80 & .71 & .66
& L & 1. & .98 & .76 & .61 & .52 & 1. & 1. & .99 & .97 & .97 & 1. & 1. & .99 & .97 &
.96 & 1. & 1. & .91 & .80 & .73 & 1. & 1. & .98 & .93 & .88 & 1. & .96 & .50 & .69 &
.77 & 1. & 1. & .94 & .84 & .75 & .97 & .90 & .65 & .55 & .50
& S & .81 & .62 & .84 & .89 & .90 & .93 & .95 & .75 & .65 & .64 &
.90 & .87 & .69 & .78 & .79 & .99 & .97 & .68 & .53 & .59 & 1. & 1. &
.75 & .52 & .54 & 1. & .97 & .55 & .73 & .74 & .99 & .99 & .89 & .76 &
.72 & 1. & .99 & .90 & .81 & .80
& E & 8.5 & 1.2 & 0.3 & 0.2 & 0.3 & 20.2 & 12.8 & 6.1 & 5.1 & 4.2 &
38.0 & 19.3 & 9.1 & 7.7 & 7.0 & 34.2 & 16.8 & 5.2 & 2.8 & 1.5 & 26.2 & 6.7 & 1.2 & 0.7 & 0.6 &
9.6 & 1.4 & 0.3 & 0.1 & 0.1 & 26.4 & 9.6 & 3.3 & 2.6 & 2.4 & 8.1 & 2.8 & 1.5 & 1.3 &
1.2
& C & .99 & .92 & .54 & .65 & .75 & .99 & .99 & .97 & .94 & .91 &
1. & 1. & .93 & .84 & .76 & 1. & .99 & .86 & .69 & .55 & 1. & 1. &
.97 & .89 & .78 & 1. & .99 & .62 & .61 & .73 & 1. & 1. & .99 &
.92 & .82 & .99 & .95 & .78 & .68 & .61
& \(\kappa\) & .83 & .63 & .78 & .87 & .90 & .88 & .96 & .78 & .68 &
.61 & .98 & .91 & .52 & .67 & .73 & .98 & .95 & .61 & .60 & .69 & 1. &
1. & .91 & .72 & .58 & 1. & .99 & .65 & .59 & .69 & 1. & 1. &
.96 & .89 & .79 & .99 & .97 & .83 & .75 & .68
& L & 1. & .98 & .80 & .67 & .55 & 1. & 1. & .99 & .98 & .97 &
1. & 1. & .99 & .98 & .96 & 1. & 1. & .93 & .84 & .75 & 1. &
1. & .98 & .94 & .88 & 1. & .97 & .59 & .61 & .75 & 1. & 1.
& .96 & .89 & .78 & .97 & .91 & .70 & .60 & .52
& S & .82 & .62 & .82 & .89 & .88 & .94 & .94 & .72 & .61 & .58 & .88 & .84 & .72 & .81 & .82 & .99 & .97 & .66 &
.57 & .64 & 1. & 1. & .71 & .55 & .63 & 1. & .98 & .57 & .66 & .70 & .99 & .99 & .90 & .78 & .70 &
1. & 1. & .97 & .91 & .87
(l)2-42 & E & 8.7 & 1.1 & 0.3 & 0.3 & 0.1 & 21.8 & 13.5 & 7.2 & 5.2 & 4.2 & 38.7 & 19.6 & 9.1 & 7.5 & 7.0 & 34.2 & 16.4 &
5.0 & 2.5 & 1.6 & 28.6 & 7.7 & 1.4 & 0.8 & 0.5 & 10.9 & 1.4 & 0.2 & 0.2 & 0.1 & 26.7 & 9.7 & 3.4 & 2.7
& 2.5 & 8.5 & 2.9 & 1.6 & 1.3 & 1.3
& C & .98 & .89 & .51 & .66 & .75 & .97 & .99 & .97 & .93 & .90 & 1. & 1. & .92 & .84 & .76 & 1. & .99 & .83 & .68 &
.56 & 1. & 1. & .96 & .88 & .77 & 1. & .98 & .56 & .64 & .74 & 1. & 1. & .98 & .92 & .82 &
.97 & .92 & .74 & .66 & .59
& \(\kappa\) & .69 & .53 & .82 & .88 & .90 & .71 & .90 & .71 & .62 & .57 & .90 & .84 & .62 &
.72 & .76 & .96 & .91 & .54 & .63 & .71 & 1. & 1. & .87 & .67 & .55 & 1. & .98 & .58 & .61 & .70 & .99
& .99 & .95 & .88 & .79 & .98 & .95 & .80 & .73 & .67
& L & 1. & .97 & .77 & .65 & .54 & 1. & 1. & .99 & .98 & .97 & 1. & 1. & .99 & .97 & .96 & 1. & .99 & .92 & .84 & .75 & 1. & 1. & .97
& .94 & .88 & .99 & .94 & .54 & .64 & .75 & 1. & 1. & .96 & .88 & .78 & .94 & .88 & .66 & .58 & .51
& S & .72 & .56 & .84 & .89 & .88 & .95 & .92 & .72 & .60 & .60 & .66 & .72 & .77 & .83 & .82
& .98 & .96 & .66 & .56 & .61 & 1. & .99 & .62 & .61 & .65 & 1. & .95 & .51 & .69 & .71 & .98 & .97 & .88 & .77 & .69 &
.99 & 1. & .98 & .92 & .90
(l)2-42 & E & 8.8 & 1.1 & 0.3 & 0.1 & 0.1 & 22.0 & 12.9 & 6.4 & 5.1 & 4.3 & 40.0 & 20.3 &
8.9 & 7.9 & 7.0 & 33.8 & 16.3 & 5.0 & 3.0 & 1.4 & 27.9 & 7.0 & 1.2 & 0.7 & 0.7 & 10.1 &
1.3 & 0.3 & 0.2 & 0.1 & 26.2 & 9.0 & 3.0 & 2.7 & 2.4 & 8.5 & 2.7 & 1.4 &
1.3 & 1.2
& C & .99 & .91 & .51 & .65 & .75 & .99 & 1. & .97 & .94 & .91 & 1. & 1. & .92 & .84 & .76 & 1. & .99 &
.83 & .68 & .55 & 1. & 1. & .96 & .88 & .78 & 1. & .99 & .57 & .63 & .74 & 1. & 1. & .98 & .92 & .80
& .99 & .94 & .76 & .68 & .61
& \(\kappa\) & .81 & .58 & .81 & .87 & .90 & .86 & .94 & .76 & .67 & .60 & .96 & .89 & .55 & .67 & .73 & .98 & .93 & .57 & .61 & .70 &
1. & 1. & .89 & .71 & .58 & 1. & .98 & .59 & .61 & .70 & 1. & .99 & .95 & .88 & .78 & .99 & .96 & .82 & .74 & .68
& L & 1. & .98 & .78 & .66 & .54 & 1. & 1. & .99 & .98 & .97 & 1. & 1. & .99 & .98 & .96 & 1. & 1. &
.92 & .83 & .74 & 1. & 1. & .98 & .94 & .88 & 1. & .96 & .54 & .63 & .75 & 1. & 1. & .96 & .88 & .78 & .97 & .89 & .67 & .59 & .52
& S & .83 & .61 & .82 & .88 & .89 & .94 & .95 & .78 & .69 & .65 & .91 &
.87 & .67 & .77 & .79 & .99 & .97 & .71 & .51 & .59 & 1. & 1. & .76 & .52 & .56 &
1. & .97 & .51 & .68 & .71 & .99 & .99 & .90 & .81 & .74 & .99 & .99 & .93 & .86 & .83
& E & 9.7 & 1.8 & 0.3 & 0.3 & 0.3 & 19.2 & 10.0 & 5.2 & 5.2 & 4.2 & 38.1 & 24.7 & 11.8 & 9.8 & 8.8 & 34.1 & 19.8 & 6.6 & 4.5 & 3.1 & 28.9 & 13.9 & 2.7
& 1.9 & 1.3 & 11.7 & 2.3 & 0.4 & 0.3 & 0.2 & 29.2 & 13.3 & 4.3 & 3.6 & 3.0 & 8.7 & 3.9 & 1.7 & 1.6 & 1.4
& C & .98 & .80 & .65 & .70 & .76 & .98 & .98 & .90 & .88 & .85 & 1. & .99 & .88 & .84 & .77 & 1. & .97 & .72 & .65 & .55 & 1. & 1. & .95 & .90
& .82 & 1. & .92 & .57 & .63 & .71 & 1. & 1. & .95 & .91 & .83 & .96 & .88 & .65 & .62 & .57
& \(\kappa\) & .62 & .65 & .88 & .89 & .90 & .67 & .67 & .53 & .52 & .51 & .78 & .65 & .68 & .71 & .74 & .94 & .79 & .60 & .65 & .70 &
1. & 1. & .80 & .72 & .61 & 1. & .92 & .55 & .61 & .67 & .99 & .98 & .90 & .87 & .79 & .97 & .92 & .72 & .69 & .65
& L & 1. & .93 & .66 & .60 & .52 & 1. & .99 & .97 & .96 & .95 & 1. & 1. & .98 & .97 & .96 & 1. & .99 & .86 & .82 & .75 & 1. & 1. & .97 & .95 &
.91 & 1. & .87 & .58 & .64 & .73 & 1. & 1. & .92 & .87 & .80 & .93 & .81 & .57 & .54 & .51
& S & .67 & .58 & .88 & .89 & .89 & .85 & .76 & .57 & .54 & .54 & .62 & .51 & .79 & .82 & .80 & .97 & .93 & .58 & .54 & .60 & .99 & .99 & .65 & .50
& .56 & .99 & .96 & .57 & .67 & .69 & .98 & .96 & .84 & .76 & .70 & .99 & 1. & .92 & .87 & .86

Figure 2: PairCos (\(\bar{s}\)) kernel density estimates. The bold black curve (with shading) shows member identities (present in the training set). All other curves show non-member identities: the nine thin solid lines correspond to cross-distribution benchmarks whose domain differs from the training data (one colour per dataset; see legend), while the dashed line (“Within-dist”) shows non-members from the held-out WebFace4M split, which shares the training domain (no domain shift). (a, b) vary training epochs at fixed \(n_{\text{ids}}\); (c) varies \(n_{\text{ids}}\) at epoch 40.. a — IR-100 / MagFace / 100K identities., b — IR-100 / MagFace / 1K identities., c — IR-100 / MagFace / epoch 40.
Figure 2 presents the main finding of this study. Panel (c) shows that the PairCos distributions of members and non-members converge as \(n_{\text{ids}}\) increases from 1K to all identities. At \(n_{\text{ids}}{=}1\)K the distributions are well separated, whereas at \(n_{\text{ids}}{=}100\)K and beyond the separation narrows, though it remains above chance even on the held-out WebFace4M split (Table 3, discussed in Section 5.4).
Panels (a) and (b) of Figure 2 isolate the effect of training duration at two extremes of identity coverage. At \(n_{\text{ids}}{=}1\)K, the PairCos separation between members and non-members increases with the number of training epochs. At \(n_{\text{ids}}{=}100\)K the same trend is present but attenuated.
Table 2 reports the ANOVA results (\(\hat{\eta}^2_p\), the proportion of variance explained after accounting for all other factors). Training-set size explains the most variance (\(\hat{\eta}^2_p \geq .95\)), followed by training duration (\(\hat{\eta}^2_p \geq .89\)). The strong \(n_{\text{ids}}\!\times\!\text{Epoch}\) interaction (\(\hat{\eta}^2_p \geq .92\)) further shows that the epoch effect is moderated by identity coverage. The \(n_{\text{ids}}\!\times\!\text{Epoch}\) cell means on the same-domain WebFace4M split show that the mean PairCos AUC increase from epoch 5 to epoch 40 ranges from \(+.38\) at \(n_{\text{ids}}{=}10\)K down to \(+.16\) at \(n_{\text{ids}}{=}100\)K, with \(n_{\text{ids}}{=}1\)K reaching a ceiling near 1.0 already at epoch 10. Models with small identity pools thus show the largest separability gains over training, whereas models trained on \(100\)K identities have only a modest increase.
Backbone and loss head. The ANOVA also shows that backbone architecture and loss head explain far less variance in threshold-classifier AUC, the loss-head factor reaching at most \(\hat{\eta}^2_p = .04\) and backbone depth at most \(.09\) across all four statistics — an order of magnitude below the contributions of \(n_{\text{ids}}\) and number of epochs. Beyond the IResNet grid, a ViT-T [33], [34] backbone (CosFace) trained under the identical protocol [35] reproduces the same \(n_{\text{ids}}\) trend in both EER and membership AUC (Table ¿tbl:tab:eer?, bottom block), so the effect is not specific to convolutional backbones.
The benchmarks used in this study differ in pose, quality, age, and demographics, which affect embedding-cluster geometry independently of membership. The distributional-shift rows of Table 2 quantify these effects, and we discuss each factor below. For matched-pair benchmarks, we also translate each effect into the AUC change it would produce at the same-domain training-attributable baseline at \(n_{\text{ids}}{=}100\)K (where the AUC is \(\approx 0.71\)). We do so by passing each AUC through the inverse standard normal cumulative distribution function (a probit transform), on which equal shifts in the underlying non-member distribution correspond to equal numerical changes regardless of baseline. We take the matched-pair difference on this transformed scale and convert it back to AUC at the chosen baseline. This step is needed because the AUC scale saturates near 1, so a fixed shift in the underlying distribution produces a smaller AUC change near the ceiling than in the middle.
Pose. Profile-view faces produce large shifts in cluster-statistic distributions. The CFP-FP vs.CFP-FF comparison is the cleanest controlled experiment: both sets share the same 500 identities, so the observed Cohen’s \(d'\) (standardised mean difference) is attributable to pose alone (Table 2). CPLFW vs.LFW yields even larger \(|d'|\) for PairCos and PenLogit, but those two sets are not identity-matched, so other factors may contribute. In both comparisons, profile views widen the intra-identity distribution regardless of membership. This pushes statistic values towards the non-member range and inflates separability. At the \(n_{\text{ids}}{=}100\)K baseline, the pose contribution from CFP-FF to CFP-FP corresponds to an AUC change of \(+0.26\), matching or exceeding the training contribution above chance (\(+0.21\)).
Image quality. XQLFW shares LFW’s identity set and pair protocol but degrades one image per pair. For PairCos and PenLogit, the resulting Cohen’s \(d'\) is of the same order as that of cross-pose variation though somewhat smaller; \(\kappa\) shows a smaller quality effect (\(|d'|{=}0.48\) vs. \(1.13\) for pose). Lower image quality increases intra-identity spread in embedding space and shifts cluster-statistic distributions in the same direction as pose variation. At the same baseline, the quality contribution from LFW to XQLFW corresponds to an AUC change of \(+0.22\), comparable to the pose contribution.
Age. Within AgeDB-30, the Spearman rank correlation \(\rho\) (a non-parametric measure of monotonic association) between within-identity age spread and each cluster statistic (Table 2) is weak and varies in sign across statistics. Only \(\kappa\) exhibits the expected negative association (\(\rho = -0.26\)), under which identities spanning wider age ranges form less concentrated clusters, whereas PairCos is essentially uncorrelated (\(\rho = +0.01\)) and both PenLogit and ProtoCE carry small positive correlations (\(\rho = +0.17\) and \(\rho = +0.33\)). Age-range variation therefore exerts the weakest and least consistent influence on cluster geometry among the confounds examined here.
Ethnicity. A Kruskal–Wallis test (non-parametric one-way ANOVA) across the four RFW demographic groups (African, Asian, Caucasian, Indian) yields \(\epsilon^2 \approx 0.02\) (small) for PairCos, \(\kappa\), and PenLogit, and \(\epsilon^2 = 0.12\) (medium) for ProtoCE. Cluster-statistic values therefore vary across demographic groups within the non-member pool, with ProtoCE again being the most sensitive statistic. Although these effect sizes are modest relative to the \(n_{\text{ids}}\) factor, they indicate that the demographic composition of a non-member benchmark affects the baseline against which membership is measured.
Pose and quality produce the largest shifts in cluster-statistic distributions, with demographic composition and age contributing smaller and, for age, less consistent effects. The matched-pair pose and quality contributions show that the domain-induced component can match or exceed the training contribution itself at higher \(n_{\text{ids}}\), so the threshold-classifier AUC on any non-WebFace4M benchmark cannot be interpreted as membership signal alone.
To test whether combining statistics recovers more membership information, we train a two-hidden-layer MLP (\(64{\to}32\) units, ReLU, BatchNorm, dropout 0.3; Adam with \(\eta=10^{-3}\)) on all four cluster statistics. For each epoch-40 configuration we perform five-fold stratified cross-validation with an 80/20 identity split on same-domain WebFace4M identities, reserving 15% of each training fold as a validation set for early stopping.
| \(\nids\) | |||||
|---|---|---|---|---|---|
| 3-6 Backbone | Loss | 1K | 10K | 50K | 100K |
| IR-34 | Arc | 1.00 / 1.00 | .98 / 1.00 | .80 / .87 | .65 / .69 |
| Cos | .99 / 1.00 | .98 / .99 | .79 / .88 | .65 / .71 | |
| Mag | 1.00 / 1.00 | .98 / 1.00 | .79 / .85 | .64 / .68 | |
| 1-6 | Arc | 1.00 / 1.00 | .99 / 1.00 | .84 / .91 | .68 / .73 |
| Cos | .99 / 1.00 | .98 / .99 | .82 / .91 | .68 / .75 | |
| Mag | 1.00 / 1.00 | .98 / 1.00 | .82 / .89 | .67 / .72 | |
| 1-6 | Arc | 1.00 / 1.00 | .99 / 1.00 | .86 / .92 | .73 / .79 |
| Cos | .99 / 1.00 | .98 / .99 | .84 / .92 | .71 / .80 | |
| Mag | 1.00 / 1.00 | .98 / 1.00 | .85 / .91 | .72 / .79 | |
3pt
Table 3 reports the best single-statistic AUC alongside the fold-averaged MLP AUC. The MLP consistently matches or outperforms the best individual statistic, with the largest gains at \(n_{\text{ids}}{=}50\text{K}\) and \(100\text{K}\) where individual statistics are weakest (one-sided Wilcoxon signed-rank: \(W{=}666\), \(p<10^{-10}\)).
Generalisation across training data. All experiments use a single training source (WebFace4M). Whether the trends generalise to other training datasets is an open question, as comparably large datasets are not publicly available or have been retracted.
Experimental controls. We do not evaluate defences such as DP-SGD [23] or knowledge distillation [36]. The number of probe images \(k\) per identity varies across benchmarks and is not controlled for, which may affect estimates of the cluster statistics non-uniformly. The held-out split is constructed by random partitioning; in practice, membership boundaries may correlate with identity-level attributes.
Statistical assumptions and scope. Because our four statistics capture only cluster geometry, the measured separability is a lower bound on the membership information available from embeddings; a richer model operating on raw embedding vectors could extract more. The probit-scale translation in Section 5.3 assumes approximately Gaussian member and non-member statistic distributions and stable variances across matched pairs.
We evaluated how much membership information is retained in the embedding geometry of 180 open-set face recognition models across nine benchmarks. The results lead to three practical conclusions. First, training-set size has the largest effect on member/non-member separability, while backbone architecture and loss head contribute far less. Training on more identities is therefore the only design choice we tested that substantially reduces the geometric membership signal. Second, cross-domain non-member benchmarks inflate the apparent signal by conflating domain shift with training-attributable differences. Privacy audits should therefore use same-domain held-out references, since cross-domain non-members overestimate the apparent membership signal. Third, even at high \(n_{\text{ids}}\) where individual statistics are weak, fusing all statistics with a learned classifier recovers additional membership information (Table 3). As our statistics capture only cluster geometry, the measured separability is a lower bound on the membership information available from embeddings. Future work includes extending these benchmarks to other backbones and evaluating whether defences such as [23] can reduce the geometric signal without degrading performance.
This work has received funding from the European Union’s Horizon Europe research and innovation programme under Grant Agreement No. 101189650 (CERTAIN: Certification for Ethical and Regulatory Transparency in Artificial Intelligence), and the Swiss State Secretariat for Education, Research and Innovation (SERI).
This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.↩︎