June 23, 2026
Federated unlearning (FU) is critical for complying with legal mandates like the right to be forgotten in decentralized systems, yet current methods face a persistent dilemma between non-target knowledge loss and high request latency. To resolve these issues, we propose FedUP, a one-shot federated unlearning framework utilizing lightweight pluggable filters that act as a “knowledge funnel" to screen out target data while preserving original model performance. By freezing original model parameters and training filters at the server side using differentially private (DP)-protected class centroid samples, FedUP bypasses the need for multi-round client-server communication and complex retraining, reducing unlearning latency from minutes to mere seconds. Additionally, the framework’s pluggable architecture ensures inherent reversibility, enabling the seamless restoration of forgotten knowledge by simply removing the filters. Extensive experiments on diverse image and text tasks demonstrate that FedUP effectively reduces non-target knowledge loss and achieves superior unlearning precision and efficiency across various scenarios. Code is available at: https://github.com/suows/FedUP-code.
Background. As a distributed machine learning paradigm, federated learning (FL) [1]–[7] has gained significant attention in privacy-sensitive scenarios recently because it eliminates the need for centralizing raw data during training. However, during the training process, the global model internalizes client information into its parameters through multiple rounds of parameter aggregation, which exposes novel privacy risks when the model is confronted with legal mandates such as the right to be forgotten under GDPR [8]. To address this, researchers have focused on Federated Unlearning (FU), aiming to remove specific knowledge without retraining the global model from scratch. Existing FU methods generally follow two main technical paradigms: server-side FU methods [9]–[11], which implement approximate unlearning [12] on the global model at the server side; and client-side FU methods [13]–[16], which achieve exact unlearning [17] through iterative client-side model retraining procedures.
Figure 1: FedEraser, a representative server-side method, shows noticeable non-target knowledge loss after unlearning across multiple tasks. Client-side methods like FUSED experience significantly longer convergence times as task difficulty increases, with responses up to 5 minutes indicating severe unlearning request latency.. a — Non-target knowledge loss of server-side FU method., b — Unlearning request latency of client-side FU method.
Existing Challenges. While current Federated Unlearning (FU) methods facilitate knowledge removal to various extents, challenges regarding non-target knowledge loss and unlearning request latency still remain. As depicted in Figure 1, server-side FU methods lack precise control over the unlearning scope, frequently incurring non-target knowledge loss, where the model performance unrelated to the unlearning request is inevitably impaired during knowledge removal. Conversely, although client-side FU methods achieve exact unlearning via retraining-based approaches, they are heavily constrained by multi-round training and communication, leading to high request latency. To date, it is hard to find a solution that simultaneously ensures rapid response and prevents non-target knowledge loss. Moreover, both server-side and client-side FU paradigms typically rely on direct modification of original model parameters, which leads to the irreversibility of unlearning. Once the unlearning process is finalized, restoring previously removed knowledge necessitates further parameter tuning or retraining, thereby imposing complexity and additional overhead on practical deployment.
Proposed Solution. To this end, we propose FedUP, a one-shot federated unlearning framework based on lightweight pluggable filters. In terms of mitigating non-target knowledge loss, the framework avoids performing direct, large-scale parameter updates on the global model; instead, it implements unlearning by introducing independent pluggable filters while keeping original model parameters frozen (shown in Figure 2). These filters serve as a “knowledge funnel" that screens out target knowledge and permits only non-target knowledge to pass, thereby reducing the interference of unlearning on non-target knowledge. To lower unlearning request latency, FedUP only needs to perform a few rounds of fine-tuning on the lightweight filters at the server side using differential privacy (DP)-protected class centroid samples to complete the unlearning task. This bypasses multi-round retraining and frequent communication with clients, significantly accelerating the response speed. Moreover, since the filters are pluggable, the framework inherently supports reversibility. As depicted in Figure 2, when restoration of forgotten knowledge is required, it can be achieved simply by removing the filters. Overall, FedUP provides a solution that balances precision, efficiency, and recoverability.
Contributions. The main contributions are as follows:
We design FedUP, a one-shot FU framework utilizing DP-protected class centroids, mitigating non-target knowledge loss and reducing unlearning request latency from minutes to seconds.
We propose a reversible unlearning mechanism via lightweight pluggable filters without altering original model parameters while ensuring rapid transitions between unlearned and pre-unlearning states.
We conduct differential privacy analysis showing that appropriate noise protects class centroid samples while minimally impacting data utility. Extensive experiments on image and text tasks confirm FedUP’s effectiveness across various scenarios.
Machine Unlearning. Simply removing training data from storage fails to purge its influence on deployed models. Machine Unlearning (MU) [18] is thus introduced to erase such learned knowledge. Existing methods are categorized as either exact [19] or approximate unlearning [20]. Exact unlearning mandates that the post-forgetting model be statistically indistinguishable from one retrained from scratch without the deleted data. It retains the guarantees of complete retraining but lowers its cost via algorithmic shortcuts. For classical models, the training process can be expressed in an invertible additive form, enabling point removal by subtracting its closed-form term [19]. For complex architectures, exact unlearning employs data sharding [21], localized retraining [22], or intermediate checkpointing [23] to reduce computation while preserving retraining-level guarantees. Approximate unlearning, in contrast, substitutes full retraining with lightweight fine-tuning, permitting a bounded, residual influence from the deleted data. Related work is typically grouped into data-driven and model-driven approaches [24]. Data-driven methods involve relabeling retained samples [25] or partitioning data shards [26] before fine-tuning. Model-driven approaches adjust parameters directly via influence functions [27], fisher-based regularization [20], or knowledge distillation [28], countering the gradient contributions of the data to be erased.
Federated Unlearning. Though federated learning maintains client privacy through local retention, knowledge from distributed datasets persists in the aggregated global model. This requires integrating MU techniques into FL, named federated unlearning (FU) [29]. Based on the execution location, FU methods can be categorized into server-side and client-side methods [30], [31]. Server-side methods are intrinsically based on approximate unlearning to expunge client contributions without client involvement. FedEraser [32] calibrates update trajectories but still requires auxiliary retraining. Conversely, Wu et al. [9] bypass client-side computation by subtracting historical updates and employing knowledge distillation. To optimize efficiency, Huynh et al. [10] utilize selective retention of influential updates to reduce memory overhead, while Pan et al. [11] resolve parameter conflicts via Orthogonal Steepest Descent to accelerate the unlearning process. Client-side FU methods perform unlearning locally, striving to achieve exact unlearning while balancing computational efficiency with global model utility. Mora et al. [33] propose FedUNRAN, utilizing local random label perturbations to disperse target gradients and attenuate their influence without server-side intervention. Wang et al. [13] employ variational Bayesian inference for parameter self-sharing to erase target data while preserving performance. To accelerate unlearning, Liu et al. [14] approximate the Hessian via a diagonal empirical Fisher Information Matrix for quasi-Newton optimization, while Zhu et al. [15] combine inverse perturbation with passive decay, propagating updates via knowledge distillation. Furthermore, Deng et al. [34] introduce model contrastive unlearning (MCU) to regularize the feature space. Zhong et al. [16] implement reversible unlearning through local fine-tuning.
In summary, existing unlearning approaches typically involve trade-offs between unlearning precision and computational efficiency, and often suffer from limited generalization. Moreover, reversibility is rarely considered in current studies. Achieving a unified balance among unlearning precision, efficiency, and reversibility therefore remains an open challenge.
In the FL framework, a set of clients denoted as \(\mathcal{C} = \{C_1, C_2, \ldots, C_N\}\) collaboratively train a global model \(\mathcal{M}_G\) without sharing local data. Each client \(C_n\) trains a local model \(\mathcal{M}_n\) using its local dataset \(\mathcal{D}_n\) and uploads the model parameters to a server. The server aggregates these local models via weighted averaging based on dataset sizes to obtain a global model \(\mathcal{M}_G\). This process iterates over \(W\) global rounds, where each global round comprises \(e\) local training epochs on the clients with a local learning rate \(l_c\). \(\mathcal{M}_G\) is structurally decomposed into a feature extractor \(\mathcal{M}_E\) and a classifier \(\mathcal{M}_{cl}\). After aggregation, \(\mathcal{M}_G\) is distributed back to all clients, where the feature extractor \(\mathcal{M}_E\) is leveraged to extract features from local data. Additionally, a pluggable filter \(\mathcal{M}_{Fi}\) is constructed to implement the filtering of knowledge that needs to be forgotten. The overall dataset is denoted as \(\mathcal{D} = \{\mathcal{D}_n\}_{n=1}^{N}\), where \(\mathcal{D}_n\) represents the local dataset of client \(C_n\). The total data volume across the federation is \(\| \mathcal{D} \|= \sum_{n=1}^{N} \| \mathcal{D}_n \|\). To facilitate data management, particularly for future unlearning operations, each local dataset \(\mathcal{D}_n\) is partitioned into two disjoint subsets: a retained subset \(\mathcal{D}_n^R\) for model training, and an unlearning subset \(\mathcal{D}_n^U\) designated to be removed in compliance with privacy or regulatory constraints. Let \(k \in \{1, 2, \ldots, K\}\) index the data categories, where \(K\) is the total number of classes.
In the FL phase, the training objective is defined as: \[\min_{\theta_{\mathcal{M}_G}} F(\theta_{\mathcal{M}_G}) = \sum_{n=1}^{N} \frac{|\mathcal{D}_n|}{|\mathcal{D}|} \sum_{(x_i^k, y_i^k) \in \mathcal{D}_n} \mathcal{L}\bigl(f(x_i^k; \theta_{\mathcal{M}_G}), y_i^k\bigr),\] where \(\mathcal{L}\) denotes the loss function. During the FU phase, the objective is to maximize the loss on the unlearning set, minimize the loss on the retained set, and keep the training cost low, described as: \[\min_{\theta_{\mathcal{M}_G'}} F^R(\theta_{\mathcal{M}_G'}) = \sum_{n=1}^{N} \frac{|\mathcal{D}_n^R|}{|\mathcal{D}^R|} \sum_{(x_i^k, y_i^k) \in \mathcal{D}_n^R} \mathcal{L}\bigl(f(x_i^k; \theta_{\mathcal{M}_G'}), y_i^k\bigr),\]
\[\max_{\theta_{\mathcal{M}_G'}} F^U(\theta_{\mathcal{M}_G'}) = \sum_{n=1}^{N} \frac{|\mathcal{D}_n^U|}{|\mathcal{D}^U|} \sum_{(x_i^k, y_i^k) \in \mathcal{D}_n^U} \mathcal{L}\bigl(f(x_i^k; \theta_{\mathcal{M}_G'}), y_i^k\bigr),\]
\[\min F^T(\theta_{\mathcal{M}_G'}) = \sum_{w=1}^{W} \text{Comm}^{(w)}(\mathcal{M}_G, \mathcal{C}\bigr),\] where Comm represents the communication resource between the server and the clients \(\mathcal{C}\) during each round.
As shown in Figure 3, our method comprises four phases: federated learning, generation of class centroid samples, one-shot federated unlearning, and inference, with each stage’s main procedures and key formulas shown in the diagram.
Each client \(C_n\) performs \(e\) rounds of training on its local dataset \(\mathcal{D}_n\), resulting in an updated local model \(\mathcal{M}_n^{w,e}\). The local training process can be expressed as: \[\theta_{\mathcal{M}_n}^{w,e} = \theta_{\mathcal{M}_n}^{w,e-1} - l_c \cdot \nabla_{\theta} \mathcal{L}(\mathcal{M}_n^{w,e-1}, \mathcal{D}_n),\] where \(\mathcal{L}\) represents the local loss function and \(\nabla_{\theta} \mathcal{L}\) denotes its gradient with respect to the model parameters.Subsequently, the server collects the local models \(\mathcal{M}_n^{w,e}\) from selected clients \(C_n\) and updates the global model through a weighted average: \[\theta_{\mathcal{M}_G}^{w+1} = \sum_{n=1}^{N} \frac{|\mathcal{D}_n|}{|\mathcal{D}|} \theta_{\mathcal{M}_n}^{w,e}.\] The trained global model \(\mathcal{M}_G\) is regarded as a combination of a feature extractor \(\mathcal{M}_E\) and a classifier \(\mathcal{M}_{cl}\). After freezing this composite model, it is deployed to the clients. The features of all data \(\mathcal{D}_n\) from each client \(C_n\) are extracted by the feature extractor \(\mathcal{M}_E\), which can be represented as follows: \[V_n^k = \left\{ \mathcal{M}_E(x_i^k) \mid x_i^k \in \mathcal{D}_n^k \right\},\] where \(V_n^k\) denotes the feature embedding of the data belonging to class \(k\) on client \(C_n\).
Upon receiving an unlearning request, client \(C_n\) removes the features associated with the unlearning data. The remaining feature embeddings belonging to class \(k\) are denoted as: \[V_n^{R,k} = V_n^k \setminus \left\{ \mathcal{M}_E(x_i) \mid x_i \in \mathcal{D}_n^{U,k} \right\}.\]
To reduce communication cost, the retained features are compressed via KMeans clustering with \[K_{n,k} = \lceil \rho \, |V_n^{R,k}| \rceil,\] the number of clusters is set proportionally according to the scenario by \(\rho\), yielding a compact set of centroids: \[\mu_n^{R,k} = \mathrm{KMeans}(V_n^{R,k}, K_{n,k}).\]
For privacy preservation, add noise to each centroid: \[\label{noise} \tilde{\mu}_{n}^{R,k} = \left\{ \mu + \mathbf{z} \;\middle|\; \mu \in \mu_{n}^{R,k}, \; \mathbf{z} \sim \mathcal{N}(0,\sigma^{2}I) \right\}.\tag{1}\] Here, \(\mathcal{N}(0, \sigma^2 \mathbf{I})\) represents a multivariate Gaussian distribution. These mechanisms ensure that the release of \(\tilde{\mu}_n^{R,k}\) does not reveal excessive information about any individual data point. Client \(C_n\) then uploads all perturbed class centroids \(\tilde{\mu}_n^{R,k}\) to the server. Using the local class centroids \(\tilde{\mu}_n^{R,k}\), the server aggregates them to generate the global class centroids \(\tilde{\mu}_\mathcal{G}^{R,k}\). The global class centroid is denoted as: \[\tilde{\mu}_\mathcal{G}^{R,k}= \left[ \tilde{\mu}_1^{R,k}; \tilde{\mu}_2^{R,k}; \ldots; \tilde{\mu}_N^{R,k}\right].\]
The objective of the FU phase is to train the filter \(\mathcal{M}_{Fi}\) such that it effectively blocks the flow of unlearning knowledge while allowing the retained learning knowledge to pass. The specific steps and formulas are as follows.
A filter \(\mathcal{M}_{Fi}\) is inserted into the global model \(\mathcal{M}_G=\mathcal{M}_E \oplus \mathcal{M}_{cl}\), yielding a new model structure: \[\mathcal{M}_G' = \mathcal{M}_E \oplus \mathcal{M}_{Fi} \oplus \mathcal{M}_{cl}.\]
Subsequently, the parameters of the global model \(\mathcal{M}_{G}\) including both \(\mathcal{M}_{E}\) and \(\mathcal{M}_{cl}\) are frozen: \[\theta_{\mathcal{M_G}} = \theta_{\mathcal{M_G}}^{*}.\]
The filter \(\mathcal{M}_{Fi}\) is trained using the global class centroids \(\tilde{\mu}_\mathcal{G}^{R,k}\). The filtered class centroid prediction is defined as: \[\{\hat{y_k}\} = \{\mathcal{M}_{\text{cl}}(\mathcal{M}_{\text{Fi}}(\tilde{\mu})) \mid \tilde{\mu} \in \tilde{\mu}_\mathcal{G}^{R,k}, k \in \mathcal{K}_R\}.\]
We propose a composite loss function for the filter consisting of cross-entropy and reconstruction losses. It preserves discriminative capability on non-target knowledge while enforcing structural consistency in the feature space, thereby improving the stability of the unlearning process without compromising knowledge selectivity. The cross-entropy loss, measuring the difference between predicted class probabilities and true labels, is defined as: \[\mathcal{L}_{\text{CE}} =-\sum_{k} y_k \log(\hat{y}_k),\] where \(y_k\) represents the true label while \(\hat{y}_k\) denotes the corresponding predicted probability for class \(k\). The reconstruction loss, which measures how well the filter reconstructs its input, is defined as: \[\mathcal{L}_{\text{RE}} = \frac{1}{d} \sum_{i=1}^{d} {|\tilde{\mu} - \mathcal{M}_{Fi}(\tilde{\mu})|^2 },\] where \(d\) is the input and output dimensionality of the filter. The total loss function used to train the filter \(\mathcal{M}_{Fi}\) is a weighted sum of these two losses: \[\label{filter32loss} \mathcal{L}_{\text{total}} = \alpha \mathcal{L}_{\text{CE}} + (1 - \alpha) \mathcal{L}_{\text{RE}},\tag{2}\] where \(\alpha\) is a weighting factor that balances the importance of the cross-entropy loss and the reconstruction loss. Train the filter for \(W_a\) rounds with a learning rate \(l_a\). The update process is as follows: \[\label{update32parameters} \theta_{\mathcal{M}_{Fi}}^w = \theta_{\mathcal{M}_{Fi}}^{w-1} - {l_a} \cdot \nabla_{\theta} \mathcal{L}_{\text{total}}\left( \theta_{\mathcal{M}_{Fi}}^{w-1}; \tilde{\mu}_\mathcal{G}^{R,k} \right) \quad w = 1, \dots, W_a.\tag{3}\]
When a client requests the restoration of forgotten knowledge, the filter \(\mathcal{M}_{Fi}\) can be removed, thereby reverting to the original global model \(\mathcal{M}_G\): \[\mathcal{M}_G = \mathcal{M}_E \oplus \mathcal{M}_{cl}.\]
Our framework injects Gaussian noise for privacy guarantee during client-side centroid uploads, satisfying (\(\varepsilon,\delta\))-DP. The key is to calibrate the noise scale \(\sigma\) according to the desired privacy budget and the \(\ell_2\)-sensitivity, which is bounded by: \[\Delta_2 f_i = \max_{p,q} \frac{ \| d_p^{(i)} - d_q^{(i)} \|_{2} }{n_i}.\]
| \(\sigma\) | |||||
| 0.001 | 0.005 | 0.01 | 0.05 | 0.1 | |
| MNIST | 11.09 | 2.22 | 1.11 | 0.22 | 0.11 |
| CIFAR-10 | 16.73 | 3.35 | 1.67 | 0.33 | 0.17 |
| AG News | 12.27 | 2.45 | 1.23 | 0.25 | 0.12 |
| CIFAR-100 | 141.40 | 28.28 | 14.14 | 2.83 | 1.41 |
9pt
We set \(\sigma \geq \frac{\sqrt{2\ln(1.25/\delta_i)} \cdot \Delta_2 f_i}{\varepsilon_i}\) with \(\delta_i = 1/n_i\). Empirical \(\varepsilon_i\) values (Table 1) confirm moderate privacy budgets \(\varepsilon \approx 10\) [35] for all datasets under specified \(\sigma\). We calculate our differentially private guarantee as:
\[\varepsilon_{i} = \frac{\sqrt{2 \ln(1.25 n_{i})} \cdot \Delta_{2} f_{i}}{ \sigma}.\] The final choice of the optimal \(\sigma\) is determined via \(\sigma = q \cdot s\), with the corresponding values of \(q\) and \(s\) established accordingly. Detailed analysis and supporting experiments are relegated to Section 2.1 and 2.2 of the supplementary material.
The algorithm, as illustrated in Algorithm 4, consists of three stages: feature extraction, generation of class centroid samples \(\mu_n^{R,k}\) and one-shot federated unlearning. Upon receiving an unlearning request, each client first removes the corresponding data points from its local feature set and computes class centroid samples using the retained data. These centroids are then protected by a differential privacy mechanism and uploaded to the server. The server aggregates the uploaded centroids from all clients to obtain global retained class centroids as \(\tilde{\mu}_\mathcal{G}^{R,k}\). Then the server freezes the parameters of the original global model \(\mathcal{M}_\mathcal{G}\) and randomly initializes a lightweight, pluggable filter \(\mathcal{M}_{Fi}\), which is trained solely using the global differentially private centroids. By jointly optimizing a cross-entropy loss and a reconstruction loss \(\mathcal{L}_{\text{total}}\), the filter blocks the propagation of forgotten knowledge while preserving the discriminative capability of retained knowledge. Finally, when restoration of forgotten knowledge is required, the filter can be removed to revert to the original global model structure, enabling efficient and reversible federated unlearning.
We conducted experiments on diverse datasets, including MNIST[36], CIFAR-10, CIFAR-100 [37], and AG News[38]. We partitioned the dataset using the Dirichlet distribution [39] with a concentration parameter of 0.5. These experiments use various network architectures, including LeNet-5, ResNet-18 [40], ResNet-34, Transformer [41] and TinyBert[42], covering three unlearning scenarios: client unlearning, class unlearning and sample unlearning. We evaluated five baselines: EraseClient [43], Federaser [32], Exact-Fun [44], Retrain and FUSED [16]. The experimental framework is implemented using PyTorch 2.3.1 and CUDA 12.1. For hardware acceleration, an NVIDIA RTX 3080 Ti GPU is utilized. We employ SGD and Adam optimizers. The balancing factor \(\alpha\) for the loss of filters is set to 0.5. During the generation of sample centroids via clustering, we set different sampling ratios \(\rho\) to accommodate various scenarios. In the client, class, and sample scenarios, the sampling ratios are 0.8, 0.1, and 0.5, respectively.
| Metric | Description |
|---|---|
| R-A(%)\(\uparrow\) | Accuracy of retained knowledge. |
| F-A(%)\(\downarrow\) | Accuracy of unlearning knowledge. |
| 0A(%)\(\uparrow\) | Accuracy of class 0. |
| PS(%)\(\uparrow\) | Prediction precision of class 0. |
| MIA(%)\(\downarrow\) | Post-unlearning privacy risk quantified via membership inference attacks. |
| Time(s)\(\downarrow\) | Time required to achieve the desired effect. |
| Comm(MB)\(\downarrow\) | Communication resource consumption between client and server. |
0pt 0pt 1pt
0pt 0pt 1pt
c c c c c c c c c c c c c c c c c c & & & &
& & E-C & Federaser & E-F & FUSED & Retrain & FedUP & FUSED & Retrain &
FedUP & & FUSED & Retrain & FedUP
& R-A & 0.99 & 0.97 & 0.97 & 0.99 & 0.99 & 0.99 & 0.97 & 0.99 & 1.00 & 0A & 0.95 & 1.00 & 0.95
& F-A & 0.00 & 0.00 & 0.00 & 0.00 & 0.00 & 0.00 & 0.00 & 0.00 & 0.00 & PS & 1.00 & 1.00 & 1.00
& MIA & 0.80 & 0.60 & 0.76 & 0.69 & 0.67 & 0.70 & 0.99 & 0.97 & 0.99 & MIA & 0.96 & 0.68 & 0.97
& R-A & 0.57 & 0.46 & 0.56 & 0.56 & 0.58 & 0.57 & 0.71 & 0.72 & 0.71 & 0A & 0.60 & 0.73 & 0.70
& F-A & 0.04 & 0.08 & 0.04 & 0.05 & 0.03 & 0.03 & 0.00 & 0.00 & 0.00 & PS & 0.52 & 0.60 & 0.53
& MIA & 0.67 & 0.70 & 0.62 & 0.68 & 0.57 & 0.63 & 0.81 & 0.63 & 0.78 & MIA & 0.95 & 0.98 & 0.96
& R-A & 0.48 & 0.51 & 0.50 & 0.49 & 0.52 & 0.52 & 0.59 & 0.60 & 0.59 & 0A & 0.35 & 0.64 & 0.57
& F-A & 0.05 & 0.05 & 0.05 & 0.05 & 0.04 & 0.04 & 0.00 & 0.00 & 0.00 & PS & 0.50 & 0.63 & 0.52
& MIA & 0.54 & 0.35 & 0.59 & 0.56 & 0.64 & 0.52 & 0.81 & 0.78 & 0.90 & MIA & 0.96 & 0.91 & 0.76
& R-A & 0.89 & 0.88 & 0.88 & 0.89 & 0.89 & 0.89 & 0.89 & 0.93 & 0.93 & 0A & 0.92 & 0.93 & 0.93
& F-A & 0.03 & 0.05 & 0.04 & 0.04 & 0.03 & 0.05 & 0.00 & 0.00 & 0.00 & PS & 0.87 & 0.89 & 0.88
& MIA & 0.63 & 0.63 & 0.79 & 0.68 & 0.68 & 0.77 & 0.44 & 0.63 & 0.50 & MIA & 0.77 & 0.54 & 0.70
& R-A & 0.13 & 0.15 & 0.27 & 0.26 & 0.27 & 0.26 & 0.36 & 0.37 & 0.37 & 0A & 0.50 & 0.57 & 0.55
& F-A & 0.01 & 0.01 & 0.01 & 0.01 & 0.01 & 0.02 & 0.00 & 0.00 & 0.00 & PS & 0.55 & 0.56 & 0.54
& MIA & 0.32 & 0.49 & 0.43 & 0.40 & 0.26 & 0.69 & 0.91 & 0.38 & 0.67 & MIA & 0.98 & 0.98 & 0.99
Evaluation Metrics. As shown in Table 2, we evaluate our proposed method and the baselines using multiple metrics.
Main Results. In the client unlearning setting, Byzantine attacks are introduced, where label flipping is applied to construct inverse feature prototypes [45], [46]. Class unlearning randomly remaps the feature labels of the target class to several classes. Sample unlearning is implemented via backdoor attacks: fixed triggers are injected into the bottom-right pixels of images or appended as specific trigger tokens to the end of text sequences. During the unlearning process, all triggered samples are consistently predicted as class 0, thereby yielding accuracy of class 0 (0A) on the forgotten samples.
As shown in Table ¿tbl:tab:results?, our method achieves superior performance across a wide range of datasets, model architectures, and unlearning scenarios. Among them, CIFAR-100 is the most challenging benchmark, while MNIST is the least. Benefiting from the pluggable filters, FedUP reduces the accuracy on the unlearning knowledge to a minimal level (F-A) while maintaining high accuracy on the retained data (R-A). Owing to its centroid-based unlearning mechanism, FedUP exhibits particularly strong performance in class unlearning scenarios. Overall, the proposed method supports single-round communication and reversible unlearning, achieving the desirable triad of high retention accuracy, low forgetting accuracy, and privacy guarantees across all evaluated datasets.
Analysis of Non-target Knowledge Loss. To validate the effectiveness of our method in mitigating non-target knowledge loss, we compare the accuracy on retained knowledge before and after executing the unlearning operation. As shown in Figure 5, FedUP achieves an effect comparable to the Retrain method, characterized by minimal non-target knowledge loss and stable model performance before and after the unlearning operation.
Unlearning Response Time. As shown in Table 3, when achieving the desired forgetting effect, our method requires only 6.33 seconds for image unlearning and 6.45 seconds for text unlearning, which is less than 10% of the latency of the second-best method and less than 1% of that of the Retrain approach. It can promptly respond to diverse unlearning requests across different modalities while maintaining consistently low latency, effectively minimizing processing delays.
Communication Cost. Communication cost is defined as the wireless resources consumed for transmitting model parameters or gradients between local clients and the central server. As shown in Table 3, although the centroid uploading phase introduces additional communication overhead, FedUP requires only a single communication round, meaning its communication cost does not accumulate with rounds. In contrast, model transmission methods gradually converge as rounds increase, causing the communication overhead to escalate to the order of \(10^3\). Specifically, the lightweight filter incurs an overhead of merely 11.50MB for image unlearning, approximately one-third of that of the second-best method and 6.73MB for text unlearning, which is second only to the FUSED method.
| CIFAR10 | AG News | |||
| Time | Comm | Time | Comm | |
| E-C | 1627.76 | 1110.98 | 567.56 | 268.64 |
| Federaser | 1586.59 | 2820.18 | 634.67 | 738.76 |
| E-F | 898.70 | 1452.82 | 629.83 | 369.38 |
| FUSED | 151.21 | 31.36 | 88.31 | 0.55 |
| Retrain | 935.54 | 1538.25 | 859.28 | 537.28 |
| FedUP | 6.33 | 11.50 | 6.45 | 6.73 |
0pt 0pt 1pt
Bottleneck Dimension. In this section, we analyse the dimension of the filter. The pluggable filter adopts a straightforward encoder-decoder architecture. Its input dimension is configured to match the output dimension of the feature extractor. Our empirical investigation focuses on determining the optimal bottleneck dimension within this encoder-decoder structure. As shown in Table 4, the remember accuracy consistently reaches its optimum across both image and text datasets when the bottleneck dimension is set to 32.
| 4 | 8 | 16 | 32 | 64 | ||||||
| R-A | M | R-A | M | R-A | M | R-A | M | R-A | M | |
| MNIST | 0.98 | 16 | 0.99 | 32 | 0.99 | 64 | 0.99 | 128 | 0.99 | 256 |
| CIFAR10 | 0.55 | 16 | 0.58 | 32 | 0.58 | 64 | 0.60 | 128 | 0.59 | 256 |
| CIFAR10-T | 0.43 | 16 | 0.47 | 32 | 0.46 | 64 | 0.51 | 128 | 0.46 | 256 |
| AG News | 0.89 | 4 | 0.90 | 8 | 0.90 | 16 | 0.91 | 32 | 0.90 | 64 |
| CIFAR100 | 0.06 | 16 | 0.18 | 32 | 0.24 | 64 | 0.25 | 128 | 0.23 | 256 |
0pt 0pt 1pt
Sensitivity Analysis of the Loss Function. The filter is trained with joint reconstruction and cross-entropy losses. We performed a sensitivity analysis on \(\alpha\) of \(\mathcal{L}_{\text{total}}\) to determine its optimal value. As shown in Figure 6, using either loss alone hinders convergence and degrades the accuracy of retained knowledge, whereas their combination maximizes the reduction of non-target knowledge loss. This is further illustrated in Table 5, which shows that increasing the weight of the cross-entropy loss lowers the recognition accuracy on retained data, while assigning a higher weight to the reconstruction loss deteriorates the unlearning effectiveness.
| \(\alpha\)-CE | 0.10 | 0.30 | 0.50 | 0.70 | 0.90 | |||||
| metric | R-A | F-A | R-A | F-A | R-A | F-A | R-A | F-A | R-A | F-A |
| MNIST | 0.99 | 0.67 | 1.00 | 0.00 | 0.99 | 0.00 | 0.99 | 0.00 | 0.99 | 0.00 |
| CIFAR10 | 0.70 | 0.00 | 0.68 | 0.00 | 0.71 | 0.00 | 0.68 | 0.00 | 0.64 | 0.00 |
| CIFAR10-T | 0.51 | 0.22 | 0.56 | 0.00 | 0.59 | 0.00 | 0.51 | 0.00 | 0.45 | 0.00 |
| AG News | 0.92 | 0.00 | 0.93 | 0.00 | 0.93 | 0.00 | 0.92 | 0.00 | 0.87 | 0.00 |
| CIFAR100 | 0.35 | 0.11 | 0.37 | 0.00 | 0.37 | 0.00 | 0.31 | 0.00 | 0.30 | 0.00 |
0pt 0pt 1pt
Analysis of Differential Privacy Effects. To ensure privacy preservation, Gaussian noise is injected into the generated class centroids. We conducted ablation studies to evaluate its impact, as shown in Table 6. The results demonstrate that an appropriately calibrated noise scale does not significantly impede the model’s performance on non-target knowledge. Instead, it enhances model robustness, achieving a synergistic improvement in both privacy protection and model utility.
| CIFAR10 | AG News | |||
| DP | w/o DP | DP | w/o DP | |
| R-A | 0.5814 | 0.5833 | 0.9132 | 0.9114 |
| F-A | 0.0438 | 0.0460 | 0.0490 | 0.0303 |
| MIA | 0.5239 | 0.6373 | 0.5654 | 0.6196 |
0pt 0pt 1pt
| R-A | F-A | |||
| Feature | Centroid | Feature | Centroid | |
| MNIST | 0.9900 | 0.9921 | 0.0001 | 0.0005 |
| CIFAR10 | 0.6011 | 0.6020 | 0.0498 | 0.0331 |
| CIFAR10-T | 0.5264 | 0.5671 | 0.0632 | 0.0490 |
| AG News | 0.9153 | 0.9118 | 0.0331 | 0.0482 |
| CIFAR100 | 0.2537 | 0.2508 | 0.0076 | 0.0172 |
0pt 0pt 1pt
Effectiveness of Class Centroid Samples. Class centroids are obtained by clustering original features and privately aggregated on the server side. To validate the effectiveness of class centroid samples, we additionally conduct experiments in which the filter is trained using only the original features. As illustrated in Table 7, the model trained on class centroid samples achieves comparable performance on the accuracy of retained and unlearning knowledge to that trained on the original features, demonstrating that class centroid samples are as effective as the original features.
Conclusion. In the field of FU, server-side methods face non-target knowledge loss, whereas client‑side methods incur high request latency. Both are limited by the irreversibility of the unlearning operation. To address these issues, this paper proposes a one-shot federated unlearning framework. By employing differentially private class centroid samples on the server, our approach achieves approximate unlearning that surpasses exact unlearning in effect, reducing non‑target knowledge loss and high resource overhead. Through fine-tuning a lightweight plug-in filter in a single round, the desired unlearning effect is achieved, significantly reducing the latency of unlearning responses. Removing the filter allows the model to revert to its pre-unlearning state, thus realizing the reversibility of unlearning. Extensive experiments across diverse datasets, scenarios, and models demonstrate that FedUP demonstrates excellent performance.
Discussion. While FedUP advances unlearning efficiency, reversibility, and responsiveness, two fundamental limitations still persist: Class centroid fidelity depends on federated feature extraction. Low-quality global models propagate bias into aggregated class centroids. It is expected that these limitations can be effectively addressed by utilizing more powerful pre-trained backbones for local feature extraction to ensure reliable class centroids.
Feihong Nan and Zhengyi Zhong contributed equally.
Corresponding Author↩︎