July 02, 2026
The rise of customized diffusion models has fueled a boom in personalized visual content creation, but it also introduces serious risks of malicious misuse, thereby posing threats to personal privacy. Image aesthetics are strongly correlated with human perception of image quality. Motivated by this observation, we address facial privacy protection from a novel aesthetic perspective by degrading the generation quality of maliciously customized models, thus reducing facial identity leakage. Specifically, we propose a Hierarchical Anti-Aesthetics (HAA) framework that exploits aesthetic cues at multiple perceptual levels. HAA consists of two key branches: (1) Global Anti-Aesthetics, which degrades overall aesthetics and generation quality by constructing a global anti-aesthetic reward mechanism and a corresponding loss; and (2) Local Anti-Aesthetics, which disrupts facial identity by using a local anti-aesthetic reward mechanism and loss to guide adversarial perturbations toward facial regions. By integrating both branches, HAA achieves anti-aesthetic degradation from a global to a local level during customized generation. Extensive experiments show that HAA outperforms existing methods in identity removal, providing an effective tool for protecting facial privacy.
Facial Privacy Protection, Diffusion Model, Text-to-Image Synthesis, Adversarial Attacks
Diffusion models (DMs) have achieved significant breakthroughs in the field of text-to-image (T2I) generation, markedly enhancing the quality and realism of image generation by executing diffusion processes in latent spaces [1]–[7]. These models diversify and enhance visual effects in image generation while ensuring a high degree of consistency between the generated images and textual descriptions. To meet personalized needs and improve fine-tuning efficiency, researchers have developed various DM fine-tuning methods, such as Textual Inversion [8], DreamBooth [9], Custom Diffusion [10], and SVDiff [11]. These fine-tuning methods offer stronger customization capabilities, allowing users to generate high-quality images of specific themes with few reference images.
Despite the immense potential of these technologies, they also harbor significant safety risks that cannot be ignored. Malicious users may exploit these models to generate forged images or deepfake content, infringing on personal privacy and intellectual property rights, and even creating fake news to mislead the public [12]–[16].
To address these threats, anti-customization methods are primarily based on adversarial attacks, such as Mist [17], ASPL [18], and CAAT [19]. These methods introduce adversarial noise to interfere with the customized fine-tuning process, preventing the malicious misuse of user images by customized diffusion models. However, as illustrated in Fig. 1, existing methods exhibit limitations in facial privacy protection and copyright preservation. These approaches primarily rely on a straightforward end-to-end paradigm that degrades overall image quality by maximizing the original training loss, yet they critically overlook the design of identity elimination for local facial regions. This oversight substantially undermines their effectiveness in removing recognizable facial features, leading to privacy leakage and the misuse of user portraits. Recent methods refine this pipeline by attacking cross-attention, timestep/frequency features, or global-local feature/attribute constraints [19]–[21]. In contrast, HAA introduces a preference-level view of protection: malicious customization is weakened by suppressing the human-preferred realism and facial details required for usable identity reconstruction. Under the same constrained perturbation interface, frozen reward models impose explicit anti-aesthetic targets at whole-image and face-local levels, providing a direct signal beyond model-internal feature disruption.
To delve into this issue, we make the following key observations and reflections from a novel aesthetic perspective: 1) There is a close connection between the aesthetic attributes of an image and human perception of image quality [22]. Specifically, the higher the conformity of an image to aesthetic standards, the higher its perceived quality; conversely, the greater the deviation from aesthetic standards, the lower the perceived quality. 2) Human aesthetic perception is not one-dimensional but hierarchical, encompassing both the perception of global harmony and the perception of local details. These two levels together form a comprehensive judgment of image aesthetics [23], [24]. 3) The findings in Fig. 2 demonstrate that aligning with human aesthetic preferences can effectively improve the quality of generated images and facial details [25], [26]. This naturally raises the question: If we adopt an anti-aesthetic alignment approach, can it achieve the goal of reducing the quality of generated images while removing facial identity?
Inspired by the above observations and reflections, we propose the Hierarchical Anti-Aesthetics (HAA) framework. This framework is designed to degrade image quality by hierarchically exploring aesthetic cues from global to local scales, thereby safeguarding users’ facial privacy and copyright. Specifically, HAA includes the following two key branches: 1) Global Anti-Aesthetics, which weakens the overall aesthetics of images generated by malicious fine-tuners and reduces overall generation quality by constructing a global anti-aesthetic reward mechanism and designing a global anti-aesthetic loss; and 2) Local Anti-Aesthetics, which guides adversarial noise to force customized diffusion models to perform local anti-facial aesthetic alignment by establishing a local anti-aesthetic reward mechanism and designing a local anti-aesthetic loss, thereby reducing their ability to reconstruct facial details. By seamlessly integrating these branches, we form a joint hierarchical anti-aesthetics framework that fully explores aesthetic cues from global to local, thereby achieving the goal of anti-aesthetics. Implementing anti-aesthetics substantially reduces the generation quality of customized DMs, thereby reducing facial identity leakage and strengthening the protection of personal privacy and copyright. Our main contributions are as follows:
From a novel aesthetic perspective, we propose a joint Hierarchical Anti-Aesthetics (HAA) framework that seamlessly integrates global and local anti-aesthetic branches. This integration effectively reduces aesthetic quality at both global and local levels, thereby achieving the goal of anti-aesthetics.
We introduce Global Anti-Aesthetics by constructing a global anti-aesthetic reward mechanism and designing a global anti-aesthetic loss, thereby reducing the overall generation quality of customized generative models.
We propose Local Anti-Aesthetics, construct a local anti-aesthetic reward mechanism, and design a local anti-aesthetic loss to guide malicious customized diffusion models to oppose alignment with human facial aesthetics.
Extensive experiments across multiple datasets and generative models show that our approach outperforms existing methods by a large margin, supporting the effectiveness and generalization capability of our method.
The remainder of this paper is organized as follows. Section II reviews related work on customized diffusion models and image cloaking-based privacy protection. Section III introduces the problem definition and DreamBooth preliminaries. Section IV details the proposed HAA framework, including Global and Local Anti-Aesthetics branches. Section V presents extensive experimental results and analysis. Finally, Section VI concludes the paper.
Recent years have witnessed the continued development of diffusion models [4]–[7], [27], [28], and these models have shown great potential in guiding the generation process through text input. Exemplar models, including DALL-E 2 [3], Stable Diffusion [2], and Stable Diffusion XL [29], have received widespread attention for their high sample quality. Using the latent diffusion method and incorporating CLIP-based [30] text encoders, these models employ a two-stage process: encoding the input image or text into a latent representation and denoising within this compact latent space to bridge the gap between textual descriptions and visual content, thereby significantly improving generation quality while maintaining high fidelity and consistency in text-to-image (T2I) generation tasks.
DreamBooth [9] further addresses the challenge of customized generation, thereby leading to significant improvements in generating personalized content from limited data while preserving diversity and flexibility. However, DreamBooth faces challenges with more complex customizations, particularly when personalization requires a diverse and extensive dataset. In contrast, Custom Diffusion [10] can first fine-tune each concept model individually and then merge them into one through constrained optimization, which enhances the ability to generate diverse outputs while maintaining coherence across different concepts. SVDiff [11] optimizes all singular values of the weight matrix and leverages spectral shifts to introduce a small number of trainable parameters into diffusion models, enabling them to effectively capture variations in specific themes. Textual Inversion [8] teaches a text model a new word using example images and trains its embedding to align with the corresponding visual representation by adding a new token to the vocabulary and optimizing the embedding using representative images.
Recent breakthroughs in generative models, represented by diffusion models, have achieved paradigm-shifting advancements. Building on this foundation, the development of customization technologies has greatly facilitated personal visual creation. However, the misuse of personalized custom generation has raised significant concerns about copyright protection, political stability, and personal privacy. To address these issues, image cloaking methods have been developed, which involve adding specific perturbations to the original images to prevent misuse by customized generative models.
AdvDM [31] performs Monte Carlo sampling on latent variables in the diffusion model’s hidden space and generates perturbations specifically targeting and misleading the model’s feature extraction process during the sampling time steps, ultimately leading the model to produce erroneous outputs or reducing its overall performance. ASPL [18] employs an Alternating Surrogate and Perturbation Learning (ASPL) strategy to enhance counterattacks against the DreamBooth method. Mist [17] is specifically designed for copyright protection of artwork and can effectively defend against malicious use. CAAT [19] effectively disrupts the text-to-image mapping by introducing subtle perturbations in the cross-attention layers of customized diffusion models, thereby protecting users’ portrait rights from infringement. SimAC [20] explores the limitations of the internal properties of diffusion models, investigates the relationship between time step selection and frequency domain perception as well as the roles of hierarchical features in the denoising process, and proposes an adaptive greedy time step search and a feature-based optimization framework, which enhances the anti-interference effect. GoodAC [21] further introduces a global-local anti-customization strategy by disrupting global perceptual feature correlations and local facial attributes.
Existing methods mainly enhance privacy protection via model-internal objectives, including reconstruction loss, cross-attention disruption, timestep/frequency manipulation, and feature- or attribute-level discrepancy. These objectives may incidentally degrade visual quality, but they do not explicitly introduce human-preference reward feedback or decompose aesthetic guidance into global image-level and local face-level cues. This underexplored use of aesthetic feedback may limit facial de-identification. To complement this line of research, HAA introduces a hierarchical anti-aesthetics perspective that explores aesthetic cues from global appearance to local facial regions, thereby improving the suppression of facial identity information and strengthening facial privacy protection.
We study facial privacy protection against personalized diffusion models. Under a black-box assumption, an attacker can collect a small number of the user’s publicly available face images and personalize (fine-tune) a diffusion model to generate forged portraits resembling the victim, enabling deepfakes and privacy abuse. Our goal is to add a magnitude-bounded, nearly imperceptible perturbation \(\delta\) to each image to be shared (i.e., \(x_{adv}=\mathrm{clip}(x+\delta)\) with \(\|\delta\|_\infty\le\eta\)), such that any model fine-tuned on these protected images produces degraded outputs (i.e., worse generation quality implies better protection) with reduced identity consistency, thereby reducing the utility of forgeries and protecting privacy.
DreamBooth is a classic fine-tuning technique for text-to-image diffusion models that enables personalized generation. Specifically, given a small number of subject images (3–5), it fine-tunes a pre-trained diffusion model to generate new images of that subject in various contexts. DreamBooth binds the target subject to a rare token identifier (e.g., sks), and the input prompt follows the format of “a sks [class noun]”, where [class noun] represents the category (e.g., person). DreamBooth combines two objective terms: a personalization reconstruction loss and a prior preservation loss. The personalization reconstruction loss encourages the model to reconstruct the subject identity from the provided reference images, while the prior preservation loss mitigates overfitting and language drift in the few-shot setting. The overall objective is formulated as: \[\begin{align} \mathcal{L}_{co}(\theta) &=\mathbb{E}_{v_0,\,c,\,t,\,\epsilon}\!\left[\left\|\epsilon-\epsilon_{\theta}\!\left(v_t,\,t,\,c\right)\right\|_2^2\right] \\ &\quad + \gamma\cdot \mathbb{E}_{v_0',\,c_{\text{pr}},\,t',\,\epsilon'}\!\left[\left\|\epsilon'-\epsilon_{\theta}\!\left(v'_{t'},\,t',\,c_{\text{pr}}\right)\right\|_2^2\right], \label{eq:train} \end{align}\tag{1}\] where \(\epsilon_{\theta}(\cdot)\) is the noise prediction network parameterized by \(\theta\) (e.g., the U-Net in latent diffusion models). \(v_0\) and \(v_0'\) denote the clean latents of a subject image and a class (prior) image, respectively. \(c\) is the subject prompt containing the rare identifier token, and \(c_{\text{pr}}\) is the class prior prompt. \(t\) and \(t'\) are diffusion timesteps sampled from a predefined schedule, and \(\epsilon,\epsilon' \sim \mathcal{N}(0,I)\) are i.i.d.Gaussian noise variables. The noisy latents are constructed as: \[v_t = \alpha_t v_0 + \sigma_t \epsilon,\quad v'_{t'} = \alpha_{t'} v_0' + \sigma_{t'} \epsilon',\] with \((\alpha_t,\sigma_t)\) determined by the noise scheduler. \(\gamma\) balances the prior preservation term. \(\|\cdot\|_2\) denotes the Euclidean norm (and \(\|\cdot\|_2^2\) its squared form).
We start from a novel perspective—the relationship between aesthetics and quality—and introduce a novel Global Anti-Aesthetics Algorithm (GAA), which focuses on mining global aesthetic cues to degrade the overall quality of images. Specifically, we construct a global anti-aesthetic reward mechanism and design a global anti-aesthetic loss function in conjunction with a reconstruction loss to train adversarial noise.
Global anti-aesthetic reward mechanism. Given a sample \(x\), we perform diffusion steps on it using a pre-trained surrogate DM. We then employ a Monte Carlo sampling strategy to sample the diffusion timestep (and the corresponding noise) and conduct a single-step prediction (i.e., one forward pass of the denoiser followed by VAE decoding) for global anti-aesthetic alignment. This process can be formalized as follows: \[x_t' = G_{\theta} (N(E(x^{adv}_t),B(T)), C), \label{eq:sample}\tag{2}\] where \(x^{adv}_t\) denotes the adversarial sample at the current optimization step (we use \(t\) to match the notation in Eq. 2 ), \(B(\cdot)\) denotes a Monte Carlo sampler, and \(T\) stands for the set of diffusion timesteps. Concretely, \(B(T)\) samples one timestep \(t\in T\) (and the associated Gaussian noise used by the scheduler) for each image in the mini-batch, yielding a noisy latent through the noise scheduler \(N(\cdot)\) after VAE encoding \(E(\cdot)\). Equivalently, for image \(i\), \(v_i^{adv}=E(x_i^{adv})\) and \(v_{t_i}^{adv}=\alpha_{t_i}v_i^{adv}+\sigma_{t_i}\epsilon_i\) before \(G_\theta(\cdot)\). Then, \(G_{\theta}(\cdot)\) performs a direct one-step prediction and decodes the predicted latent back to the image domain. This pipeline is efficient and end-to-end differentiable with respect to \(x^{adv}\) because it only consists of differentiable components (VAE encoder/decoder, noise scheduler, and the denoiser forward pass), without running the full reverse diffusion chain (see Eq. 15 in [4]).
To implement GAA, we construct a global anti-aesthetic reward model to obtain global anti-aesthetic rewards. For this purpose, we utilize the publicly available human aesthetic preference dataset VisionRewardDB-Image [32] to train a Global Aesthetic Reward Model \(RM_g\) with BLIP as the backbone. Then, \(RM_g\) is frozen. Through \(RM_g\), we can conduct a global aesthetic evaluation of the generated images to obtain global aesthetic scores. Subsequently, we construct global anti-aesthetic rewards \(R_{i}^{g}\) based on these global aesthetic scores. The formulas can be expressed as follows: \[r_i^g = - RM_g(x_{ti}', C_g), \label{eq:r}\tag{3}\] \[R_{i}^{g}=\frac{r_{i}^{g}-\frac{1}{n}\sum_{i = 1}^{n}r_{i}^{g}}{\sqrt{\frac{1}{n}\sum_{i = 1}^{n}(r_{i}^{g}-\frac{1}{n}\sum_{i = 1}^{n}r_{i}^{g})^{2}}}, \label{eq:rig}\tag{4}\] Let \(x_{ti}'\) denote the \(i\)-th generated sample in the batch at the \(t\)-th iteration, where \(n\) denotes the batch size. \(C_g\) represents the global prompt. Eq. 3 is used to calculate the global anti-aesthetic reward, aiming to measure deviations in aesthetics. This mechanism guides the model to perform “reverse optimization” for global aesthetics. Eq. 4 serves to normalize the global anti-aesthetic reward. Through normalization, we adjust the reward value to an appropriate range and avoid excessively large or small numerical values. The larger the value of \(R_{i}^{g}\), the higher the global anti-aesthetic score and the poorer the overall quality of the generated image.
Global anti-aesthetic loss. In our approach, we utilize the global anti-aesthetic reward \(R_{i}^{g}\) obtained from Eq. 4 to calculate the global anti-aesthetic loss \(\mathcal{L}_{rg}\). This loss is then employed to update the adversarial noise, thereby undermining the overall generation quality of malicious fine-tuners. The specific formula is as follows: \[L_{rg}=-\frac{1}{n}\sum_{i=1}^{n} \log\!\Big(1+\exp\!\big(-R^{g}_{i}\big)\Big), \label{eq:6}\tag{5}\] \[x_{adv}=Proj(\arg\max_{x_{adv}}(\mathcal{L}_{co}+\lambda \cdot \mathcal{L}_{rg})),\] where \(\exp(\cdot)\) denotes the exponential function, \(\mathcal{L}_{co}\) is the reconstruction loss defined in Eq. 1 , and \(\lambda\) is a balancing coefficient. \(Proj(\cdot)\) ensures that the adversarial noise remains within a prescribed \(\ell_{\infty}\) budget. Specifically, the projection is applied to the perturbation \(\delta\), while clipping is only used to map \(x+\delta\) back to the valid image range. Inspired by reinforcement learning, where the optimization process is differentiable [33], we compute \(\mathcal{L}_{rg}\) based on the rewards \(R_i^{g}\) provided by \(RM_g\). Specifically, Eq. 5 applies a softplus function to smoothly map the global anti-aesthetic rewards \(R_i^{g}\) into an averaged, differentiable objective. Maximizing this objective encourages larger \(R_i^{g}\), thereby promoting misalignment with human aesthetic preferences and guiding adversarial perturbations to degrade the generation quality of the T2I model. In practice, gradients are back-propagated from \(RM_g(\cdot)\) to \(x^{adv}\) through \(x_t'\) in Eq. 2 , enabling direct optimization of the perturbation under the \(\ell_\infty\) constraint.
Inspired by adversarial learning, we adopt an adversarial game strategy. On one hand, we train a surrogate diffusion model (SDM) with \(x_{adv}\) generated by GAA to enhance robustness. On the other hand, we attack the robust SDM to generate stronger attacks. The training process of the SDM can be represented as follows: \[\theta'= \arg\min_{\theta'} \mathcal{L}_{co}(x_{adv}, \theta'), \label{eq:8}\tag{6}\] where \(\theta'\) denotes the parameters of the SDM. Through this adversarial mechanism, adversarial samples with stronger attack effects can be generated, thereby improving the overall attack capability. This strategy is also employed in the subsequent methods.
Human perception of aesthetics is multidimensional, encompassing not only global aesthetic perception but also local aesthetic perception. Inspired by this, we propose a novel Local Anti-Aesthetics Algorithm (LAA), where we construct a local anti-aesthetic reward mechanism and design a local anti-aesthetic loss. This approach forces maliciously fine-tuned DMs to perform anti-facial aesthetic alignment, thereby enhancing the ability to eliminate facial identity cues.
Local anti-aesthetic reward mechanism. Specifically, a lightweight face detection model \(F\) is applied to the \(x'_{t}\) defined in Eq. 2 to obtain face bounding boxes \(Box_t\) and confidence scores \(S\). Similar to the global aesthetic reward model, to mine local facial aesthetic cues, we use the same BLIP architecture to construct a low-cost local facial aesthetic reward model \(RM_l\). Using the detected boxes and the generated images, we extract the corresponding facial regions and use \(RM_l\) to provide local aesthetic feedback. We then utilize the feedback to calculate the local anti-aesthetic rewards \(R^l_i\) for updating the adversarial noise. The lower the quality of the generated face, the higher the \(R^l_i\). Let \(C_l\) represent the local prompt; this process can be formalized as follows: \[S_{ti}, Box_{ti} = F(x_{ti}'), \label{eq:face}\tag{7}\] \[r_i^l = - RM_l(x_{ti}',Box_{ti}, C_l), \label{eq:10}\tag{8}\] \[R_{i}^{l}=\frac{r_{i}^{l}-\frac{1}{n}\sum_{i = 1}^{n}r_{i}^{l}}{\sqrt{\frac{1}{n}\sum_{i = 1}^{n}(r_{i}^{l}-\frac{1}{n}\sum_{i = 1}^{n}r_{i}^{l})^{2}}}. \label{eq:11}\tag{9}\]
Local anti-aesthetic loss. Based on \(R^l_i\), we further calculate the reward loss \(\mathcal{L}_{rl}\) to update the adversarial noise. This process can be formalized as follows: \[L_{rl}=-\frac{1}{n}\sum_{i=1}^{n} \log\!\Big(1+\exp\!\big(-R^{l}_{i}\big)\Big), \label{eq:12}\tag{10}\] \[x_{adv}=Proj(\arg\max_{x_{adv}}(\mathcal{L}_{co}+\beta \cdot \mathcal{L}_{rl})), \label{eq:13}\tag{11}\] where \(\beta\) is a balancing coefficient. Similarly, the parameters of the SDM are updated during the iterative attack. By attacking a more robust surrogate diffusion model, the attack strength is enhanced, thereby achieving the local anti-aesthetic objective.
| CelebA-HQ | ||||
|---|---|---|---|---|
| Method | “a photo of sks person” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 1.000 | 0.498 | 0.599 | 112.3 |
| Mist | 0.969 | 0.375 | 0.191 | 281.2 |
| CAAT | 0.813 | 0.339 | 0.326 | 271.7 |
| ASPL | 0.750 | 0.332 | 0.006 | 359.5 |
| SimAC | 0.438 | 0.321 | -0.188 | 388.9 |
| HAA | 0.281 | 0.117 | -0.894 | 471.1 |
| Method | “a dslr portrait of sks person” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.859 | 0.346 | 0.706 | 175.1 |
| Mist | 0.859 | 0.308 | 0.249 | 271.0 |
| CAAT | 0.766 | 0.260 | 0.236 | 251.8 |
| ASPL | 0.656 | 0.251 | -0.183 | 384.0 |
| SimAC | 0.500 | 0.239 | -0.202 | 369.0 |
| HAA | 0.266 | 0.115 | -0.861 | 461.4 |
| Method | “a close-up photo of sks person, high details” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.625 | 0.276 | 0.281 | 200.8 |
| Mist | 0.594 | 0.268 | -0.128 | 268.4 |
| CAAT | 0.510 | 0.177 | -0.419 | 305.9 |
| ASPL | 0.438 | 0.167 | -0.882 | 433.3 |
| SimAC | 0.333 | 0.162 | -0.779 | 401.7 |
| HAA | 0.177 | 0.077 | -1.334 | 476.0 |
| Method | “a photo of sks person looking at the mirror” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.672 | 0.276 | 0.022 | 232.9 |
| Mist | 0.695 | 0.257 | -0.289 | 287.0 |
| CAAT | 0.602 | 0.171 | -0.545 | 316.1 |
| ASPL | 0.523 | 0.157 | -0.854 | 422.4 |
| SimAC | 0.562 | 0.156 | -0.773 | 405.6 |
| HAA | 0.289 | 0.085 | -1.296 | 473.5 |
| VGGFace2 | ||||
|---|---|---|---|---|
| Method | “a photo of sks person” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.875 | 0.421 | 0.625 | 199.1 |
| Mist | 0.906 | 0.227 | 0.248 | 355.6 |
| CAAT | 0.719 | 0.145 | 0.371 | 349.8 |
| ASPL | 0.750 | 0.193 | 0.373 | 388.6 |
| SimAC | 0.500 | 0.118 | 0.143 | 435.6 |
| HAA | 0.125 | 0.031 | -0.252 | 468.6 |
| Method | “a dslr portrait of sks person” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.813 | 0.379 | 0.717 | 227.6 |
| Mist | 0.813 | 0.240 | 0.409 | 357.2 |
| CAAT | 0.750 | 0.179 | 0.559 | 332.8 |
| ASPL | 0.703 | 0.173 | 0.351 | 387.9 |
| SimAC | 0.422 | 0.102 | -0.018 | 429.3 |
| HAA | 0.109 | 0.028 | -0.608 | 462.4 |
| Method | “a close-up photo of sks person, high details” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.688 | 0.376 | 0.720 | 232.3 |
| Mist | 0.573 | 0.215 | 0.279 | 343.2 |
| CAAT | 0.521 | 0.155 | 0.353 | 341.1 |
| ASPL | 0.490 | 0.150 | 0.039 | 381.0 |
| SimAC | 0.292 | 0.078 | -0.581 | 434.5 |
| HAA | 0.083 | 0.019 | -1.162 | 458.4 |
| Method | “a photo of sks person looking at the mirror” | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.750 | 0.355 | 0.482 | 246.0 |
| Mist | 0.641 | 0.188 | 0.005 | 354.9 |
| CAAT | 0.633 | 0.139 | 0.145 | 334.0 |
| ASPL | 0.602 | 0.153 | -0.055 | 380.6 |
| SimAC | 0.414 | 0.092 | -0.616 | 429.0 |
| HAA | 0.258 | 0.031 | -1.119 | 446.0 |
Building upon the complementary strengths of GAA and LAA, we propose a unified HAA framework (Fig. 3) that integrates global and local aesthetic degradation within a single optimization objective. This design disrupts image generation quality at multiple perceptual levels, achieving anti-aesthetic effects in both global composition and local facial details. Specifically, the hierarchy in HAA is realized through the reward-model inputs and the joint objective, rather than a separate sampling schedule: GAA evaluates the full decoded image with the global prompt \(C_g\), while LAA evaluates detected facial regions with the local prompt \(C_l\). Compared with methods relying solely on \(\mathcal{L}_{co}\) or feature-level discrepancy objectives, HAA further incorporates two frozen reward-guided terms, \(\mathcal{L}_{rg}\) and \(\mathcal{L}_{rl}\), while using the same timestep sampler to maintain efficient differentiable feedback. Accordingly, the final training objective is defined as:
HAA uses the standard projected adversarial-perturbation loop as the basis for controlled comparison. The technical change is the optimization signal: HAA injects frozen human-preference feedback into both whole-image and face-local perturbation updates. These terms are not generic auxiliary losses; they define external preference targets that penalize preference-aligned realism and facial detail quality. Thus, the contribution lies in reward-guided hierarchical anti-aesthetic supervision, rather than in a new outer projected-gradient optimizer.
\[\mathcal{L}_{total} = \mathcal{L}_{co}+\lambda \cdot \mathcal{L}_{rg}+\beta \cdot \mathcal{L}_{rl}. \label{eq:total95loss}\tag{12}\] \[x_{adv}=Proj(\arg\max_{x_{adv}}(\mathcal{L}_{total})). \label{eq:15}\tag{13}\]
In each iteration, the overall loss is maximized to optimize the adversarial samples, enabling the adversarial noise to poison malicious fine-tuners and disrupt the global and local aesthetics of the generated images, thereby degrading overall generation quality. While the adversarial samples are iteratively optimized, the parameters of the SDM are fine-tuned in an adversarial training-like manner. Ultimately, through iterative attacks, the final adversarial samples can effectively remove facial identity information, thereby protecting facial privacy and copyright.
Datasets. Following previous work [19], we conduct experiments on the CelebA-HQ [34] and VGGFace2 [35] datasets. Both are large, classic datasets specifically for facial privacy research. Due to the wide range of comparison metrics, we select 10 individuals from each dataset, covering different genders, ages, and ethnicities, with at least 15 images for each individual. Unless explicitly stated, we conduct experimental tests uniformly on CelebA-HQ.
Models. Our experiments primarily utilize the popular open-source diffusion model SD-v2.1 [36]. Furthermore, to extensively validate the effectiveness of our approach, we also conduct experiments on SD-v1.4 and SD-v1.5. In addition, extended experiments on mainstream SD-V3.0 [37] and FLUX [38] are also conducted in the subsequent experiments.
Baselines and comparisons. We compare our method with various state-of-the-art methods designed to prevent the misuse of user images by DMs, including ASPL [18], Mist [17], CAAT [19], SimAC [20], and GoodAC [21].
Metrics. Following previous work [19], we employ a classic face detection model to detect the generated images and calculate Facial Detection Success Rate (FDSR) [39]. For identity consistency, we calculate Face Similarity [40] between the detected faces and clean images. FDSR is the proportion of generated images with at least one detected face. Face Similarity is the average cosine similarity between \(\ell_2\)-normalized ArcFace embeddings; if no face is detected, the per-image similarity is set to 0. Moreover, we use Image Reward [32] to assess the quality of the generated images, and report its score as IR. Finally, Fréchet Inception Distance (FID) [41] is used to assess the similarity between the generated images and clean images. Lower FDSR and Face Similarity indicate weaker face detectability and identity consistency, while lower Image Reward and higher FID indicate poorer preference-aligned generation quality; together, these metrics indicate stronger privacy protection.
Implementation Details. To ensure a fair comparison, we adopt a unified experimental setup. We use the latest Stable Diffusion (v2.1) as the pre-trained model and fine-tune the text encoder and UNet models using the DreamBooth method for 1000 training steps with a batch size of 4 and a learning rate of \(5 \times 10^{-7}\). We set \(\alpha=5\times 10^{-3}\) as the step size, along with a perturbation budget \(\eta=0.05\). Moreover, in the LAA module, we employ RetinaFace [42] for face detection because of its high efficiency and accuracy. For computing Face Similarity, we use ArcFace [40] for its high accuracy and robustness. All experiments in the table are run three times, and the results are averaged to ensure stability.
Our evaluation is designed to test both effectiveness and generalization. In addition to the main CelebA-HQ and VGGFace2 comparisons under four prompts, we include component ablations, prompt mismatch, varying numbers of protected images, perturbation-budget studies, image-transformation and purification defenses, transfer across Stable Diffusion versions, transfer to SD-V3.0 and FLUX, and three fine-tuning pipelines (SVDiff, Textual Inversion, and Custom Diffusion).
To more thoroughly assess effectiveness, we perform quantitative comparisons on two datasets (CelebA-HQ and VGGFace2) under four prompt settings, three of which remain unseen during training, and report results against recent state-of-the-art methods in Tab. 1 and Tab. 2. For each prompt, we randomly generate 32 images and report the mean of four metrics. On CelebA-HQ with the prompt “a photo of sks person”, HAA reduces FDSR from 96.9% to 28.1% and Face Similarity from 37.5% to 11.7%, lowers Image Reward from 0.191 to -0.894, and increases FID from 281.2 to 471.1 relative to Mist. On VGGFace2 with the prompt “a close-up photo of sks person, high details”, HAA reduces FDSR from 52.1% to 8.3% and Face Similarity from 15.5% to 1.9%, lowers Image Reward from 0.353 to -1.162, and increases FID from 341.1 to 458.4 relative to CAAT.
| SD-V1.4 | ||||
|---|---|---|---|---|
| attack | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.938 | 0.439 | 0.440 | 148.4 |
| Mist | 0.906 | 0.332 | -0.027 | 237.8 |
| CAAT | 0.125 | 0.035 | -1.083 | 508.8 |
| ASPL | 0.375 | 0.076 | -0.728 | 458.4 |
| SimAC | 0.186 | 0.067 | -1.079 | 508.1 |
| HAA | 0.062 | 0.030 | -1.577 | 540.4 |
| SD-V1.5 | ||||
| attack | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.969 | 0.454 | 0.577 | 140.4 |
| Mist | 1.000 | 0.383 | -0.165 | 225.9 |
| CAAT | 0.312 | 0.036 | -1.335 | 466.5 |
| ASPL | 0.250 | 0.059 | -0.904 | 479.4 |
| SimAC | 0.219 | 0.149 | -1.230 | 452.8 |
| HAA | 0.000 | 0.003 | -1.563 | 523.7 |
| Different fine-tuning scenarios | ||||
|---|---|---|---|---|
| Method | SVDiff | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.914 | 0.335 | 0.538 | 149.2 |
| Mist | 0.875 | 0.202 | 0.274 | 212.9 |
| CAAT | 0.688 | 0.199 | 0.121 | 337.4 |
| ASPL | 0.516 | 0.168 | 0.103 | 388.3 |
| SimAC | 0.789 | 0.205 | 0.095 | 289.6 |
| HAA | 0.492 | 0.133 | 0.002 | 401.2 |
| Method | Textual Inversion | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.969 | 0.258 | 0.914 | 140.8 |
| Mist | 1.000 | 0.212 | 0.432 | 185.7 |
| CAAT | 0.969 | 0.172 | 0.543 | 191.5 |
| ASPL | 1.000 | 0.178 | 0.535 | 152.9 |
| SimAC | 1.000 | 0.234 | 0.415 | 148.1 |
| HAA | 0.938 | 0.194 | 0.030 | 213.5 |
| Method | Custom Diffusion | |||
| 2-5 | FDSR↓ | Face Similarity↓ | Image Reward↓ | FID↑ |
| Clean | 0.938 | 0.437 | 0.567 | 120.0 |
| Mist | 0.906 | 0.330 | 0.173 | 194.8 |
| CAAT | 0.875 | 0.308 | 0.177 | 271.3 |
| ASPL | 0.938 | 0.319 | 0.315 | 271.2 |
| SimAC | 0.969 | 0.364 | 0.480 | 202.0 |
| HAA | 0.781 | 0.254 | 0.152 | 275.4 |
Across datasets and prompt variations, these results suggest that HAA more effectively suppresses identity-related cues in the generated outputs, as indicated by the concurrent decreases in FDSR and Face Similarity. We also observe that stronger identity removal coincides with degraded perceptual quality (higher FID) and lower preference-aligned scores (lower Image Reward), which is consistent with the goal of discouraging high-quality personalized reconstruction under malicious fine-tuning. We attribute this behavior to the hierarchical anti-aesthetics design: the global branch broadly disrupts aesthetics that correlate with overall realism, while the local branch further targets facial regions where identity information concentrates, and their combination yields a more consistent de-identification effect under prompts not seen during training.
| “a photo of [v] person” | |||||
|---|---|---|---|---|---|
| Train [v] | Test [v] | FDSR \(\downarrow\) | FS \(\downarrow\) | IR \(\downarrow\) | FID \(\uparrow\) |
| sks | sks | 0.281 | 0.117 | -0.894 | 471.1 |
| sks | t@t | 0.438 | 0.316 | 0.017 | 344.6 |
| “a dslr portrait of [v] person” | |||||
| Train [v] | Test [v] | FDSR \(\downarrow\) | FS \(\downarrow\) | IR \(\downarrow\) | FID \(\uparrow\) |
| sks | sks | 0.266 | 0.115 | -0.861 | 461.4 |
| sks | t@t | 0.578 | 0.221 | 0.197 | 276.5 |
| “a close-up photo of [v] person, high details” | |||||
| Train [v] | Test [v] | FDSR \(\downarrow\) | FS \(\downarrow\) | IR \(\downarrow\) | FID \(\uparrow\) |
| sks | sks | 0.177 | 0.077 | -1.334 | 476.0 |
| sks | t@t | 0.385 | 0.148 | -0.610 | 325.1 |
| “a photo of [v] person looking at the mirror” | |||||
| Train [v] | Test [v] | FDSR \(\downarrow\) | FS \(\downarrow\) | IR \(\downarrow\) | FID \(\uparrow\) |
| sks | sks | 0.289 | 0.085 | -1.296 | 473.5 |
| sks | st@t | 0.523 | 0.138 | -0.678 | 322.0 |
In our formulation, Eq. 12 introduces two hyperparameters, \(\lambda\) and \(\beta\), which weight the global and local anti-aesthetic losses, respectively. We perform parameter tuning on CelebA-HQ using SD-2.1 and examine how these weights affect FID.
As shown in Fig. 5, \(\lambda\) exerts a pronounced influence on FID. When \(\lambda\) increases from small values, FID rises accordingly, which is consistent with the global anti-aesthetic objective increasingly dominating optimization and thereby pushing generated samples farther from the clean distribution. However, beyond \(\lambda=0.2\), FID begins to decrease, suggesting diminishing returns—and potentially a change in the optimization regime in which overly strong global degradation does not further widen the distributional gap captured by FID. Based on this empirical trend, we set \(\lambda=0.2\) in the remaining experiments. We observe that \(\beta\) also materially affects overall performance. Following the same tuning protocol, we set \(\beta=0.2\), which provides a favorable operating point under our evaluation setting by balancing the contribution of local facial degradation against the other objectives in Eq. 12 .
| “a photo of sks person” | |||||
|---|---|---|---|---|---|
| Perturbed | Clean | FDSR↓ | FS↓ | IR↓ | FID↑ |
| 4 | 0 | 0.281 | 0.117 | -0.894 | 471.1 |
| 3 | 1 | 0.625 | 0.290 | -0.164 | 338.8 |
| 2 | 2 | 0.688 | 0.390 | 0.210 | 267.3 |
| 1 | 3 | 0.719 | 0.426 | 0.273 | 213.6 |
| 0 | 4 | 1.000 | 0.498 | 0.599 | 112.3 |
| “a dslr portrait of sks person” | |||||
| Perturbed | Clean | FDSR↓ | FS↓ | IR↓ | FID↑ |
| 4 | 0 | 0.266 | 0.115 | -0.861 | 461.4 |
| 3 | 1 | 0.641 | 0.244 | 0.055 | 320.1 |
| 2 | 2 | 0.703 | 0.310 | 0.262 | 269.4 |
| 1 | 3 | 0.688 | 0.330 | 0.230 | 215.1 |
| 0 | 4 | 0.859 | 0.346 | 0.706 | 175.1 |
| “a close-up photo of sks person, high details” | |||||
| Perturbed | Clean | FDSR↓ | FS↓ | IR↓ | FID↑ |
| 4 | 0 | 0.177 | 0.077 | -1.334 | 476.0 |
| 3 | 1 | 0.427 | 0.164 | -0.663 | 373.5 |
| 2 | 2 | 0.479 | 0.223 | -0.357 | 319.0 |
| 1 | 3 | 0.458 | 0.262 | -0.138 | 238.8 |
| 0 | 4 | 0.625 | 0.276 | 0.281 | 200.8 |
| “a photo of sks person looking at the mirror” | |||||
| Perturbed | Clean | FDSR↓ | FS↓ | IR↓ | FID↑ |
| 4 | 0 | 0.289 | 0.085 | -1.296 | 473.5 |
| 3 | 1 | 0.547 | 0.164 | -0.668 | 378.0 |
| 2 | 2 | 0.602 | 0.227 | -0.404 | 320.7 |
| 1 | 3 | 0.594 | 0.274 | -0.256 | 265.9 |
| 0 | 4 | 0.672 | 0.276 | 0.022 | 232.9 |
Effectiveness Analysis of HAA Components. We conduct systematic ablation studies to examine the individual roles of Global Anti-Aesthetics (GAA) and Local Anti-Aesthetics (LAA), as well as their combined effect. As reported in Fig. 6, relative to the Baseline (FDSR \(=0.719\)), adding GAA reduces FDSR to \(0.562\) (a \(21.8\%\) relative reduction) and decreases Face Similarity to \(0.261\) (a \(32.7\%\) relative reduction). This behavior is consistent with GAA primarily acting at the image-level distribution: by discouraging globally preferred aesthetic configurations, it tends to suppress holistic realism cues that also correlate with identity consistency.
When we instead incorporate the local anti-facial aesthetics loss (LAA) into the Baseline, LAA yields stronger privacy-oriented degradation than GAA in this setting: FDSR further drops to \(0.406\) (a \(27.8\%\) relative reduction), and Image Reward decreases from \(-0.339\) to \(-0.605\). These results suggest that explicitly targeting facial regions provides a more direct lever for weakening identity-relevant evidence, because many recognition pipelines depend on localized facial details that may persist even when global image quality degrades.
Finally, the full HAA model, which combines both global and local anti-aesthetic objectives, shows the most pronounced effect: FDSR decreases to \(0.281\) (a \(60.9\%\) relative reduction vs.the Baseline), Face Similarity decreases to \(0.117\) (a \(69.8\%\) relative reduction), Image Reward reaches \(-0.894\), and FID increases to \(471.1\) (a \(46.5\%\) relative increase). The stepwise trend across variants indicates that GAA and LAA are complementary rather than redundant: the global branch broadly shifts samples away from high-fidelity, preference-aligned generations, while the local branch concentrates the disruption on facial regions where identity information is densest. Their combination yields a more consistent degradation of both detection success and identity similarity, which aligns with the goal of reducing malicious fine-tuners’ ability to reconstruct recognizable faces under our experimental setting.
Although LAA directly targets detected faces, identity leakage in customized diffusion models is not restricted to the detected face crop. Hair, head contour, pose, clothing, lighting, and scene co-occurrence can still support subject-token binding and identity reconstruction. In addition, LAA depends on reliable face detection. GAA therefore provides a principled full-image anti-aesthetic objective for complementary image-level suppression and remains effective when local detection is unreliable.
We first mask all detected face regions and compute CLIP image-embedding cosine similarity between the masked generated images and clean references. Clean denotes unprotected training, and Base denotes reconstruction-loss-only perturbation without anti-aesthetic losses. As shown in Fig. 7, GAA reduces residual non-face similarity more than LAA (0.403 vs. 0.430), and HAA achieves the lowest value (0.381). This indicates that identity-supporting information also exists outside the face crop and that the global branch helps remove it. The quantitative ablation in Fig. 6 further shows that GAA alone improves over the baseline, LAA provides stronger face-local degradation, and HAA achieves the best overall protection. Fig. 8 visualizes the corresponding Grad-CAM response regions, with response strength following Clean \(<\) Base \(<\) GAA \(<\) LAA \(<\) HAA. LAA focuses on facial details, GAA covers broader image-level structures, and HAA yields the broadest response.
Finally, we evaluate a face-detection-failure stress setting. In Fig. 9, LAA (0%) corresponds to the case where no valid face box is available, so the local reward cannot provide effective face-local guidance. Adding GAA improves LAA at both 0% and 60% valid-local-detection availability. When 60% of local detections are available, adding GAA reduces FDSR from 0.531 to 0.375 and Face Similarity from 0.299 to 0.152, while increasing FID from 391.9 to 453.2. These results indicate that GAA provides complementary feedback when the local branch is weakened, rather than merely increasing perturbation strength.
To better approximate practical deployments where the target customized T2I model differs from the surrogate used for perturbation crafting, we evaluate HAA under black-box model-version mismatch by transferring attacks across different Stable Diffusion releases. As shown in Tab. 3, HAA maintains strong effectiveness on both SD-v1.4 and SD-v1.5. Concretely, on SD-v1.5, it further reduces FDSR to \(0.000\) and Face Similarity to \(0.003\), with Image Reward \(-1.563\) and FID \(523.7\). Compared to baselines, these results indicate that perturbations transfer across model versions and suppress both face detectability and identity consistency. We view this stability as evidence that HAA primarily disrupts version-invariant cues shared across model releases, rather than fragile version-specific artifacts. In particular, degrading global appearance statistics while targeting facial regions reduces downstream fine-tuning’s reliance on semantically meaningful identity features. This helps preserve protection strength when the attacker’s model differs from the one used during perturbation optimization.
| \(\eta\) | FDSR \(\downarrow\) | Face Similarity \(\downarrow\) | Image Reward \(\downarrow\) | FID \(\uparrow\) |
|---|---|---|---|---|
| 0.00 | 1.000 | 0.498 | 0.599 | 112.3 |
| 0.02 | 0.425 | 0.220 | -0.428 | 379.8 |
| 0.05 | 0.281 | 0.117 | -0.894 | 471.1 |
| 0.10 | 0.156 | 0.054 | -0.999 | 462.2 |
| 0.15 | 0.000 | 0.000 | -1.364 | 497.0 |
| Model | Attack | Face Similarity↓ | Image Reward↓ | FID↑ |
|---|---|---|---|---|
| SD-V3.0 | CAAT | 0.357 | 0.725 | 170.4 |
| SD-V3.0 | SimAC | 0.353 | 0.988 | 188.5 |
| SD-V3.0 | HAA | 0.343 | 0.561 | 197.0 |
| FLUX | CAAT | 0.399 | 0.630 | 123.1 |
| FLUX | SimAC | 0.397 | 0.453 | 117.7 |
| FLUX | HAA | 0.375 | 0.336 | 148.5 |
Tab. 4 evaluates HAA’s transferability across three customization pipelines (SVDiff, Textual Inversion, Custom Diffusion) on CelebA-HQ. Overall, HAA consistently shifts fine-tuned generators to weaker identity evidence (lower Face Similarity/FDSR) and reduced preference-aligned quality (lower Image Reward), supporting perturbation effectiveness beyond the DreamBooth-style surrogate used in optimization.
In the SVDiff setting, HAA outperforms other methods, achieving the best protection metrics: lowest FDSR (\(0.492\)), Face Similarity (\(0.133\)), Image Reward (\(0.002\)), and highest FID (\(401.2\)), indicating a stronger ability to disrupt face detection and identity consistency. Tab. 4 supports HAA’s generalization across heterogeneous fine-tuning strategies. Its impact is most pronounced in pipelines directly affecting facial detail generation (SVDiff, Custom Diffusion), while Textual Inversion is intrinsically harder to defend and CAAT obtains the lowest Face Similarity in that setting. Even so, HAA reliably degrades preference-aligned quality and increases distributional divergence, fulfilling the goal of reducing maliciously customized models’ utility for privacy-invasive face reconstruction.
| \(Method\) | FDSR \(\downarrow\) | Face Similarity \(\downarrow\) | Image Reward \(\downarrow\) | FID \(\uparrow\) |
|---|---|---|---|---|
| Clean | 1.000 | 0.498 | 0.599 | 112.3 |
| HAA | 0.425 | 0.220 | -0.428 | 379.8 |
| Gaussian blur k=3 | 0.644 | 0.302 | 0.315 | 215.1 |
| Gaussian blur k=5 | 0.875 | 0.326 | 0.377 | 215.6 |
| Gaussian blur k=7 | 0.906 | 0.349 | 0.410 | 209.1 |
| JPEG q=30 | 0.938 | 0.269 | -0.337 | 257.2 |
| JPEG q=50 | 0.869 | 0.262 | -0.288 | 263.2 |
| JPEG q=70 | 0.813 | 0.254 | -0.173 | 267.0 |
Tab. 5 evaluates robustness to prompt mismatch on CelebA-HQ. We optimize perturbations using a fixed training prompt (“a photo of sks person”) and then test at inference time with prompt templates that vary in style and, importantly, replace the identifier token with rare alternatives (e.g., “t@t” or “st@t”). This setting probes whether the protection hinges on a particular token string or instead transfers to semantically equivalent prompts. Across all prompts, HAA effectively reduces identity-related information. As shown in Tab. 5, when the test identifier matches training (\([v]=\)“sks”), HAA achieves strong protection, with low FDSR (e.g., \(0.281\), \(0.177\)), low Face Similarity (down to \(0.077\)), and negative Image Rewards (as low as \(-1.334\)). These results offer two complementary insights. First, HAA’s effect is partly token-agnostic, as the perturbations consistently weaken identity signals. Second, the performance gap between matched and mismatched identifiers reveals that fine-tuning relies on the specific token-to-identity link; altering the token weakens this association and, thus, the protection.
Tab. 6 evaluates HAA’s protection against the number of perturbed training images, assessing its effectiveness and sample efficiency. Across four prompts, 4 perturbed images consistently yield the strongest protection across all metrics. For example, with “a photo of sks person”, 4 images achieve FDSR=0.281, Face Similarity=0.117, Image Reward=-0.894, and FID=471.1. Notably, HAA remains effective with just 1 perturbed image: for “a dslr portrait of sks person”, 1 perturbed image reduces FDSR to 0.688 and Face Similarity to 0.330 vs.the clean baseline; for “a close-up photo of sks person, high details”, it lowers Image Reward to -0.138 and raises FID to 238.8, indicating limited poisoning degrades image quality and identity consistency. Moreover, protection strength scales predictably with perturbed image count. For “a photo of sks person looking at the mirror”, increasing perturbed images from 0 to 4 steadily improves metrics: FDSR drops from \(0.672\) to \(0.289\), Face Similarity from \(0.276\) to \(0.085\), and FID rises from \(232.9\) to \(473.5\), suggesting more poisoned images prevent perturbation averaging and reduce identity reconstruction reliability. HAA also performs well with detail-rich prompts: for “a close-up photo of sks person, high details”, 2 perturbed images achieve FDSR=0.479 and Image Reward=-0.357. Overall, Tab. 6 supports HAA’s practicality for constrained budgets and predictable protection scaling with perturbed image count.
We study how the perturbation budget affects the generation quality of customized models when applying HAA. Tab. 7 reports results for \(\eta \in \{0.0, 0.02, 0.05, 0.1, 0.15\}\). As \(\eta\) increases, the generated outputs degrade more noticeably (e.g., FDSR drops from \(1.000\) to \(0.000\) and Face Similarity from \(0.498\) to \(0.000\)), indicating stronger suppression of identity cues. At the same time, overly large budgets are less practical because they can visibly affect the protected images. Notably, even under a tight budget of \(\eta=0.02\), HAA remains effective (FDSR \(0.425\), Face Similarity \(0.220\), Image Reward \(-0.428\), FID \(379.8\)), supporting its applicability under realistic constraints.
Tab. 8 evaluates black-box transfer to SD-V3.0 and FLUX on CelebA-HQ using Face Similarity, Image Reward, and FID. On SD-V3.0, HAA attains the lowest Face Similarity (\(0.343\)) and the lowest Image Reward (\(0.561\)) among the compared methods, while its FID (\(197.0\)) is higher than CAAT and SimAC. On FLUX, HAA again yields the lowest Face Similarity (\(0.375\)) and the lowest Image Reward (\(0.336\)), and it also produces the highest FID (\(148.5\)).
Overall, these results indicate that HAA remains effective under black-box transfer to recent generative models, but the absolute transfer strength is weaker than what we observe on SD-v1.4/v1.5 in our other experiments. A plausible explanation is architectural mismatch: our surrogate is built on SD-2.1 (similar to SD-v1.4/v1.5), which uses a U-Net denoiser, whereas SD-V3.0 and FLUX adopt transformer-based denoisers. This shift likely reduces gradient and feature alignment across models, thereby weakening perturbation transfer. We also observe a consistent trade-off: stronger identity suppression and lower preference-aligned quality (lower Face Similarity and Image Reward) coincide with larger distributional deviation (higher FID), suggesting that on these newer backbones, protection gains are more tightly coupled with perceptible degradation in generation quality under our evaluation protocol.
Tab. ¿tbl:tab:def? evaluates HAA’s robustness to defensive processing with a \(\eta=0.02\) budget. Compared to clean images (FDSR \(=1.000\), Face Similarity \(=0.498\), FID \(=112.3\)), HAA degrades identity and quality signals: FDSR drops to \(0.425\), Face Similarity to \(0.220\), Image Reward becomes negative (\(-0.428\)), and FID increases to \(379.8\). We first test common non-adaptive edits, including Gaussian blur and JPEG compression. Although these operations weaken the perturbation and increase FDSR/Face Similarity, Image Reward remains lower and FID remains higher than the clean baseline, indicating that HAA retains a protective effect under common image processing.
We further evaluate DiffPure as a purification-based post-processing defense. The experimental results indicate that DiffPure weakens HAA’s protection in this setting, increasing FDSR to \(0.906\) and Face Similarity to \(0.375\). Nevertheless, it does not fully restore the generation behavior under the clean-image condition: Face Similarity remains lower than that of the clean setting (\(0.375\) vs.\(0.498\)), Image Reward is lower (\(0.451\) vs.\(0.599\)), and FID remains substantially higher (\(235.8\) vs.\(112.3\)). Therefore, under our experimental setting, the protective effect of HAA remains observable.
Fig. 10 reports the relationship between the anti-aesthetic score (\(r_i^g\) in Eq. 3 ; higher values indicate stronger aesthetic degradation), Face Similarity, and FID on VGGFace2 and CelebA-HQ. The samples cover four prompts and six comparison methods.
The anti-aesthetic score is negatively correlated with Face Similarity and positively correlated with FID, suggesting that lower preference-aligned quality often coincides with weaker identity preservation. This correlation motivates HAA but does not by itself rule out generic image degradation as a confounder. We therefore provide the controlled analyses in the following subsection to separate the proposed aesthetic mechanism from generic low-level distortion.
Fig. 11 visualizes the perturbations generated by different protection methods on the CelebA-HQ dataset, complemented by quantitative imperceptibility metrics (LPIPS [43] and SSIM [44]) to objectively evaluate the visual naturalness of protected images. Compared with Mist, CAAT, ASPL, and SimAC, the HAA method introduces fewer visually salient artifacts while still providing effective protection, indicating that it achieves a more favorable balance between preserving the natural appearance of shared images and reducing their utility for malicious customization.
Fig. 12 further compares the qualitative protection outcomes on VGGFace2 for two identities under four prompts. Across prompts, models fine-tuned on HAA-protected images tend to generate faces with substantially weaker identity-consistent details. In contrast, competing methods more often retain recognizable facial structure, which visually aligns with their higher identity leakage in our quantitative evaluation. Moreover, various figures in Appendix visually show the effectiveness of HAA in protecting individuals across different genders and ethnicities.
We further evaluate HAA under challenging input conditions and discuss the fallback mechanism used when face detection fails.
Robustness under extreme scenarios. We evaluate common input variations, including occlusion, rotation, and motion blur. As shown in Fig. 13, HAA continues to reduce identity-consistent generation under these conditions, indicating that the perturbation remains effective beyond the standard clean-input setting.
We further examine whether HAA’s identity-removal effect is attributable to generic image degradation or to the proposed aesthetic mechanism. We use four complementary analyses: quantitative and qualitative LPIPS-controlled comparisons with generic degradations, a reward-control study, and ArcFace Grad-CAM visualization.
Tab. ¿tbl:tab:degradation95control? compares HAA with five generic degradations whose LPIPS values are all higher than HAA’s (0.709–0.725 vs. 0.706). Despite larger perceptual distortion, these degradations retain substantially more identity evidence: Defocus Blur, Gaussian Blur, and Resolution Degradation still yield FDSR \(=1.000\), while Gaussian Noise and Salt-Pepper Noise retain much higher Face Similarity (0.482 and 0.476). By contrast, HAA reduces FDSR to 0.281 and Face Similarity to 0.117.
The first two rows of Fig. 14 provide a qualitative LPIPS-controlled comparison. Generic degradations introduce blur, resolution loss, or noise, but still preserve recognizable facial layouts and local identity details. HAA causes stronger identity-oriented disruption, consistent with Tab. ¿tbl:tab:degradation95control?.
Tab. ¿tbl:tab:reward95control? assesses aesthetic-mechanism-driven optimization. Replacing the aesthetic reward with a random reward does not improve over the baseline (FDSR 0.844 vs. 0.719; Face Similarity 0.394 vs. 0.388), whereas HAA achieves the best privacy metrics. Thus, meaningful aesthetic feedback, rather than a random reward signal, provides important evidence for disrupting identity-relevant cues.
The third row of Fig. 14 shows ArcFace Grad-CAM responses with respect to the original identity. Resolution Degradation remains bright around key facial regions, indicating residual original-identity evidence. HAA exhibits almost no high-response regions, indicating a larger distance from the original identity in ArcFace embedding space. Overall, generic degradation lacks an effective design for suppressing identity cues, whereas HAA achieves stronger de-identification through its aesthetics-driven mechanism.
| Method | VRAM (MiB)↓ | Training Time (s)↓ | FDSR↓ | FID↑ |
|---|---|---|---|---|
| ASPL | 20417 | 152 | 0.750 | 359.5 |
| SimAC | 30327 | 327 | 0.438 | 388.9 |
| HAA | 30647 | 486 | 0.281 | 471.1 |
Tab. 9 compares computational cost and VRAM usage on CelebA-HQ. HAA consumes \(30647\) MiB VRAM, which is essentially on par with SimAC (\(30327\) MiB), suggesting that the hierarchical anti-aesthetic design does not materially increase the memory footprint. Although HAA increases training time from 327 s to 486 s, it delivers a noticeably larger gain in protection: FDSR decreases from \(0.438\) to \(0.281\) and FID increases from \(388.9\) to \(471.1\). Given our evaluation protocol, where lower FDSR and higher FID indicate stronger protection, these results suggest that HAA offers a favorable efficiency-effectiveness trade-off. It enhances privacy protection while incurring only modest additional VRAM usage and a relatively small increase in computational cost, which may be acceptable for resource-constrained deployments that prioritize protection strength.
Since \(RM_g\) and \(RM_l\) provide the optimization feedback for HAA, we validate them along three axes: effectiveness, identity specificity, and robustness/reliability.
We further test downstream utility using the random-reward control in Tab. ¿tbl:tab:reward95control?. Replacing the aesthetic reward with a random signal does not improve protection over the baseline, whereas HAA substantially reduces FDSR and Face Similarity. Thus, HAA benefits from structured aesthetic feedback rather than arbitrary rewards.


Figure 15: Identity-specificity diagnostics of \(RM_l\). Top: different identities with similar aesthetic quality, where each tile reports the \(RM_l\) score / absolute difference from the Aesthetic Scorer. Bottom: same-identity groups under background and expression changes, where each tile reports the mean / standard deviation..
Overall, these results evaluate the reward models at the scorer, downstream-optimization, and diagnostic-behavior levels, supporting their use as frozen preference scorers in HAA.
Both reward models use a BLIP-style preference-scorer architecture with a ViT-L image encoder, a 12-layer transformer text encoder, cross-modal fusion, and an MLP scalar head. For \(RM_g\), we use the public ImageRewardDB dataset for prompt-conditioned image-level rankings. The preference data contain 8,878 prompts and 136,892 pairwise comparisons, constructed from ranked groups of 4–9 generated images for each prompt. The reported preference-accuracy test split contains 466 prompts and 6,399 comparison pairs, while Recall/Filter are evaluated on another 371 prompts with 8 images per prompt; the remaining annotated prompts are used for training. Given a prompt \(C_g\), a preferred image \(x^w\), and a less preferred image \(x^l\), the global scorer is optimized with a pairwise logistic ranking loss, i.e., it is encouraged to satisfy \(RM_g(x^w,C_g)>RM_g(x^l,C_g)\). This makes \(RM_g\) sensitive to global human-preference factors such as prompt alignment, fidelity, and overall aesthetics.
For \(RM_l\), we construct face-level preference pairs and apply the same preference-scorer principle at the face level. Facial regions are detected from LAION natural images, masked, and inpainted with a diffusion inpainting pipeline to create degraded face regions. The original face crop is treated as \(x^w_l\), while the inpainted/degraded face is treated as \(x^l_l\), yielding about 46k face-quality pairs from 23k natural images. The local prompt is fixed to the face-level prompt \(C_l\), so the scorer focuses on facial plausibility and aesthetic detail rather than full-image context. \(RM_l\) is optimized with the same pairwise logistic ranking objective, encouraging \(RM_l(x^w_l,C_l)>RM_l(x^l_l,C_l)\). No identity labels are used in this training; the supervision is quality/preference ordering. After training, both \(RM_g\) and \(RM_l\) are frozen and only provide differentiable feedback for HAA perturbation optimization. We also report the training protocol: 70% of BLIP Transformer layers are frozen, the learning rate is set to \(1\times 10^{-5}\), the batch size is 64, a cosine learning-rate schedule is used, and training is conducted on four RTX 5880 Ada GPUs. Training can also be conducted on a single RTX 5880 Ada GPU, although this increases the training time.
Existing anti-customization methods often overlook key aesthetic cues, limiting their ability to remove identities. To address this issue, we propose a novel Hierarchical Anti-Aesthetics (HAA) framework, which includes both Global Anti-Aesthetics and Local Anti-Aesthetics branches. These branches work together to leverage aesthetic cues from global to local and encourage anti-alignment with human aesthetic preferences, thereby reducing customized DMs’ ability to reconstruct facial details. Extensive experiments show that HAA outperforms existing methods and provides a useful tool for protecting facial privacy and copyright.
Songping Wang, Yueming Lyu, Shiqi Liu, Chen Zhao, Ziyuan Chen, Caifeng Shan are with the School of Intelligence Science and Technology, Nanjing University, Suzhou 215163, China (email: theone@buaa.edu.cn).↩︎
Ning Li is with the China Mobile Information Technology Co., Ltd., Beijing 100037, China.↩︎
Jing Dong is with the State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Center for Research on Intelligent Perception and Computing (CRIPAC), Institute of Automation Chinese Academy of Sciences (CASIA), Beijing 100190, China.↩︎
\(^\dagger\)Corresponding author↩︎